---
title: "Run Multi-Modal Inference with TensorRT — Troubleshooting"
canonical: "https://build.nvidia.com/playbooks/multi-modal-inference/troubleshooting.md"
---

| Symptom | Cause | Fix |
|---------|-------|-----|
| "CUDA out of memory" error | Insufficient memory for the selected model or precision | Use FP8/FP4 quantization or a smaller model variant |
| "Invalid HF token" error | Missing or expired Hugging Face token | Set a valid token: `export HF_TOKEN=<YOUR_TOKEN>` |
| Cannot access gated repo for URL | Certain Hugging Face models have restricted access | Regenerate your [Hugging Face token](https://huggingface.co/docs/hub/en/security-tokens); request access to the [gated model](https://huggingface.co/docs/hub/en/models-gated#customize-requested-information) in your browser |
| Model download timeouts | Network issues or rate limiting | Retry the command or pre-download models |

> [!NOTE]
> Some hardware platforms use Unified Memory Architecture (UMA), which enables dynamic memory sharing between the GPU and CPU. With many applications still updating to take advantage of UMA, you may encounter memory issues even when within rated capacity. If that happens, manually flush the buffer cache with:
```bash
sudo sh -c 'sync; echo 3 > /proc/sys/vm/drop_caches'
```

For latest known issues, see the documentation linked under **Resources** for your hardware platform.