---
title: "Serve LLMs with vLLM — Troubleshooting"
canonical: "https://build.nvidia.com/rtx/vllm/troubleshooting.md"
---

# Common issues

The **Hardware platform** column shows where an issue is most relevant. "All hardware platforms" applies to every platform.

| Symptom | Hardware platform | Cause | Fix |
|---------|-------------------|-------|-----|
| "permission denied" when running docker | All hardware platforms | User not in docker group | Run `sudo usermod -aG docker $USER && newgrp docker` |
| Container fails to start with GPU error | All hardware platforms | NVIDIA Container Toolkit not configured | Run `nvidia-ctk runtime configure --runtime=docker` and restart Docker |
| HuggingFace authentication failure, gated model access denied, or model download hangs/fails | All hardware platforms | Missing/invalid token, restricted model access, or network issue | Export `HF_TOKEN` before running docker; regenerate your [HuggingFace token](https://huggingface.co/docs/hub/en/security-tokens) and request access to the [gated model](https://huggingface.co/docs/hub/en/models-gated) if needed; check internet connection and verify the token is valid |
| CUDA out of memory | All hardware platforms | Context length too large / model too big | Reduce `--max-model-len` and `--max-num-seqs`, or lower `--gpu-memory-utilization` |
| Port 8000 is in use | All hardware platforms | Another app is using the port | Follow **Use another API port** below |
| Laptop cannot reach the API, but `/health` returns HTTP `200` on the DGX device | All hardware platforms | NVIDIA Sync is not forwarding the expected port, or the test is running in WSL on Windows | Keep the Sync custom app running and match its port to Docker's left-hand port. Use that port in the laptop URL; on Windows, test in PowerShell, not WSL |
| PowerShell prompts for `Uri` after `curl -i` | All hardware platforms | `curl` is a PowerShell alias for `Invoke-WebRequest` | Use `curl.exe` for the Step 6 health and model checks |
| Chat request returns `The model 'unknown' does not exist` | All hardware platforms | The request omits `model` | Use the Step 6 command for your DGX device; its model ID must match `/v1/models` |
| NGC authentication fails | All hardware platforms | Invalid or missing credentials | Run `docker login nvcr.io` with your NGC API key |
| `rm: cannot remove '.../.cache/huggingface/hub/models--...': Permission denied` | All hardware platforms | The container downloads weights as root into the mounted hub cache, so cached model files are root-owned | Remove with `sudo rm -rf $HOME/.cache/huggingface/hub/"<downloaded model name>"` |
| Memory pressure within capacity | DGX Spark | UMA buffer cache not released | See UMA note below |
| Container startup fails / missing ARM64 image | DGX Spark | Image not built for ARM64 | Use the default NGC image for your hardware platform from the Instructions tab |
| Model runs on wrong GPU | DGX Station | Default GPU selection with two GPUs | Use `--gpus '"device=N"'` to pin the GB300 (`N` from `nvidia-smi`) |
| EngineCore failed / FlashInfer "Buffer overflow when allocating memory for batch_prefill_tmp_v" | DGX Station | CUDA graph capture failure during batch prefill | Use the recommended container image: `nvcr.io/nvidia/vllm:26.01-py3` |
| Chat completion returns `content: null` with `finish_reason: length` | All hardware platforms | `max_tokens` was exhausted during reasoning | Raise `max_tokens` in the request (Step 6 uses `4096`) so the model can finish with a visible answer |

# Use another API port

If port 8000 is in use, you can use 8001 without changing the port inside the container:

1. Set the NVIDIA Sync custom app port to **8001**.
2. In its launch script, change the Docker mapping to `-p 8001:8000`. Leave `vllm serve --port 8000` unchanged.
3. Restart the custom app. In every **Step 6** URL, use `http://localhost:8001` instead of `http://localhost:8000`.

Docker's left-hand port is on the remote device; the right-hand port is inside the container. NVIDIA Sync forwards your laptop's port 8001 to port 8001 on the remote device.

> [!NOTE]
> **Unified memory (UMA).** On hardware platforms with unified memory, GPU and CPU share memory dynamically. Some applications have not yet been updated for UMA, so you may hit memory issues even within capacity. If that happens, manually flush the buffer cache:
> ```bash
> sudo sh -c 'sync; echo 3 > /proc/sys/vm/drop_caches'
> ```

> [!NOTE]
> **Monitoring GPU memory with UMA.** Because of unified memory, `nvidia-smi --query-gpu` memory fields report `N/A`. Use plain `nvidia-smi` instead.