High-throughput serving for 30+ models, with continuous batching and an OpenAI-compatible API
The Hardware platform column shows where an issue is most relevant. "All hardware platforms" applies to every platform.
| Symptom | Hardware platform | Cause | Fix |
|---|---|---|---|
| "permission denied" when running docker | All hardware platforms | User not in docker group | Run sudo usermod -aG docker $USER && newgrp docker |
| Container fails to start with GPU error | All hardware platforms | NVIDIA Container Toolkit not configured | Run nvidia-ctk runtime configure --runtime=docker and restart Docker |
| HuggingFace authentication failure, gated model access denied, or model download hangs/fails | All hardware platforms | Missing/invalid token, restricted model access, or network issue | Export HF_TOKEN before running docker; regenerate your HuggingFace token and request access to the gated model if needed; check internet connection and verify the token is valid |
| CUDA out of memory | All hardware platforms | Context length too large / model too big | Reduce --max-model-len and --max-num-seqs, or lower --gpu-memory-utilization |
| Port 8000 is in use | All hardware platforms | Another app is using the port | Follow Use another API port below |
Laptop cannot reach the API, but /health returns HTTP 200 on the DGX device | All hardware platforms | NVIDIA Sync is not forwarding the expected port, or the test is running in WSL on Windows | Keep the Sync custom app running and match its port to Docker's left-hand port. Use that port in the laptop URL; on Windows, test in PowerShell, not WSL |
PowerShell prompts for Uri after curl -i | All hardware platforms | curl is a PowerShell alias for Invoke-WebRequest | Use curl.exe for the Step 6 health and model checks |
Chat request returns The model 'unknown' does not exist | All hardware platforms | The request omits model | Use the Step 6 command for your DGX device; its model ID must match /v1/models |
| NGC authentication fails | All hardware platforms | Invalid or missing credentials | Run docker login nvcr.io with your NGC API key |
rm: cannot remove '.../.cache/huggingface/hub/models--...': Permission denied | All hardware platforms | The container downloads weights as root into the mounted hub cache, so cached model files are root-owned | Remove with sudo rm -rf $HOME/.cache/huggingface/hub/"<downloaded model name>" |
| Memory pressure within capacity | DGX Spark | UMA buffer cache not released | See UMA note below |
| Container startup fails / missing ARM64 image | DGX Spark | Image not built for ARM64 | Use the default NGC image for your hardware platform from the Instructions tab |
| Model runs on wrong GPU | DGX Station | Default GPU selection with two GPUs | Use --gpus '"device=N"' to pin the GB300 (N from nvidia-smi) |
| EngineCore failed / FlashInfer "Buffer overflow when allocating memory for batch_prefill_tmp_v" | DGX Station | CUDA graph capture failure during batch prefill | Use the recommended container image: nvcr.io/nvidia/vllm:26.01-py3 |
Chat completion returns content: null with finish_reason: length | All hardware platforms | max_tokens was exhausted during reasoning | Raise max_tokens in the request (Step 6 uses 4096) so the model can finish with a visible answer |
If port 8000 is in use, you can use 8001 without changing the port inside the container:
-p 8001:8000. Leave vllm serve --port 8000 unchanged.http://localhost:8001 instead of http://localhost:8000.Docker's left-hand port is on the remote device; the right-hand port is inside the container. NVIDIA Sync forwards your laptop's port 8001 to port 8001 on the remote device.
NOTE
Unified memory (UMA). On hardware platforms with unified memory, GPU and CPU share memory dynamically. Some applications have not yet been updated for UMA, so you may hit memory issues even within capacity. If that happens, manually flush the buffer cache:
sudo sh -c 'sync; echo 3 > /proc/sys/vm/drop_caches'
NOTE
Monitoring GPU memory with UMA. Because of unified memory, nvidia-smi --query-gpu memory fields report N/A. Use plain nvidia-smi instead.