High-throughput serving for 30+ models, with continuous batching and an OpenAI-compatible API
Follow the Connect to Your Spark playbook.
First, use NVIDIA Sync to open a terminal on the remote device.
Next, check that your user is in the Docker group.
docker ps > /dev/null
If it returns a blank line, then the group is already configured.
If it reports a permission error, add your user to the docker group as follows:
sudo usermod -aG docker "$USER"newgrp dockerdocker ps > /dev/nullSuccess: It returns a blank line.
Finally, check that the Hugging Face CLI is installed and authenticated.
hf auth whoami
If this returns command not found or Not logged in, then install (see here) and/or authenticate the CLI (see here).
NOTE
If you install the CLI, you need to refresh the terminal with the command source "$HOME/.bashrc".
If hf is still not found after refreshing the terminal, add the local user binary directory to your current PATH:
export PATH="$HOME/.local/bin:$PATH"
Success: The command hf auth whoami returns your username and org.
First, pull the appropriate container and follow progress in the terminal.
| Device | Docker Command to Run |
|---|---|
| DGX Spark | docker pull vllm/vllm-openai:qwen38 |
| DGX Station | docker pull vllm/vllm-openai:qwen38-flash-next |
Success: The Docker CLI reports that the download has succeeded.
Then, download the appropriate model and follow progress in the terminal.
| Device | HF CLI Command to Run |
|---|---|
| Single DGX Spark | hf download nvidia/Qwen3.8-27B-NVFP4 --cache-dir "$HOME/.cache/huggingface/hub" |
| Single DGX Station | hf download Inferact/Qwen3.8-Flash-Next-NVFP4 --cache-dir "$HOME/.cache/huggingface/hub" |
Success: The command output will stop and print a path under $HOME/.cache/huggingface/hub.
First, open the launch script for your device.
| Device | Launch Script to Copy |
|---|---|
| Single DGX Spark | open the DGX Spark Qwen3.8 27B script |
| Single DGX Station | open the DGX Station Qwen3.8 Flash script |
Next, create the custom app in NVIDIA Sync.
Success: The application name shows up in the Custom section.
| Device | Docker Command to See Logs |
|---|---|
| DGX Spark | docker logs --tail 50 --follow vllm-qwen38 |
| DGX Station | docker logs --tail 50 --follow vllm-qwen38-flash |
Success: The logs show OpenAI server is ready to accept requests or Application startup complete. You can close the log terminal without stopping vLLM.
IMPORTANT
GPU activity in the Resource Monitor shows that vLLM is working but does not mean the API is ready. Wait for a readiness message before you test the endpoint.
IMPORTANT
Run these commands on the laptop where NVIDIA Sync is running, not in the remote terminal. On Windows, use PowerShell, not a WSL terminal.
Use the custom app port in every URL. The examples use 8000; if you assigned 8001, change every localhost:8000 to localhost:8001.
First, check the endpoint health and list the models.
| Laptop terminal | Health check command | List models command |
|---|---|---|
| Windows (PowerShell) | curl.exe -i http://localhost:8000/health | curl.exe -sS http://localhost:8000/v1/models |
| macOS or Linux (terminal) | curl -i http://localhost:8000/health | curl -sS http://localhost:8000/v1/models |
Success: The health check returns HTTP 200, and the model list shows the model for your device.
Next, send a chat request using the command for your laptop terminal and DGX device.
| Laptop terminal | DGX Spark | DGX Station |
|---|---|---|
| Windows (PowerShell) | Invoke-RestMethod -Uri http://localhost:8000/v1/chat/completions -Method Post -ContentType application/json -Body '{"model":"nvidia/Qwen3.8-27B-NVFP4","messages":[{"role":"user","content":"Write a haiku about a GPU."}],"max_tokens":4096}' | ConvertTo-Json -Depth 8 | Invoke-RestMethod -Uri http://localhost:8000/v1/chat/completions -Method Post -ContentType application/json -Body '{"model":"Inferact/Qwen3.8-Flash-Next-NVFP4","messages":[{"role":"user","content":"Write a haiku about a GPU."}],"max_tokens":4096}' | ConvertTo-Json -Depth 8 |
| macOS or Linux (terminal) | curl -sS http://localhost:8000/v1/chat/completions -H "Content-Type: application/json" -d '{"model":"nvidia/Qwen3.8-27B-NVFP4","messages":[{"role":"user","content":"Write a haiku about a GPU."}],"max_tokens":4096}' | curl -sS http://localhost:8000/v1/chat/completions -H "Content-Type: application/json" -d '{"model":"Inferact/Qwen3.8-Flash-Next-NVFP4","messages":[{"role":"user","content":"Write a haiku about a GPU."}],"max_tokens":4096}' |
Success: The response should contain a choices array and the model's answer.