---
title: "Serve LLMs with vLLM — Single device"
canonical: "https://build.nvidia.com/rtx/vllm/instructions.md"
---

## Step 1. Install NVIDIA Sync locally and add the DGX Spark or DGX Station device (one time)

Follow the [Connect to Your Spark](https://build.nvidia.com/spark/connect-to-your-spark) playbook.

# Step 2. Open a remote terminal to check Docker and the Hugging Face CLI (one time)

**First, use NVIDIA Sync to open a terminal on the remote device.**

1. On your computer, open NVIDIA Sync and select the device.
3. Then select **Connect**.
4. After the device connects, open **Terminal**.

**Next, check that your user is in the Docker group.**

```bash
docker ps > /dev/null
```

If it returns a blank line, then the group is already configured.
If it reports a permission error, add your user to the `docker` group as follows:

1. Add your user: `sudo usermod -aG docker "$USER"`
2. Activate the group: `newgrp docker`
3. Check status again: `docker ps > /dev/null`

**Success**: It returns a blank line.

**Finally, check that the Hugging Face CLI is installed and authenticated.**

```bash
hf auth whoami
```

If this returns `command not found` or `Not logged in`, then install ([see here](https://huggingface.co/docs/huggingface_hub/main/en/guides/cli#standalone-installer-recommended)) and/or authenticate the CLI ([see here](https://huggingface.co/docs/huggingface_hub/main/en/quick-start#authentication)).

> [!NOTE]
> If you install the CLI, you need to refresh the terminal with the command `source "$HOME/.bashrc"`.

If `hf` is still not found after refreshing the terminal, add the local user binary directory to your current `PATH`:

```bash
export PATH="$HOME/.local/bin:$PATH"
```

**Success**: The command `hf auth whoami` returns your username and org.

# Step 3. Download the vLLM container and model for the recipe (one time)

**First, pull the appropriate container and follow progress in the terminal.**

| Device | Docker Command to Run |
| --- | --- |
| DGX Spark | `docker pull vllm/vllm-openai:qwen38` |
| DGX Station | `docker pull vllm/vllm-openai:qwen38-flash-next` |

**Success**: The Docker CLI reports that the download has succeeded.

**Then, download the appropriate model and follow progress in the terminal**.

| Device | HF CLI Command to Run |
| --- | --- |
| Single DGX Spark | `hf download nvidia/Qwen3.8-27B-NVFP4 --cache-dir "$HOME/.cache/huggingface/hub"` |
| Single DGX Station | `hf download Inferact/Qwen3.8-Flash-Next-NVFP4 --cache-dir "$HOME/.cache/huggingface/hub"` |

**Success**: The command output will stop and print a path under `$HOME/.cache/huggingface/hub`.

# Step 4. Add the vLLM launch script as a custom app (one time)

**First, open the launch script for your device.**

| Device | Launch Script to Copy |
| --- | --- |
| Single DGX Spark | [open the DGX Spark Qwen3.8 27B script](https://raw.githubusercontent.com/NVIDIA/dgx-spark-playbooks/refs/heads/main/nvidia/playbook-vllm/assets/sync-vllm-single-spark.sh) |
| Single DGX Station | [open the DGX Station Qwen3.8 Flash script](https://raw.githubusercontent.com/NVIDIA/dgx-spark-playbooks/refs/heads/main/nvidia/playbook-vllm/assets/sync-vllm-single-station.sh) |

**Next, create the custom app in NVIDIA Sync.**

1. Select **Custom > Add New** to open the form.
2. Name the app, e.g "vLLM-Qwen3.827B" for DGX Spark or "vllm-Qwen3.8Flash" for DGX Station.
3. Enter **8000** for the port. If it is already in use, follow [Use another API port](troubleshooting.md#use-another-api-port) before copying the script.
4. Leave **Auto open in browser** turned off because vLLM serves an API, not a web app.
5. Copy the appropriate script into the **Launch Script** field.
6. Select **Add**.

**Success**: The application name shows up in the **Custom** section.

# Step 5. Launch vLLM and watch the model load (repeat use)

1. In NVIDIA Sync, select the custom app you added in Step 4.
2. Then open **Resource Monitor** to watch GPU activity and memory use as the model loads.
3. Follow the container logs in an NVIDIA Sync terminal with the command for your device:

| Device | Docker Command to See Logs |
| --- | --- |
| DGX Spark | `docker logs --tail 50 --follow vllm-qwen38` |
| DGX Station | `docker logs --tail 50 --follow vllm-qwen38-flash` |

**Success**: The logs show `OpenAI server is ready to accept requests` or `Application startup complete`. You can close the log terminal without stopping vLLM.

> [!IMPORTANT]
> GPU activity in the Resource Monitor shows that vLLM is working but does not mean the API is ready. Wait for a readiness message before you test the endpoint.

# Step 6. Once vLLM is ready, test the API from your laptop (one time)

> [!IMPORTANT]
> Run these commands on the laptop where NVIDIA Sync is running, not in the remote terminal. On Windows, use PowerShell, not a WSL terminal.
> Use the custom app port in every URL. The examples use 8000; if you assigned 8001, change every `localhost:8000` to `localhost:8001`.

**First, check the endpoint health and list the models.**

| Laptop terminal | Health check command | List models command |
| --- | --- | --- |
| Windows (PowerShell) | `curl.exe -i http://localhost:8000/health` | `curl.exe -sS http://localhost:8000/v1/models` |
| macOS or Linux (terminal) | `curl -i http://localhost:8000/health` | `curl -sS http://localhost:8000/v1/models` |

**Success**: The health check returns HTTP `200`, and the model list shows the model for your device.

**Next, send a chat request using the command for your laptop terminal and DGX device.**

| Laptop terminal | DGX Spark | DGX Station |
| --- | --- | --- |
| Windows (PowerShell) | `Invoke-RestMethod -Uri http://localhost:8000/v1/chat/completions -Method Post -ContentType application/json -Body '{"model":"nvidia/Qwen3.8-27B-NVFP4","messages":[{"role":"user","content":"Write a haiku about a GPU."}],"max_tokens":4096}' \| ConvertTo-Json -Depth 8` | `Invoke-RestMethod -Uri http://localhost:8000/v1/chat/completions -Method Post -ContentType application/json -Body '{"model":"Inferact/Qwen3.8-Flash-Next-NVFP4","messages":[{"role":"user","content":"Write a haiku about a GPU."}],"max_tokens":4096}' \| ConvertTo-Json -Depth 8` |
| macOS or Linux (terminal) | `curl -sS http://localhost:8000/v1/chat/completions -H "Content-Type: application/json" -d '{"model":"nvidia/Qwen3.8-27B-NVFP4","messages":[{"role":"user","content":"Write a haiku about a GPU."}],"max_tokens":4096}'` | `curl -sS http://localhost:8000/v1/chat/completions -H "Content-Type: application/json" -d '{"model":"Inferact/Qwen3.8-Flash-Next-NVFP4","messages":[{"role":"user","content":"Write a haiku about a GPU."}],"max_tokens":4096}'` |

**Success**: The response should contain a `choices` array and the model's answer.

# Next steps

- To serve the model across two DGX Sparks, go to [Multi-node DGX Spark](multi-node-spark.md).
- To choose another model, return to [Choose a Recipe](choose-recipe.md).
- If vLLM does not start or respond, see [Troubleshooting](troubleshooting.md).