---
title: "Serve LLMs with vLLM — Multi-node DGX Spark"
canonical: "https://build.nvidia.com/rtx/vllm/multi-node.md"
---

## Step 1. Install NVIDIA Sync on your laptop and add the two DGX Spark devices

Follow the [Connect to Your Spark](https://build.nvidia.com/spark/connect-to-your-spark) playbook.
Estimate time is 5 minutes. 

## Step 2. Physically connect the two Sparks and configure the ConnectX-7 network using NVIDIA Sync

Follow the [Connect Multiple Sparks](https://build.nvidia.com/playbooks/connect-multiple-sparks) playbook.

It shows you how to connect the devices with a cable or switch and then walks you the NVIDIA Sync Clustering Assistant ([see demo video](https://www.youtube.com/watch?v=MehBUQtb9qM)). 

## Step 2. Check the Docker group on each device and configure if needed

We will be using a vLLM container on both devices.
This simplifies things in a variety of ways, a major way being the elimination of installing and configuring NCCL on the two devices.

Docker commands 

If you have not used Cluster Assistant, follow [Configure Manually](https://build.nvidia.com/playbooks/connect-multiple-sparks/manual) in the Connect Multiple Sparks playbook to set up the two-node cluster: physical cabling, network configuration, passwordless SSH, and connectivity verification.

> **Manual setup only:** the connectivity script writes its SSH key to `~/.ssh/` and fails if the directory does not exist. Run `mkdir -p ~/.ssh && chmod 700 ~/.ssh` on both nodes first if you have never used SSH on them.

## Step 2. Prepare the environment

On the **`head node`** (first node in your cluster), clone the DGX Spark community container repo [**spark-vllm-docker**](https://github.com/eugr/spark-vllm-docker)

```bash
git clone https://github.com/eugr/spark-vllm-docker.git
cd spark-vllm-docker
```

Install `uv` if not present

```bash
curl -LsSf https://astral.sh/uv/install.sh | sh
```

Download the model you want to serve on all nodes in the cluster eg. **nvidia/Qwen3.8-27B-NVFP4**

```bash
./hf-download.sh nvidia/Qwen3.8-27B-NVFP4 -c --copy-parallel
```

## Step 3. Serve the model

To serve the model across the cluster, you can use two options

**Option 1**: Serve using a pre-defined recipe

If a pre-defined recipe exists in the repo then you can run it directly like below. This runs the **nvidia/Qwen3.8-27B-NVFP4** model across a two node cluster.

```bash
./run-recipe.sh recipes/qwen3.8-27b-nvfp4-dflash2.yaml --setup
```

> !NOTE
> See ./run-recipe.sh --help for full usage

**Option 2**: Serve with manual command

If a pre-defined recipe does not exist or if you want to run with your own custom arguments, you can run it like below. 

```bash
./launch-cluster.sh -t vllm/vllm-openai:v0.28.0 \
--earlyoom \
-e VLLM_USE_V2_MODEL_RUNNER="1" \
-e VLLM_FLOAT32_MATMUL_PRECISION=high \
-e PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True \
exec \
vllm serve nvidia/Qwen3.8-27B-NVFP4 \
--host 0.0.0.0 \
--port 8000 \
--trust-remote-code \
--kv-cache-dtype fp8 \
--gpu-memory-utilization 0.8 \
--max-model-len 262144 \
--max-num-seqs 8 \
--max-num-batched-tokens 16384 \
--enable-chunked-prefill \
--async-scheduling \
--enable-prefix-caching \
--speculative-config '{"method":"dflash","model": "z-lab/Qwen3.8-27B-DFlash2", "num_speculative_tokens":8, "draft_tensor_parallel_size": 2}' \
--load-format safetensors \
--reasoning-parser qwen3 \
--tool-call-parser qwen3_xml \
--enable-auto-tool-choice \
--tensor-parallel-size 2
```

> [!NOTE]
> You can specify a different VLLM container image using the -t flag. Eg. eugr/spark-vllm-b12x:latest, vllm/vllm-openai:latest etc.
> --earlyoom helps detect OOM early and kills the VLLM process to avoid system hang due to OOM
> See ./launch-cluster.sh --help for full usage.

## Step 4. Test inference

Run on **`head node`**. If you want to run from an external client, replace `localhost` with head node's reachable IP.

```bash
curl -s http://localhost:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "nvidia/Qwen3.8-27B-NVFP4",
"messages": [
{"role": "user", "content": "Write a haiku about a GPU"}
],
"max_tokens": 64
}'
```

# Next steps

Consider for production:

- Health checks and automatic restarts  
- Log rotation for long-running services  
- Persistent model caching across restarts