High-throughput serving for 30+ models, with continuous batching and an OpenAI-compatible API
Follow the Connect to Your Spark playbook. Estimate time is 5 minutes.
Follow the Connect Multiple Sparks playbook.
It shows you how to connect the devices with a cable or switch and then walks you the NVIDIA Sync Clustering Assistant (see demo video).
We will be using a vLLM container on both devices. This simplifies things in a variety of ways, a major way being the elimination of installing and configuring NCCL on the two devices.
Docker commands
If you have not used Cluster Assistant, follow Configure Manually in the Connect Multiple Sparks playbook to set up the two-node cluster: physical cabling, network configuration, passwordless SSH, and connectivity verification.
Manual setup only: the connectivity script writes its SSH key to
~/.ssh/and fails if the directory does not exist. Runmkdir -p ~/.ssh && chmod 700 ~/.sshon both nodes first if you have never used SSH on them.
On the head node (first node in your cluster), clone the DGX Spark community container repo spark-vllm-docker
git clone https://github.com/eugr/spark-vllm-docker.git
cd spark-vllm-docker
Install uv if not present
curl -LsSf https://astral.sh/uv/install.sh | sh
Download the model you want to serve on all nodes in the cluster eg. nvidia/Qwen3.8-27B-NVFP4
./hf-download.sh nvidia/Qwen3.8-27B-NVFP4 -c --copy-parallel
To serve the model across the cluster, you can use two options
Option 1: Serve using a pre-defined recipe
If a pre-defined recipe exists in the repo then you can run it directly like below. This runs the nvidia/Qwen3.8-27B-NVFP4 model across a two node cluster.
./run-recipe.sh recipes/qwen3.8-27b-nvfp4-dflash2.yaml --setup
!NOTE See ./run-recipe.sh --help for full usage
Option 2: Serve with manual command
If a pre-defined recipe does not exist or if you want to run with your own custom arguments, you can run it like below.
./launch-cluster.sh -t vllm/vllm-openai:v0.28.0 \
--earlyoom \
-e VLLM_USE_V2_MODEL_RUNNER="1" \
-e VLLM_FLOAT32_MATMUL_PRECISION=high \
-e PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True \
exec \
vllm serve nvidia/Qwen3.8-27B-NVFP4 \
--host 0.0.0.0 \
--port 8000 \
--trust-remote-code \
--kv-cache-dtype fp8 \
--gpu-memory-utilization 0.8 \
--max-model-len 262144 \
--max-num-seqs 8 \
--max-num-batched-tokens 16384 \
--enable-chunked-prefill \
--async-scheduling \
--enable-prefix-caching \
--speculative-config '{"method":"dflash","model": "z-lab/Qwen3.8-27B-DFlash2", "num_speculative_tokens":8, "draft_tensor_parallel_size": 2}' \
--load-format safetensors \
--reasoning-parser qwen3 \
--tool-call-parser qwen3_xml \
--enable-auto-tool-choice \
--tensor-parallel-size 2
NOTE
You can specify a different VLLM container image using the -t flag. Eg. eugr/spark-vllm-b12x
Run on head node. If you want to run from an external client, replace localhost with head node's reachable IP.
curl -s http://localhost:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "nvidia/Qwen3.8-27B-NVFP4",
"messages": [
{"role": "user", "content": "Write a haiku about a GPU"}
],
"max_tokens": 64
}'
Consider for production: