High-throughput serving for 30+ models, with continuous batching and an OpenAI-compatible API
Choose the configuration that matches the hardware you will use:
| Configuration | How vLLM will run |
|---|---|
| One DGX Spark | On the GB10 GPU in one device |
| One DGX Station | On the GB300 GPU in one device |
| Two DGX Sparks | Across one GB10 GPU in each device |
If you are unsure which NVIDIA GPU is available, connect to the device with NVIDIA Sync, open Terminal, and run:
nvidia-smi --query-gpu=name --format=csv,noheader
Select the row for your configuration. These recipes were chosen to provide reasoning and tool calling while making effective use of the available hardware.
| Configuration | Recommended model | Why this recipe | Continue with |
|---|---|---|---|
| One DGX Spark | Qwen3.8-27B NVFP4 | The quantized model fits one Spark and has a hardware-specific vLLM configuration. | Single device |
| One DGX Station | Qwen3.8-Flash-Next NVFP4 | The larger mixture-of-experts model takes advantage of the Station's greater memory and uses a dedicated container. | Single device |
| Two DGX Sparks | Qwen3.8-27B NVFP4 | The predefined cluster recipe uses tensor parallelism across both Sparks and adds DFlash2 speculative decoding. | Multi-node DGX Spark |
The launch tabs provide the complete model, container, and serving configuration for these recommended recipes. You do not need to translate the recipe commands yourself.
For the simplest path, keep the recommended recipe and continue to the tab shown in the table.
If you are comfortable selecting a different model, container, and launch configuration, continue to Step 4.
IMPORTANT
The copy-and-paste launch configurations in this playbook are tested for the recommended recipes. Another recipe may require different model-download, container, environment, memory, parser, or parallelism settings.
To use another recipe:
vllm serve command from the generated recipe.spark-vllm-docker. Do not assume a single-device recipe can run across two devices.NOTE
Model size is not the only compatibility requirement. The container architecture, vLLM version, quantization format, parsers, and parallel configuration must also match the selected model and hardware.