---
title: "Serve LLMs with vLLM — Choose a Recipe"
canonical: "https://build.nvidia.com/spark/vllm/choose-recipe.md"
---

# Step 1. Identify your configuration

Choose the configuration that matches the hardware you will use:

| Configuration | How vLLM will run |
| --- | --- |
| **One DGX Spark** | On the GB10 GPU in one device |
| **One DGX Station** | On the GB300 GPU in one device |
| **Two DGX Sparks** | Across one GB10 GPU in each device |

If you are unsure which NVIDIA GPU is available, connect to the device with NVIDIA Sync, open **Terminal**, and run:

```bash
nvidia-smi --query-gpu=name --format=csv,noheader
```

# Step 2. Use the recommended recipe

Select the row for your configuration. These recipes were chosen to provide reasoning and tool calling while making effective use of the available hardware.

| Configuration | Recommended model | Why this recipe | Continue with |
| --- | --- | --- | --- |
| **One DGX Spark** | [Qwen3.8-27B NVFP4](https://recipes.vllm.ai/Qwen/Qwen3.8-27B?hardware=dgx_spark_gb10) | The quantized model fits one Spark and has a hardware-specific vLLM configuration. | **Single device** |
| **One DGX Station** | [Qwen3.8-Flash-Next NVFP4](https://recipes.vllm.ai/Qwen/Qwen3.8-Flash-Next?hardware=dgx_station_gb300) | The larger mixture-of-experts model takes advantage of the Station's greater memory and uses a dedicated container. | **Single device** |
| **Two DGX Sparks** | [Qwen3.8-27B NVFP4](https://github.com/eugr/spark-vllm-docker/blob/main/recipes/qwen3.8-27b-nvfp4-dflash2.yaml) | The predefined cluster recipe uses tensor parallelism across both Sparks and adds DFlash2 speculative decoding. | **Multi-node DGX Spark** |

The launch tabs provide the complete model, container, and serving configuration for these recommended recipes. You do not need to translate the recipe commands yourself.

# Step 3. Decide whether to use another recipe

For the simplest path, keep the recommended recipe and continue to the tab shown in the table.

If you are comfortable selecting a different model, container, and launch configuration, continue to Step 4.

> [!IMPORTANT]
> The copy-and-paste launch configurations in this playbook are tested for the recommended recipes. Another recipe may require different model-download, container, environment, memory, parser, or parallelism settings.

# Step 4. Generate another recipe

To use another recipe:

1. Open the filtered recipe catalog for your hardware:
- [DGX Spark recipes](https://recipes.vllm.ai/browse?panel=open&hw=dgx_spark_gb10)
- [DGX Station recipes](https://recipes.vllm.ai/browse?panel=open&hw=dgx_station_gb300)
2. Select a model that supports your hardware configuration.
3. Select the model variant and precision.
4. Enable the capabilities you need, such as tool calling or reasoning.
5. Copy the model ID, container image, environment variables, and complete `vllm serve` command from the generated recipe.
6. Keep all values from the same generated recipe. Do not combine settings from different model variants or hardware configurations.
7. For two-Spark serving, confirm that a matching recipe exists in [`spark-vllm-docker`](https://github.com/eugr/spark-vllm-docker/tree/main/recipes). Do not assume a single-device recipe can run across two devices.

> [!NOTE]
> Model size is not the only compatibility requirement. The container architecture, vLLM version, quantization format, parsers, and parallel configuration must also match the selected model and hardware.

# Step 5. Continue to the launch instructions

- For a recommended one-Spark or one-Station recipe, continue with **Single device**.
- For the recommended two-Spark recipe, continue with **Multi-node DGX Spark**.
- For another recipe, use its generated installation and serving commands as the manual path. Substitute its settings only where the following launch instructions explicitly tell you to do so.