vLLM is a highly performant inference engine for serving large language models through an OpenAI-compatible API.
It uses efficient memory management and continuous batching to increase serving throughput.
You can learn more about vLLM in their blog.
This playbook walks you through two situations for configuring a vLLM container to serve a model.
Single remote device: How to configure and access a model in a vLLM container on a DGX Spark or DGX Station
DGX Spark cluster: How to serve a model across vLLM containers running on two clustered DGX Sparks
The playbook focuses on the recommended paths using NVIDIA Sync.
Advanced users can take the manual path to get under the hood for details.
What you'll accomplish
Choose a model and vLLM recipe for your hardware, launch the model, and send a test request to its OpenAI-compatible API.
How to use NVIDIA Sync Cluster Assistant (see here and the demo video)
How to run shell scripts
Suggested:
How to use the NVIDIA Sync Custom App feature (see here)
How to download and run containers
How to edit and run shell scripts
Supported hardware platforms
Check the table below to confirm which path this playbook supports for your hardware.
Hardware platform
OS
Memory
One device
Clustered Sparks
DGX Spark
DGX OS (Linux)
128 GB Unified Memory
✅
✅ Two devices
DGX Station
DGX OS (Linux)
Large HBM + Grace DRAM
✅
—
NOTE
Cluster Assistant can configure two, three, or four DGX Sparks, but model sharding depends on the model and cluster configuration. For simplicity, this playbook limits cluster serving to two Sparks.
Prerequisites
Hardware requirements
Single device: a DGX Spark or DGX Station device — see Supported hardware platforms matrix above
Estimated time: 30 minutes for one device; longer for a cluster or the first model download
Risk level: Low for one device; medium when using a cluster
Rollback: For one device, stop the vLLM custom application in NVIDIA Sync or stop the container you launched manually. For a cluster, stop vLLM on both devices before deleting or changing the cluster.
Last Updated: 09/14/2026
Reorganized the playbook around the recommended NVIDIA Sync paths for one device and DGX Spark clusters.