---
title: "Serve LLMs with vLLM — Overview"
canonical: "https://build.nvidia.com/rtx/vllm/overview.md"
---

# Basic idea

vLLM is a highly performant inference engine for serving large language models through an OpenAI-compatible API.
It uses efficient memory management and continuous batching to increase serving throughput.
You can learn more about vLLM [in their blog](https://vllm.ai/blog).

This playbook walks you through two situations for configuring a vLLM container to serve a model.

- **Single remote device**: How to configure and access a model in a vLLM container on a DGX Spark or DGX Station
- **DGX Spark cluster**: How to serve a model across vLLM containers running on two clustered DGX Sparks

The playbook focuses on the recommended paths using NVIDIA Sync.
Advanced users can take the manual path to get under the hood for details.

# What you'll accomplish

Choose a model and vLLM recipe for your hardware, launch the model, and send a test request to its OpenAI-compatible API.

**Recommended path:** Use [NVIDIA Sync](https://docs.nvidia.com/sync/latest/index.html)

- One DGX Spark or Station: Use NVIDIA Sync to start/stop the remote container and handle port forwarding for the API
- One time: Download the recommended container and model
- One time: Save the launch script as an NVIDIA Sync custom application
- Repeat use: Start/stop the vLLM container from NVIDIA Sync

- Two DGX Sparks: Use NVIDIA Sync to set up the cluster, then run a recipe from the head Spark to serve the model across both devices
- One time: Physically connect the Sparks and use the Cluster Assistant to configure and test the high-speed network and interdevice SSH
- One time: Prepare the recommended container and model on both Sparks
- Repeat use: Run the multi-node recipe from the head Spark to start vLLM

**Advanced path:** Run the setup and recipe commands directly on the devices when you need to customize the deployment.

# What to know before starting

**Required**:

- Familiarity with [NVIDIA Sync](https://docs.nvidia.com/sync/latest/index.html) and adding devices ([see steps here](https://docs.nvidia.com/sync/latest/direct-connections.html))
- How to run simple terminal commands in Linux
- Basic familiarity with [Hugging Face](https://huggingface.co/) and the [CLI](https://huggingface.co/docs/huggingface_hub/main/en/guides/cli#standalone-installer-recommended)
- If using a DGX Spark cluster:
- How to physically connect the devices ([see here](https://docs.nvidia.com/dgx/dgx-spark/spark-clustering.html#the-qsfp-ports-and-cables))
- How to use NVIDIA Sync Cluster Assistant (see [here](https://build.nvidia.com/playbooks/connect-multiple-sparks) and the [demo video](https://www.youtube.com/watch?v=MehBUQtb9qM))
- How to run shell scripts

**Suggested**:

- How to use the NVIDIA Sync Custom App feature ([see here](https://docs.nvidia.com/sync/latest/applications.html#adding-and-editing-a-custom-script-to-a-remote-device))
- How to download and run containers
- How to edit and run shell scripts

# Supported hardware platforms

Check the table below to confirm which path this playbook supports for your hardware.

| Hardware platform | OS | Memory | One device | Clustered Sparks |
| :---- | :---- | :---- | :----: | :----: |
| **DGX Spark** | DGX OS (Linux) | 128 GB Unified Memory | ✅ | ✅ Two devices |
| **DGX Station** | DGX OS (Linux) | Large HBM + Grace DRAM | ✅ | — |

> [!NOTE]
> Cluster Assistant can configure two, three, or four DGX Sparks, but model sharding depends on the model and cluster configuration. For simplicity, this playbook limits cluster serving to two Sparks.

# Prerequisites

**Hardware requirements**

- Single device: a DGX Spark or DGX Station device — see Supported hardware platforms matrix above
- DGX Spark cluster: Two DGX Sparks and a single QSFP cable ([see here for cabling two devices](https://docs.nvidia.com/dgx/dgx-spark/spark-clustering.html#the-qsfp-ports-and-cables))

**Software requirements**

- NVIDIA Sync installed on your laptop ([see installation instructions here](https://docs.nvidia.com/sync/latest/getting-started.html#installation-and-onboarding))
- Each remote device added to NVIDIA Sync ([see how to do this here](https://docs.nvidia.com/sync/latest/direct-connections.html#adding-a-device-for-a-direct-connection))
- Each remote device has Docker installed and the user in the Docker group ([see here for how](https://docs.docker.com/engine/install/linux-postinstall/#add-your-user-to-the-docker-group))
- The Hugging Face CLI installed and authenticated on each device
- [See here](https://huggingface.co/docs/huggingface_hub/main/en/guides/cli#standalone-installer-recommended) on Hugging Face for CLI installation
- [See here](https://huggingface.co/docs/huggingface_hub/main/en/quick-start#authentication) on Hugging Face for token authentication

# Time & risk

- **Estimated time:** 30 minutes for one device; longer for a cluster or the first model download
- **Risk level:** Low for one device; medium when using a cluster
- **Rollback:** For one device, stop the vLLM custom application in NVIDIA Sync or stop the container you launched manually. For a cluster, stop vLLM on both devices before deleting or changing the cluster.
- **Last Updated:** 09/14/2026
- Reorganized the playbook around the recommended NVIDIA Sync paths for one device and DGX Spark clusters.