High-throughput serving with RadixAttention, structured output, and an OpenAI-compatible API
NOTE
These instructions target Linux (containerized SGLang). WSL and Windows Native are not applicable to the containerized SGLang workflow at this time.
To manage containers without sudo, add your user to the docker group. Open a terminal and test Docker access:
docker ps
If you see a permission-denied error, add your user to the docker group (skip if it already works):
sudo usermod -aG docker $USER
newgrp docker
Pick a model for your hardware platform (see Overview → Supported models). Set these so the SGLang container can download and serve your model:
# HuggingFace token (required for gated / private models)
# Get a token from https://huggingface.co/settings/tokens
# Leave empty for public models such as Qwen/Qwen3-8B
export HF_TOKEN=""
# Model to serve (HuggingFace handle)
# Fast first-run validation default:
export MODEL_HANDLE="Qwen/Qwen3-8B"
# Maximum context length (prompt + output). Size to your workload and memory.
export MAX_MODEL_LEN=8192
# SGLang container image (CUDA 13.0 — required for Blackwell)
export SGLANG_IMAGE="lmsysorg/sglang:latest-cu130"
MODEL_HANDLE)Use any Hugging Face text-generation or chat checkpoint your SGLang build supports. Common starting points:
| Model ID | Notes |
|---|---|
Qwen/Qwen3-8B | Default. Dense 8B; fast warmup for validating the workflow end-to-end. |
Qwen/Qwen3.6-35B-A3B | Qwen3.6 MoE (~3B active); strong quality per GPU hour. Hybrid mamba/SSM — prefix-cache validation in Step 6 does not apply (see note there). |
Qwen/Qwen3.6-27B | Dense Qwen3.6; higher memory than the MoE row at equal batch settings. |
google/gemma-3-12b-it / google/gemma-3-27b-it | Gemma 3 instruct variants. |
meta-llama/Llama-3.3-70B-Instruct | Gated on Hugging Face — accept the license before download. |
deepseek-ai/DeepSeek-V2-Lite | Small DeepSeek path used in Spark validation examples. |
| Spark NVFP4 / FP8 handles | See Overview → Supported models → DGX Spark (for example nvidia/Qwen3-32B-FP4); add --quantization modelopt_fp4 for NVFP4. |
Heavyweight MoE (confirm SGLang version + memory before serving):
| Model ID | Notes |
|---|---|
deepseek-ai/DeepSeek-V4-Flash | Large local MoE for high-memory hardware platforms; long download; may need lower --mem-fraction-static / --context-length. |
deepseek-ai/DeepSeek-V4-Pro | Larger V4 variant — only with sufficient memory and a supported SGLang build. |
docker pull "$SGLANG_IMAGE"
# Optional: verify GPU access inside the image
docker run --rm --gpus all "$SGLANG_IMAGE" nvidia-smi
Container flags differ slightly by hardware platform. --gpus all is correct on all supported hardware platforms unless noted below. Apply the note for your hardware platform:
| Hardware platform | Launch notes |
|---|---|
| DGX Station | Add --ipc host and --cap-add SYS_NICE. --gpus all uses the GB300 when it is the only GPU; if multiple GPUs are present, pin with --gpus '"device=N"' where N is the GB300 index from nvidia-smi --query-gpu=index,name --format=csv,noheader. Use --attention-backend flashinfer (required on GB300 / SM103). |
| DGX Spark | Unified memory (UMA). Prefer a slightly lower --mem-fraction-static (for example 0.75) under memory pressure. If you hit memory issues even within capacity, flush the buffer cache (see Troubleshooting). |
Identify the GPU index when needed (DGX Station with more than one GPU):
nvidia-smi --query-gpu=index,name --format=csv,noheader
Recommended starting point for any model that fits in memory on a single node:
docker run -d \
--name sglang-server \
--gpus all \
--ipc host \
--cap-add SYS_NICE \
--ulimit memlock=-1 \
--ulimit stack=67108864 \
-p 30000:30000 \
-e HF_TOKEN="$HF_TOKEN" \
-v "$HOME/.cache/huggingface/hub:/root/.cache/huggingface/hub" \
"$SGLANG_IMAGE" \
sglang serve --model-path "$MODEL_HANDLE" \
--host 0.0.0.0 \
--port 30000 \
--context-length $MAX_MODEL_LEN \
--mem-fraction-static 0.85 \
--attention-backend flashinfer \
--enable-cache-report \
--trust-remote-code
Settings used:
--context-length — maximum context length per request; larger values reserve more memory for the KV cache.--mem-fraction-static — fraction of GPU memory reserved for weights and KV cache. Use 0.85 as a starting point; lower toward 0.7–0.75 on unified-memory hardware or when hitting OOM.--attention-backend flashinfer — validated attention backend for Blackwell. Prefer this over auto-selected backends that fail CUDA-graph capture on GB300.--enable-cache-report — populates usage.prompt_tokens_details.cached_tokens in OpenAI-style responses for prefix-cache checks.--cap-add SYS_NICE — allows NUMA affinity; avoids repeated permission warnings in logs.--trust-remote-code — required for some model families with custom modeling code.NVFP4 models (Spark-validated NVIDIA FP4 checkpoints): add --quantization modelopt_fp4 to the sglang serve arguments.
docker logs -f sglang-server
Expected output includes model download (first run), CUDA-graph capture, then readiness messages such as Uvicorn listening on port 30000 and the server ready to accept requests.
Or wait for the health endpoint:
timeout 900 bash -c 'until curl -sf http://localhost:30000/health > /dev/null 2>&1; do sleep 10; done' \
|| { echo "Server failed to start within 900s"; docker logs sglang-server | tail -50; exit 1; }
NOTE
First launch downloads weights and captures CUDA graphs. Plan for ~10–15 min for Qwen/Qwen3-8B and longer for large MoE models before the first successful request. Subsequent starts are faster thanks to cached weights.
Send a chat completion request:
curl http://localhost:30000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "'"$MODEL_HANDLE"'",
"messages": [{"role": "user", "content": "Explain quantum computing in simple terms."}],
"max_tokens": 256
}'
The response should contain a choices array with the model's answer.
Optional native /generate check:
curl -f -X POST http://localhost:30000/generate \
-H "Content-Type: application/json" \
-d '{
"text": "What does NVIDIA love?",
"sampling_params": {
"temperature": 0.7,
"max_new_tokens": 100
}
}'
SGLang's RadixAttention caches KV entries for processed tokens. Follow-up messages that share the same conversation prefix reuse those entries and skip repeated prefill for previously seen tokens.
# Turn 1
curl -s http://localhost:30000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "'"$MODEL_HANDLE"'",
"messages": [
{"role": "system", "content": "You are an expert physics tutor who explains concepts clearly and concisely. You use real-world analogies and everyday examples to make abstract ideas concrete. When answering, first state the key concept in one sentence, then give a short explanation with an example."},
{"role": "user", "content": "What is the difference between speed and velocity?"}
],
"max_tokens": 256
}' | python3 -m json.tool
# Turn 2 — extends the same conversation
curl -s http://localhost:30000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "'"$MODEL_HANDLE"'",
"messages": [
{"role": "system", "content": "You are an expert physics tutor who explains concepts clearly and concisely. You use real-world analogies and everyday examples to make abstract ideas concrete. When answering, first state the key concept in one sentence, then give a short explanation with an example."},
{"role": "user", "content": "What is the difference between speed and velocity?"},
{"role": "assistant", "content": "Speed is a scalar quantity that measures how fast an object moves, while velocity is a vector quantity that includes both speed and direction. For example, a car driving at 60 km/h has a speed of 60 km/h regardless of where it is headed. But if that car is driving 60 km/h north, that is its velocity — change direction to south and the velocity changes even though the speed stays the same."},
{"role": "user", "content": "Can you give me another example that shows why the distinction matters in real physics problems?"}
],
"max_tokens": 256
}' | python3 -m json.tool
Check cache reuse in the server logs:
docker logs sglang-server 2>&1 | grep "cached-token" | tail -10
Look for #cached-token values greater than 0 on later turns. Treat that as the primary signal of prefix caching; wall-clock curl latency alone can be misleading.
NOTE
This prefix-cache check does not apply to hybrid mamba/SSM models such as Qwen/Qwen3.6-35B-A3B. Cross-request prefix reuse is skipped for these architectures and #cached-token / cached_tokens stay 0 even when radix cache is enabled. To validate prefix caching, use a standard-attention model such as Qwen/Qwen3-8B.
Generate a schema-constrained response:
curl -s http://localhost:30000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "'"$MODEL_HANDLE"'",
"messages": [
{"role": "user", "content": "List three programming languages with their primary use case and year created."}
],
"max_tokens": 512,
"response_format": {
"type": "json_schema",
"json_schema": {
"name": "languages",
"schema": {
"type": "object",
"properties": {
"languages": {
"type": "array",
"items": {
"type": "object",
"properties": {
"name": {"type": "string"},
"primary_use": {"type": "string"},
"year_created": {"type": "integer"}
},
"required": ["name", "primary_use", "year_created"]
}
}
},
"required": ["languages"]
}
}
}
}' | python3 -m json.tool
Parse choices[0].message.content — it should be well-formed JSON matching the schema.
This step uses assets/benchmark_multiturn.py, which ships with this playbook. Steps 1–7 run entirely in the container, so clone the playbook repository now if you have not already:
git clone https://github.com/NVIDIA/dgx-spark-playbooks
cd dgx-spark-playbooks/nvidia/playbook-sglang
That directory — the one containing assets/ — is the playbook root for the commands below. Run them from there, in a shell where MODEL_HANDLE is exported (re-export it as in Step 2 if you opened a new terminal):
sudo apt update && sudo apt install -y python3-venv
python3 -m venv .venv && source .venv/bin/activate
pip install requests
python3 assets/benchmark_multiturn.py \
--base-url http://localhost:30000 \
--model "$MODEL_HANDLE" \
--num-conversations 20 \
--turns-per-conversation 5 \
--cache-detail-file ./sglang_benchmark_cache_details.log
To isolate prefix-cache behavior from multi-client contention, rerun with --num-conversations 1. Always correlate with docker logs (#cached-token lines).
docker stop sglang-server
docker rm sglang-server
Optionally remove the image and cached model. The container downloads weights as root into the mounted hub cache, so the cached model files are root-owned and need sudo to delete:
docker rmi "$SGLANG_IMAGE"
sudo rm -rf $HOME/.cache/huggingface/hub/"<downloaded model name>"
--mem-fraction-static, --context-length, and concurrency for your workloadassets/offline-inference.py for an in-process Engine example (clone the repository as shown in Step 8 to get it)