Skip to main content
NVIDIA
Explore
Models
Skills
Blueprints
GPUs
Docs
Help Center
Getting Started
  1. Create and verify your account to unlock full access to NVIDIA NIM APIs.
ResourcesDeveloper ForumsContact Support
FAQs
  • Serve LLMs with SGLang

    30 MIN

    High-throughput serving with RadixAttention, structured output, and an OpenAI-compatible API

    • DGX Spark
    • DGX Station
    • Inference
    • SGLang
    View on GitHub
    OverviewOverviewInstructionsInstructionsTroubleshootingTroubleshooting

    NOTE

    These instructions target Linux (containerized SGLang). WSL and Windows Native are not applicable to the containerized SGLang workflow at this time.

    Step 1
    Set up Docker permissions

    To manage containers without sudo, add your user to the docker group. Open a terminal and test Docker access:

    docker ps
    

    If you see a permission-denied error, add your user to the docker group (skip if it already works):

    sudo usermod -aG docker $USER
    newgrp docker
    

    Step 2
    Set up environment variables

    Pick a model for your hardware platform (see Overview → Supported models). Set these so the SGLang container can download and serve your model:

    # HuggingFace token (required for gated / private models)
    # Get a token from https://huggingface.co/settings/tokens
    # Leave empty for public models such as Qwen/Qwen3-8B
    export HF_TOKEN=""
    
    # Model to serve (HuggingFace handle)
    # Fast first-run validation default:
    export MODEL_HANDLE="Qwen/Qwen3-8B"
    
    # Maximum context length (prompt + output). Size to your workload and memory.
    export MAX_MODEL_LEN=8192
    
    # SGLang container image (CUDA 13.0 — required for Blackwell)
    export SGLANG_IMAGE="lmsysorg/sglang:latest-cu130"
    

    Example model IDs (MODEL_HANDLE)

    Use any Hugging Face text-generation or chat checkpoint your SGLang build supports. Common starting points:

    Model IDNotes
    Qwen/Qwen3-8BDefault. Dense 8B; fast warmup for validating the workflow end-to-end.
    Qwen/Qwen3.6-35B-A3BQwen3.6 MoE (~3B active); strong quality per GPU hour. Hybrid mamba/SSM — prefix-cache validation in Step 6 does not apply (see note there).
    Qwen/Qwen3.6-27BDense Qwen3.6; higher memory than the MoE row at equal batch settings.
    google/gemma-3-12b-it / google/gemma-3-27b-itGemma 3 instruct variants.
    meta-llama/Llama-3.3-70B-InstructGated on Hugging Face — accept the license before download.
    deepseek-ai/DeepSeek-V2-LiteSmall DeepSeek path used in Spark validation examples.
    Spark NVFP4 / FP8 handlesSee Overview → Supported models → DGX Spark (for example nvidia/Qwen3-32B-FP4); add --quantization modelopt_fp4 for NVFP4.

    Heavyweight MoE (confirm SGLang version + memory before serving):

    Model IDNotes
    deepseek-ai/DeepSeek-V4-FlashLarge local MoE for high-memory hardware platforms; long download; may need lower --mem-fraction-static / --context-length.
    deepseek-ai/DeepSeek-V4-ProLarger V4 variant — only with sufficient memory and a supported SGLang build.

    Step 3
    Pull the SGLang container image

    docker pull "$SGLANG_IMAGE"
    
    # Optional: verify GPU access inside the image
    docker run --rm --gpus all "$SGLANG_IMAGE" nvidia-smi
    

    Step 4
    Start the SGLang server

    Hardware platform launch notes

    Container flags differ slightly by hardware platform. --gpus all is correct on all supported hardware platforms unless noted below. Apply the note for your hardware platform:

    Hardware platformLaunch notes
    DGX StationAdd --ipc host and --cap-add SYS_NICE. --gpus all uses the GB300 when it is the only GPU; if multiple GPUs are present, pin with --gpus '"device=N"' where N is the GB300 index from nvidia-smi --query-gpu=index,name --format=csv,noheader. Use --attention-backend flashinfer (required on GB300 / SM103).
    DGX SparkUnified memory (UMA). Prefer a slightly lower --mem-fraction-static (for example 0.75) under memory pressure. If you hit memory issues even within capacity, flush the buffer cache (see Troubleshooting).

    Identify the GPU index when needed (DGX Station with more than one GPU):

    nvidia-smi --query-gpu=index,name --format=csv,noheader
    

    Base configuration (most models)

    Recommended starting point for any model that fits in memory on a single node:

    docker run -d \
      --name sglang-server \
      --gpus all \
      --ipc host \
      --cap-add SYS_NICE \
      --ulimit memlock=-1 \
      --ulimit stack=67108864 \
      -p 30000:30000 \
      -e HF_TOKEN="$HF_TOKEN" \
      -v "$HOME/.cache/huggingface/hub:/root/.cache/huggingface/hub" \
      "$SGLANG_IMAGE" \
      sglang serve --model-path "$MODEL_HANDLE" \
        --host 0.0.0.0 \
        --port 30000 \
        --context-length $MAX_MODEL_LEN \
        --mem-fraction-static 0.85 \
        --attention-backend flashinfer \
        --enable-cache-report \
        --trust-remote-code
    

    Settings used:

    • --context-length — maximum context length per request; larger values reserve more memory for the KV cache.
    • --mem-fraction-static — fraction of GPU memory reserved for weights and KV cache. Use 0.85 as a starting point; lower toward 0.7–0.75 on unified-memory hardware or when hitting OOM.
    • --attention-backend flashinfer — validated attention backend for Blackwell. Prefer this over auto-selected backends that fail CUDA-graph capture on GB300.
    • --enable-cache-report — populates usage.prompt_tokens_details.cached_tokens in OpenAI-style responses for prefix-cache checks.
    • --cap-add SYS_NICE — allows NUMA affinity; avoids repeated permission warnings in logs.
    • --trust-remote-code — required for some model families with custom modeling code.

    NVFP4 models (Spark-validated NVIDIA FP4 checkpoints): add --quantization modelopt_fp4 to the sglang serve arguments.

    Watch startup

    docker logs -f sglang-server
    

    Expected output includes model download (first run), CUDA-graph capture, then readiness messages such as Uvicorn listening on port 30000 and the server ready to accept requests.

    Or wait for the health endpoint:

    timeout 900 bash -c 'until curl -sf http://localhost:30000/health > /dev/null 2>&1; do sleep 10; done' \
      || { echo "Server failed to start within 900s"; docker logs sglang-server | tail -50; exit 1; }
    

    NOTE

    First launch downloads weights and captures CUDA graphs. Plan for ~10–15 min for Qwen/Qwen3-8B and longer for large MoE models before the first successful request. Subsequent starts are faster thanks to cached weights.

    Step 5
    Test the API

    Send a chat completion request:

    curl http://localhost:30000/v1/chat/completions \
      -H "Content-Type: application/json" \
      -d '{
        "model": "'"$MODEL_HANDLE"'",
        "messages": [{"role": "user", "content": "Explain quantum computing in simple terms."}],
        "max_tokens": 256
      }'
    

    The response should contain a choices array with the model's answer.

    Optional native /generate check:

    curl -f -X POST http://localhost:30000/generate \
      -H "Content-Type: application/json" \
      -d '{
          "text": "What does NVIDIA love?",
          "sampling_params": {
              "temperature": 0.7,
              "max_new_tokens": 100
          }
      }'
    

    Step 6
    Multi-turn conversation with prefix caching

    SGLang's RadixAttention caches KV entries for processed tokens. Follow-up messages that share the same conversation prefix reuse those entries and skip repeated prefill for previously seen tokens.

    # Turn 1
    curl -s http://localhost:30000/v1/chat/completions \
      -H "Content-Type: application/json" \
      -d '{
        "model": "'"$MODEL_HANDLE"'",
        "messages": [
          {"role": "system", "content": "You are an expert physics tutor who explains concepts clearly and concisely. You use real-world analogies and everyday examples to make abstract ideas concrete. When answering, first state the key concept in one sentence, then give a short explanation with an example."},
          {"role": "user", "content": "What is the difference between speed and velocity?"}
        ],
        "max_tokens": 256
      }' | python3 -m json.tool
    
    # Turn 2 — extends the same conversation
    curl -s http://localhost:30000/v1/chat/completions \
      -H "Content-Type: application/json" \
      -d '{
        "model": "'"$MODEL_HANDLE"'",
        "messages": [
          {"role": "system", "content": "You are an expert physics tutor who explains concepts clearly and concisely. You use real-world analogies and everyday examples to make abstract ideas concrete. When answering, first state the key concept in one sentence, then give a short explanation with an example."},
          {"role": "user", "content": "What is the difference between speed and velocity?"},
          {"role": "assistant", "content": "Speed is a scalar quantity that measures how fast an object moves, while velocity is a vector quantity that includes both speed and direction. For example, a car driving at 60 km/h has a speed of 60 km/h regardless of where it is headed. But if that car is driving 60 km/h north, that is its velocity — change direction to south and the velocity changes even though the speed stays the same."},
          {"role": "user", "content": "Can you give me another example that shows why the distinction matters in real physics problems?"}
        ],
        "max_tokens": 256
      }' | python3 -m json.tool
    

    Check cache reuse in the server logs:

    docker logs sglang-server 2>&1 | grep "cached-token" | tail -10
    

    Look for #cached-token values greater than 0 on later turns. Treat that as the primary signal of prefix caching; wall-clock curl latency alone can be misleading.

    NOTE

    This prefix-cache check does not apply to hybrid mamba/SSM models such as Qwen/Qwen3.6-35B-A3B. Cross-request prefix reuse is skipped for these architectures and #cached-token / cached_tokens stay 0 even when radix cache is enabled. To validate prefix caching, use a standard-attention model such as Qwen/Qwen3-8B.

    Step 7
    Structured JSON output

    Generate a schema-constrained response:

    curl -s http://localhost:30000/v1/chat/completions \
      -H "Content-Type: application/json" \
      -d '{
        "model": "'"$MODEL_HANDLE"'",
        "messages": [
          {"role": "user", "content": "List three programming languages with their primary use case and year created."}
        ],
        "max_tokens": 512,
        "response_format": {
          "type": "json_schema",
          "json_schema": {
            "name": "languages",
            "schema": {
              "type": "object",
              "properties": {
                "languages": {
                  "type": "array",
                  "items": {
                    "type": "object",
                    "properties": {
                      "name": {"type": "string"},
                      "primary_use": {"type": "string"},
                      "year_created": {"type": "integer"}
                    },
                    "required": ["name", "primary_use", "year_created"]
                  }
                }
              },
              "required": ["languages"]
            }
          }
        }
      }' | python3 -m json.tool
    

    Parse choices[0].message.content — it should be well-formed JSON matching the schema.

    Step 8
    (Optional) Benchmark multi-turn throughput

    This step uses assets/benchmark_multiturn.py, which ships with this playbook. Steps 1–7 run entirely in the container, so clone the playbook repository now if you have not already:

    git clone https://github.com/NVIDIA/dgx-spark-playbooks
    cd dgx-spark-playbooks/nvidia/playbook-sglang
    

    That directory — the one containing assets/ — is the playbook root for the commands below. Run them from there, in a shell where MODEL_HANDLE is exported (re-export it as in Step 2 if you opened a new terminal):

    sudo apt update && sudo apt install -y python3-venv
    python3 -m venv .venv && source .venv/bin/activate
    pip install requests
    
    python3 assets/benchmark_multiturn.py \
      --base-url http://localhost:30000 \
      --model "$MODEL_HANDLE" \
      --num-conversations 20 \
      --turns-per-conversation 5 \
      --cache-detail-file ./sglang_benchmark_cache_details.log
    

    To isolate prefix-cache behavior from multi-client contention, rerun with --num-conversations 1. Always correlate with docker logs (#cached-token lines).

    Step 9
    Stop the container

    docker stop sglang-server
    docker rm sglang-server
    

    Optionally remove the image and cached model. The container downloads weights as root into the mounted hub cache, so the cached model files are root-owned and need sudo to delete:

    docker rmi "$SGLANG_IMAGE"
    sudo rm -rf $HOME/.cache/huggingface/hub/"<downloaded model name>"
    

    Next steps

    • More models: see Overview → Supported models
    • Production: tune --mem-fraction-static, --context-length, and concurrency for your workload
    • Offline inference: see assets/offline-inference.py for an in-process Engine example (clone the repository as shown in Step 8 to get it)

    Resources

    • SGLang Documentation
    • SGLang Cookbook
    • SGLang Supported Models
    • SGLang Cookbook (source)
    • SGLang OpenAI API Reference
    • SGLang (GitHub)
    • DGX Spark Documentation
    • DGX Spark Forum
    • DGX Station Support
    • NVIDIA Developer Forums
    Terms of Use
    Privacy Policy
    Your Privacy Choices
    Contact

    Copyright © 2026 NVIDIA Corporation