Skip to main content
NVIDIA
Explore
Models
Skills
Blueprints
GPUs
Docs
Help Center
Getting Started
  1. Create and verify your account to unlock full access to NVIDIA NIM APIs.
ResourcesDeveloper ForumsContact Support
FAQs
  • Serve LLMs with vLLM

    30 MIN

    High-throughput serving for 30+ models, with continuous batching and an OpenAI-compatible API

    • DGX Spark
    • DGX Station
    • Inference
    • RTX PRO
    • vLLM
    View on GitHub
    OverviewOverviewInstructionsInstructionsAgent-ready ModelsAgent-ready ModelsMulti-node servingMulti-node servingTroubleshootingTroubleshooting

    Agent-ready models

    Agent-ready models are tuned for agentic workloads — tool calling, reasoning traces, and long multi-turn sessions. Use this tab to pick a recommended model for your hardware platform and follow the launch guidance below.

    Follow the launch recipe below for your platform to set up the recommended model.

    Recommendations by hardware platform

    Hardware platformRecommended agent-ready modelMODEL_HANDLERecipe
    DGX SparkAgent-ready Qwen3.6-35B-A3B (NVFP4)nvidia/Qwen3.6-35B-A3B-NVFP4Launch recipe
    DGX StationDeepSeek-V4-Flashdeepseek-ai/DeepSeek-V4-FlashLaunch recipe
    RTX PROQwen3.6 27Bnvidia/Qwen3.6-27B-NVFP4Launch recipe

    Verify your server

    After you launch a recipe above, confirm startup using Watch startup in the Instructions tab, then test with Step 5. Test the API.

    Next steps

    • General serving workflow: Docker setup, health checks, and API testing — see the Instructions tab
    • Scale out: multi-node serving on multi-node capable hardware — see the Multi-node serving tab

    Resources

    • vLLM Recipes
    • vLLM Recipes — DGX Spark
    • vLLM Recipes — DGX Station
    • vLLM Recipes — RTX PRO
    • vLLM Documentation
    • NGC vLLM Container
    • DGX Spark Documentation
    • DGX Spark Forum
    • DGX Station Support
    • RTX PRO Support
    • NVIDIA Developer Forums
    • DGX Spark User Performance Guide
    Terms of Use
    Privacy Policy
    Your Privacy Choices
    Contact

    Copyright © 2026 NVIDIA Corporation