High-throughput serving for 30+ models, with continuous batching and an OpenAI-compatible API
Agent-ready models are tuned for agentic workloads — tool calling, reasoning traces, and long multi-turn sessions. Use this tab to pick a recommended model for your hardware platform and follow the launch guidance below.
Follow the launch recipe below for your platform to set up the recommended model.
| Hardware platform | Recommended agent-ready model | MODEL_HANDLE | Recipe |
|---|---|---|---|
| DGX Spark | Agent-ready Qwen3.6-35B-A3B (NVFP4) | nvidia/Qwen3.6-35B-A3B-NVFP4 | Launch recipe |
| DGX Station | DeepSeek-V4-Flash | deepseek-ai/DeepSeek-V4-Flash | Launch recipe |
| RTX PRO | Qwen3.6 27B | nvidia/Qwen3.6-27B-NVFP4 | Launch recipe |
After you launch a recipe above, confirm startup using Watch startup in the Instructions tab, then test with Step 5. Test the API.