Models
Deploy and scale models on your GPU infrastructure of choice with NVIDIA NIM inference microservices
models
552B MoE, 8B active params with native multimodal support and lower API cost using smaller KV cacheFree Endpointdeepseek-v4.1-flash
284B MoE (13B active) model ideal for long-context workloads optimized for coding, chat, and agentic workflowsDeprecatedFree Endpointdeepseek-v4-flash-0731
Open, efficient hybrid Mamba-Transformer MoE with 1M context, excelling in agentic reasoning, coding, planning, tool calling, and moreDownloadableFree Endpointnemotron-3-ultra-550b-a55b
Open, efficient hybrid Mamba-Transformer MoE with 1M context, excelling in agentic reasoning, coding, planning, tool calling, and moreDownloadableFree Endpointnemotron-3-super-120b-a12b
Items per page
of 1 pages