Skip to main content
Explore
Models
Skills
Blueprints
GPUs
Docs
Search
⌘K
Ctrl+K
?
Help Center
Getting Started
1
Set up your account
Create and verify your account to unlock full access to NVIDIA NIM APIs.
Create an Account
2
Generate API Key
3
Make your first API call
4
Prototype in your environment
5
Connect to inference partners
Resources
Developer Forums
Contact Support
FAQs
Login
Models
Deploy and scale models on your GPU infrastructure of choice with NVIDIA NIM inference microservices
Optimized by NVIDIA
Launch from Hugging Face
Beta
Filters (1)
4 models
Sort By
Most Recent
Select item
Most Recent
Most Popular
Most Downloaded
Alphabetical (A-Z)
Alphabetical (Z-A)
Most Recent
Most Popular
Most Downloaded
Alphabetical (A-Z)
Alphabetical (Z-A)
models
DeepSeek AI
Free Endpoint
deepseek-v4-flash-0731
284B MoE (13B active) model ideal for long-context workloads optimized for coding, chat, and agentic workflows
MoE
+3
Reasoning
Long Context
Hybrid Attention
7d
Last updated on August 19, 2026
Items per page
24
12
24
48
96
1
1
of 1 pages
NVIDIA
Downloadable
Free Endpoint
nemotron-3-ultra-550b-a55b
Open, efficient hybrid Mamba-Transformer MoE with 1M context, excelling in agentic reasoning, coding, planning, tool calling, and more
Agent
+4
MoE
Frontier
Reasoning
Long Context
52M
52M API calls in the last 30 days
2mo
Last updated on June 4, 2026
NVIDIA
Downloadable
Free Endpoint
nemotron-3-super-120b-a12b
Open, efficient hybrid Mamba-Transformer MoE with 1M context, excelling in agentic reasoning, coding, planning, tool calling, and more
MoE
+4
Reasoning
Chat
Long Context
Instruction Following
65M
65M API calls in the last 30 days
5mo
Last updated on March 11, 2026
NVIDIA
Downloadable
Free Endpoint
nemotron-3-nano-30b-a3b
Open, efficient MoE model with 1M context, excelling in coding, reasoning, instruction following, tool calling, and more
MoE
+3
Reasoning
Long Context
Instruction Following
12M
12M API calls in the last 30 days
8mo
Last updated on December 15, 2025