Skip to main content
Explore
Models
Skills
Blueprints
GPUs
Docs
Search
⌘K
Ctrl+K
?
Forums
Support
Login
Models
Deploy and scale models on your GPU infrastructure of choice with NVIDIA NIM inference microservices
Optimized by NVIDIA
Launch from Hugging Face
Beta
Filters (1)
5 models
Sort By
Most Recent
Select item
Most Recent
Most Popular
Most Downloaded
Alphabetical (A-Z)
Alphabetical (Z-A)
Most Recent
Most Popular
Most Downloaded
Alphabetical (A-Z)
Alphabetical (Z-A)
NVIDIA
Downloadable
Free Endpoint
nemotron-3-ultra-550b-a55b
Open, efficient hybrid Mamba-Transformer MoE with 1M context, excelling in agentic reasoning, coding, planning, tool calling, and more
Agent
+4
MoE
Frontier
Reasoning
Long Context
Items per page
24
12
24
48
96
1
1
of 1 pages
52M
52M API calls in the last 30 days
1mo
Last updated on June 4, 2026
DeepSeek AI
Downloadable
Free Endpoint
deepseek-v4-flash
DeepSeek V4 Flash is a 284B MoE model with 1M-token context optimized for fast coding and agents.
coding
+3
MoE
fast
agentic
17M
17M API calls in the last 30 days
3mo
Last updated on April 24, 2026
DeepSeek AI
Downloadable
Free Endpoint
deepseek-v4-pro
DeepSeek V4 scales to 1M-token context windows with efficient MoE architecture for coding tasks.
Moe
+3
reasoning
coding
agentic
7M
7M API calls in the last 30 days
3mo
Last updated on April 24, 2026
NVIDIA
Downloadable
Free Endpoint
nemotron-3-super-120b-a12b
Open, efficient hybrid Mamba-Transformer MoE with 1M context, excelling in agentic reasoning, coding, planning, tool calling, and more
MoE
+4
Reasoning
Chat
Long Context
Instruction Following
65M
65M API calls in the last 30 days
4mo
Last updated on March 11, 2026
NVIDIA
Downloadable
Free Endpoint
nemotron-3-nano-30b-a3b
Open, efficient MoE model with 1M context, excelling in coding, reasoning, instruction following, tool calling, and more
MoE
+3
Reasoning
Long Context
Instruction Following
12M
12M API calls in the last 30 days
7mo
Last updated on December 15, 2025