Skip to main content
Explore
Models
Skills
Blueprints
GPUs
Docs
Search
⌘K
Ctrl+K
?
Help Center
Getting Started
1
Set up your account
Create and verify your account to unlock full access to NVIDIA NIM APIs.
Create an Account
2
Generate API Key
3
Make your first API call
4
Prototype in your environment
5
Connect to inference partners
Resources
Developer Forums
Contact Support
FAQs
Login
Models
Deploy and scale models on your GPU infrastructure of choice with NVIDIA NIM inference microservices
Optimized by NVIDIA
Launch from Hugging Face
Beta
Filters (1)
14 models
Sort By
Most Recent
Select item
Most Recent
Most Popular
Most Downloaded
Alphabetical (A-Z)
Alphabetical (Z-A)
Most Recent
Most Popular
Most Downloaded
Alphabetical (A-Z)
Alphabetical (Z-A)
models
Moonshotai
Downloadable
Free Endpoint
kimi-k3
~2.8T hybrid KDA+MLA multimodal MoE for long-horizon coding, agentic tool use, and image understanding.
Multimodal
+3
Mixture-of-Experts
Reasoning
Image-to-Text
6d
Last updated on August 27, 2026
Items per page
24
12
24
48
96
1
1
of 1 pages
DeepSeek AI
Free Endpoint
deepseek-v4-pro-0813
DeepSeek V4 scales to 1M-token context windows with efficient MoE architecture for coding tasks.
coding
+3
Moe
reasoning
agentic
7d
Last updated on August 26, 2026
DeepSeek AI
Free Endpoint
deepseek-v4-flash-0731
284B MoE (13B active) model ideal for long-context workloads optimized for coding, chat, and agentic workflows
MoE
+3
Reasoning
Long Context
Hybrid Attention
15d
Last updated on August 19, 2026
Meta
Downloadable
Free Endpoint
muse-glimmer-30b
Muse Glimmer 30B is a multimodal reasoning model accepting text and images, with native tool-calling and separate reasoning output.
Multimodal
+5
Image-to-Text
Reasoning
Chat
Text-to-Text
Large Language Models
23d
Last updated on August 10, 2026
Poolside
Free Endpoint
laguna-xs-2.1
Efficient 33B MoE for local, long-horizon agentic coding and terminal tasks
Agentic AI
+3
Coding
Reasoning
Tool Use
1mo
Last updated on July 15, 2026
Minimaxai
Deprecation in 7d
Free Endpoint
minimax-m3
MiniMax M3 Preview is a multimodal MoE vision-language model with strong reasoning, coding, and tool-calling capabilities.
coding
+2
text-to-text
reasoning
10M
10M API calls in the last 30 days
2mo
Last updated on June 12, 2026
Google
Downloadable
Free Endpoint
diffusiongemma-26b-a4b-it
Diffusion-based 26B parameter LLM enabling parallel token generation for real-time text apps
diffusion-llm
+2
text-to-text
reasoning
4M
4M API calls in the last 30 days
2mo
Last updated on June 10, 2026
NVIDIA
Downloadable
Free Endpoint
nemotron-3-ultra-550b-a55b
Open, efficient hybrid Mamba-Transformer MoE with 1M context, excelling in agentic reasoning, coding, planning, tool calling, and more
Agent
+4
MoE
Frontier
Reasoning
Long Context
52M
52M API calls in the last 30 days
3mo
Last updated on June 4, 2026
NVIDIA
Downloadable
Free Endpoint
cosmos3-nano-reasoner
Vision language model that excels in understanding the physical world using structured reasoning on videos or images.
video understanding
+8
autonomous vehicles
industrial
Physical AI
vision language model
reasoning
robotics
smart cities
Synthetic Data Generation
2K
2K API calls in the last 30 days
3mo
Last updated on June 1, 2026
NVIDIA
Downloadable
Free Endpoint
ising-calibration-1-35b-a3b
Open VLM for quantum computer calibration chart understanding across a range of qubit modalities.
Quantum
+3
reasoning
Vision Language Model
calibration
442K
442K API calls in the last 30 days
4mo
Last updated on April 14, 2026
Google
Downloadable
Free Endpoint
gemma-4-31b-it
Dense 31B model delivering frontier reasoning for coding, agentic workflows, and fine-tuning.
coding
+3
text-to-text
reasoning
agentic
6M
6M API calls in the last 30 days
5mo
Last updated on April 2, 2026
NVIDIA
Downloadable
Free Endpoint
nemotron-3-super-120b-a12b
Open, efficient hybrid Mamba-Transformer MoE with 1M context, excelling in agentic reasoning, coding, planning, tool calling, and more
MoE
+4
Reasoning
Chat
Long Context
Instruction Following
65M
65M API calls in the last 30 days
5mo
Last updated on March 11, 2026
OpenAI
Downloadable
Free Endpoint
gpt-oss-20b
Smaller Mixture of Experts (MoE) text-only LLM for efficient AI reasoning and math
reasoning
+3
text-to-text
chat
math
19M
19M API calls in the last 30 days
1y
Last updated on August 5, 2025
OpenAI
Downloadable
Free Endpoint
gpt-oss-120b
Mixture of Experts (MoE) reasoning LLM (text-only) designed to fit within 80GB GPU.
reasoning
+3
text-to-text
chat
math
45M
45M API calls in the last 30 days
1y
Last updated on August 5, 2025