Skip to main content
Explore
Models
Skills
Blueprints
GPUs
Docs
Search
⌘K
Ctrl+K
?
Help Center
Getting Started
1
Set up your account
Create and verify your account to unlock full access to NVIDIA NIM APIs.
Create an Account
2
Generate API Key
3
Make your first API call
4
Prototype in your environment
5
Connect to inference partners
Resources
Developer Forums
Contact Support
FAQs
Login
Models
Deploy and scale models on your GPU infrastructure of choice with NVIDIA NIM inference microservices
Optimized by NVIDIA
Launch from Hugging Face
Beta
Filters (1)
8 models
Sort By
Most Recent
Select item
Most Recent
Most Popular
Most Downloaded
Alphabetical (A-Z)
Alphabetical (Z-A)
Most Recent
Most Popular
Most Downloaded
Alphabetical (A-Z)
Alphabetical (Z-A)
models
DeepSeek AI
Free Endpoint
deepseek-v4-flash-0731
284B MoE (13B active) model ideal for long-context workloads optimized for coding, chat, and agentic workflows
MoE
+3
Reasoning
Long Context
Hybrid Attention
10d
Last updated on August 19, 2026
Items per page
24
12
24
48
96
1
1
of 1 pages
Meta
Downloadable
Free Endpoint
muse-glimmer-30b
Muse Glimmer 30B is a multimodal reasoning model accepting text and images, with native tool-calling and separate reasoning output.
Multimodal
+5
Image-to-Text
Reasoning
Chat
Text-to-Text
Large Language Models
19d
Last updated on August 10, 2026
NVIDIA
Downloadable
Free Endpoint
nemotron-3-nano-omni-30b-a3b-reasoning
Nemotron 3 Nano Omni is an omni-modal reasoning model that understands images, video, speech, text.
Image-to-Text
+4
VLM
Video
Omni
OCR
8M
8M API calls in the last 30 days
4mo
Last updated on April 28, 2026
NVIDIA
Downloadable
Free Endpoint
ising-calibration-1-35b-a3b
Open VLM for quantum computer calibration chart understanding across a range of qubit modalities.
Quantum
+3
reasoning
Vision Language Model
calibration
442K
442K API calls in the last 30 days
4mo
Last updated on April 14, 2026
NVIDIA
Free Endpoint
nemotron-voicechat
Nemotron 3 Voicechat
English
+2
voice chat
NVIDIA NIM
1K
1K API calls in the last 30 days
5mo
Last updated on March 16, 2026
NVIDIA
Downloadable
Free Endpoint
nemotron-3-super-120b-a12b
Open, efficient hybrid Mamba-Transformer MoE with 1M context, excelling in agentic reasoning, coding, planning, tool calling, and more
MoE
+4
Reasoning
Chat
Long Context
Instruction Following
65M
65M API calls in the last 30 days
5mo
Last updated on March 11, 2026
OpenAI
Downloadable
Free Endpoint
gpt-oss-20b
Smaller Mixture of Experts (MoE) text-only LLM for efficient AI reasoning and math
reasoning
+3
text-to-text
chat
math
19M
19M API calls in the last 30 days
1y
Last updated on August 5, 2025
OpenAI
Downloadable
Free Endpoint
gpt-oss-120b
Mixture of Experts (MoE) reasoning LLM (text-only) designed to fit within 80GB GPU.
reasoning
+3
text-to-text
chat
math
45M
45M API calls in the last 30 days
1y
Last updated on August 5, 2025