Skip to main content
Explore
Models
Skills
Blueprints
GPUs
Docs
Search
⌘K
Ctrl+K
?
Help Center
Getting Started
1
Set up your account
Create and verify your account to unlock full access to NVIDIA NIM APIs.
Create an Account
2
Generate API Key
3
Make your first API call
4
Prototype in your environment
5
Connect to inference partners
Resources
Developer Forums
Contact Support
FAQs
Login
14 results for
Filters (1)
Models (14)
Blueprints (0)
Skills (0)
Other (0)
Sort By
Best Match
Select item
Best Match
Most Popular
Most Downloaded
Alphabetical (A-Z)
Alphabetical (Z-A)
Best Match
Most Popular
Most Downloaded
Alphabetical (A-Z)
Alphabetical (Z-A)
Search results
Google
Downloadable
Free Endpoint
diffusiongemma-26b-a4b-it
Diffusion-based 26B parameter LLM enabling parallel token generation for real-time text apps
Model
diffusion-llm
+2
text-to-text
reasoning
Items per page
24
12
24
48
96
1
1
of 1 pages
4M
4M API calls in the last 30 days
2mo
Last updated on June 10, 2026
OpenAI
Downloadable
Free Endpoint
gpt-oss-20b
Smaller Mixture of Experts (MoE) text-only LLM for efficient AI reasoning and math
Model
text-to-text
+3
chat
reasoning
math
19M
19M API calls in the last 30 days
1y
Last updated on August 5, 2025
NVIDIA
Deprecation in 14d
Free Endpoint
nemotron-mini-4b-instruct
Optimized SLM for on-device inference and fine-tuned for roleplay, RAG and function calling
Model
Chat
+2
Text-to-Text
Language Generation
3M
3M API calls in the last 30 days
1y
Last updated on August 26, 2024
Google
Downloadable
Free Endpoint
gemma-4-31b-it
Dense 31B model delivering frontier reasoning for coding, agentic workflows, and fine-tuning.
Model
coding
+3
text-to-text
reasoning
agentic
6M
6M API calls in the last 30 days
4mo
Last updated on April 2, 2026
Thinkingmachines
Downloadable
Free Endpoint
inkling
Inkling is a multimodal (text + image) reasoning model from Thinking Machines — a Mamba-hybrid, 256-expert Mixture-of-Experts architecture with tool use and switchable reasoning.
Model
text-to-text
+3
reasoning
image-to-text
multimodal
27d
Last updated on July 16, 2026
Meta
Downloadable
Free Endpoint
llama-3.1-70b-instruct
Powers complex conversations with superior contextual understanding, reasoning and text generation.
Model
Chat
+3
Text-to-Text
Language Generation
Code Generation
5M
5M API calls in the last 30 days
1y
Last updated on June 12, 2025
Meta
Downloadable
Free Endpoint
llama-3.1-8b-instruct
Advanced state-of-the-art model with language understanding, superior reasoning, and text generation.
Model
Chat
+4
Text-to-Text
Language Generation
Run-on-RTX
Code Generation
19M
19M API calls in the last 30 days
1y
Last updated on July 9, 2025
Meta
Downloadable
Free Endpoint
llama-3.2-1b-instruct
Advanced state-of-the-art small language model with language understanding, superior reasoning, and text generation.
Model
chat
+3
Text-to-Text
Language Generation
Code Generation
40K
40K downloads in the last 30 days
545K
545K API calls in the last 30 days
1y
Last updated on May 21, 2025
Meta
Downloadable
Free Endpoint
llama-3.2-3b-instruct
Advanced state-of-the-art small language model with language understanding, superior reasoning, and text generation.
Model
Chat
+3
Text-to-Text
Language Generation
Code Generation
27K
27K downloads in the last 30 days
1M
1M API calls in the last 30 days
1y
Last updated on May 22, 2025
Meta
Downloadable
Free Endpoint
llama-3.3-70b-instruct
Advanced LLM for reasoning, math, general knowledge, and function calling
Model
Reasoning
+4
Text-to-Text
Instruction following
Math
Code Generation
27M
27M API calls in the last 30 days
1y
Last updated on June 12, 2025
Minimaxai
Free Endpoint
minimax-m3
MiniMax M3 Preview is a multimodal MoE vision-language model with strong reasoning, coding, and tool-calling capabilities.
Model
coding
+2
text-to-text
reasoning
10M
10M API calls in the last 30 days
2mo
Last updated on June 12, 2026
NVIDIA
Downloadable
Free Endpoint
nemotron-3.5-lightning-30b-a3b
Fastest 30B A3B MoE model with leading domain accuracy for specialized agentic tasks
Model
Customization
+3
Text-to-Text
Long-running agents
Open
1d
Last updated on August 11, 2026
OpenAI
Downloadable
Free Endpoint
gpt-oss-120b
Mixture of Experts (MoE) reasoning LLM (text-only) designed to fit within 80GB GPU.
Model
reasoning
+3
text-to-text
chat
math
45M
45M API calls in the last 30 days
1y
Last updated on August 5, 2025
Meta
Downloadable
Free Endpoint
muse-glimmer-30b
Muse Glimmer 30B is a multimodal reasoning model accepting text and images, served on vLLM with native Onyx tool-calling and reasoning parsers.
Model
Multimodal
+5
Image-to-Text
Reasoning
Chat
Text-to-Text
Large Language Models
1d
Last updated on August 10, 2026