Skip to main content
Explore
Models
Skills
Blueprints
GPUs
Docs
Search
⌘K
Ctrl+K
?
Forums
Support
Login
Models
Deploy and scale models on your GPU infrastructure of choice with NVIDIA NIM inference microservices
Optimized by NVIDIA
Launch from Hugging Face
Beta
Filters
14 models
Sort By
Most Recent
Select item
Most Recent
Most Popular
Most Downloaded
Alphabetical (A-Z)
Alphabetical (Z-A)
Most Recent
Most Popular
Most Downloaded
Alphabetical (A-Z)
Alphabetical (Z-A)
Minimaxai
Free Endpoint
minimax-m3
MiniMax M3 Preview is a multimodal MoE vision-language model with strong reasoning, coding, and tool-calling capabilities.
coding
+2
text-to-text
reasoning
Items per page
24
12
24
48
96
1
1
of 1 pages
10M
10M API calls in the last 30 days
1mo
Last updated on June 12, 2026
NVIDIA
Downloadable
Free Endpoint
cosmos3-nano-reasoner
Vision language model that excels in understanding the physical world using structured reasoning on videos or images.
video understanding
+8
autonomous vehicles
industrial
Physical AI
vision language model
reasoning
robotics
smart cities
Synthetic Data Generation
2K
2K API calls in the last 30 days
1mo
Last updated on June 1, 2026
Stepfun-ai
Downloadable
Free Endpoint
step-3.7-flash
A sparse MoE multimodal reasoning model good for enterprise, agentic and coding tasks.
Coding
+2
Vision
Agents
7M
7M API calls in the last 30 days
1mo
Last updated on May 29, 2026
NVIDIA
Downloadable
Free Endpoint
ising-calibration-1-35b-a3b
Open VLM for quantum computer calibration chart understanding across a range of qubit modalities.
Quantum
+3
reasoning
Vision Language Model
calibration
442K
442K API calls in the last 30 days
3mo
Last updated on April 14, 2026
Qwen
Downloadable
Free Endpoint
qwen3.5-397b-a17b
Next-gen Qwen 3.5 VLM (400B MoE) brings advanced vision, chat, RAG, and agentic capabilities.
MoE
+3
image-to-image
VLM
agentic
16M
16M API calls in the last 30 days
5mo
Last updated on February 16, 2026
NVIDIA
Downloadable
cosmos-reason2-8b
Vision language model that excels in understanding the physical world using structured reasoning on videos or images.
video understanding
+8
autonomous vehicles
industrial
Physical AI
vision language model
reasoning
robotics
smart cities
Synthetic Data Generation
191K
191K API calls in the last 30 days
6mo
Last updated on December 27, 2025
NVIDIA
Downloadable
nemotron-parse
Cutting-edge vision-language model exceling in retrieving text and metadata from images.
text and table extraction
+2
document parsing
supported language - english
1M
1M API calls in the last 30 days
8mo
Last updated on October 28, 2025
NVIDIA
Downloadable
Free Endpoint
nemotron-nano-12b-v2-vl
Nemotron Nano 12B v2 VL enables multi-image and video understanding, along with visual Q&A and summarization capabilities.
language generation
+3
vision assistant
visual question answering
Image-to-Text
5M
5M API calls in the last 30 days
8mo
Last updated on October 28, 2025
NVIDIA
Downloadable
Free Endpoint
llama-3.1-nemotron-nano-vl-8b-v1
Multi-modal vision-language model that understands text/img and creates informative responses
doc intelligence
+2
multiple image understanding
OCR
14M
14M API calls in the last 30 days
1y
Last updated on July 1, 2025
Meta
Deprecation in 5d
Free Endpoint
llama-4-maverick-17b-128e-instruct
A general purpose multimodal, multilingual 128 MoE model with 17B parameters.
language generation
+3
vision assistant
visual question answering
Image-to-Text
16M
16M API calls in the last 30 days
1y
Last updated on July 17, 2025
NVIDIA
Downloadable
nemoretriever-parse
Cutting-edge vision-language model exceling in retrieving text and metadata from images.
optical character recognition
+4
nemo retriever
data ingestion
table extraction
supported language - english
386K
386K API calls in the last 30 days
1y
Last updated on June 6, 2025
Meta
Downloadable
Free Endpoint
llama-3.2-11b-vision-instruct
Cutting-edge vision-language model exceling in high-quality reasoning from images.
Image-Text Retrieval
+4
Visual QA
Image Captioning
Visual Grounding
Image-to-Text
3M
3M API calls in the last 30 days
1y
Last updated on May 30, 2025
Meta
Downloadable
Free Endpoint
llama-3.2-90b-vision-instruct
Cutting-edge vision-Language model exceling in high-quality reasoning from images.
Image-Text Retrieval
+4
Visual QA
image captioning
Visual Grounding
Image-to-Text
4M
4M API calls in the last 30 days
1y
Last updated on May 30, 2025
Google
Free Endpoint
paligemma
Vision language model adept at comprehending text and visual inputs to produce informative responses
image
+8
cv
Vision Assistant
vlm
Visual Question Answering
computer vision
Language Generation
video
Image-to-Text
12K
12K API calls in the last 30 days
1y
Last updated on August 26, 2024