Skip to main content
Explore
Models
Skills
Blueprints
GPUs
Docs
Search
⌘K
Ctrl+K
?
Forums
Support
Login
Models
Deploy and scale models on your GPU infrastructure of choice with NVIDIA NIM inference microservices
Optimized by NVIDIA
Launch from Hugging Face
Beta
Filters (1)
7 models
Sort By
Most Recent
Select item
Most Recent
Most Popular
Most Downloaded
Alphabetical (A-Z)
Alphabetical (Z-A)
Most Recent
Most Popular
Most Downloaded
Alphabetical (A-Z)
Alphabetical (Z-A)
NVIDIA
Downloadable
Free Endpoint
cosmos3-nano-reasoner
Vision language model that excels in understanding the physical world using structured reasoning on videos or images.
video understanding
+8
autonomous vehicles
industrial
Physical AI
vision language model
reasoning
robotics
smart cities
Synthetic Data Generation
Items per page
24
12
24
48
96
1
1
of 1 pages
2K
2K API calls in the last 30 days
1mo
Last updated on June 1, 2026
Stepfun-ai
Downloadable
Free Endpoint
step-3.7-flash
A sparse MoE multimodal reasoning model good for enterprise, agentic and coding tasks.
Coding
+2
Vision
Agents
7M
7M API calls in the last 30 days
1mo
Last updated on May 29, 2026
NVIDIA
Downloadable
Free Endpoint
ising-calibration-1-35b-a3b
Open VLM for quantum computer calibration chart understanding across a range of qubit modalities.
Quantum
+3
reasoning
Vision Language Model
calibration
442K
442K API calls in the last 30 days
3mo
Last updated on April 14, 2026
NVIDIA
Downloadable
cosmos-reason2-8b
Vision language model that excels in understanding the physical world using structured reasoning on videos or images.
video understanding
+8
autonomous vehicles
industrial
Physical AI
vision language model
reasoning
robotics
smart cities
Synthetic Data Generation
191K
191K API calls in the last 30 days
6mo
Last updated on December 27, 2025
NVIDIA
Downloadable
Free Endpoint
nemotron-nano-12b-v2-vl
Nemotron Nano 12B v2 VL enables multi-image and video understanding, along with visual Q&A and summarization capabilities.
language generation
+3
vision assistant
visual question answering
Image-to-Text
5M
5M API calls in the last 30 days
8mo
Last updated on October 28, 2025
Meta
Deprecation in 7d
Free Endpoint
llama-4-maverick-17b-128e-instruct
A general purpose multimodal, multilingual 128 MoE model with 17B parameters.
language generation
+3
vision assistant
visual question answering
Image-to-Text
16M
16M API calls in the last 30 days
1y
Last updated on July 17, 2025
Google
Free Endpoint
paligemma
Vision language model adept at comprehending text and visual inputs to produce informative responses
image
+8
cv
Vision Assistant
vlm
Visual Question Answering
computer vision
Language Generation
video
Image-to-Text
12K
12K API calls in the last 30 days
1y
Last updated on August 26, 2024