Skip to main content
Explore
Models
Skills
Blueprints
GPUs
Docs
Search
⌘K
Ctrl+K
?
Forums
Support
Login
Models
Deploy and scale models on your GPU infrastructure of choice with NVIDIA NIM inference microservices
Optimized by NVIDIA
Launch from Hugging Face
Beta
Filters (3)
3 models
Sort By
Most Recent
Select item
Most Recent
Most Popular
Most Downloaded
Alphabetical (A-Z)
Alphabetical (Z-A)
Most Recent
Most Popular
Most Downloaded
Alphabetical (A-Z)
Alphabetical (Z-A)
NVIDIA
Downloadable
Free Endpoint
nemotron-nano-12b-v2-vl
Nemotron Nano 12B v2 VL enables multi-image and video understanding, along with visual Q&A and summarization capabilities.
language generation
+3
vision assistant
visual question answering
Image-to-Text
Items per page
24
12
24
48
96
1
1
of 1 pages
5M
5M API calls in the last 30 days
8mo
Last updated on October 28, 2025
Meta
Deprecation in 7d
Free Endpoint
llama-4-maverick-17b-128e-instruct
A general purpose multimodal, multilingual 128 MoE model with 17B parameters.
language generation
+3
vision assistant
visual question answering
Image-to-Text
16M
16M API calls in the last 30 days
1y
Last updated on July 17, 2025
Google
Free Endpoint
paligemma
Vision language model adept at comprehending text and visual inputs to produce informative responses
image
+8
cv
Vision Assistant
vlm
Visual Question Answering
computer vision
Language Generation
video
Image-to-Text
12K
12K API calls in the last 30 days
1y
Last updated on August 26, 2024