Skip to main content
Explore
Models
Skills
Blueprints
GPUs
Docs
Search
⌘K
Ctrl+K
?
Help Center
Getting Started
1
Set up your account
Create and verify your account to unlock full access to NVIDIA NIM APIs.
Create an Account
2
Generate API Key
3
Make your first API call
4
Prototype in your environment
5
Connect to inference partners
Resources
Developer Forums
Contact Support
FAQs
Login
14 results for
Filters (1)
Models (10)
Blueprints (1)
Skills (0)
Other (3)
Sort By
Best Match
Select item
Best Match
Most Popular
Most Downloaded
Alphabetical (A-Z)
Alphabetical (Z-A)
Best Match
Most Popular
Most Downloaded
Alphabetical (A-Z)
Alphabetical (Z-A)
Search results
Playbooks
45 MIN
Generate Images and Videos with ComfyUI
Node-based diffusion workflows for images and videos with FLUX, Wan, HunyuanVideo, and Stable Diffusion
Playbook
Content Creation
+5
Image Generation
ComfyUI
DGX Spark
Docker
DGX Station
8d
Last updated on August 4, 2026
Items per page
24
12
24
48
96
1
1
of 1 pages
DGX Spark
1 HR
FLUX.1 Dreambooth LoRA Fine-tuning
Fine-tune FLUX.1-dev 12B model using Dreambooth LoRA for custom image generation
Playbook
Image Generation
+6
ComfyUI
DGX
LoRA
Spark
Fine-tuning
Text-to-Image
10mo
Last updated on October 7, 2025
Meta
Downloadable
Free Endpoint
llama-3.2-11b-vision-instruct
Cutting-edge vision-language model exceling in high-quality reasoning from images.
Model
Image-Text Retrieval
+4
Visual QA
Image Captioning
Visual Grounding
Image-to-Text
3M
3M API calls in the last 30 days
1y
Last updated on May 30, 2025
Meta
Downloadable
Free Endpoint
llama-3.2-90b-vision-instruct
Cutting-edge vision-Language model exceling in high-quality reasoning from images.
Model
Image-Text Retrieval
+4
Visual QA
image captioning
Visual Grounding
Image-to-Text
4M
4M API calls in the last 30 days
1y
Last updated on May 30, 2025
NVIDIA
Downloadable
Free Endpoint
llama-3.1-nemotron-nano-vl-8b-v1
Multi-modal vision-language model that understands text/img and creates informative responses
Model
doc intelligence
+2
multiple image understanding
OCR
14M
14M API calls in the last 30 days
1y
Last updated on July 1, 2025
DGX Spark
1 HR
Vision-Language Model Fine-tuning
Fine-tune Vision-Language Models for image and video understanding tasks using Qwen2.5-VL and InternVL3
Playbook
DGX
+6
Image Understanding
Vision-Language Models
GRPO
Spark
Fine-tuning
Video Analysis
10mo
Last updated on October 7, 2025
Black-forest-labs
Downloadable
flux.2-klein-4b
FLUX.2-klein-4B is a distilled image generation and editing model, producing outputs at lighting speed
Model
image editing
+3
Run-on-RTX
Text-to-Image
Image Generation
338K
338K API calls in the last 30 days
5mo
Last updated on March 13, 2026
Microsoft
Downloadable
TRELLIS
MSFT TRELLIS is a 3D AI model that generates high-quality 3D assets from text or image inputs.
Model
text-to-3d
+2
Run-on-RTX
image-to-3d
18K
18K API calls in the last 30 days
11mo
Last updated on September 3, 2025
Thinkingmachines
Downloadable
Free Endpoint
inkling
Inkling is a multimodal (text + image) reasoning model from Thinking Machines — a Mamba-hybrid, 256-expert Mixture-of-Experts architecture with tool use and switchable reasoning.
Model
text-to-text
+3
reasoning
image-to-text
multimodal
28d
Last updated on July 16, 2026
Google
Free Endpoint
paligemma
Vision language model adept at comprehending text and visual inputs to produce informative responses
Model
image
+8
cv
Vision Assistant
vlm
Visual Question Answering
computer vision
Language Generation
video
Image-to-Text
12K
12K API calls in the last 30 days
1y
Last updated on August 26, 2024
NVIDIA
Downloadable
vista-3d
VISTA-3D is a specialized interactive foundation model for segmenting and anotating human anatomies.
Model
Interactive Annotation
+3
Image Segmentation
Non-Commercial Use Only
Medical Imaging
587
587 API calls in the last 30 days
1y
Last updated on April 21, 2025
NVIDIA
Downloadable
Free Endpoint
cosmos3-nano
Generates physics-aware videos from text prompts or an image prompt for physical AI development.
Model
autonomous vehicles
+5
Physical AI
robotics
text-to-world
image-to-world
Synthetic Data Generation
2K
2K API calls in the last 30 days
2mo
Last updated on June 1, 2026
Meta
Downloadable
Free Endpoint
muse-glimmer-30b
Muse Glimmer 30B is a multimodal reasoning model accepting text and images, served on vLLM with native Onyx tool-calling and reasoning parsers.
Model
Multimodal
+5
Image-to-Text
Reasoning
Chat
Text-to-Text
Large Language Models
3d
Last updated on August 10, 2026
Robotics
Enterprise
Synthetic Manipulation Motion Generation for Robotics
Generate exponentially large amounts of synthetic motion trajectories for robot manipulation from just a few human demonstrations.
Blueprint
synthetic data
+9
robotics
physical ai
robot learning
Humanoids
text-to-world
image-to-world
teleop
NVIDIA Isaac GR00T
NVIDIA Omniverse
5mo
Last updated on February 17, 2026