Models
Deploy and scale models on your GPU infrastructure of choice with NVIDIA NIM inference microservices
models
A relational foundation model for prediction over structured, multi-table data.DownloadableFree EndpointKumo Relational
Multimodal 320B-total / 18B-active MoE with hybrid KDA and sparse MLA attention, native FP8 weights, reasoning and tool calling.Free Endpointglm-5-3-flash
Cutting-edge vision-language model excelling in retrieving text and metadata from images.Downloadablenemotron-parse-2.0
Multilingual ASR across 25 European languages with punctuation, capitalization, and word timestampsDownloadableparakeet-tdt-0.6b
~2.8T hybrid KDA+MLA multimodal MoE for long-horizon coding, agentic tool use, and image understanding.DownloadableFree Endpointkimi-k3
Wan2.2-Animate-2 is a novel end-to-end character animation frameworkDownloadablewan2.2-animate-2-14b
284B MoE (13B active) model ideal for long-context workloads optimized for coding, chat, and agentic workflowsDeprecation in 4dFree Endpointdeepseek-v4-flash-0731
Fastest 30B A3B MoE model with leading domain accuracy for specialized agentic tasksDownloadableFree Endpointnemotron-3.5-lightning-30b-a3b
Muse Glimmer 30B is a multimodal reasoning model accepting text and images, with native tool-calling and separate reasoning output.DownloadableFree Endpointmuse-glimmer-30b
Translation model in 37 languages with few-shots example prompts capability.Free Endpointriva-translate-4b-instruct-v2
NVIDIA-Ising-Calibration-1.5 is a dense multimodal vision-language model built on Gemma 4 31B. It analyzes quantum computing calibration experiment plots and generates structured technical text.Free Endpointising-calibration-1.5-31b
Upscale encoded or ST 2110 video to higher resolutions with NVIDIA Video Super Resolution.DownloadableVideo Super Resolution NIM
1B embedding model for semantic search, retrieval, and RAG applications.Free Endpointnemotron-3-embed-1b
Efficient 33B MoE for local, long-horizon agentic coding and terminal tasksFree Endpointlaguna-xs-2.1
An image edit model specialized for Omniverse synthetic to photographic solder-light style captured at NVIDIA PCB inspection stationsDownloadableqwen-image-edit-nvpcb-ovsl2sl
Nemotron OCR v2 is a state-of-the-art multilingual text recognition model designed for robust end-to-end optical character recognition (OCR) on complex real-world images.Downloadablenemotron-ocr-v2
Diffusion-based 26B parameter LLM enabling parallel token generation for real-time text appsDownloadableFree Endpointdiffusiongemma-26b-a4b-it
Open, efficient hybrid Mamba-Transformer MoE with 1M context, excelling in agentic reasoning, coding, planning, tool calling, and moreDownloadableFree Endpointnemotron-3-ultra-550b-a55b
Natural and expressive voices in 23 languages. For voice agents and brand ambassadors.Downloadablechatterbox-multilingual-tts
Multilingual, multimodal model for detecting unsafe and toxic content.DownloadableFree Endpointnemotron-3.5-content-safety
Generates physics-aware videos from text prompts or an image prompt for physical AI development.DownloadableFree Endpointcosmos3-nano
Vision language model that excels in understanding the physical world using structured reasoning on videos or images.DownloadableFree Endpointcosmos3-nano-reasoner
Qwen-Image is a text-to-image foundation model with advanced multilingual text rendering.Downloadableqwen-image
Items per page
of 5 pages