Models
Deploy and scale models on your GPU infrastructure of choice with NVIDIA NIM inference microservices
models
552B MoE, 8B active params with native multimodal support and lower API cost using smaller KV cacheFree Endpointdeepseek-v4.1-flash
A relational foundation model for prediction over structured, multi-table data.DownloadableFree EndpointKumo Relational
Multimodal 320B-total / 18B-active MoE with hybrid KDA and sparse MLA attention, native FP8 weights, reasoning and tool calling.Free Endpointglm-5-3-flash
Cutting-edge vision-language model excelling in retrieving text and metadata from images.Downloadablenemotron-parse-2.0
Multilingual ASR across 25 European languages with punctuation, capitalization, and word timestampsDownloadableparakeet-tdt-0.6b
~2.8T hybrid KDA+MLA multimodal MoE for long-horizon coding, agentic tool use, and image understanding.DownloadableFree Endpointkimi-k3
Wan2.2-Animate-2 is a novel end-to-end character animation frameworkDownloadablewan2.2-animate-2-14b
284B MoE (13B active) model ideal for long-context workloads optimized for coding, chat, and agentic workflowsDeprecatedFree Endpointdeepseek-v4-flash-0731
Fastest 30B A3B MoE model with leading domain accuracy for specialized agentic tasksDownloadableFree Endpointnemotron-3.5-lightning-30b-a3b
Muse Glimmer 30B is a multimodal reasoning model accepting text and images, with native tool-calling and separate reasoning output.DownloadableFree Endpointmuse-glimmer-30b
Translation model in 37 languages with few-shots example prompts capability.Free Endpointriva-translate-4b-instruct-v2
NVIDIA-Ising-Calibration-1.5 is a dense multimodal vision-language model built on Gemma 4 31B. It analyzes quantum computing calibration experiment plots and generates structured technical text.Free Endpointising-calibration-1.5-31b
Upscale encoded or ST 2110 video to higher resolutions with NVIDIA Video Super Resolution.DownloadableVideo Super Resolution NIM
1B embedding model for semantic search, retrieval, and RAG applications.Free Endpointnemotron-3-embed-1b
Efficient 33B MoE for local, long-horizon agentic coding and terminal tasksFree Endpointlaguna-xs-2.1
An image edit model specialized for Omniverse synthetic to photographic solder-light style captured at NVIDIA PCB inspection stationsDownloadableqwen-image-edit-nvpcb-ovsl2sl
Nemotron OCR v2 is a state-of-the-art multilingual text recognition model designed for robust end-to-end optical character recognition (OCR) on complex real-world images.Downloadablenemotron-ocr-v2
Diffusion-based 26B parameter LLM enabling parallel token generation for real-time text appsDownloadableFree Endpointdiffusiongemma-26b-a4b-it
Open, efficient hybrid Mamba-Transformer MoE with 1M context, excelling in agentic reasoning, coding, planning, tool calling, and moreDownloadableFree Endpointnemotron-3-ultra-550b-a55b
Natural and expressive voices in 23 languages. For voice agents and brand ambassadors.Downloadablechatterbox-multilingual-tts
Multilingual, multimodal model for detecting unsafe and toxic content.DownloadableFree Endpointnemotron-3.5-content-safety
Generates images, videos and action predictions from text, visual inputs and spatial controls, with mandatory input screening and visual SynthID watermarking.DownloadableFree Endpointcosmos3-nano
Vision language model that excels in understanding the physical world using structured reasoning on videos or images.DownloadableFree Endpointcosmos3-nano-reasoner
Items per page
of 5 pages