Models
Deploy and scale models on your GPU infrastructure of choice with NVIDIA NIM inference microservices
models
A relational foundation model for prediction over structured, multi-table data.DownloadableFree EndpointKumo Relational
Cutting-edge vision-language model excelling in retrieving text and metadata from images.Downloadablenemotron-parse-2.0
Multilingual ASR across 25 European languages with punctuation, capitalization, and word timestampsDownloadableparakeet-tdt-0.6b
Fastest 30B A3B MoE model with leading domain accuracy for specialized agentic tasksDownloadableFree Endpointnemotron-3.5-lightning-30b-a3b
Translation model in 37 languages with few-shots example prompts capability.Free Endpointriva-translate-4b-instruct-v2
NVIDIA-Ising-Calibration-1.5 is a dense multimodal vision-language model built on Gemma 4 31B. It analyzes quantum computing calibration experiment plots and generates structured technical text.Free Endpointising-calibration-1.5-31b
Upscale encoded or ST 2110 video to higher resolutions with NVIDIA Video Super Resolution.DownloadableVideo Super Resolution NIM
1B embedding model for semantic search, retrieval, and RAG applications.Free Endpointnemotron-3-embed-1b
An image edit model specialized for Omniverse synthetic to photographic solder-light style captured at NVIDIA PCB inspection stationsDownloadableqwen-image-edit-nvpcb-ovsl2sl
Nemotron OCR v2 is a state-of-the-art multilingual text recognition model designed for robust end-to-end optical character recognition (OCR) on complex real-world images.Downloadablenemotron-ocr-v2
Open, efficient hybrid Mamba-Transformer MoE with 1M context, excelling in agentic reasoning, coding, planning, tool calling, and moreDownloadableFree Endpointnemotron-3-ultra-550b-a55b
Multilingual, multimodal model for detecting unsafe and toxic content.DownloadableFree Endpointnemotron-3.5-content-safety
Generates images, videos and action predictions from text, visual inputs and spatial controls, with mandatory input screening and visual SynthID watermarking.DownloadableFree Endpointcosmos3-nano
Vision language model that excels in understanding the physical world using structured reasoning on videos or images.DownloadableFree Endpointcosmos3-nano-reasoner
Nemotron 3 Nano Omni is an omni-modal reasoning model that understands images, video, speech, text.DownloadableFree Endpointnemotron-3-nano-omni-30b-a3b-reasoning
Re-illuminate people in video to match target lighting from a 360 HDRI environment map.DownloadableRelighting
NVIDIA Synthetic Video Detector is an AI-powered micro-service for detecting AI‑generated (synthetic) videos.DownloadableFree Endpointsynthetic-video-detector
Detect and track speaker identities across video frames.DownloadableFree EndpointActive Speaker Detection
Open VLM for quantum computer calibration chart understanding across a range of qubit modalities.DownloadableFree Endpointising-calibration-1-35b-a3b
GPU-accelerated model optimized for providing a probability score that a given passage contains the information to answer a question.Downloadablellama-nemotron-rerank-vl-1b-v2
Powerful OCR model for fast, accurate real-world image text extraction, layout, and structure analysis.Downloadablenemotron-ocr-v1
Items per page
of 3 pages