NVIDIA
93 resultsModels (61)
A relational foundation model for prediction over structured, multi-table data.DownloadableFree EndpointKumo Relational
Cutting-edge vision-language model excelling in retrieving text and metadata from images.Downloadablenemotron-parse-2.0
Multilingual ASR across 25 European languages with punctuation, capitalization, and word timestampsDownloadableparakeet-tdt-0.6b
Fastest 30B A3B MoE model with leading domain accuracy for specialized agentic tasksDownloadableFree Endpointnemotron-3.5-lightning-30b-a3b
Translation model in 37 languages with few-shots example prompts capability.Free Endpointriva-translate-4b-instruct-v2
NVIDIA-Ising-Calibration-1.5 is a dense multimodal vision-language model built on Gemma 4 31B. It analyzes quantum computing calibration experiment plots and generates structured technical text.Free Endpointising-calibration-1.5-31b
Upscale encoded or ST 2110 video to higher resolutions with NVIDIA Video Super Resolution.DownloadableVideo Super Resolution NIM
1B embedding model for semantic search, retrieval, and RAG applications.Free Endpointnemotron-3-embed-1b
An image edit model specialized for Omniverse synthetic to photographic solder-light style captured at NVIDIA PCB inspection stationsDownloadableqwen-image-edit-nvpcb-ovsl2sl
Nemotron OCR v2 is a state-of-the-art multilingual text recognition model designed for robust end-to-end optical character recognition (OCR) on complex real-world images.Downloadablenemotron-ocr-v2
Open, efficient hybrid Mamba-Transformer MoE with 1M context, excelling in agentic reasoning, coding, planning, tool calling, and moreDownloadableFree Endpointnemotron-3-ultra-550b-a55b
Multilingual, multimodal model for detecting unsafe and toxic content.DownloadableFree Endpointnemotron-3.5-content-safety
Generates physics-aware videos from text prompts or an image prompt for physical AI development.DownloadableFree Endpointcosmos3-nano
Vision language model that excels in understanding the physical world using structured reasoning on videos or images.DownloadableFree Endpointcosmos3-nano-reasoner
Nemotron 3 Nano Omni is an omni-modal reasoning model that understands images, video, speech, text.DownloadableFree Endpointnemotron-3-nano-omni-30b-a3b-reasoning
Re-illuminate people in video to match target lighting from a 360 HDRI environment map.DownloadableRelighting
NVIDIA Synthetic Video Detector is an AI-powered micro-service for detecting AI‑generated (synthetic) videos.DownloadableFree Endpointsynthetic-video-detector
Detect and track speaker identities across video frames.DownloadableFree EndpointActive Speaker Detection
Open VLM for quantum computer calibration chart understanding across a range of qubit modalities.DownloadableFree Endpointising-calibration-1-35b-a3b
GPU-accelerated model optimized for providing a probability score that a given passage contains the information to answer a question.Downloadablellama-nemotron-rerank-vl-1b-v2
Powerful OCR model for fast, accurate real-world image text extraction, layout, and structure analysis.Downloadablenemotron-ocr-v1
Open, efficient hybrid Mamba-Transformer MoE with 1M context, excelling in agentic reasoning, coding, planning, tool calling, and moreDownloadableFree Endpointnemotron-3-super-120b-a12b
Model for object detection, fine-tuned to detect charts, tables, and titles in documents.Downloadablenemotron-table-structure-v1
Model for object detection, fine-tuned to detect charts, tables, and titles in documents.Downloadablenemotron-page-elements-v3
Model for object detection, fine-tuned to detect charts, tables, and titles in documents.Downloadablenemotron-graphic-elements-v1
Generates physics-aware video world states for physical AI development using text prompts and multiple spatial control inputs derived from real-world data or simulation.Free Endpointcosmos-transfer2.5-2b
Multimodal question-answer retrieval representing user queries as text and documents as images.Downloadablellama-nemotron-embed-vl-1b-v2
Translation model in 12 languages with few-shots example prompts capability.Free Endpointriva-translate-4b-instruct-v1_1
StreamPETR offers efficient 3D object detection for autonomous driving by propagating sparse object queries temporally.Free Endpointstreampetr
Cutting-edge vision-language model exceling in retrieving text and metadata from images.Downloadablenemotron-parse
Leading multilingual content safety model for enhancing the safety and moderation capabilities of LLMsFree Endpointllama-3.1-nemotron-safety-guard-8b-v3
Record-setting accuracy and performance for Mandarin Taiwanese English transcriptions.Downloadableparakeet-ctc-0.6b-zh-tw
Record-setting accuracy and performance for Mandarin English transcriptions.Downloadableparakeet-ctc-0.6b-zh-cn
Accurate and optimized Spanish English transcriptions with punctuation and word timestamps.Downloadableparakeet-ctc-0.6b-es
Accurate and optimized Vietnamese-English transcriptions with punctuation and word timestamps.Downloadableparakeet-ctc-0.6b-vi
Accurate and optimized English transcriptions with punctuation and word timestampsDownloadableparakeet-tdt-0.6b-v2
Powerful OCR model for fast, accurate real-world image text extraction, layout, and structure analysis.Downloadablenemoretriever-ocr
Removes unwanted noises from audio improving speech intelligibility.DownloadableFree EndpointBackground Noise Removal
Expressive and engaging text-to-speech, generated from a short audio sample.Free Endpointmagpie-tts-zeroshot
High accuracy and optimized performance for transcription in 25 languagesDownloadableparakeet-1.1b-rnnt-multilingual-asr
End-to-end autonomous driving stack integrating perception, prediction, and planning with sparse scene representations for efficiency and safety.Free Endpointsparsedrive
Natural and expressive voices in multiple languages. For voice agents and brand ambassadors.Downloadablemagpie-tts-multilingual
Multi-lingual model supporting speech-to-text recognition and translation.Downloadablecanary-1b-asr
Topic control model to keep conversations focused on approved topics, avoiding inappropriate content.Downloadablellama-3.1-nemoguard-8b-topic-control
Industry leading jailbreak classification model for protection from adversarial attemptsDownloadablenemoguard-jailbreak-detect
Leading content safety model for enhancing the safety and moderation capabilities of LLMsDownloadablellama-3.1-nemoguard-8b-content-safety
Automatic speech recognition model that transcribes speech in lower case Spanish with record-setting accuracy and performanceDownloadableconformer-ctc-asr
FourCastNet predicts global atmospheric dynamics of various weather / climate variables.Downloadablefourcastnet
Enhance input speech recorded with low-quality microphones in noisy or reverberant environments, producing studio-quality speech.DownloadableFree EndpointStudio Voice
Record-setting accuracy and performance for English transcription.Downloadableparakeet-ctc-1.1b-asr
State-of-the-art accuracy and speed for English transcriptions.Downloadableparakeet-ctc-0.6b-asr
Estimate gaze angles of a person in a video and redirect to make it frontal.Downloadableeyecontact