Skip to main content
NVIDIA
Explore
Models
Skills
Blueprints
GPUs
Docs
Help Center
Getting Started
  1. Create and verify your account to unlock full access to NVIDIA NIM APIs.
ResourcesDeveloper ForumsContact Support
FAQs
  • DiscoverModelsSkillsBlueprintsGPUsDocsForums

    workstations

    • Run on RTX
    • Run on Spark
    • Run on Station

    models

    • Reasoning
    • Vision
    • Visual Design
    • Retrieval
    • Speech
    • Biology
    • Simulation
    • Climate & Weather
    • Safety & Moderation

    industries

    • Automotive
    • Financial Services
    • Gaming
    • Healthcare
    • Industrial
    • Robotics

    Retrieval

    Terms of Use
    Privacy Policy
    Your Privacy Choices
    Contact

    Copyright © 2026 NVIDIA Corporation

    Deploy Models Now with NVIDIA NIM

    Optimized inference for the world’s leading models
    Free serverless APIs for development
    Accelerated by DGX Cloud
    Self-Host on your GPU infrastructure
    Continuous vulnerability fixes

    Embedding Models

    Power enterprise search and question answering with NVIDIA Nemotron RAG models by embedding text and scanned documents, including PDFs and images, for fast, multilingual, multimodal data retrieval.

    NVIDIA
    Downloadable

    llama-nemotron-embed-vl-1b-v2

    Multimodal question-answer retrieval representing user queries as text and documents as images.
    • embedding
    • nemo retriever
    16M API calls in the last 30 days
    Last updated on February 10, 2026
    NVIDIA
    Deprecation in 1dDownloadable

    llama-nemotron-embed-1b-v2

    Multilingual, cross-lingual embedding model for long-document QA retrieval, supporting 26 languages.
    • NeMo Retriever
    6M API calls in the last 30 days
    Last updated on March 4, 2026

    Reranking Models

    Improve information retrieval accuracy with world-class NVIDIA Nemotron RAG models by reranking retrieved enterprise data to improve answer relevancy.

    NVIDIA
    Downloadable

    llama-nemotron-rerank-vl-1b-v2

    GPU-accelerated model optimized for providing a probability score that a given passage contains the information to answer a question.
    • nemo retriever
    • reranking
    1M API calls in the last 30 days
    Last updated on March 31, 2026
    NVIDIA
    Deprecation in 1dDownloadable

    llama-nemotron-rerank-1b-v2

    GPU-accelerated model optimized for providing a probability score that a given passage contains the information to answer a question.
    • nemo retriever
    • reranking
    632K API calls in the last 30 days
    Last updated on March 6, 2026
    NVIDIA
    Deprecation in 1dDownloadable

    nv-yolox-page-elements-v1

    Model for object detection, fine-tuned to detect charts, tables, and titles in documents.
    • Chart Detection
    • Data ingestion
    • Object Detection
    • Table Detection
    • extraction
    • nemo retriever
    • run-on-rtx
    455 API calls in the last 30 days
    Last updated on July 9, 2025
    NVIDIA
    Deprecation in 1dDownloadable

    nemoretriever-parse

    Cutting-edge vision-language model exceling in retrieving text and metadata from images.
    • data ingestion
    • nemo retriever
    • optical character recognition
    • supported language - english
    • table extraction
    386K API calls in the last 30 days
    Last updated on June 6, 2025
    NVIDIA
    Deprecation in 1dDownloadable

    nemoretriever-page-elements-v2

    Model for object detection, fine-tuned to detect charts, tables, and titles in documents.
    • Chart Detection
    • Object Detection
    • Table Detection
    • data ingestion
    • nemo retriever
    • run-on-rtx
    78K API calls in the last 30 days
    Last updated on March 17, 2025

    Extraction Models

    Accelerate large-scale extraction from massive collections of multimodal data - text, images, and complex documents - with NVIDIA Nemotron RAG models for rapid, context-aware insights across your enterprise.

    Explore NVIDIA Blueprints

    Accelerate AI application development with ready-to-use workflows powered by NVIDIA NIM and NeMo microservices for RAG, AI agents, video search and summarization, and more.

    nvidiaBuild an Enterprise RAG Pipeline Blueprint

    Power fast, accurate semantic search across multimodal enterprise data with NVIDIA’s RAG Blueprint—built on NeMo Retriever and Nemotron models—to connect your agents to trusted, authoritative sources of knowledge.

    • NIM
    • NeMo Retriever
    • Nemotron
    • Retrieval-Augmented Generation

    nvidiaBuild an AI Virtual Assistant

    Create intelligent virtual assistants for customer service across every industry

    • Customer Service
    • Retrieval-augmented generation
    • contact center
    • llm

    nvidiaBuild a Video Search and Summarization (VSS) Agent

    Ingest massive volumes of live or archived videos and extract insights for summarization and interactive Q&A

    • chat
    • generative AI
    • video-to-text
    • vision
    Enterprise

    nvidiaNVIDIA AI-Q Blueprint for intelligent agents

    AI agents that connect, retrieve, and reason on enterprise data—making information accessible, actionable, and intelligent.

    • Agents
    • Enterprise
    • NIM
    • NeMo
    • Nemotron

    crewaiCode Documentation for Software Development

    Document your github repositories with AI Agents using CrewAI and Llama3.3 70B NIM.

    • AI Agents
    • Code Documentation
    • CrewAI
    • Partner