NVIDIA
Explore
Models
Blueprints
GPUs
Docs
⌘KCtrl+K
Terms of Use
Privacy Policy
Your Privacy Choices
Contact

Copyright © 2026 NVIDIA Corporation

16 results for

Filters

  • Free Endpoint
    2
  • Partner Endpoint
    4
  • Download Available
    12
  • Enterprise
    0
  • Launchable
    0
  • Retrieval Augmented Generation
    4
  • Text-to-Embedding
    2
  • Deep Infra
    3
  • Together AI
    1
  • NVIDIA
    16
  • NVIDIA AI
    0
  • NVIDIA
    Free Endpoint

    llama-3.1-nemotron-70b-reward

    Leaderboard topping reward model supporting RLHF for better alignment with human preferences.
    Model
    Text-to-text
    100K
    1y
    NVIDIA
    Downloadable

    llama-nemotron-embed-1b-v2

    Multilingual, cross-lingual embedding model for long-document QA retrieval, supporting 26 languages.
    Model
    Text-to-Embedding
    1.12M
    1mo
    NVIDIA
    Downloadable

    llama-nemotron-rerank-1b-v2

    GPU-accelerated model optimized for providing a probability score that a given passage contains the information to answer a question.
    Model
    nemo retriever
    100K
    1mo
    NVIDIA
    Deprecation in 2dDownloadable

    llama-3.1-nemotron-nano-4b-v1.1

    State-of-the-art open model for reasoning, code, math, and tool calling - suitable for edge agents
    Model
    chat
    80.37K
    9mo
    NVIDIA
    Downloadable

    llama-3.1-nemotron-nano-8b-v1

    Leading reasoning and agentic AI accuracy model for PC and edge.
    Model
    chat
    514K
    9mo
    NVIDIA
    Downloadable

    llama-3.1-nemotron-nano-vl-8b-v1

    Multi-modal vision-language model that understands text/img and creates informative responses
    Model
    chat
    8.91M
    9mo
    NVIDIA
    Free Endpoint

    llama-3.1-nemotron-safety-guard-8b-v3

    Leading multilingual content safety model for enhancing the safety and moderation capabilities of LLMs
    Model
    content moderation
    98.03K
    5mo
    NVIDIA
    Downloadable

    llama-3.1-nemotron-ultra-253b-v1

    Superior inference efficiency with highest accuracy for scientific and complex math reasoning, coding, tool calling, and instruction following.
    Model
    chat
    4.41M
    9mo
    NVIDIA
    Deprecation in 2dDownloadable

    llama-3.3-nemotron-super-49b-v1

    High efficiency model with leading accuracy for reasoning, tool calling, chat, and instruction following.
    Model
    chat
    1.22M
    8mo
    NVIDIA
    Downloadable

    llama-3.3-nemotron-super-49b-v1.5

    High efficiency model with leading accuracy for reasoning, tool calling, chat, and instruction following.
    Model
    chat
    2.37M
    8mo
    NVIDIA
    Downloadable

    llama-nemotron-embed-vl-1b-v2

    Multimodal question-answer retrieval representing user queries as text and documents as images.
    Model
    nemo retriever
    8.02M
    2mo
    NVIDIA
    Downloadable

    llama-nemotron-rerank-vl-1b-v2

    GPU-accelerated model optimized for providing a probability score that a given passage contains the information to answer a question.
    Model
    nemo retriever
    522
    1w
    DGX Spark
    30 MIN

    Nemotron-3-Nano with llama.cpp

    Run Nemotron-3-Nano-30B model using llama.cpp on DGX Spark
    Playbook
    Nemotron
    3mo
    NVIDIA
    Downloadable

    llama-3.1-nemoguard-8b-content-safety

    Leading content safety model for enhancing the safety and moderation capabilities of LLMs
    Model
    nemo guardrails
    96.21K
    1y
    NVIDIA
    Downloadable

    llama-3.1-nemoguard-8b-topic-control

    Topic control model to keep conversations focused on approved topics, avoiding inappropriate content.
    Model
    nemo guardrails
    63.98K
    1y
    DGX Spark
    30 MINS

    NemoClaw with Nemotron-3-Super and Telegram on DGX Spark

    Install NemoClaw on DGX Spark with local Ollama inference and Telegram bot integration
    Playbook
    Telegram
    1w
    Items per page
    of 1 pages