NVIDIA
Explore
Models
Blueprints
GPUs
Docs
⌘KCtrl+K
Terms of Use
Privacy Policy
Your Privacy Choices
Contact

Copyright © 2026 NVIDIA Corporation

68 results for

Filters

  • Free Endpoint
    30
  • Partner Endpoint
    45
  • Download Available
    38
  • Enterprise
    0
  • Launchable
    0
  • Code Generation
    16
  • Image-to-Text
    7
  • Synthetic Data Generation
    1
  • Deep Infra
    33
  • Together AI
    29
  • GMI Cloud
    20
  • Bitdeer AI
    12
  • Fireworks AI
    12
  • NVIDIA
    10
  • Meta
    9
  • Microsoft
    9
  • DeepSeek AI
    7
  • Mistral AI
    7
  • NVIDIA AI
    0
  • NVIDIA
    Free Endpoint

    nemotron-content-safety-reasoning-4b

    A context‑aware safety model that applies reasoning to enforce domain‑specific policies.
    Model
    NeMo Guardrails
    163K
    2mo
    Microsoft
    Deprecation in 6dFree Endpoint

    phi-4-mini-flash-reasoning

    Lightweight reasoning model for applications in latency bound, memory/compute constrained environments
    Model
    chat
    120K
    8mo
    IBM
    Free Endpoint

    granite-3.3-8b-instruct

    Small language model fine-tuned for improved reasoning, coding, and instruction-following
    Model
    coding
    41.3K
    9mo
    Mistral AI
    Free Endpoint

    magistral-small-2506

    High performance reasoning model optimized for efficiency and edge deployment
    Model
    coding
    1.14M
    9mo
    NVIDIA
    Downloadable

    llama-3.1-nemotron-nano-8b-v1

    Leading reasoning and agentic AI accuracy model for PC and edge.
    Model
    chat
    485K
    9mo
    Marin
    Free Endpoint

    marin-8b-instruct

    State-of-the-art open model trained on open datasets, excelling in reasoning, math, and science.
    Model
    chat
    113K
    10mo
    NVIDIA
    Downloadable

    nvidia-nemotron-nano-9b-v2

    High‑efficiency LLM with hybrid Transformer‑Mamba design, excelling in reasoning and agentic tasks.
    Model
    chat
    309K
    7mo
    Qwen
    Downloadable

    qwen3-next-80b-a3b-thinking

    80B parameter AI model with hybrid reasoning, MoE architecture, support for 119 languages.
    Model
    chat
    1.87M
    6mo
    ByteDance
    Free Endpoint

    seed-oss-36b-instruct

    ByteDance open-source LLM with long-context, reasoning, and agentic intelligence.
    Model
    chat
    1.28M
    7mo
    Stepfun-ai
    Free Endpoint

    step-3.5-flash

    200B open-source reasoning engine with sparse MoE powering frontier agentic AI.
    Model
    chat
    8.73M
    2mo
    NVIDIA
    Deprecation in 6dDownloadable

    llama-3.1-nemotron-nano-4b-v1.1

    State-of-the-art open model for reasoning, code, math, and tool calling - suitable for edge agents
    Model
    chat
    111K
    9mo
    NVIDIA
    Downloadable

    llama-3.1-nemotron-ultra-253b-v1

    Superior inference efficiency with highest accuracy for scientific and complex math reasoning, coding, tool calling, and instruction following.
    Model
    chat
    4.72M
    9mo
    Mistral AI
    Downloadable

    mistral-small-24b-instruct

    Latency-optimized language model excelling in code, math, general knowledge, and instruction-following.
    Model
    chat
    203K
    9mo
    Sarvamai
    Downloadable

    sarvam-m

    Multilingual, hybrid-reasoning model optimized for Indian language tasks, programming, mathematical reasoning capabilities.
    Model
    coding
    153K
    8mo
    Qwen
    Deprecation in 6dFree Endpoint

    qwq-32b

    Powerful reasoning model capable of thinking and reasoning, can achieve significantly enhanced performance in downstream tasks, especially hard problems.
    Model
    coding
    1.22M
    9mo
    DeepSeek AI
    Downloadable

    deepseek-r1-distill-llama-8b

    Distilled version of Llama 3.1 8B using reasoning data generated by DeepSeek R1 for enhanced performance.
    Model
    Distillation
    1.36M
    9mo
    DeepSeek AI
    Downloadable

    deepseek-r1-distill-qwen-14b

    Distilled version of Qwen 2.5 14B using reasoning data generated by DeepSeek R1 for enhanced performance.
    Model
    coding
    2.29K1.24M
    10mo
    DeepSeek AI
    Downloadable

    deepseek-r1-distill-qwen-32b

    Distilled version of Qwen 2.5 32B using reasoning data generated by DeepSeek R1 for enhanced performance.
    Model
    coding
    7.76K1.64M
    10mo
    DeepSeek AI
    Free Endpoint

    deepseek-v3.2

    State-of-the-art 685B reasoning LLM with sparse attention, long context, and integrated agentic tools.
    Model
    chat
    15.06M
    3mo
    Mistral AI
    Free Endpoint

    devstral-2-123b-instruct-2512

    State-of-the-art open code model with deep reasoning, 256k context, and unmatched efficiency.
    Model
    coding
    3.61M
    4mo
    Tiiuae
    Free Endpoint

    falcon3-7b-instruct

    Instruction tuned LLM achieving SoTA performance on reasoning, math and general knowledge capabilities
    Model
    chat
    421K
    10mo
    Google
    Downloadable

    gemma-4-31b-it

    Dense 31B model delivering frontier reasoning for coding, agentic workflows, and fine-tuning.
    Model
    reasoning
    437K
    6d
    Z.ai
    Free Endpoint

    glm-4.7

    GLM-4.7 is a multilingual agentic coding partner with stronger reasoning, tool use, and UI skills.
    Model
    Tool Calling
    15.62M
    2mo
    OpenAI
    Downloadable

    gpt-oss-20b

    Smaller Mixture of Experts (MoE) text-only LLM for efficient AI reasoning and math
    Model
    reasoning
    7.78M
    8mo
    Items per page
    of 3 pages