NVIDIA
Explore
Models
Blueprints
GPUs
Docs
⌘KCtrl+K
Terms of Use
Privacy Policy
Your Privacy Choices
Contact

Copyright © 2026 NVIDIA Corporation

62 results for

Filters

  • Free Endpoint
    28
  • Partner Endpoint
    41
  • Download Available
    34
  • Enterprise
    0
  • Launchable
    0
  • Code Generation
    14
  • Image-to-Text
    7
  • Synthetic Data Generation
    1
  • Deep Infra
    31
  • Together AI
    27
  • GMI Cloud
    16
  • Bitdeer AI
    12
  • CoreWeave
    9
  • NVIDIA
    11
  • Microsoft
    9
  • Meta
    8
  • Mistral AI
    7
  • Qwen
    6
  • NVIDIA AI
    0
  • NVIDIA
    Free Endpoint

    nemotron-content-safety-reasoning-4b

    A context‑aware safety model that applies reasoning to enforce domain‑specific policies.
    Model
    NeMo Guardrails
    207K
    2mo
    Microsoft
    DeprecatedFree Endpoint

    phi-4-mini-flash-reasoning

    Lightweight reasoning model for applications in latency bound, memory/compute constrained environments
    Model
    edge
    162K
    9mo
    Mistral AI
    Free Endpoint

    magistral-small-2506

    High performance reasoning model optimized for efficiency and edge deployment
    Model
    coding
    1.18M
    9mo
    Marin
    DeprecatedFree Endpoint

    marin-8b-instruct

    State-of-the-art open model trained on open datasets, excelling in reasoning, math, and science.
    Model
    Reasoning
    148K
    10mo
    Qwen
    Downloadable

    qwen3-next-80b-a3b-thinking

    80B parameter AI model with hybrid reasoning, MoE architecture, support for 119 languages.
    Model
    Reasoning
    1.86M
    7mo
    ByteDance
    Free Endpoint

    seed-oss-36b-instruct

    ByteDance open-source LLM with long-context, reasoning, and agentic intelligence.
    Model
    thinking budget
    1.36M
    7mo
    NVIDIA
    Deprecation in 1dDownloadable

    llama-3.1-nemotron-nano-4b-v1.1

    State-of-the-art open model for reasoning, code, math, and tool calling - suitable for edge agents
    Model
    edge
    82.25K
    9mo
    NVIDIA
    Downloadable

    ising-calibration-1-35b-a3b

    Open VLM for quantum computer calibration chart understanding across a range of qubit modalities.
    Model
    Quantum
    1d
    Moonshotai
    Free Endpoint

    kimi-k2-instruct

    State-of-the-art open mixture-of-experts model with strong reasoning, coding, and agentic capabilities
    Model
    coding
    21.04M
    9mo
    NVIDIA
    Downloadable

    llama-3.1-nemotron-nano-8b-v1

    Leading reasoning and agentic AI accuracy model for PC and edge.
    Model
    math
    622K
    9mo
    NVIDIA
    Downloadable

    nvidia-nemotron-nano-9b-v2

    High‑efficiency LLM with hybrid Transformer‑Mamba design, excelling in reasoning and agentic tasks.
    Model
    thinking budget
    285K
    8mo
    Stepfun-ai
    Free Endpoint

    step-3.5-flash

    200B open-source reasoning engine with sparse MoE powering frontier agentic AI.
    Model
    Agentic
    9.4M
    2mo
    NVIDIA
    Deprecation in 6dDownloadable

    llama-3.1-nemotron-ultra-253b-v1

    Superior inference efficiency with highest accuracy for scientific and complex math reasoning, coding, tool calling, and instruction following.
    Model
    math
    5.15M
    9mo
    Mistral AI
    DeprecatedDownloadable

    mistral-small-24b-instruct

    Latency-optimized language model excelling in code, math, general knowledge, and instruction-following.
    Model
    code
    198K
    9mo
    Qwen
    DeprecatedFree Endpoint

    qwq-32b

    Powerful reasoning model capable of thinking and reasoning, can achieve significantly enhanced performance in downstream tasks, especially hard problems.
    Model
    coding
    1.13M
    9mo
    Z.ai
    Free Endpoint

    glm-4.7

    GLM-4.7 is a multilingual agentic coding partner with stronger reasoning, tool use, and UI skills.
    Model
    Tool Calling
    16.27M
    2mo
    Z.ai
    Deprecation in 4dDownloadable

    glm-5

    GLM-5 744B MoE enables efficient reasoning for complex systems and long-horizon agentic tasks.
    Model
    MoE
    41.67M
    2mo
    Moonshotai
    Free Endpoint

    kimi-k2-instruct-0905

    Follow-on version of Kimi-K2-Instruct with longer context window and enhanced reasoning capabilities.
    Model
    long-context
    14.63M
    6mo
    Moonshotai
    Free Endpoint

    kimi-k2-thinking

    Open reasoning model with 256K context window, native INT4 quantization and enhanced tool use.
    Model
    Conversational
    4.09M
    4mo
    NVIDIA
    Downloadable

    llama-3.3-nemotron-super-49b-v1

    High efficiency model with leading accuracy for reasoning, tool calling, chat, and instruction following.
    Model
    math
    1.39M
    9mo
    NVIDIA
    Downloadable

    llama-3.3-nemotron-super-49b-v1.5

    High efficiency model with leading accuracy for reasoning, tool calling, chat, and instruction following.
    Model
    math
    2.54M
    8mo
    Qwen
    Downloadable

    qwen3.5-122b-a10b

    122B MoE LLM (10B active) for coding, reasoning, multimodal chat. Agent-ready.
    Model
    tool calling
    8.34M
    1mo
    Mistral AI
    Downloadable

    mixtral-8x7b-instruct-v0.1

    An MOE LLM that follows instructions, completes requests, and generates creative text.
    Model
    Advanced Reasoning
    353K
    9mo
    Sarvamai
    Downloadable

    sarvam-m

    Multilingual, hybrid-reasoning model optimized for Indian language tasks, programming, mathematical reasoning capabilities.
    Model
    coding
    134K
    8mo
    Items per page
    of 3 pages