NVIDIA
Explore
Models
Blueprints
GPUs
Docs
⌘KCtrl+K
Terms of Use
Privacy Policy
Your Privacy Choices
Contact

Copyright © 2026 NVIDIA Corporation

Models

Deploy and scale models on your GPU infrastructure of choice with NVIDIA NIM inference microservices

Optimized by NVIDIALaunch from Hugging FaceBeta

Filters (2)

  • Free Endpoint
    9
  • Partner Endpoint
    12
  • Download Available
    5
  • Code Generation
    1
  • Drug Discovery
    0
  • Retrieval Augmented Generation
    0
  • Image-to-Text
    0
  • Object Detection
    0
  • Fireworks AI
    8
  • Deep Infra
    7
  • Together AI
    6
  • GMI Cloud
    5
  • Bitdeer AI
    3
  • DeepSeek AI
    3
  • Mistral AI
    2
  • Moonshotai
    2
  • Qwen
    1
  • IBM
    1
  • Enterprise
    0
  • NVIDIA BioNemo
    0
  • reasoning
  • coding
  • 14 models
    Minimaxai
    Downloadable

    minimax-m2.5

    MiniMax M2.5 is a 230B-parameter text-to-text AI model excelling in coding, reasoning, and office tasks.
    reasoning
    8.25M
    3w
    Stepfun-ai
    Free Endpoint

    step-3.5-flash

    200B open-source reasoning engine with sparse MoE powering frontier agentic AI.
    chat
    9.09M
    1mo
    Z.ai
    Free Endpoint

    glm-4.7

    GLM-4.7 is a multilingual agentic coding partner with stronger reasoning, tool use, and UI skills.
    Tool Calling
    15.18M
    1mo
    Mistral AI
    Free Endpoint

    devstral-2-123b-instruct-2512

    State-of-the-art open code model with deep reasoning, 256k context, and unmatched efficiency.
    coding
    6.1M
    3mo
    Moonshotai
    Free Endpoint

    kimi-k2-instruct-0905

    Follow-on version of Kimi-K2-Instruct with longer context window and enhanced reasoning capabilities.
    long-context
    11.44M
    6mo
    Sarvamai
    Downloadable

    sarvam-m

    Multilingual, hybrid-reasoning model optimized for Indian language tasks, programming, mathematical reasoning capabilities.
    coding
    586K
    8mo
    Moonshotai
    Free Endpoint

    kimi-k2-instruct

    State-of-the-art open mixture-of-experts model with strong reasoning, coding, and agentic capabilities
    coding
    20.28M
    8mo
    Mistral AI
    Free Endpoint

    magistral-small-2506

    High performance reasoning model optimized for efficiency and edge deployment
    coding
    4.17M
    8mo
    IBM
    Free Endpoint

    granite-3.3-8b-instruct

    Small language model fine-tuned for improved reasoning, coding, and instruction-following
    coding
    113K
    8mo
    Qwen
    Free Endpoint

    qwq-32b

    Powerful reasoning model capable of thinking and reasoning, can achieve significantly enhanced performance in downstream tasks, especially hard problems.
    coding
    4.27M
    9mo
    DeepSeek AI
    Downloadable

    deepseek-r1-distill-llama-8b

    Distilled version of Llama 3.1 8B using reasoning data generated by DeepSeek R1 for enhanced performance.
    Distillation
    4.81M
    8mo
    DeepSeek AI
    Downloadable

    deepseek-r1-distill-qwen-32b

    Distilled version of Qwen 2.5 32B using reasoning data generated by DeepSeek R1 for enhanced performance.
    coding
    2.29K5.12M
    10mo
    DeepSeek AI
    Downloadable

    deepseek-r1-distill-qwen-14b

    Distilled version of Qwen 2.5 14B using reasoning data generated by DeepSeek R1 for enhanced performance.
    coding
    1.87K4.52M
    10mo
    Tiiuae
    Free Endpoint

    falcon3-7b-instruct

    Instruction tuned LLM achieving SoTA performance on reasoning, math and general knowledge capabilities
    chat
    2.02M
    10mo
    Items per page
    of 1 pages