Skip to main content
NVIDIA
Explore
Models
Skills
Blueprints
GPUs
Docs
Help Center
Getting Started
  1. Create and verify your account to unlock full access to NVIDIA NIM APIs.
ResourcesDeveloper ForumsContact Support
FAQs
  • Terms of Use
    Privacy Policy
    Your Privacy Choices
    Contact

    Copyright © 2026 NVIDIA Corporation

    Models

    Deploy and scale models on your GPU infrastructure of choice with NVIDIA NIM inference microservices

    Optimized by NVIDIALaunch from Hugging FaceBeta

    Filters (1)

    Use Case
    Inference Providers
    Publisher
    NIM Container GPUs
    Labels (1)
    14 models

    models

    • Moonshotai
      DownloadableFree Endpoint

      kimi-k3

      ~2.8T hybrid KDA+MLA multimodal MoE for long-horizon coding, agentic tool use, and image understanding.
      • Multimodal
      • Mixture-of-Experts
      • Reasoning
      • Image-to-Text
      Last updated on August 27, 2026
    Items per page
    of 1 pages
  • DeepSeek AI
    Free Endpoint

    deepseek-v4-pro-0813

    DeepSeek V4 scales to 1M-token context windows with efficient MoE architecture for coding tasks.
    • coding
    • Moe
    • reasoning
    • agentic
    Last updated on August 26, 2026
  • DeepSeek AI
    Free Endpoint

    deepseek-v4-flash-0731

    284B MoE (13B active) model ideal for long-context workloads optimized for coding, chat, and agentic workflows
    • MoE
    • Reasoning
    • Long Context
    • Hybrid Attention
    Last updated on August 19, 2026
  • Meta
    DownloadableFree Endpoint

    muse-glimmer-30b

    Muse Glimmer 30B is a multimodal reasoning model accepting text and images, with native tool-calling and separate reasoning output.
    • Multimodal
    • Image-to-Text
    • Reasoning
    • Chat
    • Text-to-Text
    • Large Language Models
    Last updated on August 10, 2026
  • Poolside
    Free Endpoint

    laguna-xs-2.1

    Efficient 33B MoE for local, long-horizon agentic coding and terminal tasks
    • Agentic AI
    • Coding
    • Reasoning
    • Tool Use
    Last updated on July 15, 2026
  • Minimaxai
    Deprecation in 7dFree Endpoint

    minimax-m3

    MiniMax M3 Preview is a multimodal MoE vision-language model with strong reasoning, coding, and tool-calling capabilities.
    • coding
    • text-to-text
    • reasoning
    10M API calls in the last 30 days
    Last updated on June 12, 2026
  • Google
    DownloadableFree Endpoint

    diffusiongemma-26b-a4b-it

    Diffusion-based 26B parameter LLM enabling parallel token generation for real-time text apps
    • diffusion-llm
    • text-to-text
    • reasoning
    4M API calls in the last 30 days
    Last updated on June 10, 2026
  • NVIDIA
    DownloadableFree Endpoint

    nemotron-3-ultra-550b-a55b

    Open, efficient hybrid Mamba-Transformer MoE with 1M context, excelling in agentic reasoning, coding, planning, tool calling, and more
    • Agent
    • MoE
    • Frontier
    • Reasoning
    • Long Context
    52M API calls in the last 30 days
    Last updated on June 4, 2026
  • NVIDIA
    DownloadableFree Endpoint

    cosmos3-nano-reasoner

    Vision language model that excels in understanding the physical world using structured reasoning on videos or images.
    • video understanding
    • autonomous vehicles
    • industrial
    • Physical AI
    • vision language model
    • reasoning
    • robotics
    • smart cities
    • Synthetic Data Generation
    2K API calls in the last 30 days
    Last updated on June 1, 2026
  • NVIDIA
    DownloadableFree Endpoint

    ising-calibration-1-35b-a3b

    Open VLM for quantum computer calibration chart understanding across a range of qubit modalities.
    • Quantum
    • reasoning
    • Vision Language Model
    • calibration
    442K API calls in the last 30 days
    Last updated on April 14, 2026
  • Google
    DownloadableFree Endpoint

    gemma-4-31b-it

    Dense 31B model delivering frontier reasoning for coding, agentic workflows, and fine-tuning.
    • coding
    • text-to-text
    • reasoning
    • agentic
    6M API calls in the last 30 days
    Last updated on April 2, 2026
  • NVIDIA
    DownloadableFree Endpoint

    nemotron-3-super-120b-a12b

    Open, efficient hybrid Mamba-Transformer MoE with 1M context, excelling in agentic reasoning, coding, planning, tool calling, and more
    • MoE
    • Reasoning
    • Chat
    • Long Context
    • Instruction Following
    65M API calls in the last 30 days
    Last updated on March 11, 2026
  • OpenAI
    DownloadableFree Endpoint

    gpt-oss-20b

    Smaller Mixture of Experts (MoE) text-only LLM for efficient AI reasoning and math
    • reasoning
    • text-to-text
    • chat
    • math
    19M API calls in the last 30 days
    Last updated on August 5, 2025
  • OpenAI
    DownloadableFree Endpoint

    gpt-oss-120b

    Mixture of Experts (MoE) reasoning LLM (text-only) designed to fit within 80GB GPU.
    • reasoning
    • text-to-text
    • chat
    • math
    45M API calls in the last 30 days
    Last updated on August 5, 2025