Skip to main content
NVIDIA
Explore
Models
Skills
Blueprints
GPUs
Docs
Help Center
Getting Started
  1. Create and verify your account to unlock full access to NVIDIA NIM APIs.
ResourcesDeveloper ForumsContact Support
FAQs
  • Terms of Use
    Privacy Policy
    Your Privacy Choices
    Contact

    Copyright © 2026 NVIDIA Corporation

    Models

    Deploy and scale models on your GPU infrastructure of choice with NVIDIA NIM inference microservices

    Optimized by NVIDIALaunch from Hugging FaceBeta

    Filters (1)

    Use Case
    Inference Providers
    Publisher
    NIM Container GPUs
    Labels (1)
    8 models

    models

    • DeepSeek AI
      Free Endpoint

      deepseek-v4-flash-0731

      284B MoE (13B active) model ideal for long-context workloads optimized for coding, chat, and agentic workflows
      • MoE
      • Reasoning
      • Long Context
      • Hybrid Attention
      Last updated on August 19, 2026
    Items per page
    of 1 pages
  • Meta
    DownloadableFree Endpoint

    muse-glimmer-30b

    Muse Glimmer 30B is a multimodal reasoning model accepting text and images, with native tool-calling and separate reasoning output.
    • Multimodal
    • Image-to-Text
    • Reasoning
    • Chat
    • Text-to-Text
    • Large Language Models
    Last updated on August 10, 2026
  • NVIDIA
    DownloadableFree Endpoint

    nemotron-3-nano-omni-30b-a3b-reasoning

    Nemotron 3 Nano Omni is an omni-modal reasoning model that understands images, video, speech, text.
    • Image-to-Text
    • VLM
    • Video
    • Omni
    • OCR
    8M API calls in the last 30 days
    Last updated on April 28, 2026
  • NVIDIA
    DownloadableFree Endpoint

    ising-calibration-1-35b-a3b

    Open VLM for quantum computer calibration chart understanding across a range of qubit modalities.
    • Quantum
    • reasoning
    • Vision Language Model
    • calibration
    442K API calls in the last 30 days
    Last updated on April 14, 2026
  • NVIDIA
    Free Endpoint

    nemotron-voicechat

    Nemotron 3 Voicechat
    • English
    • voice chat
    • NVIDIA NIM
    1K API calls in the last 30 days
    Last updated on March 16, 2026
  • NVIDIA
    DownloadableFree Endpoint

    nemotron-3-super-120b-a12b

    Open, efficient hybrid Mamba-Transformer MoE with 1M context, excelling in agentic reasoning, coding, planning, tool calling, and more
    • MoE
    • Reasoning
    • Chat
    • Long Context
    • Instruction Following
    65M API calls in the last 30 days
    Last updated on March 11, 2026
  • OpenAI
    DownloadableFree Endpoint

    gpt-oss-20b

    Smaller Mixture of Experts (MoE) text-only LLM for efficient AI reasoning and math
    • reasoning
    • text-to-text
    • chat
    • math
    19M API calls in the last 30 days
    Last updated on August 5, 2025
  • OpenAI
    DownloadableFree Endpoint

    gpt-oss-120b

    Mixture of Experts (MoE) reasoning LLM (text-only) designed to fit within 80GB GPU.
    • reasoning
    • text-to-text
    • chat
    • math
    45M API calls in the last 30 days
    Last updated on August 5, 2025