Skip to main content
NVIDIA
Explore
Models
Skills
Blueprints
GPUs
Docs
Help Center
Getting Started
  1. Create and verify your account to unlock full access to NVIDIA NIM APIs.
ResourcesDeveloper ForumsContact Support
FAQs
  • Terms of Use
    Privacy Policy
    Your Privacy Choices
    Contact

    Copyright © 2026 NVIDIA Corporation

    Models

    Deploy and scale models on your GPU infrastructure of choice with NVIDIA NIM inference microservices

    Optimized by NVIDIALaunch from Hugging FaceBeta

    Filters (1)

    Use Case
    Inference Providers
    Publisher
    NIM Container GPUs
    Labels (1)
    7 models

    models

    • NVIDIA
      DownloadableFree Endpoint

      nemotron-3.5-lightning-30b-a3b

      Fastest 30B A3B MoE model with leading domain accuracy for specialized agentic tasks
      • Customization
      • Text-to-Text
      • Long-running agents
      • Open
      Last updated on August 11, 2026
    Items per page
    of 1 pages
  • Meta
    DownloadableFree Endpoint

    muse-glimmer-30b

    Muse Glimmer 30B is a multimodal reasoning model accepting text and images, with native tool-calling and separate reasoning output.
    • Multimodal
    • Reasoning
    • Chat
    • Text-to-Text
    • Image-to-Text
    • Large Language Models
    Last updated on August 10, 2026
  • Minimaxai
    Deprecation in 8dFree Endpoint

    minimax-m3

    MiniMax M3 Preview is a multimodal MoE vision-language model with strong reasoning, coding, and tool-calling capabilities.
    • coding
    • text-to-text
    • reasoning
    10M API calls in the last 30 days
    Last updated on June 12, 2026
  • Google
    DownloadableFree Endpoint

    diffusiongemma-26b-a4b-it

    Diffusion-based 26B parameter LLM enabling parallel token generation for real-time text apps
    • diffusion-llm
    • text-to-text
    • reasoning
    4M API calls in the last 30 days
    Last updated on June 10, 2026
  • Google
    DownloadableFree Endpoint

    gemma-4-31b-it

    Dense 31B model delivering frontier reasoning for coding, agentic workflows, and fine-tuning.
    • coding
    • text-to-text
    • reasoning
    • agentic
    6M API calls in the last 30 days
    Last updated on April 2, 2026
  • OpenAI
    DownloadableFree Endpoint

    gpt-oss-20b

    Smaller Mixture of Experts (MoE) text-only LLM for efficient AI reasoning and math
    • reasoning
    • text-to-text
    • chat
    • math
    19M API calls in the last 30 days
    Last updated on August 5, 2025
  • OpenAI
    DownloadableFree Endpoint

    gpt-oss-120b

    Mixture of Experts (MoE) reasoning LLM (text-only) designed to fit within 80GB GPU.
    • reasoning
    • text-to-text
    • chat
    • math
    45M API calls in the last 30 days
    Last updated on August 5, 2025