Skip to main content
NVIDIA
Explore
Models
Skills
Blueprints
GPUs
Docs
Help Center
Getting Started
  1. Create and verify your account to unlock full access to NVIDIA NIM APIs.
ResourcesDeveloper ForumsContact Support
FAQs
    Terms of Use
    Privacy Policy
    Your Privacy Choices
    Contact

    Copyright © 2026 NVIDIA Corporation

    14 results for

    Filters (1)

    Use Case
    Inference Providers
    Publisher
    Audience
    Blueprint Type
    Domain
    Library
    Labels (1)

    Search results

    • Google
      DownloadableFree Endpoint

      diffusiongemma-26b-a4b-it

      Diffusion-based 26B parameter LLM enabling parallel token generation for real-time text apps
      Model
      • diffusion-llm
      • text-to-text
      • reasoning
    Items per page
    of 1 pages
    4M API calls in the last 30 days
    Last updated on June 10, 2026
  • OpenAI
    DownloadableFree Endpoint

    gpt-oss-20b

    Smaller Mixture of Experts (MoE) text-only LLM for efficient AI reasoning and math
    Model
    • text-to-text
    • chat
    • reasoning
    • math
    19M API calls in the last 30 days
    Last updated on August 5, 2025
  • NVIDIA
    Deprecation in 14dFree Endpoint

    nemotron-mini-4b-instruct

    Optimized SLM for on-device inference and fine-tuned for roleplay, RAG and function calling
    Model
    • Chat
    • Text-to-Text
    • Language Generation
    3M API calls in the last 30 days
    Last updated on August 26, 2024
  • Google
    DownloadableFree Endpoint

    gemma-4-31b-it

    Dense 31B model delivering frontier reasoning for coding, agentic workflows, and fine-tuning.
    Model
    • coding
    • text-to-text
    • reasoning
    • agentic
    6M API calls in the last 30 days
    Last updated on April 2, 2026
  • Thinkingmachines
    DownloadableFree Endpoint

    inkling

    Inkling is a multimodal (text + image) reasoning model from Thinking Machines — a Mamba-hybrid, 256-expert Mixture-of-Experts architecture with tool use and switchable reasoning.
    Model
    • text-to-text
    • reasoning
    • image-to-text
    • multimodal
    Last updated on July 16, 2026
  • Meta
    DownloadableFree Endpoint

    llama-3.1-70b-instruct

    Powers complex conversations with superior contextual understanding, reasoning and text generation.
    Model
    • Chat
    • Text-to-Text
    • Language Generation
    • Code Generation
    5M API calls in the last 30 days
    Last updated on June 12, 2025
  • Meta
    DownloadableFree Endpoint

    llama-3.1-8b-instruct

    Advanced state-of-the-art model with language understanding, superior reasoning, and text generation.
    Model
    • Chat
    • Text-to-Text
    • Language Generation
    • Run-on-RTX
    • Code Generation
    19M API calls in the last 30 days
    Last updated on July 9, 2025
  • Meta
    DownloadableFree Endpoint

    llama-3.2-1b-instruct

    Advanced state-of-the-art small language model with language understanding, superior reasoning, and text generation.
    Model
    • chat
    • Text-to-Text
    • Language Generation
    • Code Generation
    40K downloads in the last 30 days
    545K API calls in the last 30 days
    Last updated on May 21, 2025
  • Meta
    DownloadableFree Endpoint

    llama-3.2-3b-instruct

    Advanced state-of-the-art small language model with language understanding, superior reasoning, and text generation.
    Model
    • Chat
    • Text-to-Text
    • Language Generation
    • Code Generation
    27K downloads in the last 30 days
    1M API calls in the last 30 days
    Last updated on May 22, 2025
  • Meta
    DownloadableFree Endpoint

    llama-3.3-70b-instruct

    Advanced LLM for reasoning, math, general knowledge, and function calling
    Model
    • Reasoning
    • Text-to-Text
    • Instruction following
    • Math
    • Code Generation
    27M API calls in the last 30 days
    Last updated on June 12, 2025
  • Minimaxai
    Free Endpoint

    minimax-m3

    MiniMax M3 Preview is a multimodal MoE vision-language model with strong reasoning, coding, and tool-calling capabilities.
    Model
    • coding
    • text-to-text
    • reasoning
    10M API calls in the last 30 days
    Last updated on June 12, 2026
  • NVIDIA
    DownloadableFree Endpoint

    nemotron-3.5-lightning-30b-a3b

    Fastest 30B A3B MoE model with leading domain accuracy for specialized agentic tasks
    Model
    • Customization
    • Text-to-Text
    • Long-running agents
    • Open
    Last updated on August 11, 2026
  • OpenAI
    DownloadableFree Endpoint

    gpt-oss-120b

    Mixture of Experts (MoE) reasoning LLM (text-only) designed to fit within 80GB GPU.
    Model
    • reasoning
    • text-to-text
    • chat
    • math
    45M API calls in the last 30 days
    Last updated on August 5, 2025
  • Meta
    DownloadableFree Endpoint

    muse-glimmer-30b

    Muse Glimmer 30B is a multimodal reasoning model accepting text and images, served on vLLM with native Onyx tool-calling and reasoning parsers.
    Model
    • Multimodal
    • Image-to-Text
    • Reasoning
    • Chat
    • Text-to-Text
    • Large Language Models
    Last updated on August 10, 2026