Skip to main content
NVIDIA
Explore
Models
Skills
Blueprints
GPUs
Docs
Help Center
Getting Started
  1. Create and verify your account to unlock full access to NVIDIA NIM APIs.
ResourcesDeveloper ForumsContact Support
FAQs
  • Models

    Deploy and scale models on your GPU infrastructure of choice with NVIDIA NIM inference microservices

    Filters

    Use Case
    Inference Providers
    Publisher
    NIM Container GPUs
    97 models
    Items per page
    of 5 pages
    Terms of Use
    Privacy Policy
    Your Privacy Choices
    Contact

    Copyright © 2026 NVIDIA Corporation

    models

    • Moonshotai
      DownloadableFree Endpoint

      kimi-k3

      ~2.8T hybrid KDA+MLA multimodal MoE for long-horizon coding, agentic tool use, and image understanding.
      • Multimodal
      • Mixture-of-Experts
      • Reasoning
      • Image-to-Text
      Last updated on August 27, 2026
    • DeepSeek AI
      Free Endpoint

      deepseek-v4-pro-0813

      DeepSeek V4 scales to 1M-token context windows with efficient MoE architecture for coding tasks.
      • coding
      • Moe
      • reasoning
      • agentic
      Last updated on August 26, 2026
    • Wan-ai
      Downloadable

      wan2.2-animate-2-14b

      Wan2.2-Animate-2 is a novel end-to-end character animation framework
      • video editing
      • character animation
      Last updated on August 19, 2026
    • DeepSeek AI
      Free Endpoint

      deepseek-v4-flash-0731

      284B MoE (13B active) model ideal for long-context workloads optimized for coding, chat, and agentic workflows
      • MoE
      • Reasoning
      • Long Context
      • Hybrid Attention
      Last updated on August 19, 2026
    • NVIDIA
      DownloadableFree Endpoint

      nemotron-3.5-lightning-30b-a3b

      Fastest 30B A3B MoE model with leading domain accuracy for specialized agentic tasks
      • Customization
      • Text-to-Text
      • Long-running agents
      • Open
      Last updated on August 11, 2026
    • Meta
      DownloadableFree Endpoint

      muse-glimmer-30b

      Muse Glimmer 30B is a multimodal reasoning model accepting text and images, with native tool-calling and separate reasoning output.
      • Multimodal
      • Image-to-Text
      • Reasoning
      • Chat
      • Text-to-Text
      • Large Language Models
      Last updated on August 10, 2026
    • NVIDIA
      Free Endpoint

      riva-translate-4b-instruct-v2

      Translation model in 37 languages with few-shots example prompts capability.
      • nvidia nim
      • neural machine translation
      • Text Translation
      Last updated on July 27, 2026
    • NVIDIA
      Free Endpoint

      ising-calibration-1.5-31b

      NVIDIA-Ising-Calibration-1.5 is a dense multimodal vision-language model built on Gemma 4 31B. It analyzes quantum computing calibration experiment plots and generates structured technical text.
      • Quantum Computing
      • Calibration
      • NVIDIA NIM
      • Vision Language Model
      Last updated on July 23, 2026
    • NVIDIA
      Downloadable

      Video Super Resolution NIM

      Upscale encoded or ST 2110 video to higher resolutions with NVIDIA Video Super Resolution.
      • broadcast
      • video upscaling
      • streaming
      • nvidia ai for media
      • video super resolution
      Last updated on July 22, 2026
    • NVIDIA
      Free Endpoint

      nemotron-3-embed-1b

      1B embedding model for semantic search, retrieval, and RAG applications.
      • Nemotron Retriever
      • Agentic Retrieval
      • Code Retrieval
      • Text-to-Embedding
      • Retrieval Augmented Generation
      Last updated on July 16, 2026
    • Poolside
      Free Endpoint

      laguna-xs-2.1

      Efficient 33B MoE for local, long-horizon agentic coding and terminal tasks
      • Agentic AI
      • Coding
      • Reasoning
      • Tool Use
      Last updated on July 15, 2026
    • NVIDIA
      Downloadable

      qwen-image-edit-nvpcb-ovsl2sl

      An image edit model specialized for Omniverse synthetic to photographic solder-light style captured at NVIDIA PCB inspection stations
      • Synthetic Data Generation
      • Image Generation
      • Physical AI
      Last updated on July 3, 2026
    • NVIDIA
      Downloadable

      nemotron-ocr-v2

      Nemotron OCR v2 is a state-of-the-art multilingual text recognition model designed for robust end-to-end optical character recognition (OCR) on complex real-world images.
      • Table Extraction
      • nemo retriever
      • data ingestion
      • extraction
      • Optical Character Recognition
      338K API calls in the last 30 days
      Last updated on June 24, 2026
    • Minimaxai
      Deprecation in 2dFree Endpoint

      minimax-m3

      MiniMax M3 Preview is a multimodal MoE vision-language model with strong reasoning, coding, and tool-calling capabilities.
      • coding
      • text-to-text
      • reasoning
    • Google
      DownloadableFree Endpoint

      diffusiongemma-26b-a4b-it

      Diffusion-based 26B parameter LLM enabling parallel token generation for real-time text apps
      • diffusion-llm
      • text-to-text
      • reasoning
      4M API calls in the last 30 days
      Last updated on June 10, 2026
    • NVIDIA
      DownloadableFree Endpoint

      nemotron-3-ultra-550b-a55b

      Open, efficient hybrid Mamba-Transformer MoE with 1M context, excelling in agentic reasoning, coding, planning, tool calling, and more
      • Agent
      • MoE
      • Frontier
      • Reasoning
      • Long Context
    • Resemble.AI
      Downloadable

      chatterbox-multilingual-tts

      Natural and expressive voices in 23 languages. For voice agents and brand ambassadors.
      • TTS
      • Chatterbox
      • Speech Generation
      • multilingual
      • Text-to-Speech
      22K API calls in the last 30 days
      Last updated on June 3, 2026
    • NVIDIA
      DownloadableFree Endpoint

      nemotron-3.5-content-safety

      Multilingual, multimodal model for detecting unsafe and toxic content.
      • llm safety
      • safety and moderation
      • multilingual content safety
      • ai safety nemo guardrails
    • NVIDIA
      DownloadableFree Endpoint

      cosmos3-nano

      Generates physics-aware videos from text prompts or an image prompt for physical AI development.
      • autonomous vehicles
      • Physical AI
      • robotics
      • text-to-world
      • image-to-world
      • Synthetic Data Generation
      2K API calls in the last 30 days
      Last updated on June 1, 2026
    • NVIDIA
      DownloadableFree Endpoint

      cosmos3-nano-reasoner

      Vision language model that excels in understanding the physical world using structured reasoning on videos or images.
      • video understanding
      • autonomous vehicles
      • industrial
      • Physical AI
      • vision language model
      • reasoning
      • robotics
      • smart cities
      • Synthetic Data Generation
    • Qwen
      Downloadable

      qwen-image

      Qwen-Image is a text-to-image foundation model with advanced multilingual text rendering.
      • Text-to-Image
      • Image Generation
      Last updated on May 1, 2026
    • Qwen
      Downloadable

      qwen-image-edit

      Qwen-Image-Edit is an image editing model with multilingual text editing and strong subject consistency.
      • Text-to-Image
      • Image Generation
      Last updated on May 1, 2026
    • NVIDIA
      DownloadableFree Endpoint

      nemotron-3-nano-omni-30b-a3b-reasoning

      Nemotron 3 Nano Omni is an omni-modal reasoning model that understands images, video, speech, text.
      • Image-to-Text
      • VLM
      • Video
      • Omni
      • OCR
      8M API calls in the last 30 days
      Last updated on April 28, 2026
    • NVIDIA
      Downloadable

      Relighting

      Re-illuminate people in video to match target lighting from a 360 HDRI environment map.
      • HDRI
      • remote contribution
      • lighting
      • nvidia ai for media
      242 API calls in the last 30 days
      Last updated on April 17, 2026
    Optimized by NVIDIALaunch from Hugging FaceBeta
    10M API calls in the last 30 days
    Last updated on June 12, 2026
    52M API calls in the last 30 days
    Last updated on June 4, 2026
    2M API calls in the last 30 days
    Last updated on June 2, 2026
    2K API calls in the last 30 days
    Last updated on June 1, 2026