Skip to main content
NVIDIA
Explore
Models
Skills
Blueprints
GPUs
Docs
Help Center
Getting Started
  1. Create and verify your account to unlock full access to NVIDIA NIM APIs.
ResourcesDeveloper ForumsContact Support
FAQs
  • Terms of Use
    Privacy Policy
    Your Privacy Choices
    Contact

    Copyright © 2026 NVIDIA Corporation

    24 results for

    Filters

    Use Case
    Inference Providers
    Publisher
    Audience
    Domain
    Library

    Search results

    • Meta
      DownloadableFree Endpoint

      llama-3.2-11b-vision-instruct

      Cutting-edge vision-language model exceling in high-quality reasoning from images.
      Model
      • Image-Text Retrieval
      • Visual QA
      • Image Captioning
      • Visual Grounding
      • Image-to-Text
    Items per page
    of 1 pages
    3M API calls in the last 30 days
    Last updated on May 30, 2025
  • Meta
    DownloadableFree Endpoint

    llama-3.2-90b-vision-instruct

    Cutting-edge vision-Language model exceling in high-quality reasoning from images.
    Model
    • Image-Text Retrieval
    • Visual QA
    • image captioning
    • Visual Grounding
    • Image-to-Text
    4M API calls in the last 30 days
    Last updated on May 30, 2025
  • NVIDIA
    Downloadable

    llama-3.1-nemoguard-8b-content-safety

    Leading content safety model for enhancing the safety and moderation capabilities of LLMs
    Model
    • nemo guardrails
    • LLM safety
    • Safety and moderation
    • dialogue safety
    • nemotron
    342K API calls in the last 30 days
    Last updated on March 25, 2025
  • NVIDIA
    Downloadable

    llama-3.1-nemoguard-8b-topic-control

    Topic control model to keep conversations focused on approved topics, avoiding inappropriate content.
    Model
    • nemo guardrails
    • LLM safety
    • Safety and moderation
    • dialogue safety
    • nemotron
    491K API calls in the last 30 days
    Last updated on March 25, 2025
  • NVIDIA
    Free Endpoint

    llama-3.1-nemotron-safety-guard-8b-v3

    Leading multilingual content safety model for enhancing the safety and moderation capabilities of LLMs
    Model
    • content moderation
    • llm safety
    • multilingual guard model
    • multilingual content safety
    • nemoguard
    344K API calls in the last 30 days
    Last updated on October 28, 2025
  • Google
    DownloadableFree Endpoint

    gemma-4-31b-it

    Dense 31B model delivering frontier reasoning for coding, agentic workflows, and fine-tuning.
    Model
    • reasoning
    • coding
    • text-to-text
    • agentic
    6M API calls in the last 30 days
    Last updated on April 2, 2026
  • NVIDIA
    DownloadableFree Endpoint

    ising-calibration-1-35b-a3b

    Open VLM for quantum computer calibration chart understanding across a range of qubit modalities.
    Model
    • Quantum
    • reasoning
    • Vision Language Model
    • calibration
    442K API calls in the last 30 days
    Last updated on April 14, 2026
  • NVIDIA
    Free Endpoint

    ising-calibration-1.5-31b

    NVIDIA-Ising-Calibration-1.5 is a dense multimodal vision-language model built on Gemma 4 31B. It analyzes quantum computing calibration experiment plots and generates structured technical text.
    Model
    • Quantum Computing
    • Calibration
    • NVIDIA NIM
    • Vision Language Model
    Last updated on July 23, 2026
  • Meta
    DownloadableFree Endpoint

    muse-glimmer-30b

    Muse Glimmer 30B is a multimodal reasoning model accepting text and images, with native tool-calling and separate reasoning output.
    Model
    • Multimodal
    • Image-to-Text
    • Reasoning
    • Chat
    • Text-to-Text
    • Large Language Models
    Last updated on August 10, 2026
  • NVIDIA
    Free Endpoint

    nemotron-3-embed-1b

    1B embedding model for semantic search, retrieval, and RAG applications.
    Model
    • Nemotron Retriever
    • Agentic Retrieval
    • Code Retrieval
    • Text-to-Embedding
    • Retrieval Augmented Generation
    Last updated on July 16, 2026
  • NVIDIA
    DownloadableFree Endpoint

    nemotron-3-nano-omni-30b-a3b-reasoning

    Nemotron 3 Nano Omni is an omni-modal reasoning model that understands images, video, speech, text.
    Model
    • Image-to-Text
    • VLM
    • Video
    • Omni
    • OCR
    8M API calls in the last 30 days
    Last updated on April 28, 2026
  • NVIDIA
    DownloadableFree Endpoint

    nemotron-3-super-120b-a12b

    Open, efficient hybrid Mamba-Transformer MoE with 1M context, excelling in agentic reasoning, coding, planning, tool calling, and more
    Model
    • MoE
    • Reasoning
    • Chat
    • Long Context
    • Instruction Following
    65M API calls in the last 30 days
    Last updated on March 11, 2026
  • NVIDIA
    DownloadableFree Endpoint

    nemotron-3-ultra-550b-a55b

    Open, efficient hybrid Mamba-Transformer MoE with 1M context, excelling in agentic reasoning, coding, planning, tool calling, and more
    Model
    • Agent
    • MoE
    • Frontier
    • Reasoning
    • Long Context
    52M API calls in the last 30 days
    Last updated on June 4, 2026
  • NVIDIA
    DownloadableFree Endpoint

    nemotron-3.5-content-safety

    Multilingual, multimodal model for detecting unsafe and toxic content.
    Model
    • llm safety
    • safety and moderation
    • multilingual content safety
    • ai safety nemo guardrails
    2M API calls in the last 30 days
    Last updated on June 2, 2026
  • NVIDIA
    DownloadableFree Endpoint

    nemotron-3.5-lightning-30b-a3b

    Fastest 30B A3B MoE model with leading domain accuracy for specialized agentic tasks
    Model
    • Customization
    • Text-to-Text
    • Long-running agents
    • Open
    Last updated on August 11, 2026
  • Stability AI
    Downloadable

    stable-diffusion-3.5-large

    Stable Diffusion 3.5 is a popular text-to-image generation model
    Model
    • Text-to-Image
    • Image Generation
    Last updated on August 12, 2025
  • Deploy and operate the RTVI-CV-3D microservice as MV3DT (`MODE=mv3dt`): per-camera DeepStream perception plus BEV Fusion over calibrated cameras. Supports the bundled sample dataset, custom video files, and RTSP streams, and chains to `vss-generate-video-
    Skill
    • Video Search and Summarization (VSS)
    • AI And Machine Learning
    • DevOps Engineer
    • Platform Engineer
    • Application Developer
    • Solutions Architect
    2K downloads in the last 30 days
    Last updated on June 13, 2026
  • Playbooks

    LLaMA Factory

    Install and fine-tune models with LLaMA Factory
    Playbook
    • DGX
    • Spark
    Last updated on October 9, 2025
  • Meta
    Free Endpoint

    llama-guard-4-12b

    Multi-modal model to classify safety for input prompts as well output responses.
    Model
    • LLM Multimodal Safety
    • Content Safety
    • Guardrail
    • Content Moderator
    357K API calls in the last 30 days
    Last updated on July 1, 2025
  • NVIDIA
    Downloadable

    llama-nemotron-embed-vl-1b-v2

    Multimodal question-answer retrieval representing user queries as text and documents as images.
    Model
    • nemo retriever
    • embedding
    • Text-to-Embedding
    • Retrieval Augmented Generation
    16M API calls in the last 30 days
    Last updated on February 10, 2026
  • NVIDIA
    Downloadable

    llama-nemotron-rerank-vl-1b-v2

    GPU-accelerated model optimized for providing a probability score that a given passage contains the information to answer a question.
    Model
    • nemo retriever
    • reranking
    • Retrieval Augmented Generation
    1M API calls in the last 30 days
    Last updated on March 31, 2026
  • Playbooks
    30 MIN

    Run models with llama.cpp on DGX Spark

    Build llama.cpp with CUDA and serve models via an OpenAI-compatible API
    Playbook
    • DGX Spark
    • Inference
    • LLM
    • llama.cpp
    Last updated on April 2, 2026
  • DGX Station
    2 HRS

    Profiler-Driven Kernel Optimization for Fine-Tuning

    Use torch.profiler to find training bottlenecks, then write custom Triton kernels to optimize LLaMA 8B fine-tuning
    Playbook
    • Training
    • Fine-Tuning
    • Performance Optimization
    • Kernel Development
    • DGX Station
    • LLaMA
    • Triton
    • GB300
    Last updated on May 26, 2026
  • DGX Station
    30 MIN

    NVFP4 Pretraining with Megatron Bridge

    Pretrain Llama 3.1 8B with NVFP4 mixed precision on DGX Station using Megatron Bridge
    Playbook
    • Training
    • NVFP4
    • Megatron Bridge
    Last updated on May 27, 2026