Skip to main content
NVIDIA
Explore
Models
Skills
Blueprints
GPUs
Docs
Help Center
Getting Started
  1. Create and verify your account to unlock full access to NVIDIA NIM APIs.
ResourcesDeveloper ForumsContact Support
FAQs
    Terms of Use
    Privacy Policy
    Your Privacy Choices
    Contact

    Copyright © 2026 NVIDIA Corporation

    14 results for

    Filters (1)

    Use Case
    Inference Providers
    Publisher
    Audience
    Blueprint Type
    Domain
    NIM Container GPUs
    Library
    Labels (1)

    Search results

    • Playbooks
      45 MIN

      Generate Images and Videos with ComfyUI

      Node-based diffusion workflows for images and videos with FLUX, Wan, HunyuanVideo, and Stable Diffusion
      Playbook
      • Content Creation
      • Image Generation
      • ComfyUI
      • DGX Spark
      • Docker
      • DGX Station
      Last updated on August 4, 2026
    Items per page
    of 1 pages
  • DGX Spark
    1 HR

    FLUX.1 Dreambooth LoRA Fine-tuning

    Fine-tune FLUX.1-dev 12B model using Dreambooth LoRA for custom image generation
    Playbook
    • Image Generation
    • ComfyUI
    • DGX
    • LoRA
    • Spark
    • Fine-tuning
    • Text-to-Image
    Last updated on October 7, 2025
  • Meta
    DownloadableFree Endpoint

    llama-3.2-11b-vision-instruct

    Cutting-edge vision-language model exceling in high-quality reasoning from images.
    Model
    • Image-Text Retrieval
    • Visual QA
    • Image Captioning
    • Visual Grounding
    • Image-to-Text
    3M API calls in the last 30 days
    Last updated on May 30, 2025
  • Meta
    DownloadableFree Endpoint

    llama-3.2-90b-vision-instruct

    Cutting-edge vision-Language model exceling in high-quality reasoning from images.
    Model
    • Image-Text Retrieval
    • Visual QA
    • image captioning
    • Visual Grounding
    • Image-to-Text
    4M API calls in the last 30 days
    Last updated on May 30, 2025
  • NVIDIA
    DownloadableFree Endpoint

    llama-3.1-nemotron-nano-vl-8b-v1

    Multi-modal vision-language model that understands text/img and creates informative responses
    Model
    • doc intelligence
    • multiple image understanding
    • OCR
    14M API calls in the last 30 days
    Last updated on July 1, 2025
  • DGX Spark
    1 HR

    Vision-Language Model Fine-tuning

    Fine-tune Vision-Language Models for image and video understanding tasks using Qwen2.5-VL and InternVL3
    Playbook
    • DGX
    • Image Understanding
    • Vision-Language Models
    • GRPO
    • Spark
    • Fine-tuning
    • Video Analysis
    Last updated on October 7, 2025
  • Black-forest-labs
    Downloadable

    flux.2-klein-4b

    FLUX.2-klein-4B is a distilled image generation and editing model, producing outputs at lighting speed
    Model
    • image editing
    • Run-on-RTX
    • Text-to-Image
    • Image Generation
    338K API calls in the last 30 days
    Last updated on March 13, 2026
  • Microsoft
    Downloadable

    TRELLIS

    MSFT TRELLIS is a 3D AI model that generates high-quality 3D assets from text or image inputs.
    Model
    • text-to-3d
    • Run-on-RTX
    • image-to-3d
    18K API calls in the last 30 days
    Last updated on September 3, 2025
  • Thinkingmachines
    DownloadableFree Endpoint

    inkling

    Inkling is a multimodal (text + image) reasoning model from Thinking Machines — a Mamba-hybrid, 256-expert Mixture-of-Experts architecture with tool use and switchable reasoning.
    Model
    • text-to-text
    • reasoning
    • image-to-text
    • multimodal
    Last updated on July 16, 2026
  • Google
    Free Endpoint

    paligemma

    Vision language model adept at comprehending text and visual inputs to produce informative responses
    Model
    • image
    • cv
    • Vision Assistant
    • vlm
    • Visual Question Answering
    • computer vision
    • Language Generation
    • video
    • Image-to-Text
    12K API calls in the last 30 days
    Last updated on August 26, 2024
  • NVIDIA
    Downloadable

    vista-3d

    VISTA-3D is a specialized interactive foundation model for segmenting and anotating human anatomies.
    Model
    • Interactive Annotation
    • Image Segmentation
    • Non-Commercial Use Only
    • Medical Imaging
    587 API calls in the last 30 days
    Last updated on April 21, 2025
  • NVIDIA
    DownloadableFree Endpoint

    cosmos3-nano

    Generates physics-aware videos from text prompts or an image prompt for physical AI development.
    Model
    • autonomous vehicles
    • Physical AI
    • robotics
    • text-to-world
    • image-to-world
    • Synthetic Data Generation
    2K API calls in the last 30 days
    Last updated on June 1, 2026
  • Meta
    DownloadableFree Endpoint

    muse-glimmer-30b

    Muse Glimmer 30B is a multimodal reasoning model accepting text and images, served on vLLM with native Onyx tool-calling and reasoning parsers.
    Model
    • Multimodal
    • Image-to-Text
    • Reasoning
    • Chat
    • Text-to-Text
    • Large Language Models
    Last updated on August 10, 2026
  • Robotics
    Enterprise

    Synthetic Manipulation Motion Generation for Robotics

    Generate exponentially large amounts of synthetic motion trajectories for robot manipulation from just a few human demonstrations.
    Blueprint
    • synthetic data
    • robotics
    • physical ai
    • robot learning
    • Humanoids
    • text-to-world
    • image-to-world
    • teleop
    • NVIDIA Isaac GR00T
    • NVIDIA Omniverse
    Last updated on February 17, 2026