Skip to main content
NVIDIA
Explore
Models
Skills
Blueprints
GPUs
Docs
ForumsSupport
Terms of Use
Privacy Policy
Your Privacy Choices
Contact

Copyright © 2026 NVIDIA Corporation

42 results for

Filters

Use Case
Inference Providers
Publisher
Audience
Blueprint Type
Domain
Library
NVIDIA
DownloadableFree Endpoint

synthetic-video-detector

NVIDIA Synthetic Video Detector is an AI-powered micro-service for detecting AI‑generated (synthetic) videos.
Model
broadcast
media2forensicsnvidia ai for mediadiffusion models
Items per page
of 2 pages
90K API calls in the last 30 days
Last updated on April 16, 2026

Use this skill when deploying, operating, or integrating the VSS 3.2 GA RT-Embed Video Embedding microservice. Covers Docker Compose bring-up, GPU and storage prerequisites, the `/v1` REST API (file uploads, text and video embeddings, live RTSP streams, h
Skill
Video Search and Summarization (VSS)
AI And Machine LearningDevOps EngineerMl EngineerPlatform EngineerApplication DeveloperSolutions Architect
2K downloads in the last 30 days
Last updated on June 13, 2026
npx skills add NVIDIA/skills --skill vss-deploy-video-embedding

Use this skill when producing a VSS analysis report — Mode A per-clip VLM, Mode B incident-range via video-analytics. Not for standalone video summarization, real-time alerts or ad-hoc Q&A.
Skill
Video Search and Summarization (VSS)
AI EngineerMl EngineerApplication DeveloperSolutions ArchitectAI And Machine Learning
2K downloads in the last 30 days
Last updated on June 13, 2026
npx skills add NVIDIA/skills --skill vss-generate-video-report
RTX Workstation
18 MIN

NVIDIA Video Generation Guide

Learn how to create videos using LTX-2 in ComfyUI, accelerated on RTX. Learn how to take control of visual generative AI, creating high resolution video on RTX.
Playbook
ComfyUI
LTX-2RTX
Last updated on June 1, 2026

Use to summarize a recorded video via the LVS summarization microservice (HITL-gated) with a VLM fallback. Not for report generation or live RTSP captioning.
Skill
Developer
AI EngineerMl EngineerApplication DeveloperSolutions ArchitectVideo Search and Summarization (VSS)AI And Machine Learning
2K downloads in the last 30 days
Last updated on June 13, 2026
npx skills add NVIDIA/skills --skill vss-summarize-video

Calibrate a new dataset from pre-recorded video files via the AutoMagicCalib REST API. Use when user has local MP4s and says 'calibrate my videos', 'run AMC on these videos', or similar. For RTSP/live streams, use amc-run-rtsp-calibration instead.
Skill
Developer
AI EngineerMl EngineerApplication DeveloperSolutions ArchitectAI And Machine LearningDeepStream SDK
519 downloads in the last 30 days
Last updated on July 6, 2026
npx skills add NVIDIA/skills --skill amc-run-video-calibration

Use this skill to ask the VSS agent's video_understanding tool a fresh visual question about a recorded clip. Not for prior tool output, search hits, or metadata-answerable questions.
Skill
Developer
AI EngineerMl EngineerApplication DeveloperVideo Search and Summarization (VSS)AI And Machine Learning
2K downloads in the last 30 days
Last updated on June 13, 2026
npx skills add NVIDIA/skills --skill vss-ask-video

Use to run AutoMagicCalib on local MP4s, RTSP, or the bundled sample dataset, and to deploy vss-auto-calibration when needed. Do not use for non-AMC calibration or runtime analytics.
Skill
Video Search and Summarization (VSS)
AI EngineerDevOps EngineerHands On BuilderApplication DeveloperSolutions ArchitectAI And Machine Learning
2K downloads in the last 30 days
Last updated on June 13, 2026
npx skills add NVIDIA/skills --skill vss-generate-video-calibration

Use when running video data augmentation and auto-labeling workflows on OSMO: flow selection, preflight, submit-time interpolation, monitoring, and output retrieval. Trigger keywords: video data augmentation, data enrichment, auto labeling, VDA demo, OSMO
Skill
Developer
AI EngineerMl EngineerPhysical AIPhysical AI Dataset
2K downloads in the last 30 days
Last updated on May 31, 2026
npx skills add NVIDIA/skills --skill physical-ai-video-data-augmentation

Use to deploy the vss-video-analytics-api REST service standalone (config-source, data-log bind, Elasticsearch, optional Kafka). Not for full warehouse deploy.
Skill
Video Search and Summarization (VSS)
AI And Machine LearningDevOps EngineerPlatform EngineerApplication DeveloperSolutions Architect
2K downloads in the last 30 days
Last updated on June 13, 2026
npx skills add NVIDIA/skills --skill vss-setup-video-analytics-api

Multi-step video annotation pipeline that turns raw videos into Chain-of-Thought training data — multi-level captions, structured descriptions, and QA pairs (MCQ, binary, open-ended) with reasoning traces, via VLM/LLM distillation. Use when the user wants
Skill
Data Engineer
AI EngineerMl EngineerApplication DeveloperAI And Machine LearningTAO Toolkit
1K downloads in the last 30 days
Last updated on June 13, 2026
npx skills add NVIDIA/skills --skill tao-generate-video-reasoning-annotations

Use to call the VIOS REST API (sensor list, timelines, clip extraction, snapshots, add/delete sensors and streams). Not for VLM inference or search.
Skill
Developer
DevOps EngineerPlatform EngineerApplication DeveloperSolutions ArchitectVideo Search and Summarization (VSS)AI And Machine Learning
2K downloads in the last 30 days
Last updated on June 13, 2026
npx skills add NVIDIA/skills --skill vss-manage-video-io-storage
DGX Station
45 MIN

Image & Video Generation with ComfyUI

Generate images and videos with FLUX, Wan 2.1, HunyuanVideo, and Cosmos on DGX Station
Playbook
Image Generation
ComfyUIDockerVideo GenerationDGX StationWan 2.1FLUXHunyuanVideoCosmosGB300
Last updated on May 26, 2026
General
LaunchableEnterprise

Build a Video Search and Summarization (VSS) Agent

Ingest massive volumes of live or archived videos and extract insights for summarization and interactive Q&A
Blueprint
NVIDIA AI
visionvideo-to-textgenerative AIchat
Last updated on February 17, 2026
DGX Spark
30 MIN

Build a Video Search and Summarization (VSS) Agent

Run the VSS Blueprint on your Spark
Playbook
DGX
Spark
Last updated on October 7, 2025
NVIDIA
DownloadableFree Endpoint

nemotron-3-nano-omni-30b-a3b-reasoning

Nemotron 3 Nano Omni is an omni-modal reasoning model that understands images, video, speech, text.
Model
Image-to-Text
VLMVideoOmniOCR
8M API calls in the last 30 days
Last updated on April 28, 2026
DGX Spark
1 HR

Vision-Language Model Fine-tuning

Fine-tune Vision-Language Models for image and video understanding tasks using Qwen2.5-VL and InternVL3
Playbook
DGX
Image UnderstandingVision-Language ModelsGRPOSparkFine-tuningVideo Analysis
Last updated on October 7, 2025
NVIDIA
Free Endpoint

cosmos-transfer1-7b

Generates physics-aware video world states for physical AI development using text prompts and multiple spatial control inputs derived from real-world data or simulation.
Model
Synthetic Data Generation
Autonomous VehiclesPhysical AIroboticsvideo-to-world
250 API calls in the last 30 days
Last updated on June 30, 2025
NVIDIA
Free Endpoint

cosmos-transfer2.5-2b

Generates physics-aware video world states for physical AI development using text prompts and multiple spatial control inputs derived from real-world data or simulation.
Model
Synthetic Data Generation
Autonomous VehiclesPhysical AIroboticsvideo-to-world
Last updated on February 26, 2026
Google
Free Endpoint

paligemma

Vision language model adept at comprehending text and visual inputs to produce informative responses
Model
image
cvVision AssistantvlmVisual Question Answeringcomputer visionLanguage GenerationvideoImage-to-Text
12K API calls in the last 30 days
Last updated on August 26, 2024
NVIDIA
Downloadable

cosmos-reason2-8b

Vision language model that excels in understanding the physical world using structured reasoning on videos or images.
Model
video understanding
autonomous vehiclesindustrialPhysical AIvision language modelreasoningroboticssmart citiesSynthetic Data Generation
191K API calls in the last 30 days
Last updated on December 27, 2025
NVIDIA
DownloadableFree Endpoint

cosmos3-nano-reasoner

Vision language model that excels in understanding the physical world using structured reasoning on videos or images.
Model
video understanding
autonomous vehiclesindustrialPhysical AIvision language modelreasoningroboticssmart citiesSynthetic Data Generation
2K API calls in the last 30 days
Last updated on June 1, 2026

Cosmos-Embed1 video-text embedding for text-to-video retrieval, video-to-video search, semantic deduplication, and fine-tuning. Use when the user asks to "fine-tune Cosmos-Embed1", "run cosmos-embed inference", "export Cosmos-Embed1", "embed videos", or "
Skill
AI Engineer
Data ScientistMl EngineerApplication DeveloperAI And Machine LearningTAO Toolkit
1K downloads in the last 30 days
Last updated on June 16, 2026
npx skills add NVIDIA/skills --skill tao-finetune-cosmos-embed
NVIDIA
DownloadableFree Endpoint

Active Speaker Detection

Detect and track speaker identities across video frames.
Model
broadcast
localizationsmptespeaker detectiondubbingnvidia ai for mediabroadcast-loggingnvidia holoscan for media
1K API calls in the last 30 days
Last updated on April 16, 2026