Search results
Use this skill to bring a supported object-detection vision model from HuggingFace or NVIDIA NGC into an NVIDIA DeepStream pipeline with end-to-end automation: ONNX download, SafeTensors export, TRT engine build, custom nvinfer bbox parser, multi-stream b Cutting-edge vision-language model exceling in high-quality reasoning from images.DownloadableFree Endpointllama-3.2-11b-vision-instruct
Cutting-edge vision-Language model exceling in high-quality reasoning from images.DownloadableFree Endpointllama-3.2-90b-vision-instruct
Fine-tune Vision-Language Models for image and video understanding tasks using Qwen2.5-VL and InternVL3 Real-time Vision Language Model interaction with webcam streamingPlaybooks20 MINLive VLM WebUI
Open VLM for quantum computer calibration chart understanding across a range of qubit modalities.DownloadableFree Endpointising-calibration-1-35b-a3b
NVIDIA-Ising-Calibration-1.5 is a dense multimodal vision-language model built on Gemma 4 31B. It analyzes quantum computing calibration experiment plots and generates structured technical text.Free Endpointising-calibration-1.5-31b
Ingest massive volumes of live or archived videos and extract insights for summarization and interactive Q&AGeneralLaunchableEnterpriseBuild a Video Search and Summarization (VSS) Agent
Vision language model that excels in understanding the physical world using structured reasoning on videos or images.DownloadableFree Endpointcosmos3-nano-reasoner
Cutting-edge vision-language model exceling in retrieving text and metadata from images.Downloadablenemotron-parse
Cutting-edge vision-language model excelling in retrieving text and metadata from images.Downloadablenemotron-parse-2.0
Grade or filter workflow HDF5 episodes with an OpenAI-compatible vision model. Use for visual success labels; do not use for replay, policy evaluation, or recordings without frames. CLIP vision-language model for image-text retrieval, zero-shot classification, embedding extraction, ONNX export, and TensorRT deployment. Use when fine-tuning or training CLIP, running zero-shot classification, computing image embeddings, or deploying CL NVDINOv2 for self-supervised visual representation learning. Trains vision transformers via self-distillation (teacher-student) without labels and produces general-purpose visual features. Use when training, exporting, or running inference for a TAO NVDIN Fine-tune any HuggingFace CV / VLM / LLM model on local NVIDIA GPUs inside an NGC PyTorch container when no dedicated TAO model skill matches. Use when the user wants to fine-tune a HuggingFace model (full or LoRA), train a vision / VLM / LLM model end-to Integrate a HuggingFace Computer Vision model into the NVIDIA TAO Toolkit ecosystem (tao-core config, tao-pytorch trainer, tao-deploy TensorRT pipeline). Use when the user asks to "integrate a HuggingFace model into TAO", "add an HF model to TAO Toolkit",
Items per page
of 1 pages