Search results
Benchmark Jetson LLM/VLM serving performance across vLLM, llama.cpp, and Ollama with structured JSON output. Stand up vLLM or SGLang serving on Jetson, using upstream vLLM on Thor and Orin JetPack 7.2+, and NVIDIA-AI-IOT vLLM on older Orin. Prebuilt, GPU-optimized model containers with a ready-to-use HTTP endpointPlaybooksIntermediate30 MINDeploy NVIDIA NIM for LLM Inference
Chat from the terminal against local vLLM with the self-improving Nous Research agent (Telegram optional) Build a local AI assistant in an OpenShell sandbox with vLLM inference and optional Telegram Install a local-first AI agent and connect it to a private OpenAI-compatible model endpoint High-throughput serving for 30+ models, with continuous batching and an OpenAI-compatible APIPlaybooksIntermediate30 MINServe LLMs with vLLM
High-throughput serving with RadixAttention, structured output, and an OpenAI-compatible API One interface for supervised, RLHF, and parameter-efficient trainingPlaybooksAdvanced60 MINFine-Tune LLMs with LLaMA Factory
LoRA, full training, and RL paths plus Nemotron open modelsPlaybooksIntermediate8 MINFine-Tune Specialized LLMs with Unsloth
A self-hosted browser interface with models running locally on your GPUPlaybooksIntermediate15 MINChat with LLMs Using Open WebUI and Ollama
Industry leading jailbreak classification model for protection from adversarial attemptsDownloadablenemoguard-jailbreak-detect
Diffusion-based 26B parameter LLM enabling parallel token generation for real-time text appsDownloadableFree Endpointdiffusiongemma-26b-a4b-it
Build advanced AI agents for providers and patients using this developer example powered by NeMo Microservices, NVIDIA Nemotron, Riva ASR and TTS, and NVIDIA LLM NIMHealthcare & Life SciencesLaunchableDeveloper ExampleAmbient Healthcare Agents
Cut memory ~3.5× vs FP16 while keeping accuracy close to FP8, then validate with an OpenAI-compatible endpointPlaybooksIntermediate60 MINQuantize Models to NVFP4 with NVIDIA Model Optimizer
Leading content safety model for enhancing the safety and moderation capabilities of LLMsDownloadablellama-3.1-nemoguard-8b-content-safety
Leading multilingual content safety model for enhancing the safety and moderation capabilities of LLMsFree Endpointllama-3.1-nemotron-safety-guard-8b-v3
Topic control model to keep conversations focused on approved topics, avoiding inappropriate content.Downloadablellama-3.1-nemoguard-8b-topic-control
Multi-modal model to classify safety for input prompts as well output responses.Free Endpointllama-guard-4-12b
Run cuTile kernel benchmarks, FMHA implementation, and LLM inference on DGX Spark and B300Playbooks60 MINcuTile Kernels
Real-time Vision Language Model interaction with webcam streamingPlaybooks20 MINLive VLM WebUI
Distill and deploy domain-specific AI models from unstructured financial data to generate market signals efficiently—scaling your workflow with the NVIDIA Data Flywheel Blueprint for high-performance, cost-efficient experimentation.Financial ServicesLaunchableDeveloper ExampleAI Model Distillation for Financial Data
Items per page
of 3 pages