Skip to main content
Search
⌘K
Ctrl+K
Login
2 results for
Filters (1)
Models (0)
Blueprints (0)
Skills (0)
Other (2)
Sort By
Best Match
Select item
Best Match
Most Popular
Most Downloaded
Alphabetical (A-Z)
Alphabetical (Z-A)
Best Match
Most Popular
Most Downloaded
Alphabetical (A-Z)
Alphabetical (Z-A)
Search results
Playbooks
30 MIN
Serve LLMs with vLLM
High-throughput serving for 30+ models, with continuous batching and an OpenAI-compatible API
Playbook
Playbooks
60 MIN
Quantize Models to NVFP4 with NVIDIA Model Optimizer
Cut memory ~3.5× vs FP16 while keeping accuracy close to FP8, then validate with an OpenAI-compatible endpoint
Playbook
Items per page
24
12
24
48
96
1
1
of 1 pages
Explore
Models
Skills
Blueprints
GPUs
Docs
Search
⌘K
Ctrl+K
?
Login
Explore
Models
Skills
Blueprints
GPUs
Docs
?
vLLM
+4
DGX Spark
DGX Station
Inference
RTX PRO
1mo
Last updated on August 6, 2026
vLLM
+5
DGX Spark
Model Optimizer
DGX Station
TensorRT-LLM
Inference
1mo
Last updated on August 4, 2026