Models
Deploy and scale models on your GPU infrastructure of choice with NVIDIA NIM inference microservices
models
Multimodal 320B-total / 18B-active MoE with hybrid KDA and sparse MLA attention, native FP8 weights, reasoning and tool calling.Free Endpointglm-5-3-flash
Diffusion-based 26B parameter LLM enabling parallel token generation for real-time text appsDownloadableFree Endpointdiffusiongemma-26b-a4b-it
Multilingual, multimodal model for detecting unsafe and toxic content.DownloadableFree Endpointnemotron-3.5-content-safety
Nemotron 3 Nano Omni is an omni-modal reasoning model that understands images, video, speech, text.DownloadableFree Endpointnemotron-3-nano-omni-30b-a3b-reasoning
Open VLM for quantum computer calibration chart understanding across a range of qubit modalities.DownloadableFree Endpointising-calibration-1-35b-a3b
Leading multilingual content safety model for enhancing the safety and moderation capabilities of LLMsFree Endpointllama-3.1-nemotron-safety-guard-8b-v3
Smaller Mixture of Experts (MoE) text-only LLM for efficient AI reasoning and mathDownloadableFree Endpointgpt-oss-20b
Multi-modal model to classify safety for input prompts as well output responses.Free Endpointllama-guard-4-12b
Topic control model to keep conversations focused on approved topics, avoiding inappropriate content.Downloadablellama-3.1-nemoguard-8b-topic-control
Industry leading jailbreak classification model for protection from adversarial attemptsDownloadablenemoguard-jailbreak-detect
Leading content safety model for enhancing the safety and moderation capabilities of LLMsDownloadablellama-3.1-nemoguard-8b-content-safety
Items per page
of 1 pages