Models
Deploy and scale models on your GPU infrastructure of choice with NVIDIA NIM inference microservices
models
Multimodal 320B-total / 18B-active MoE with hybrid KDA and sparse MLA attention, native FP8 weights, reasoning and tool calling.Free Endpointglm-5-3-flash
Multi-modal model to classify safety for input prompts as well output responses.Free Endpointllama-guard-4-12b
Items per page
of 1 pages