Models
Deploy and scale models on your GPU infrastructure of choice with NVIDIA NIM inference microservices
models
552B MoE, 8B active params with native multimodal support and lower API cost using smaller KV cacheFree Endpointdeepseek-v4.1-flash
Multimodal 320B-total / 18B-active MoE with hybrid KDA and sparse MLA attention, native FP8 weights, reasoning and tool calling.Free Endpointglm-5-3-flash
~2.8T hybrid KDA+MLA multimodal MoE for long-horizon coding, agentic tool use, and image understanding.DownloadableFree Endpointkimi-k3
Muse Glimmer 30B is a multimodal reasoning model accepting text and images, with native tool-calling and separate reasoning output.DownloadableFree Endpointmuse-glimmer-30b
Multi-modal model to classify safety for input prompts as well output responses.Free Endpointllama-guard-4-12b
Items per page
of 1 pages