Models
Deploy and scale models on your GPU infrastructure of choice with NVIDIA NIM inference microservices
models
Multimodal 320B-total / 18B-active MoE with hybrid KDA and sparse MLA attention, native FP8 weights, reasoning and tool calling.Free Endpointglm-5-3-flash
~2.8T hybrid KDA+MLA multimodal MoE for long-horizon coding, agentic tool use, and image understanding.DownloadableFree Endpointkimi-k3
Muse Glimmer 30B is a multimodal reasoning model accepting text and images, with native tool-calling and separate reasoning output.DownloadableFree Endpointmuse-glimmer-30b
Nemotron 3 Nano Omni is an omni-modal reasoning model that understands images, video, speech, text.DownloadableFree Endpointnemotron-3-nano-omni-30b-a3b-reasoning
Cutting-edge vision-language model exceling in high-quality reasoning from images.DownloadableFree Endpointllama-3.2-11b-vision-instruct
Cutting-edge vision-Language model exceling in high-quality reasoning from images.DownloadableFree Endpointllama-3.2-90b-vision-instruct
Items per page
of 1 pages