Models
Deploy and scale models on your GPU infrastructure of choice with NVIDIA NIM inference microservices
models
Multimodal 320B-total / 18B-active MoE with hybrid KDA and sparse MLA attention, native FP8 weights, reasoning and tool calling.Free Endpointglm-5-3-flash
Muse Glimmer 30B is a multimodal reasoning model accepting text and images, with native tool-calling and separate reasoning output.DownloadableFree Endpointmuse-glimmer-30b
Vision language model that excels in understanding the physical world using structured reasoning on videos or images.DownloadableFree Endpointcosmos3-nano-reasoner
Open VLM for quantum computer calibration chart understanding across a range of qubit modalities.DownloadableFree Endpointising-calibration-1-35b-a3b
Open, efficient hybrid Mamba-Transformer MoE with 1M context, excelling in agentic reasoning, coding, planning, tool calling, and moreDownloadableFree Endpointnemotron-3-super-120b-a12b
Smaller Mixture of Experts (MoE) text-only LLM for efficient AI reasoning and mathDownloadableFree Endpointgpt-oss-20b
Items per page
of 1 pages