Models
Deploy and scale models on your GPU infrastructure of choice with NVIDIA NIM inference microservices
models
Multimodal 320B-total / 18B-active MoE with hybrid KDA and sparse MLA attention, native FP8 weights, reasoning and tool calling.Free Endpointglm-5-3-flash
Fastest 30B A3B MoE model with leading domain accuracy for specialized agentic tasksDownloadableFree Endpointnemotron-3.5-lightning-30b-a3b
Muse Glimmer 30B is a multimodal reasoning model accepting text and images, with native tool-calling and separate reasoning output.DownloadableFree Endpointmuse-glimmer-30b
An image edit model specialized for Omniverse synthetic to photographic solder-light style captured at NVIDIA PCB inspection stationsDownloadableqwen-image-edit-nvpcb-ovsl2sl
Diffusion-based 26B parameter LLM enabling parallel token generation for real-time text appsDownloadableFree Endpointdiffusiongemma-26b-a4b-it
Generates images, videos and action predictions from text, visual inputs and spatial controls, with mandatory input screening and visual SynthID watermarking.DownloadableFree Endpointcosmos3-nano
Qwen-Image is a text-to-image foundation model with advanced multilingual text rendering.Downloadableqwen-image
Qwen-Image-Edit is an image editing model with multilingual text editing and strong subject consistency.Downloadableqwen-image-edit
Dense 31B model delivering frontier reasoning for coding, agentic workflows, and fine-tuning.DownloadableFree Endpointgemma-4-31b-it
Stable Diffusion 3.5 is a popular text-to-image generation modelDownloadablestable-diffusion-3.5-large
Smaller Mixture of Experts (MoE) text-only LLM for efficient AI reasoning and mathDownloadableFree Endpointgpt-oss-20b
Expressive and engaging text-to-speech, generated from a short audio sample.Free Endpointmagpie-tts-zeroshot
Items per page
of 1 pages