Models
Deploy and scale models on your GPU infrastructure of choice with NVIDIA NIM inference microservices
models
Multimodal 320B-total / 18B-active MoE with hybrid KDA and sparse MLA attention, native FP8 weights, reasoning and tool calling.Free Endpointglm-5-3-flash
Cutting-edge vision-language model excelling in retrieving text and metadata from images.Downloadablenemotron-parse-2.0
~2.8T hybrid KDA+MLA multimodal MoE for long-horizon coding, agentic tool use, and image understanding.DownloadableFree Endpointkimi-k3
Muse Glimmer 30B is a multimodal reasoning model accepting text and images, with native tool-calling and separate reasoning output.DownloadableFree Endpointmuse-glimmer-30b
An image edit model specialized for Omniverse synthetic to photographic solder-light style captured at NVIDIA PCB inspection stationsDownloadableqwen-image-edit-nvpcb-ovsl2sl
Nemotron OCR v2 is a state-of-the-art multilingual text recognition model designed for robust end-to-end optical character recognition (OCR) on complex real-world images.Downloadablenemotron-ocr-v2
Generates images, videos and action predictions from text, visual inputs and spatial controls, with mandatory input screening and visual SynthID watermarking.DownloadableFree Endpointcosmos3-nano
Vision language model that excels in understanding the physical world using structured reasoning on videos or images.DownloadableFree Endpointcosmos3-nano-reasoner
Qwen-Image is a text-to-image foundation model with advanced multilingual text rendering.Downloadableqwen-image
Qwen-Image-Edit is an image editing model with multilingual text editing and strong subject consistency.Downloadableqwen-image-edit
Nemotron 3 Nano Omni is an omni-modal reasoning model that understands images, video, speech, text.DownloadableFree Endpointnemotron-3-nano-omni-30b-a3b-reasoning
FLUX.2-klein-4B is a distilled image generation and editing model, producing outputs at lighting speedDownloadableflux.2-klein-4b
Powerful OCR model for fast, accurate real-world image text extraction, layout, and structure analysis.Downloadablenemotron-ocr-v1
Multimodal question-answer retrieval representing user queries as text and documents as images.Downloadablellama-nemotron-embed-vl-1b-v2
Cutting-edge vision-language model exceling in retrieving text and metadata from images.Downloadablenemotron-parse
Stable Diffusion 3.5 is a popular text-to-image generation modelDownloadablestable-diffusion-3.5-large
FLUX.1 Kontext is a multimodal model that enables in-context image generation and editing.DownloadableFLUX.1-Kontext-dev
Powerful OCR model for fast, accurate real-world image text extraction, layout, and structure analysis.Downloadablenemoretriever-ocr
FLUX.1-schnell is a distilled image generation model, producing high quality images at fast speedsDownloadableFLUX.1-schnell
Cutting-edge vision-language model exceling in high-quality reasoning from images.DownloadableFree Endpointllama-3.2-11b-vision-instruct
Cutting-edge vision-Language model exceling in high-quality reasoning from images.DownloadableFree Endpointllama-3.2-90b-vision-instruct
Items per page
of 2 pages