Skip to main content
NVIDIA
Explore
Models
Skills
Blueprints
GPUs
Docs
ForumsSupport
Terms of Use
Privacy Policy
Your Privacy Choices
Contact

Copyright © 2026 NVIDIA Corporation

Models

Deploy and scale models on your GPU infrastructure of choice with NVIDIA NIM inference microservices

Optimized by NVIDIALaunch from Hugging FaceBeta

Filters (1)

Use Case
Inference Providers
Publisher
NIM Container GPUs
Labels (1)
11 models
NVIDIA
Free Endpoint

nemotron-3-embed-1b

1B embedding model for semantic search, retrieval, and RAG applications.
Nemotron Retriever
Agentic RetrievalCode RetrievalText-to-EmbeddingRetrieval Augmented Generation
Last updated on July 16, 2026
Items per page
of 1 pages
Poolside
Free Endpoint

laguna-xs-2.1

Efficient 33B MoE for local, long-horizon agentic coding and terminal tasks
Agentic AI
CodingReasoningTool Use
Last updated on July 15, 2026
Z.ai
DownloadableFree Endpoint

glm-5.2

GLM-5.2 is a flagship LLM for agentic workflows, coding, and long-horizon reasoning tasks.
Agentic AI
CodingReasoningTool Use
8M API calls in the last 30 days
Last updated on July 3, 2026
Mistral AI
DownloadableFree Endpoint

mistral-medium-3.5-128b

A high performing model for text generation, coding and agentic use cases
coding
reasoningtextagentic
5M API calls in the last 30 days
Last updated on April 29, 2026
DeepSeek AI
DownloadableFree Endpoint

deepseek-v4-flash

DeepSeek V4 Flash is a 284B MoE model with 1M-token context optimized for fast coding and agents.
MoE
codingfastagentic
15M API calls in the last 30 days
Last updated on April 24, 2026
DeepSeek AI
DownloadableFree Endpoint

deepseek-v4-pro

DeepSeek V4 scales to 1M-token context windows with efficient MoE architecture for coding tasks.
Moe
reasoningcodingagentic
8M API calls in the last 30 days
Last updated on April 24, 2026
Google
DownloadableFree Endpoint

gemma-4-31b-it

Dense 31B model delivering frontier reasoning for coding, agentic workflows, and fine-tuning.
reasoning
codingtext-to-textagentic
6M API calls in the last 30 days
Last updated on April 2, 2026
Qwen
DownloadableFree Endpoint

qwen3.5-397b-a17b

Next-gen Qwen 3.5 VLM (400B MoE) brings advanced vision, chat, RAG, and agentic capabilities.
MoE
image-to-imageVLMagentic
16M API calls in the last 30 days
Last updated on February 16, 2026
Stepfun-ai
Deprecation in 7dFree Endpoint

step-3.5-flash

200B open-source reasoning engine with sparse MoE powering frontier agentic AI.
Agentic
CodingReasoning
10M API calls in the last 30 days
Last updated on February 2, 2026
Mistral AI
Downloadable

mistral-large-3-675b-instruct-2512

A state-of-the-art general purpose MoE VLM ideal for chat, agentic and instruction based use cases.
language generation
multimodalagenticImage-to-Text
2M API calls in the last 30 days
Last updated on December 2, 2025
Qwen
DownloadableFree Endpoint

qwen3-next-80b-a3b-instruct

Qwen3-Next Instruct blends hybrid attention, sparse MoE, and stability boosts for ultra-long context AI.
text-generation
agentic
25M API calls in the last 30 days
Last updated on September 22, 2025