Skip to main content
Explore
Models
Skills
Blueprints
GPUs
Docs
⌘K
Ctrl+K
?
Login
Models
Deploy and scale models on your GPU infrastructure of choice with NVIDIA NIM inference microservices
Optimized by NVIDIA
Launch from Hugging Face
Beta
Filters (2)
9 models
Sort By
dateCreated:DESC
Most Recent
Mistral AI
Downloadable
Free Endpoint
mistral-medium-3.5-128b
A high performing model for text generation, coding and agentic use cases
coding
+3
3.05M
1mo
Items per page
24
1
1
of 1 pages
DeepSeek AI
Downloadable
deepseek-v4-pro
DeepSeek V4 scales to 1M-token context windows with efficient MoE architecture for coding tasks.
B200
+4
8.1M
1mo
Z.ai
Downloadable
Free Endpoint
glm-5.1
GLM-5.1 is a flagship LLM for agentic workflows, coding, and long-horizon reasoning tasks.
B200
+5
24.9M
1mo
Minimaxai
Downloadable
Free Endpoint
minimax-m2.7
MiniMax M2.7 is a 230B-parameter text-to-text AI model excelling in coding, reasoning, and office tasks.
B200
+5
13.46M
1mo
Google
Downloadable
Free Endpoint
gemma-4-31b-it
Dense 31B model delivering frontier reasoning for coding, agentic workflows, and fine-tuning.
B200
+6
5.76M
2mo
Minimaxai
Deprecated
Downloadable
minimax-m2.5
MiniMax M2.5 is a 230B-parameter text-to-text AI model excelling in coding, reasoning, and office tasks.
B200
+7
2.5M
3mo
Stepfun-ai
Free Endpoint
step-3.5-flash
200B open-source reasoning engine with sparse MoE powering frontier agentic AI.
Agentic
+2
11.52M
4mo
Sarvamai
Downloadable
Free Endpoint
sarvam-m
Multilingual, hybrid-reasoning model optimized for Indian language tasks, programming, mathematical reasoning capabilities.
coding
+5
286K
10mo
Mistral AI
Deprecated
Free Endpoint
magistral-small-2506
High performance reasoning model optimized for efficiency and edge deployment
coding
+3
364K
10mo