Skip to main content
Explore
Models
Skills
Blueprints
GPUs
Docs
Search
⌘K
Ctrl+K
?
Login
Models
Deploy and scale models on your GPU infrastructure of choice with NVIDIA NIM inference microservices
Optimized by NVIDIA
Launch from Hugging Face
Beta
Filters (2)
8 models
Sort By
dateCreated:DESC
Most Recent
Microsoft
Downloadable
Free Endpoint
phi-4-mini-instruct
Lightweight multilingual LLM powering AI applications in latency bound, memory/compute constrained environments
Chat
+3
Items per page
24
1
1
of 1 pages
445K
1y
Meta
Downloadable
Free Endpoint
llama-3.2-3b-instruct
Advanced state-of-the-art small language model with language understanding, superior reasoning, and text generation.
Language Generation
+3
26.7K
1.22M
1y
Meta
Downloadable
Free Endpoint
llama-3.2-1b-instruct
Advanced state-of-the-art small language model with language understanding, superior reasoning, and text generation.
Language Generation
+3
45.6K
290K
1y
NVIDIA
Free Endpoint
nemotron-mini-4b-instruct
Optimized SLM for on-device inference and fine-tuned for roleplay, RAG and function calling
Chat
+2
1.53M
1y
Google
Free Endpoint
gemma-2-2b-it
Advanced small language generative AI model for edge applications
Chat
+3
1.42M
1y
Meta
Downloadable
Free Endpoint
llama-3.1-70b-instruct
Powers complex conversations with superior contextual understanding, reasoning and text generation.
Chat
+3
3.9M
1y
Meta
Downloadable
Free Endpoint
llama-3.1-8b-instruct
Advanced state-of-the-art model with language understanding, superior reasoning, and text generation.
Chat
+4
25.09M
11mo
Upstage
Free Endpoint
solar-10.7b-instruct
Excels in NLP tasks, particularly in instruction-following, reasoning, and mathematics.
Non-Commercial Use Only
+4
449K
1y