Skip to main content
Explore
Models
Skills
Blueprints
GPUs
Docs
⌘K
Ctrl+K
?
Login
Models
Deploy and scale models on your GPU infrastructure of choice with NVIDIA NIM inference microservices
Optimized by NVIDIA
Launch from Hugging Face
Beta
Filters (1)
8 models
Sort By
dateCreated:DESC
Most Recent
Microsoft
Downloadable
Free Endpoint
phi-4-mini-instruct
Lightweight multilingual LLM powering AI applications in latency bound, memory/compute constrained environments
Chat
+3
Items per page
24
1
1
of 1 pages
440K
1y
Qwen
Deprecated
Downloadable
qwen2.5-coder-32b-instruct
Advanced LLM for code generation, reasoning, and fixing across popular programming languages.
code completion
+2
834K
11mo
Meta
Downloadable
Free Endpoint
llama-3.2-3b-instruct
Advanced state-of-the-art small language model with language understanding, superior reasoning, and text generation.
B200
+13
25.28K
1.08M
1y
Meta
Downloadable
Free Endpoint
llama-3.2-1b-instruct
Advanced state-of-the-art small language model with language understanding, superior reasoning, and text generation.
B200
+17
43.13K
341K
1y
Abacus.AI
Free Endpoint
dracarys-llama-3.1-70b-instruct
Fine-tuned Llama 3.1 70B model for code generation, summarization, and multi-language tasks.
Code Generation
+1
612K
1y
Google
Free Endpoint
gemma-2-2b-it
Advanced small language generative AI model for edge applications
Chat
+3
885K
1y
Meta
Downloadable
Free Endpoint
llama-3.1-70b-instruct
Powers complex conversations with superior contextual understanding, reasoning and text generation.
B200
+18
3.74M
11mo
Meta
Downloadable
Free Endpoint
llama-3.1-8b-instruct
Advanced state-of-the-art model with language understanding, superior reasoning, and text generation.
B200
+19
30.68M
10mo