Skip to main content
Explore
Models
Skills
Blueprints
GPUs
Docs
Search
⌘K
Ctrl+K
?
Forums
Support
Login
Models
Deploy and scale models on your GPU infrastructure of choice with NVIDIA NIM inference microservices
Optimized by NVIDIA
Launch from Hugging Face
Beta
Filters (2)
3 models
Sort By
Most Recent
Select item
Most Recent
Most Popular
Most Downloaded
Alphabetical (A-Z)
Alphabetical (Z-A)
Most Recent
Most Popular
Most Downloaded
Alphabetical (A-Z)
Alphabetical (Z-A)
NVIDIA
Downloadable
llama-nemotron-embed-vl-1b-v2
Multimodal question-answer retrieval representing user queries as text and documents as images.
nemo retriever
+3
embedding
Text-to-Embedding
Retrieval Augmented Generation
Items per page
24
12
24
48
96
1
1
of 1 pages
8M
8M API calls in the last 30 days
5mo
Last updated on February 10, 2026
NVIDIA
Free Endpoint
nv-embedcode-7b-v1
The NV-EmbedCode model is a 7B Mistral-based embedding model optimized for code retrieval, supporting text, code, and hybrid queries.
nemo retriever
+2
Embedding
Retrieval Augmented Generation
1M
1M API calls in the last 30 days
1y
Last updated on May 29, 2025
NVIDIA
Downloadable
nv-embedqa-e5-v5
English text embedding model for question-answering retrieval.
Embedding
+4
run-on-rtx
Nemo retriever
Text-to-Embedding
Retrieval Augmented Generation
16M
16M API calls in the last 30 days
1y
Last updated on July 22, 2025