Skip to main content
NVIDIA
Explore
Models
Skills
Blueprints
GPUs
Docs
Help Center
Getting Started
  1. Create and verify your account to unlock full access to NVIDIA NIM APIs.
ResourcesDeveloper ForumsContact Support
FAQs
  • Terms of Use
    Privacy Policy
    Your Privacy Choices
    Contact

    Copyright © 2026 NVIDIA Corporation

    nvidia/nemotron-mini-4b-instruct

    API Reference
    This API will be deprecated on 08/25/2026. It will no longer be supported after 08/25/2026. Please transition to another model to avoid any service interruptions. For more models information, visit our API Reference.

    Prototype

    Start building with a free API endpoint.
    from openai import OpenAI
    
    client = OpenAI(
      base_url = "https://integrate.api.nvidia.com/v1",
      api_key = "$NVIDIA_API_KEY"
    )
    
    completion = client.chat.completions.create(
      model="nvidia/nemotron-mini-4b-instruct",
      messages=[{"role":"user","content":"Write a limerick about the wonders of GPU computing."}],
      temperature=0.2,
      top_p=0.7,
      max_tokens=1024,
      stream=False
    )
    
    print(completion.choices[0].message)

    Specifications

    Optimized SLM for on-device inference and fine-tuned for roleplay, RAG and function calling

    • Chat
    • Language Generation
    • Text-to-Text
    Provider
    NVIDIA
    Last Modified
    1 year ago
    Context Length
    4K
    Parameters
    4B
    Input Modalities
    Text
    Output Modalities
    Text

    Model Availability

    Free Endpoint
    Available
    Partner Endpoint
    Not available
    Download Available
    Not available
    API calls in last 30 days
    3M