Skip to main content
NVIDIA
Explore
Models
Skills
Blueprints
GPUs
Docs
Help Center
Getting Started
  1. Create and verify your account to unlock full access to NVIDIA NIM APIs.
ResourcesDeveloper ForumsContact Support
FAQs
  • Terms of Use
    Privacy Policy
    Your Privacy Choices
    Contact

    Copyright © 2026 NVIDIA Corporation

    openai/gpt-oss-120b

    API Reference

    Prototype

    Start building with a free API endpoint.
    from openai import OpenAI
    
    client = OpenAI(
      base_url = "https://integrate.api.nvidia.com/v1",
      api_key = "$NVIDIA_API_KEY"
    )
    
    completion = client.chat.completions.create(
      model="openai/gpt-oss-120b",
      messages=[{"role":"user","content":"Which number is larger, 9.11 or 9.8?"}],
      temperature=1,
      top_p=1,
      max_tokens=4096,
      stream=False
    )
    
    reasoning = getattr(completion.choices[0].message, "reasoning_content", None)
    if reasoning:
      print(reasoning)
    print(completion.choices[0].message.content)

    Deploy

    Ready to scale? Choose your deployment path.
    Deploying your application in production? Get started with a 90-day evaluation of NVIDIA AI Enterprise

    Follow the steps below to download and run the NVIDIA NIM inference microservice for this model on your infrastructure of choice.

    Step 1
    Generate API Key

    Step 2
    Pull and Run the NIM

    $ docker login nvcr.io
    Username: $oauthtoken
    Password: <PASTE_API_KEY_HERE>
    

    Pull and run the NVIDIA NIM with the command below. This will download the optimized model for your infrastructure.

    export NGC_API_KEY=<PASTE_API_KEY_HERE>
    export LOCAL_NIM_CACHE=~/.cache/nim
    mkdir -p "$LOCAL_NIM_CACHE"
    chmod -R 777 $LOCAL_NIM_CACHE
    docker run -it --rm \
        --gpus all \
        --shm-size=16GB \
        -e NGC_API_KEY \
        -v "$LOCAL_NIM_CACHE:/opt/nim/.cache" \
        -p 8000:8000 \
        nvcr.io/nim/openai/gpt-oss-120b:latest
    

    Step 3
    Test the NIM

    You can now make a local API call using this curl command:

    curl -X 'POST' \
    'http://0.0.0.0:8000/v1/chat/completions' \
    -H 'accept: application/json' \
    -H 'Content-Type: application/json' \
    -d '{
        "model": "openai/gpt-oss-120b",
        "messages": [{"role":"user", "content":"Which number is larger, 9.11 or 9.8?"}],
        "max_tokens": 64
    }'
    

    For more details on getting started with this NIM, visit the NVIDIA NIM Docs.

    Specifications

    Mixture of Experts (MoE) reasoning LLM (text-only) designed to fit within 80GB GPU.

    • chat
    • math
    • reasoning
    • text-to-text
    Provider
    OpenAI
    Last Modified
    1 year ago
    Context Length
    131K
    Parameters
    117B
    Input Modalities
    Text
    Output Modalities
    Text

    Model Availability

    Free Endpoint
    Available
    Partner Endpoint
    Available
    Download Available
    Available
    API calls in last 30 days
    45M