Skip to main content
NVIDIA
Explore
Models
Skills
Blueprints
GPUs
Docs
Help Center
Getting Started
  1. Create and verify your account to unlock full access to NVIDIA NIM APIs.
ResourcesDeveloper ForumsContact Support
FAQs
  • Terms of Use
    Privacy Policy
    Your Privacy Choices
    Contact

    Copyright © 2026 NVIDIA Corporation

    OpenAI

    whisper-large-v3

    Downloadable

    Robust Speech Recognition via Large-Scale Weak Supervision.

    • ASR
    • AST
    • Multilingual
    • NVIDIA NIM
    • NVIDIA Riva
    • OpenAI
    • batch
    • Speech-to-Text
    • whisper
    Get API Key
    API ReferenceAPI Reference
    Accelerated by DGX Cloud

    Model Overview

    Description:

    This model is used to transcribe short-form audio files and is designed to be compatible with OpenAI's sequential long-form transcription algorithm. Whisper is a pre-trained model for automatic speech recognition (ASR) and speech translation. Trained on 680k hours of labeled data, Whisper models demonstrate a strong ability to generalize to many datasets and domains without the need for fine-tuning. Whisper-large-v3 is one of the 5 configurations of the model with 1550M parameters.
    This model version is optimized to run with NVIDIA TensorRT-LLM. This model is ready for commercial use.

    Third-Party Community Consideration

    This model is not owned or developed by NVIDIA. This model has been developed and built to a third-party’s requirements for this application and use case; see the Whisper Model Card on GitHub.(https://github.com/openai/whisper/blob/main/model-card.md).

    License/Terms of Use:

    GOVERNING TERMS: Use of this model is governed by the NVIDIA Community Model License Agreement. ADDITIONAL INFORMATION: MIT License.

    References:

    Whisper website
    Whisper paper:

    @misc{radford2022robust,
          title={Robust Speech Recognition via Large-Scale Weak Supervision}, 
          author={Alec Radford and Jong Wook Kim and Tao Xu and Greg Brockman and Christine McLeavey and Ilya Sutskever},
          year={2022},
          eprint={2212.04356},
          archivePrefix={arXiv},
          primaryClass={eess.AS}
    }
    

    Model Architecture:

    Architecture Type: Transformer (Encoder-Decoder)
    Network Architecture: Whisper

    Input:

    Input Type(s): Audio, Text-Prompt
    Input Format(s): Linear PCM 16-bit 1 channel (Audio), String (Text Prompt)
    Input Parameters: One-Dimensional (1D)

    Output:

    Output Type(s): Text Output Format: String Output Parameters: 1D

    Supported Hardware Microarchitecture Compatibility:

    • NVIDIA Ampere
    • NVIDIA Blackwell

    Supported Operating System(s):

    • Linux

    Model Version(s):

    Large-v3: Whisper large-v3 has the same architecture as the previous large and large-v2 models, except for the following minor differences:

    • The spectrogram input uses 128 Mel frequency bins instead of 80.
    • A new language token for Cantonese.

    Training Dataset:

    Data Collection Method by dataset: [Hybrid: human, automatic]

    Labeling Method by dataset: [Automated]

    Dataset License(s): NA

    Inference:

    Engine: Tensor(RT)-LLM, Triton
    Test Hardware:

    • A100
    • H100

    For more detail on model usage, evaluation, training data set and implications, please refer to Whisper Model Card.

    Ethical Considerations:

    NVIDIA believes Trustworthy AI is a shared responsibility and we have established policies and practices to enable development for a wide array of AI applications. When downloaded or used in accordance with our terms of service, developers should work with their internal model team to ensure this model meets requirements for the relevant industry and use case and addresses unforeseen product misuse.
    Please report security vulnerabilities or NVIDIA AI Concerns here.

    On this page

    1. Description
    2. Third-Party Community Consideration
      1. License/Terms of Use
    3. References
    4. Model Architecture
    5. Input
    6. Output
    7. Model Version(s)
    8. Training Dataset
    9. Inference
    10. Ethical Considerations