---
title: "audio2face-2d"
publisher: "nvidia"
type: "endpoint"
updated: "2025-06-13T09:40:56.957Z"
description: "Create facial animations using a portrait photo and synchronize mouth movement with audio."
canonical: "https://build.nvidia.com/nvidia/audio2face-2d"
---

# Model Overview

## Description:

NVIDIA Maxine Audio2Face-2D is a generative model to create facial animations using a portrait photo and a driving audio such that the mouth movement in the photo synchronizes with the speech in the provided audio.

The model takes the input audio to estimate landmark motions that represent the mouth movements articulating the words in the audio. These landmarks are then encoded into latent representations that were passed to a generative model to animate the input portrait.

## Terms of use

The use of NVIDIA Maxine Audio2Face-2D is available as a demonstration of the input and output of the Live Portrait generative model. As such the user may submit a reference “driving” audio or use the sample “driving” audio and download the generated video for evaluation under the terms of the [NVIDIA MAXINE EVALUATION LICENSE AGREEMENT](https://developer.download.nvidia.com/maxine/nvidia-maxine-evaluation-license-24oct2023.pdf).

## References(s):

* [NVIDIA Maxine](https://developer.nvidia.com/maxine )
* [NVIDIA ACE](https://developer.nvidia.com/ace)

## Model Architecture:
**Architecture Type:** Recurrent Neural Network (RNN), Convolutional Neural Networks (CNNs), Generative Adversarial Networks (GANs) <br>
**Network Architecture:** Encoder-Decoder <br>

## Input:
**Input Format:** RGB image (portrait photo), float vector containing 32-bit float Pulse Code Modulation (PCM) data (driving audio) <br>
**Input Parameters:** 720p to 4K <br>
**Other Properties Related to Input:** Input images pre-processed using proprietary technique; portrait photo supports 3 channel, 32 bit images; PCM audio samples with no encoding or pre-processing; 16kHz sampling rate and mono channel is required for audio.  <br>

## Output: <br>
**Output Format:** RGB image <br>
**Output Parameters:** 512 x 512 <br>
**Other Properties Related to Output:** Input images post-processed using proprietary technique; 3 Channel, 32 bit image supported. <br>

## Software Integration:
0.8.6.0

## Supported Operating System(s):
Linux

## Model Version(s):
0.8.6.0

## Supported Hardware Microarchitecture Compatability:

* [Volta] <br>
* [Turing] <br>
* [Ampere] <br>
* [ADA] <br>
* [Blackwell] <br>

## Training and Evaluation Dataset:

**Data Collection Method by dataset:** Automated <br>

**Properties (Quantity, Dataset Descriptions, Sensor(s)):**

Datasets used in speech-live-portrait tranining are as follows:

One dataset includes 7,356 files collected of 24 professional actors (12 female and 12 male) with different expressions, head poses, and backgrounds.

One dataset consists of about 160000 videos of different speakers in different environments such as outdoor recording, indoor recording, data covering different phonemes. It is made of audio-visual data consisting of short clips of human speech, extracted from interview videos.

### Evaluation Dataset:

**Data Collection Method by dataset:** Automated, Human <br>

**Properties (Quantity, Dataset Descriptions, Sensor(s)):**

The dataset consists of 5000 samples. This data captures variety among different speakers, languages and phonemes.

NVIDIA models are trained on a diverse set of public and proprietary datasets. This model was trained on a dataset containing facial images of people covering different attributes such as expressions, head poses, backgrounds etc. NVIDIA is committed to the responsible development of AI Foundation models and conducts reviews of all datasets included in training.

# Inference:
**Engine:** [TensorRT](https://developer.nvidia.com/tensorrt), [Triton](https://developer.nvidia.com/triton-inference-server) <br>
**Test Hardware:**
* CUDA 12.1 compatible hardware versions.

## Ethical Considerations:
NVIDIA believes Trustworthy AI is a shared responsibility and we have established policies and practices to enable development for a wide array of AI applications.  When downloaded or used in accordance with our terms of service, developers should work with their supporting model team to ensure this model meets requirements for the relevant industry and use case and addresses unforeseen product misuse.  For more detailed information on ethical considerations for this model, please see the Model Card++ Explainability, Bias, Safety & Security, and Privacy Subcards.  Please report security vulnerabilities or NVIDIA AI Concerns [here](https://www.nvidia.com/en-us/support/submit-security-vulnerability/).

## Bias

Field                                                                                               |  Response
:---------------------------------------------------------------------------------------------------|:---------------------------------------------------------------------------------
Participation considerations from adversely impacted groups [protected classes](https://www.senate.ca.gov/content/protected-classes) in model design and testing:  |  Age (18+), Gender, Ethnicity
Measures taken to mitigate against unwanted bias:                                                   |  Collected data from public and private sources to better balance gender diversity.  Evaluated using internal, proprietary data mix to achieve identical key performance indicators.

## Explainability

Field                                                                                                  |  Response
:------------------------------------------------------------------------------------------------------|:---------------------------------------------------------------------------------
Intended Applications & Domains:                                                                       |  Content Creation and Video communication.
Types:                                                                                                 |  Facial Animation
Intended Users:                                                                                        |  Content creators, Video conference users
Output:                                                                                                |  Image
Describe how the model works:                                                                          |  The model takes audio input and an image as input. The audio is processed by a LSTM network to produce facial landmarks that animate the mouth articulation of the speech present in the audio input. These facial landmarks are then tranformed to a latent representation that encapsulates the mouth articulation and also the information in input image. This is then passed through a generative model, which generates photo-realistic animation of facial landmarks and mouth articulation of the input image to match the given audio.
Name the adversely impacted groups this has been tested to deliver comparable outcomes regardless of:  |  Age (18+), Gender, Ethnicity
Technical Limitations:                                                                                 |  Input image: This model may not perform well with images of persons with non-neutral expressions, non-neutral gazes, non-neutral head pose, occluded face, or audio with background noise.
Verified to have met prescribed NVIDIA quality standards:  |  Yes
Performance Metrics:                                                                                   |  Throughput, Latency, and other quality metrics including Mean Absolute Error (MAE) of face landmarks, MAE of face landmark velocity, MAE of mouth landmarks, MAE of mouth landmark velocity
Potential Known Risks:                                                                                 |  The model could be used to generate deep fakes if misused.
Licensing:                                                                                             | [NVIDIA Maxine Evaluation License Agreement](https://developer.download.nvidia.com/maxine/nvidia-maxine-evaluation-license-24oct2023.pdf)

## Privacy

Field                                                                                                                              |  Response
:----------------------------------------------------------------------------------------------------------------------------------|:-----------------------------------------------
Generatable or reverse engineerable personal data?                                                     |  None
Was consent obtained for any personal data used?                                                                                             |  Not Applicable
Protected class data used to create this model?                               |  Yes
How often is dataset reviewed?                                                                                                     |  Before Release
Is a mechanism in place to honor data subject right of access or deletion of personal data?                                        |  Not Applicable
If personal data collected for the development of the model, was it collected directly by NVIDIA?                                            |  Not Applicable
If personal data collected for the development of the model by NVIDIA, do you maintain or have access to disclosures made to data subjects?  |  Not Applicable
If personal data collected for the development of this AI model, was it minimized to only what was required?                                 |  Not Applicable
Is there provenance for all datasets used in training?                                                                                                      |  Yes
Does data labeling (annotation, metadata) comply with privacy laws?                                                                |  Not Applicable
Is data compliant with data subject requests for data correction or removal, if such a request was made?                           |  Not applicable

## Safety & Security

Field                                               |  Response
:---------------------------------------------------|:----------------------------------
Model Application(s):                               |  Video Conferencing and Content creation
Describe the life critical impact (if present).   |  Not Applicable
Use Case Restrictions:                              |  Refer to [Maxine Evaluation EULA](https://developer.download.nvidia.com/maxine/nvidia-maxine-evaluation-license-24oct2023.pdf)
Model and dataset protection:            |  The Principle of least privilege (PoLP) is applied limiting access for dataset generation and model development.  Restrictions enforce dataset access during training, and dataset license constraints adhered to.