---
title: "Run Multi-Modal Inference with TensorRT — Instructions"
canonical: "https://build.nvidia.com/playbooks/multi-modal-inference/instructions.md"
---

# Step 1. Configure Docker permissions

To manage containers without sudo, add your user to the `docker` group. If you skip this step, run Docker commands with sudo.

Open a terminal and test Docker access:

```bash
docker ps
```

If you see a permission denied error (for example, permission denied while trying to connect to the Docker daemon socket), add your user to the docker group:

```bash
sudo usermod -aG docker $USER
newgrp docker
```

# Step 2. Launch the TensorRT container environment

Start the NVIDIA PyTorch container with GPU access and Hugging Face cache mounting. This provides the TensorRT development environment with required dependencies pre-installed.

```bash
docker run --gpus all --ipc=host --ulimit memlock=-1 \
--ulimit stack=67108864 -it --rm \
-v $HOME/.cache/huggingface:/root/.cache/huggingface \
nvcr.io/nvidia/pytorch:26.03-py3
```

# Step 3. Clone and set up the TensorRT repository

Download the TensorRT repository and configure the environment for diffusion model demos.

```bash
git clone https://github.com/NVIDIA/TensorRT.git -b main --single-branch && cd TensorRT
export TRT_OSSPATH=/workspace/TensorRT/
cd $TRT_OSSPATH/demo/Diffusion
```

# Step 4. Install required dependencies

Install the system libraries used by the image-generation demo:

```bash
apt update
apt install -y libgl1 libglu1-mesa libglib2.0-0t64 libxrender1 libxext6 libx11-6 libxrandr2 libxss1 libxcomposite1 libxdamage1 libxfixes3 libxcb1
```

TensorRT organizes its diffusion dependencies by model family. Because this playbook uses Flux models, install the Flux dependencies:

```bash
python3 setup.py flux
```

To view the installed model-family dependencies and their status, run:

```bash
python3 -c "from demo_diffusion import deps; deps.print_status()"
```

Set up your Hugging Face token to access gated models:

```bash
export HF_TOKEN=<YOUR_HUGGING_FACE_TOKEN>
```

# Step 5. Run Flux.1 Dev model inference

Test multi-modal inference using the Flux.1 Dev model with different precision formats.

**Substep A. BF16 precision**

```bash
python3 demo_txt2img_flux.py "a beautiful photograph of Mt. Fuji during cherry blossom" \
--hf-token=$HF_TOKEN --download-onnx-models --bf16
```

**Substep B. FP8 quantized precision**

```bash
python3 demo_txt2img_flux.py "a beautiful photograph of Mt. Fuji during cherry blossom" \
--hf-token=$HF_TOKEN --quantization-level 4 --fp8 --download-onnx-models
```

**Substep C. FP4 quantized precision**

```bash
python3 demo_txt2img_flux.py "a beautiful photograph of Mt. Fuji during cherry blossom" \
--hf-token=$HF_TOKEN --fp4 --download-onnx-models
```

# Step 6. Run Flux.1 Schnell model inference

Test the faster Flux.1 Schnell variant with different precision formats.

> [!WARNING]
> FP16 Flux.1 Schnell requires more than 48 GB of available memory for native export.

**Substep A. FP16 precision (high memory requirement)**

```bash
python3 demo_txt2img_flux.py "a beautiful photograph of Mt. Fuji during cherry blossom" \
--hf-token=$HF_TOKEN --version="flux.1-schnell"
```

**Substep B. FP8 quantized precision**

```bash
python3 demo_txt2img_flux.py "a beautiful photograph of Mt. Fuji during cherry blossom" \
--hf-token=$HF_TOKEN --version="flux.1-schnell" \
--quantization-level 4 --fp8 --download-onnx-models
```

**Substep C. FP4 quantized precision**

```bash
python3 demo_txt2img_flux.py "a beautiful photograph of Mt. Fuji during cherry blossom" \
--hf-token=$HF_TOKEN --version="flux.1-schnell" \
--fp4 --download-onnx-models
```

# Step 7. Validate inference outputs

Confirm that the models generated images successfully and that TensorRT is available.

```bash
# Check for generated images in the output directory
ls -la *.png *.jpg 2>/dev/null || echo "No image files found"

# Verify CUDA is accessible
nvidia-smi

# Check TensorRT version
python3 -c "import tensorrt as trt; print(f'TensorRT version: {trt.__version__}')"
```

# Step 8. Cleanup and rollback

Remove downloaded models and exit the container environment to free disk space when you are finished.

> [!WARNING]
> This deletes cached models and generated images if you remove the Hugging Face cache.

```bash
# Exit container
exit

# Remove Hugging Face cache (optional)
rm -rf $HOME/.cache/huggingface/
```

# Step 9. Next steps

Use the validated setup to generate custom images or integrate multi-modal inference into your applications. Try different prompts, precision formats, or explore model fine-tuning with the established TensorRT environment.