---
title: "Run Multi-Modal Inference with TensorRT"
publisher: "nvidia"
type: "playbook"
updated: "2026-10-01T20:29:43.839Z"
description: "GPU-accelerated text-to-image generation with diffusion models"
canonical: "https://build.nvidia.com/playbooks/multi-modal-inference.md"
---

# Basic idea

Multi-modal inference combines different data types, such as **text, images, and audio**, within a single model pipeline to generate or interpret richer outputs. Instead of processing one input type at a time, multi-modal systems share representations that support **text-to-image generation**, **image captioning**, or **vision-language reasoning**.

On GPUs, this enables **parallel processing across modalities** for faster, higher-fidelity results on tasks that combine language and vision.

# What you'll accomplish

You'll deploy GPU-accelerated multi-modal inference on your **hardware platform** using torch and torch-TensorRT to run Flux.1 diffusion models with optimized performance across multiple precision formats (BF16, FP16, FP8, FP4).

# What to know before starting

**Required:**

- Working with Docker containers and GPU passthrough
- Using torch-TensorRT for model optimization
- Hugging Face model hub authentication and downloads
- Command-line tools for GPU workloads

**Optional:**

- Basic understanding of diffusion models and image generation

# Supported hardware platforms

Use the matrix below to confirm your hardware platform, default container image, and whether multi-node applies.

| Hardware platform | OS | Memory | Recommended default local settings | Multi-node capable hardware |
| :---- | :---- | :---- | :---- | :---- |
| **DGX Spark** | DGX OS (Linux) | 128 GB unified memory | `nvcr.io/nvidia/pytorch:26.03-py3` | — |

# Prerequisites

**Hardware requirements**

- Supported hardware platform — see the matrix above
- At least 48 GB available memory for FP16 Flux.1 Schnell operations
- Sufficient available storage for container images, Hugging Face model downloads, and generated outputs

**Software requirements**

- NVIDIA driver and GPU visibility: `nvidia-smi`
- Docker installed and accessible to the current user: `docker --version`
- NVIDIA Container Toolkit configured for Docker
- Hugging Face account with access to Black Forest Labs models [FLUX.1-dev](https://huggingface.co/black-forest-labs/FLUX.1-dev) and [FLUX.1-dev-onnx](https://huggingface.co/black-forest-labs/FLUX.1-dev-onnx)
- Hugging Face [token](https://huggingface.co/settings/tokens) configured with access to both FLUX.1 model repositories
- Network access to NGC and Hugging Face

Verify GPU and Docker GPU integration:

```bash
nvidia-smi
docker run --rm --gpus all nvcr.io/nvidia/pytorch:26.03-py3 nvidia-smi
```

# Find model recipes

Start with the Flux.1 Dev and Flux.1 Schnell precision examples in **Instructions**. For additional diffusion demos, scripts, and dependency files, use the torch-TensorRT documentation (Compiling LLM models from Huggingface) **Resources**.

| Hardware platform | More recipes |
| ----------------- | ------------ |
| **DGX Spark** | [Compiling LLM models from Huggingface](https://docs.pytorch.org/TensorRT/tutorials/_rendered_examples/dynamo/torch_export_flux_dev.html) |

Follow the documentation for an example inference workflow for Flux.1-dev.

> [!NOTE]
> **Memory determines what you can run.** FP16 Flux.1 Schnell needs substantially more memory than FP8 or FP4. If a model or precision is not listed for your hardware platform, confirm it fits available memory before downloading.

# Ancillary files

Alternatively, use the example hosted on [torch-TensorRT](https://github.com/pytorch/TensorRT/tree/main/examples/apps):

- [**README.md**](https://github.com/pytorch/TensorRT/blob/main/examples/apps/README.md) — Explanation how to run the example
- [**flux_demo.py**](https://github.com/pytorch/TensorRT/blob/main/examples/apps/flux_demo.py) — Flux.1 model inference script

# Time & risk

- **Estimated time:** 60 MIN (longer on first run due to model downloads and optimization)
- **Risk level:** Medium
- Large model downloads may time out
- High memory requirements may cause out-of-memory errors
- Quantized models may show quality differences versus full precision
- **Rollback:** Exit the container, then optionally remove downloaded models from the Hugging Face cache
- **Last Updated:** 08/31/2026
- Flux.1 TensorRT inference workflow via torch-TensorRT for text-to-image generation

## More

- [Instructions](/playbooks/multi-modal-inference/instructions.md)
- [Troubleshooting](/playbooks/multi-modal-inference/troubleshooting.md)