GPU-accelerated text-to-image generation with diffusion models
Multi-modal inference combines different data types, such as text, images, and audio, within a single model pipeline to generate or interpret richer outputs. Instead of processing one input type at a time, multi-modal systems share representations that support text-to-image generation, image captioning, or vision-language reasoning.
On GPUs, this enables parallel processing across modalities for faster, higher-fidelity results on tasks that combine language and vision.
You'll deploy GPU-accelerated multi-modal inference on your hardware platform using torch and torch-TensorRT to run Flux.1 diffusion models with optimized performance across multiple precision formats (BF16, FP16, FP8, FP4).
Required:
Optional:
Use the matrix below to confirm your hardware platform, default container image, and whether multi-node applies.
| Hardware platform | OS | Memory | Recommended default local settings | Multi-node capable hardware |
|---|---|---|---|---|
| DGX Spark | DGX OS (Linux) | 128 GB unified memory | nvcr.io/nvidia/pytorch:26.03-py3 | — |
Hardware requirements
Software requirements
nvidia-smidocker --versionVerify GPU and Docker GPU integration:
nvidia-smi
docker run --rm --gpus all nvcr.io/nvidia/pytorch:26.03-py3 nvidia-smi
Start with the Flux.1 Dev and Flux.1 Schnell precision examples in Instructions. For additional diffusion demos, scripts, and dependency files, use the torch-TensorRT documentation (Compiling LLM models from Huggingface) Resources.
| Hardware platform | More recipes |
|---|---|
| DGX Spark | Compiling LLM models from Huggingface |
Follow the documentation for an example inference workflow for Flux.1-dev.
NOTE
Memory determines what you can run. FP16 Flux.1 Schnell needs substantially more memory than FP8 or FP4. If a model or precision is not listed for your hardware platform, confirm it fits available memory before downloading.
Alternatively, use the example hosted on torch-TensorRT: