GPU-accelerated text-to-image generation with diffusion models
To manage containers without sudo, add your user to the docker group. If you skip this step, run Docker commands with sudo.
Open a terminal and test Docker access:
docker ps
If you see a permission denied error (for example, permission denied while trying to connect to the Docker daemon socket), add your user to the docker group:
sudo usermod -aG docker $USER
newgrp docker
Start the NVIDIA PyTorch container with GPU access and Hugging Face cache mounting. This provides the TensorRT development environment with required dependencies pre-installed.
docker run --gpus all --ipc=host --ulimit memlock=-1 \
--ulimit stack=67108864 -it --rm \
-v $HOME/.cache/huggingface:/root/.cache/huggingface \
nvcr.io/nvidia/pytorch:26.03-py3
Download the TensorRT repository and configure the environment for diffusion model demos.
git clone https://github.com/NVIDIA/TensorRT.git -b main --single-branch && cd TensorRT
export TRT_OSSPATH=/workspace/TensorRT/
cd $TRT_OSSPATH/demo/Diffusion
Install the system libraries used by the image-generation demo:
apt update
apt install -y libgl1 libglu1-mesa libglib2.0-0t64 libxrender1 libxext6 libx11-6 libxrandr2 libxss1 libxcomposite1 libxdamage1 libxfixes3 libxcb1
TensorRT organizes its diffusion dependencies by model family. Because this playbook uses Flux models, install the Flux dependencies:
python3 setup.py flux
To view the installed model-family dependencies and their status, run:
python3 -c "from demo_diffusion import deps; deps.print_status()"
Set up your Hugging Face token to access gated models:
export HF_TOKEN=<YOUR_HUGGING_FACE_TOKEN>
Test multi-modal inference using the Flux.1 Dev model with different precision formats.
Substep A. BF16 precision
python3 demo_txt2img_flux.py "a beautiful photograph of Mt. Fuji during cherry blossom" \
--hf-token=$HF_TOKEN --download-onnx-models --bf16
Substep B. FP8 quantized precision
python3 demo_txt2img_flux.py "a beautiful photograph of Mt. Fuji during cherry blossom" \
--hf-token=$HF_TOKEN --quantization-level 4 --fp8 --download-onnx-models
Substep C. FP4 quantized precision
python3 demo_txt2img_flux.py "a beautiful photograph of Mt. Fuji during cherry blossom" \
--hf-token=$HF_TOKEN --fp4 --download-onnx-models
Test the faster Flux.1 Schnell variant with different precision formats.
WARNING
FP16 Flux.1 Schnell requires more than 48 GB of available memory for native export.
Substep A. FP16 precision (high memory requirement)
python3 demo_txt2img_flux.py "a beautiful photograph of Mt. Fuji during cherry blossom" \
--hf-token=$HF_TOKEN --version="flux.1-schnell"
Substep B. FP8 quantized precision
python3 demo_txt2img_flux.py "a beautiful photograph of Mt. Fuji during cherry blossom" \
--hf-token=$HF_TOKEN --version="flux.1-schnell" \
--quantization-level 4 --fp8 --download-onnx-models
Substep C. FP4 quantized precision
python3 demo_txt2img_flux.py "a beautiful photograph of Mt. Fuji during cherry blossom" \
--hf-token=$HF_TOKEN --version="flux.1-schnell" \
--fp4 --download-onnx-models
Confirm that the models generated images successfully and that TensorRT is available.
# Check for generated images in the output directory
ls -la *.png *.jpg 2>/dev/null || echo "No image files found"
# Verify CUDA is accessible
nvidia-smi
# Check TensorRT version
python3 -c "import tensorrt as trt; print(f'TensorRT version: {trt.__version__}')"
Remove downloaded models and exit the container environment to free disk space when you are finished.
WARNING
This deletes cached models and generated images if you remove the Hugging Face cache.
# Exit container
exit
# Remove Hugging Face cache (optional)
rm -rf $HOME/.cache/huggingface/
Use the validated setup to generate custom images or integrate multi-modal inference into your applications. Try different prompts, precision formats, or explore model fine-tuning with the established TensorRT environment.