---
title: "Fine-Tune FLUX.1 for Custom Image Generation — Instructions"
canonical: "https://build.nvidia.com/playbooks/flux-finetuning/instructions.md"
---

# Step 1. Configure Docker permissions

To manage containers without `sudo`, your user must be in the `docker` group. If you skip this step, run Docker commands with `sudo`.

Open a terminal and test Docker access:

```bash
docker ps
```

If you see a permission-denied error (for example, permission denied while trying to connect to the Docker daemon socket), add your user to the docker group:

```bash
sudo usermod -aG docker $USER
newgrp docker
```

# Step 2. Clone the repository

Clone the playbook assets and open the assets directory:

```bash
git clone https://github.com/NVIDIA/dgx-spark-playbooks
cd dgx-spark-playbooks/nvidia/playbook-flux-finetuning/assets
```

# Step 3. Model download

FLUX.1-dev is gated. Open the [model card](https://huggingface.co/black-forest-labs/FLUX.1-dev), accept the terms, and gain access to the checkpoints. If you do not already have an `HF_TOKEN`, follow the [Hugging Face token guide](https://huggingface.co/docs/hub/en/security-tokens), then authenticate:

```bash
export HF_TOKEN=<YOUR_HF_TOKEN>
sh download.sh
```

The download can take about 30–45 minutes depending on network speed. It pulls approximately:

- `flux1-dev.safetensors` (~23.8 GB)
- `ae.safetensors` (~335 MB)
- `clip_l.safetensors` (~246 MB)
- `t5xxl_fp16.safetensors` (~9.8 GB)

After download, `models/` should look like:

```text
models/
├── checkpoints/
│   └── flux1-dev.safetensors
├── loras/
├── text_encoders/
│   ├── clip_l.safetensors
│   └── t5xxl_fp16.safetensors
└── vae/
└── ae.safetensors
```

If you already have fine-tuned LoRAs, place them in `models/loras/`. Otherwise continue to training in Step 6.

# Step 4. Base model inference

Generate an image with the base FLUX.1 model for the sample concepts (Toy Jensen and a custom GPU) before training.

```bash
# Build the inference Docker image (run from assets/)
docker build -f Dockerfile.inference -t flux-comfyui .

# Launch ComfyUI; you can ignore import errors for torchaudio
sh launch_comfyui.sh
```

Open ComfyUI at `http://localhost:8188` (or `http://<HARDWARE_IP>:8188` from another device). Do not select a pre-existing template.

Open the workflow panel (left side, or press `w`) and load `base_flux.json`. Enter a prompt in the **CLIP Text Encode (Prompt)** node — for example, `Toy Jensen holding a DGX Spark in a datacenter`. High-resolution 1024px generation can take about three minutes.

Next steps:

- If you already placed LoRAs in `models/loras/`, skip to Step 7.
- If you will train, stop the ComfyUI container with `Ctrl+C` first.

> [!NOTE]
> To clear buffer cache after stopping ComfyUI (outside the container):
> ```bash
> sudo sh -c 'sync; echo 3 > /proc/sys/vm/drop_caches'
> ```

# Step 5. Dataset preparation

Prepare a dataset for Dreambooth LoRA fine-tuning on FLUX.1-dev. This playbook ships a two-concept sample dataset of public-domain images. If you use those concepts as-is, you do not need to edit `data.toml`.

**TJToy concept**

- **Trigger phrase:** `tjtoy toy`
- **Training images:** 6 images of Toy Jensen figures
- **Use case:** Generate scenes featuring that toy character

**SparkGPU concept**

- **Trigger phrase:** `sparkgpu gpu`
- **Training images:** 7 images of a custom GPU design
- **Use case:** Generate scenes featuring that GPU

For your own concepts, collect about 5–10 images per concept. Create one folder per concept under `flux_data/` (this playbook uses `tjtoy` and `sparkgpu`). Update `flux_data/data.toml` so each `[[datasets.subsets]]` entry has the correct `image_dir` and `class_tokens`. Appending a class token (for example `toy` or `gpu`) usually improves fine-tuning.

# Step 6. Training

Build the training image and start Dreambooth LoRA training:

```bash
docker build -f Dockerfile.train -t flux-train .
sh launch_train.sh
```

`launch_train.sh` runs `--max_train_epochs=100` and saves a LoRA checkpoint every 25 epochs (`--save_every_n_epochs=25`) into `models/loras/`, named with the `flux_dreambooth` prefix. A complete 100-epoch run takes about four hours and gives the highest quality.

You do not have to wait for the full run. Intermediate checkpoints are usable on their own: results that capture the sample concepts often appear within roughly the first 90 minutes of training. To use an earlier checkpoint, pick the most recent file in `models/loras/` and load it in ComfyUI (Step 7). For a shorter run overall, lower the epoch count in `launch_train.sh`, for example:

```bash
--max_train_epochs=25
```

Other useful knobs in `launch_train.sh` include LoRA dimension and alpha (256), learning rate (1.0 with the Prodigy optimizer), mixed precision (`bfloat16`), and caching / torch compile options. Training resolution comes from `flux_data/data.toml` (1024×1024 by default).

# Step 7. Fine-tuned model inference

Generate images with your trained LoRAs:

```bash
# Launch ComfyUI (from assets/); you can ignore import errors for torchaudio
sh launch_comfyui.sh
```

Open `http://localhost:8188`, skip pre-existing templates, open the workflow panel (`w`), and load `finetuned_flux.json`.

Prompt with your trigger phrases — for example, `tjtoy toy holding sparkgpu gpu in a datacenter`. Expect about three minutes for 1024px generation. The fine-tuned path can combine multiple concepts in one image. Use ComfyUI nodes to adjust LoRA strength, resolution, seed, sampler, scheduler, and steps.

# Step 8. Cleanup (optional)

Stop running containers with `Ctrl+C`. Remove local images if you no longer need them:

```bash
docker rmi flux-comfyui flux-train
```

> [!WARNING]
> Removing the images deletes the local Docker builds. Rebuild them with the Dockerfiles in `assets/` before running this playbook again. Downloaded models under `models/` are separate; delete those directories only if you intend to free disk space.