---
title: "Fine-Tune LLMs with LLaMA Factory — Instructions"
canonical: "https://build.nvidia.com/playbooks/llama-factory/instructions.md"
---

> [!NOTE]
> These instructions target **Linux** with a Python virtual environment and GPU-enabled PyTorch. Run install and training steps in the activated venv unless noted otherwise.

# Step 1. Verify system prerequisites

Confirm that your hardware platform has the required components installed and accessible.

```bash
nvcc --version
nvidia-smi
python3 --version
git --version
```

Expected output should show a supported CUDA toolkit, GPU summary from `nvidia-smi`, Python 3, and Git.

# Step 2. Create and activate a Python virtual environment

```bash
python3 -m venv factoryEnv
source ./factoryEnv/bin/activate
```

# Step 3. Install PyTorch with CUDA 13 support

Install PyTorch, torchvision, and torchaudio with CUDA 13.0 support from the official PyTorch index.

```bash
pip3 install torch torchvision torchaudio --index-url https://download.pytorch.org/whl/cu130
```

# Step 4. Verify PyTorch CUDA support

Confirm that PyTorch can see the GPU.

```bash
python -c "import torch; print(f'PyTorch: {torch.__version__}, CUDA: {torch.cuda.is_available()}')"
```

Expected output should show a PyTorch version and `CUDA: True`.

# Step 5. Clone LLaMA Factory repository

```bash
git clone --depth 1 https://github.com/hiyouga/LLaMA-Factory.git
cd LLaMA-Factory
```

# Step 6. Install LLaMA Factory with dependencies

Install LLaMA Factory in editable mode with metrics support.

```bash
pip install -e ".[metrics]"
```

# Step 7. Prepare training configuration

Examine the provided LoRA fine-tuning configuration for Qwen3.

```bash
cat examples/train_lora/qwen3_lora_sft.yaml
```

# Step 8. Launch fine-tuning training

> [!NOTE]
> Log in to Hugging Face Hub to download the model if the model is gated.

```bash
hf auth login   # if the model is gated
llamafactory-cli train examples/train_lora/qwen3_lora_sft.yaml
```

Example output:

```
***** train metrics *****
epoch                    =        3.0
total_flos               = 11076559GF
train_loss               =     0.9993
train_runtime            = 0:14:32.12
train_samples_per_second =      3.749
train_steps_per_second   =      0.471
Figure saved at: saves/qwen3-4b/lora/sft/training_loss.png
```

# Step 9. Validate training completion

Verify that training completed successfully and checkpoints were saved.

```bash
ls -la saves/qwen3-4b/lora/sft/
```

Expected output should show:

- A final checkpoint directory (`checkpoint-411` or similar)
- Model configuration files such as `adapter_config.json`
- Training metrics showing decreasing loss values
- A training loss plot saved as a PNG file

# Step 10. Test inference with the fine-tuned model

```bash
llamafactory-cli chat examples/inference/qwen3_lora_sft.yaml
# Type: "Hello, how can you help me today?"
# Expect: Response showing fine-tuned behavior
```

# Step 11. Export the model for production deployment

```bash
llamafactory-cli export examples/merge_lora/qwen3_lora_sft.yaml
```

# Step 12. Cleanup (optional)

> [!WARNING]
> This will delete all training progress and checkpoints in the cloned repository and remove the virtual environment.

```bash
deactivate
cd ..
rm -rf LLaMA-Factory/
rm -rf factoryEnv/
```