---
title: "Fine-Tune LLMs with LLaMA Factory — Overview"
canonical: "https://build.nvidia.com/playbooks/llama-factory/overview.md"
---

# Basic idea

LLaMA Factory is an open-source framework that simplifies training and fine-tuning large language models. It offers a unified interface for methods such as supervised fine-tuning (SFT), RLHF, and QLoRA, and supports a wide range of LLM architectures including LLaMA, Mistral, and Qwen.

- **Unified CLI** — one interface for supervised, RLHF, and parameter-efficient training
- **Broad model support** — works with popular open LLM families and Hugging Face workflows
- **Flexible fine-tuning** — LoRA, QLoRA, and full fine-tuning with ready example configs

# What you'll accomplish

You'll set up LLaMA Factory on your **hardware platform** to fine-tune large language models with LoRA, QLoRA, and full fine-tuning methods, using the CLI and example configs from the upstream repository.

# What to know before starting

**Required:**

- Basic Python knowledge for editing config files and troubleshooting
- Command-line usage for shell commands and virtual environments
- Familiarity with PyTorch and the Hugging Face Transformers ecosystem
- Fine-tuning concepts: tradeoffs between LoRA, QLoRA, and full fine-tuning

**Optional:**

- Dataset preparation: formatting text data into JSON for instruction tuning
- Resource management: adjusting batch size and memory settings for GPU constraints

# Supported hardware platforms

Use the matrix below to confirm your hardware platform, recommended default local settings, and whether multi-node applies.

| Hardware platform | OS | Memory | Recommended default local settings | Multi-node capable hardware |
| :---- | :---- | :---- | :---- | :---- |
| **DGX Spark** | DGX OS (Linux) | 128 GB Unified Memory | Python venv + PyTorch CUDA 13.0 (`cu130`) | — |

# Prerequisites

**Hardware requirements**

- Supported hardware platform — see Supported hardware platforms matrix above
- Sufficient storage space (plan for >50 GB for models and checkpoints): `df -h`

**Software requirements**

- CUDA 12.9 or newer: `nvcc --version`
- Git: `git --version`
- Python 3 with venv and pip: `python3 --version && pip3 --version`
- Internet connection for downloading models from Hugging Face Hub

# Ancillary files

All required assets are in the [LLaMA Factory repository](https://github.com/hiyouga/LLaMA-Factory).

- `examples/train_lora/qwen3_lora_sft.yaml` — example LoRA SFT training configuration
- `examples/inference/qwen3_lora_sft.yaml` — example chat/inference configuration for the fine-tuned adapter
- `examples/merge_lora/qwen3_lora_sft.yaml` — example export/merge configuration for production use
- [Data preparation docs](https://llamafactory.readthedocs.io/en/latest/getting_started/data_preparation.html) — dataset formatting guidance

# Time & risk

- **Estimated time:** 60 MIN for initial setup; training can take 1–7 hours depending on model size and dataset
- **Risk level:** Medium
- Model downloads require significant bandwidth and storage
- Training may consume substantial GPU memory and need batch-size or accumulation tuning
- **Rollback:** Deactivate the virtual environment and remove the `factoryEnv` and `LLaMA-Factory` directories. Delete local training checkpoints to reclaim storage.
- **Last Updated:** 07/31/2026
- Set up LLaMA Factory with a venv-based PyTorch CUDA 13 workflow for Qwen3 LoRA fine-tuning, chat validation, and export