Basic idea
LLaMA Factory is an open-source framework that simplifies training and fine-tuning large language models. It offers a unified interface for methods such as supervised fine-tuning (SFT), RLHF, and QLoRA, and supports a wide range of LLM architectures including LLaMA, Mistral, and Qwen.
- Unified CLI — one interface for supervised, RLHF, and parameter-efficient training
- Broad model support — works with popular open LLM families and Hugging Face workflows
- Flexible fine-tuning — LoRA, QLoRA, and full fine-tuning with ready example configs
What you'll accomplish
You'll set up LLaMA Factory on your hardware platform to fine-tune large language models with LoRA, QLoRA, and full fine-tuning methods, using the CLI and example configs from the upstream repository.
What to know before starting
Required:
- Basic Python knowledge for editing config files and troubleshooting
- Command-line usage for shell commands and virtual environments
- Familiarity with PyTorch and the Hugging Face Transformers ecosystem
- Fine-tuning concepts: tradeoffs between LoRA, QLoRA, and full fine-tuning
Optional:
- Dataset preparation: formatting text data into JSON for instruction tuning
- Resource management: adjusting batch size and memory settings for GPU constraints
Supported hardware platforms
Use the matrix below to confirm your hardware platform, recommended default local settings, and whether multi-node applies.
| Hardware platform | OS | Memory | Recommended default local settings | Multi-node capable hardware |
|---|
| DGX Spark | DGX OS (Linux) | 128 GB Unified Memory | Python venv + PyTorch CUDA 13.0 (cu130) | — |
Prerequisites
Hardware requirements
- Supported hardware platform — see Supported hardware platforms matrix above
- Sufficient storage space (plan for >50 GB for models and checkpoints):
df -h
Software requirements
- CUDA 12.9 or newer:
nvcc --version
- Git:
git --version
- Python 3 with venv and pip:
python3 --version && pip3 --version
- Internet connection for downloading models from Hugging Face Hub
Ancillary files
All required assets are in the LLaMA Factory repository.
examples/train_lora/qwen3_lora_sft.yaml — example LoRA SFT training configuration
examples/inference/qwen3_lora_sft.yaml — example chat/inference configuration for the fine-tuned adapter
examples/merge_lora/qwen3_lora_sft.yaml — example export/merge configuration for production use
- Data preparation docs — dataset formatting guidance
Time & risk
- Estimated time: 60 MIN for initial setup; training can take 1–7 hours depending on model size and dataset
- Risk level: Medium
- Model downloads require significant bandwidth and storage
- Training may consume substantial GPU memory and need batch-size or accumulation tuning
- Rollback: Deactivate the virtual environment and remove the
factoryEnv and LLaMA-Factory directories. Delete local training checkpoints to reclaim storage.
- Last Updated: 07/31/2026
- Set up LLaMA Factory with a venv-based PyTorch CUDA 13 workflow for Qwen3 LoRA fine-tuning, chat validation, and export