---
title: "Fine-Tune LLMs with LLaMA Factory — Troubleshooting"
canonical: "https://build.nvidia.com/playbooks/llama-factory/troubleshooting.md"
---

# Common issues

The **Hardware platform** column shows where an issue is most relevant. "All hardware platforms" applies to every platform listed in this playbook.

| Symptom | Hardware platform | Cause | Fix |
|---------|-------------------|-------|-----|
| CUDA out of memory during training | All hardware platforms | Batch size too large for available GPU memory | Reduce `per_device_train_batch_size` or increase `gradient_accumulation_steps` |
| Cannot access gated repo for URL | All hardware platforms | Certain Hugging Face models have restricted access | Regenerate your [Hugging Face token](https://huggingface.co/docs/hub/en/security-tokens); request access to the [gated model](https://huggingface.co/docs/hub/en/models-gated#customize-requested-information) in your browser |
| Model download fails or is slow | All hardware platforms | Network connectivity or Hugging Face Hub issues | Check internet connection; try `HF_HUB_OFFLINE=1` for cached models |
| Training loss not decreasing | All hardware platforms | Learning rate too high/low or insufficient data | Adjust `learning_rate` or check dataset quality |
| Memory pressure within capacity | DGX Spark | UMA buffer cache not released | See UMA note below |

> [!NOTE]
> **Unified memory (UMA).** On hardware platforms with unified memory, GPU and CPU share memory dynamically. Some applications have not yet been updated for UMA, so you may hit memory issues even within capacity. If that happens, manually flush the buffer cache:
> ```bash
> sudo sh -c 'sync; echo 3 > /proc/sys/vm/drop_caches'
> ```

For latest known issues, see the documentation linked under **Resources** for your hardware platform.