One interface for supervised, RLHF, and parameter-efficient training
The Hardware platform column shows where an issue is most relevant. "All hardware platforms" applies to every platform listed in this playbook.
| Symptom | Hardware platform | Cause | Fix |
|---|---|---|---|
| CUDA out of memory during training | All hardware platforms | Batch size too large for available GPU memory | Reduce per_device_train_batch_size or increase gradient_accumulation_steps |
| Cannot access gated repo for URL | All hardware platforms | Certain Hugging Face models have restricted access | Regenerate your Hugging Face token; request access to the gated model in your browser |
| Model download fails or is slow | All hardware platforms | Network connectivity or Hugging Face Hub issues | Check internet connection; try HF_HUB_OFFLINE=1 for cached models |
| Training loss not decreasing | All hardware platforms | Learning rate too high/low or insufficient data | Adjust learning_rate or check dataset quality |
| Memory pressure within capacity | DGX Spark | UMA buffer cache not released | See UMA note below |
NOTE
Unified memory (UMA). On hardware platforms with unified memory, GPU and CPU share memory dynamically. Some applications have not yet been updated for UMA, so you may hit memory issues even within capacity. If that happens, manually flush the buffer cache:
sudo sh -c 'sync; echo 3 > /proc/sys/vm/drop_caches'
For latest known issues, see the documentation linked under Resources for your hardware platform.