Hugging Face model training from single-GPU to multi-node jobs
| Symptom | Cause | Fix |
|---|---|---|
nvcc: command not found | CUDA toolkit not in PATH | Add CUDA toolkit to PATH: export PATH=/usr/local/cuda/bin:$PATH |
pip install uv permission denied | System-level pip restrictions | Use pip3 install --user uv and update PATH |
| GPU not detected in training | CUDA driver/runtime mismatch | Verify driver compatibility with nvidia-smi and reinstall CUDA if needed |
| Out of memory during training | Model too large for available GPU memory | Reduce batch size, enable gradient checkpointing, or use model parallelism |
| Package compatibility issues on your architecture | Package not available for the host architecture | Use source installation or build from source with architecture-appropriate flags |
| Cannot access gated repo for URL | Certain Hugging Face models have restricted access | Regenerate your Hugging Face token; request access to the gated model in your browser |
| Memory pressure within capacity | Unified memory buffer cache not released | See UMA note below |
NOTE
Some hardware platforms use Unified Memory Architecture (UMA), which enables dynamic memory sharing between the GPU and CPU. If you hit memory pressure even when within rated capacity, manually flush the buffer cache with:
sudo sh -c 'sync; echo 3 > /proc/sys/vm/drop_caches'
For latest known issues, see the documentation linked under Resources for your hardware platform.