Hugging Face model training from single-GPU to multi-node jobs
NVIDIA NeMo AutoModel provides GPU-accelerated, end-to-end fine-tuning for Hugging Face large language models and vision-language models with native PyTorch support. You can start training without conversion delays, using optimized kernels and memory-efficient recipes from a single GPU through distributed setups.
You'll set up a fine-tuning environment for large language models (about 1–70B parameters) and vision-language models using NeMo AutoModel on your hardware platform. By the end, you'll have a working Docker-based installation that supports parameter-efficient fine-tuning (PEFT), supervised fine-tuning (SFT), and related training workflows with FP8 precision options, while staying compatible with the Hugging Face ecosystem.
Required:
Optional:
Use the matrix below to confirm your hardware platform, recommended default local settings, and whether multi-node applies.
| Hardware platform | OS | Memory | Recommended default local settings | Multi-node capable hardware |
|---|---|---|---|---|
| DGX Spark | DGX OS (Linux) | 128 GB Unified Memory | nvcr.io/nvidia/nemo-automodel:26.02 | — |
Hardware requirements
Software requirements
nvcc --versionpython3 --versiondocker psgit --versionAll necessary files for this playbook are in the NeMo AutoModel GitHub repository.
--rm, so exiting removes it. Optionally remove the Docker image to reclaim disk space (see Cleanup in the Instructions tab). No lasting host changes beyond optional Docker group membership.