---
title: "Fine-Tune with NVIDIA NeMo"
publisher: "nvidia"
type: "playbook"
updated: "2026-09-22T17:55:59.891Z"
description: "Hugging Face model training from single-GPU to multi-node jobs"
canonical: "https://build.nvidia.com/playbooks/nemo-fine-tune.md"
---

# Basic idea

NVIDIA NeMo AutoModel provides GPU-accelerated, end-to-end fine-tuning for Hugging Face large language models and vision-language models with native PyTorch support. You can start training without conversion delays, using optimized kernels and memory-efficient recipes from a single GPU through distributed setups.

# What you'll accomplish

You'll set up a fine-tuning environment for large language models (about 1–70B parameters) and vision-language models using NeMo AutoModel on your hardware platform. By the end, you'll have a working Docker-based installation that supports parameter-efficient fine-tuning (PEFT), supervised fine-tuning (SFT), and related training workflows with FP8 precision options, while staying compatible with the Hugging Face ecosystem.

# What to know before starting

**Required:**

- Working in Linux terminal environments and SSH connections
- Basic understanding of Python virtual environments and package management
- Familiarity with GPU computing concepts and CUDA toolkit usage
- Experience with containerized workflows and Docker operations
- Understanding of machine learning model training and fine-tuning concepts

**Optional:**

- Experience with distributed training configurations on multi-node capable hardware

# Supported hardware platforms

Use the matrix below to confirm your hardware platform, recommended default local settings, and whether multi-node applies.

| Hardware platform | OS | Memory | Recommended default local settings | Multi-node capable hardware |
| :---- | :---- | :---- | :---- | :---- |
| **DGX Spark** | DGX OS (Linux) | 128 GB Unified Memory | `nvcr.io/nvidia/nemo-automodel:26.02` | — |

# Prerequisites

**Hardware requirements**

- Supported hardware platform — see Supported hardware platforms matrix above
- Minimum 32 GB system memory for efficient model loading and training
- Active internet connection for downloading models and packages
- SSH access to your hardware platform configured

**Software requirements**

- CUDA toolkit 12.0+ installed and configured: `nvcc --version`
- Python 3.10+ environment available: `python3 --version`
- Docker installed and usable: `docker ps`
- Git installed for repository cloning: `git --version`
- Hugging Face account and access token (for gated models)

# Ancillary files

All necessary files for this playbook are in the [NeMo AutoModel GitHub repository](https://github.com/NVIDIA-NeMo/Automodel).

# Time & risk

- **Estimated time:** 45–90 MIN for complete setup and an initial fine-tuning run (longer on first run due to model download)
- **Risk level:** Medium
- Model downloads can be large (several GB)
- Package or architecture compatibility issues may require troubleshooting
- Distributed training complexity increases if you extend beyond the single-node examples in this playbook
- **Rollback:** The container was launched with `--rm`, so exiting removes it. Optionally remove the Docker image to reclaim disk space (see Cleanup in the **Instructions** tab). No lasting host changes beyond optional Docker group membership.
- **Last Updated:** 07/31/2026
- NeMo AutoModel Docker workflow for LoRA, QLoRA, and full SFT fine-tuning on supported hardware platforms

## More

- [Instructions](/playbooks/nemo-fine-tune/instructions.md)
- [Troubleshooting](/playbooks/nemo-fine-tune/troubleshooting.md)