---
title: "Set Up Claude Code with Local Inference"
publisher: "nvidia"
type: "playbook"
updated: "2026-09-22T17:55:58.025Z"
description: "A coding agent connected to a local Ollama model"
canonical: "https://build.nvidia.com/playbooks/local-coding-agent.md"
---

# Basic idea

Use [Ollama](https://ollama.com) on your **hardware platform** to run a local coding model and connect a CLI coding agent. This playbook uses **[Claude Code](https://docs.claude.com/en/docs/claude-code)** with Ollama's built-in [launch method](https://ollama.com/blog/launch) (`ollama launch claude`), so you can work without environment variables, provider config files, or external cloud APIs.

# CLI agent

This playbook uses **Claude Code** as the CLI agent, connected to a local Ollama model for inference.

# What you'll accomplish

You'll run a local coding model ([Qwen3.6](https://ollama.com/library/qwen3.6) `qwen3.6:27b`) on your **hardware platform** with Ollama, launch Claude Code against it with a single command, and complete a small coding task end-to-end.

# What to know before starting

**Required:**

- Comfort with Linux command line basics
- Experience running terminal-based tools and editors
- Familiarity with Python for the short coding task

**Optional:**

- Experience browsing models on the [Ollama library](https://ollama.com/library) (for example, [Qwen3.6](https://ollama.com/library/qwen3.6))

# Supported hardware platforms

Use the matrix below to confirm your hardware platform, recommended default local settings, and whether multi-node applies.

| Hardware platform | OS | Memory | Recommended default local settings | Multi-node capable hardware |
| :---- | :---- | :---- | :---- | :---- |
| **DGX Station** | DGX OS (Linux) | Large HBM + Grace DRAM | Ollama + `qwen3.6:27b`; Claude Code via `ollama launch` | — |

# Prerequisites

**Hardware requirements**

- Supported hardware platform — see Supported hardware platforms matrix above
- Hardware platform powered on, networked, and reachable for local or SSH terminal access
- Sufficient GPU memory for `qwen3.6:27b`
- Enough free storage for the model download

**Software requirements**

- A [current Ollama release](https://ollama.com/download) (required for [`ollama launch`](https://ollama.com/blog/launch) and the recommended model): `ollama --version`
- Internet access to download model weights
- Python 3 with `venv` support for the coding-task verification

# Find model recipes

Browse models you can pull and run with Ollama in the [Ollama library](https://ollama.com/library). Use the tags and sizes that fit your hardware platform’s memory and storage.

| Hardware platform | More recipes |
| ----------------- | ------------ |
| **DGX Station** | [Ollama library](https://ollama.com/library) · [Qwen3.6](https://ollama.com/library/qwen3.6) |

Use the **Claude Code** tab for the base workflow.

# Ancillary files

No local ancillary files are required. All steps use Ollama and Claude Code on your hardware platform.

# Time & risk

- **Estimated time:** 30 MIN (mostly model download time)
- **Risk level:** Low
- Large model downloads can fail if network connectivity is unstable
- Ollama must support [`ollama launch`](https://ollama.com/blog/launch) and the recommended model tag — verify with `ollama --version` and install from [ollama.com/download](https://ollama.com/download) if needed
- **Rollback:** Stop Ollama and delete the downloaded model from `~/.ollama/models` (see Cleanup in the Claude Code tab). Cleanup is optional and removes downloaded model files.
- **Last Updated:** 08/03/2026
- Set up Claude Code against a local Qwen3.6 model with `ollama launch` on supported hardware platforms

## More

- [Claude Code](/playbooks/local-coding-agent/claude-code.md)
- [Troubleshooting](/playbooks/local-coding-agent/troubleshooting.md)