---
title: "Set Up Claude Code with Local Inference — Claude Code"
canonical: "https://build.nvidia.com/playbooks/local-coding-agent/claude-code.md"
---

# Step 1. Confirm your environment

Verify the GPU is visible before installing anything.

```bash
nvidia-smi
```

Expected output should show a detected GPU and driver version for your hardware platform.

# Step 2. Install or update Ollama

Install [Ollama](https://ollama.com/download) or confirm your install supports [`ollama launch`](https://ollama.com/blog/launch).

```bash
curl -fsSL https://ollama.com/install.sh | sh
ollama --version
```

If Ollama is already installed, just verify the version:

```bash
ollama --version
```

Expected output should show a current Ollama release.

# Step 3. Pull a coding model

Download the [Qwen3.6](https://ollama.com/library/qwen3.6) model weights to your hardware platform.

This playbook uses **qwen3.6:27b** with Claude Code through Ollama:

```bash
ollama pull qwen3.6:27b
```

Expected output should show progress lines followed by success. Confirm the model appears in the list:

```bash
ollama list
```

```text
NAME                                ID              SIZE    MODIFIED
qwen3.6:27b                         abc123...       ...     1 minute ago
```

# Step 4. Test local inference

Run a quick prompt to confirm the model loads.

```bash
ollama run qwen3.6:27b
```

Try a prompt like:

```text
Write a short README checklist for a Python project.
```

Expected output should show the model responding in the terminal. When you are done, type `/bye` or press **Ctrl+D** to exit the interactive session before continuing.

# Step 5. Install Claude Code

Install [Claude Code](https://docs.claude.com/en/docs/claude-code), the CLI tool that will drive the local model.

```bash
curl -fsSL https://claude.ai/install.sh | bash
```

Verify the installation:

```bash
claude --version
```

Expected output should show a version string such as `claude 0.x.x`. If you see `claude: command not found`, ensure the install script added the CLI to your PATH (for example, restart the terminal or source your shell profile); see [Troubleshooting](troubleshooting.md).

# Step 6. Increase context length (optional)

Ollama defaults to a 4096 token context length. For coding agents and larger codebases, set it to 64K tokens. This increases memory usage. For more details on configuring context length and other parameters, see the [Ollama documentation](https://ollama.com/docs).

Set the context length per session in the Ollama REPL:

```bash
ollama run qwen3.6:27b
```

Then, in the Ollama prompt:

```text
/set parameter num_ctx 64000
```

Exit when done: type `/bye` or press **Ctrl+D**.

Optional method (set globally when serving Ollama):

```bash
sudo systemctl stop ollama
OLLAMA_CONTEXT_LENGTH=64000 ollama serve
```

Keep this terminal open and run the next step in a new terminal.

# Step 7. Connect Claude Code to Ollama

Launch Claude Code through Ollama with the model you pulled. No environment variables or config files are required.

```bash
ollama launch claude --model qwen3.6:27b
```

Expected output should show Claude Code starting and using the local Ollama model.

Exit Claude Code when done: type `/exit` or press **Ctrl+C**.

# Step 8. Complete a small coding task

Create a tiny repo and let Claude Code implement a function and tests.

```bash
mkdir -p ~/cli-agent-demo
cd ~/cli-agent-demo
python3 -m venv .venv
source .venv/bin/activate
python3 -m pip install -U pytest

printf 'def add(a, b):\n    """Return the sum of a and b."""\n    pass\n' > math_utils.py
printf 'import math_utils\n\n\ndef test_add():\n    assert math_utils.add(1, 2) == 3\n' > test_math_utils.py
```

If Claude Code is not already running, launch it:

```bash
ollama launch claude --model qwen3.6:27b
```

In Claude Code, enter:

```text
Please implement add() in math_utils.py and make sure the test passes.
```

Exit Claude Code when finished: type `/exit` or press **Ctrl+C**, then run the test:

```bash
python3 -m pytest -q
deactivate
```

Expected output should show the test passing.

# Step 9. Cleanup

Remove the model and stop the Ollama service if you no longer need them. Cleanup is optional. **Remove the model first** (while the Ollama server is running), then stop the service.

> [!WARNING]
> The following removes the downloaded model files from disk.

**1. Remove the model** (Ollama must be running). Use the same name you pulled:

```bash
ollama rm qwen3.6:27b
```

**2. Stop the Ollama service**:

```bash
sudo systemctl stop ollama
```

# Step 10. Next steps

- Use larger context (for example, 64K–198K) for big codebases
- Use Claude Code on multi-file refactors or test-generation tasks
- Browse additional models in the [Ollama library](https://ollama.com/library)