A coding agent connected to a local Ollama model
Verify the GPU is visible before installing anything.
nvidia-smi
Expected output should show a detected GPU and driver version for your hardware platform.
Install Ollama or confirm your install supports ollama launch.
curl -fsSL https://ollama.com/install.sh | sh
ollama --version
If Ollama is already installed, just verify the version:
ollama --version
Expected output should show a current Ollama release.
Download the Qwen3.6 model weights to your hardware platform.
This playbook uses qwen3.6:27b with Claude Code through Ollama:
ollama pull qwen3.6:27b
Expected output should show progress lines followed by success. Confirm the model appears in the list:
ollama list
NAME ID SIZE MODIFIED
qwen3.6:27b abc123... ... 1 minute ago
Run a quick prompt to confirm the model loads.
ollama run qwen3.6:27b
Try a prompt like:
Write a short README checklist for a Python project.
Expected output should show the model responding in the terminal. When you are done, type /bye or press Ctrl+D to exit the interactive session before continuing.
Install Claude Code, the CLI tool that will drive the local model.
curl -fsSL https://claude.ai/install.sh | bash
Verify the installation:
claude --version
Expected output should show a version string such as claude 0.x.x. If you see claude: command not found, ensure the install script added the CLI to your PATH (for example, restart the terminal or source your shell profile); see Troubleshooting.
Ollama defaults to a 4096 token context length. For coding agents and larger codebases, set it to 64K tokens. This increases memory usage. For more details on configuring context length and other parameters, see the Ollama documentation.
Set the context length per session in the Ollama REPL:
ollama run qwen3.6:27b
Then, in the Ollama prompt:
/set parameter num_ctx 64000
Exit when done: type /bye or press Ctrl+D.
Optional method (set globally when serving Ollama):
sudo systemctl stop ollama
OLLAMA_CONTEXT_LENGTH=64000 ollama serve
Keep this terminal open and run the next step in a new terminal.
Launch Claude Code through Ollama with the model you pulled. No environment variables or config files are required.
ollama launch claude --model qwen3.6:27b
Expected output should show Claude Code starting and using the local Ollama model.
Exit Claude Code when done: type /exit or press Ctrl+C.
Create a tiny repo and let Claude Code implement a function and tests.
mkdir -p ~/cli-agent-demo
cd ~/cli-agent-demo
python3 -m venv .venv
source .venv/bin/activate
python3 -m pip install -U pytest
printf 'def add(a, b):\n """Return the sum of a and b."""\n pass\n' > math_utils.py
printf 'import math_utils\n\n\ndef test_add():\n assert math_utils.add(1, 2) == 3\n' > test_math_utils.py
If Claude Code is not already running, launch it:
ollama launch claude --model qwen3.6:27b
In Claude Code, enter:
Please implement add() in math_utils.py and make sure the test passes.
Exit Claude Code when finished: type /exit or press Ctrl+C, then run the test:
python3 -m pytest -q
deactivate
Expected output should show the test passing.
Remove the model and stop the Ollama service if you no longer need them. Cleanup is optional. Remove the model first (while the Ollama server is running), then stop the service.
WARNING
The following removes the downloaded model files from disk.
1. Remove the model (Ollama must be running). Use the same name you pulled:
ollama rm qwen3.6:27b
2. Stop the Ollama service:
sudo systemctl stop ollama