---
title: "Set Up CLI Coding Agents with Local Inference — Troubleshooting"
canonical: "https://build.nvidia.com/playbooks/cli-coding-agent/troubleshooting.md"
---

| Symptom | Cause | Fix |
|---------|-------|-----|
| `ollama: command not found` | Ollama not installed or PATH not updated | Rerun `curl -fsSL https://ollama.com/install.sh \| sh` and open a new shell |
| `ollama launch` reports unknown command | `ollama launch` not available in this install | Install or reinstall Ollama from [ollama.com/download](https://ollama.com/download): `curl -fsSL https://ollama.com/install.sh \| sh` |
| Model load fails with version error or HTTP 412 | Installed Ollama does not accept the default MTP Q4_K_M tag | Install or reinstall Ollama from [ollama.com/download](https://ollama.com/download): `curl -fsSL https://ollama.com/install.sh \| sh` |
| `model not found` when launching an agent | Model was not pulled | Run `ollama pull qwen3.6:35b-a3b-mtp-q4_K_M` and retry |
| `connection refused` to localhost:11434 | Ollama service not running | Start with `ollama serve` or `sudo systemctl start ollama` |
| `ollama launch <agent>` exits immediately | Agent integration failed to initialize | Re-run `ollama launch <agent>`; if it persists, check `journalctl -u ollama` |
| Slow responses or OOM errors | Model variant too large for available memory | Use the default `qwen3.6:35b-a3b-mtp-q4_K_M` variant and close other GPU workloads |
| `python3 -m pip install -U pytest` reports `externally-managed-environment` | System Python environment is protected | Create and activate a virtual environment first: `python3 -m venv .venv && source .venv/bin/activate` |
| `ollama pull` reports that a model tag is a sharded GGUF | The selected model tag is not supported by Ollama | Use the Qwen3.6 commands in Step 3 instead of sharded GGUF tags |
| `ollama run` fails with `CUDA error: context is destroyed` on a multi-GPU system | Ollama is initializing across a mixed-GPU topology | Pin Ollama to one GPU. For a foreground test, run `CUDA_VISIBLE_DEVICES=0 ollama serve`; for a system service, add `Environment="CUDA_VISIBLE_DEVICES=0"` to an Ollama systemd drop-in and restart Ollama |
| A direct Claude Code setup using an Anthropic-compatible Ollama endpoint produces prose but does not edit files | Some model/server combinations do not emit tool calls reliably | Use `ollama launch claude` with Qwen3.6 as shown in this playbook |
| OpenCode launch reports `opencode is not installed` or fails while fetching its version | OpenCode is not installed, or Ollama cannot complete its installer flow | Install OpenCode directly: `curl -fsSL https://opencode.ai/install \| bash`, then run `export PATH="$HOME/.opencode/bin:$PATH"` and retry `ollama launch opencode --model qwen3.6:35b-a3b-mtp-q4_K_M` |

> [!NOTE]
> Some hardware platforms use Unified Memory Architecture (UMA), which enables dynamic memory sharing between the GPU and CPU. With many applications still updating to take advantage of UMA, you may encounter memory pressure even when within rated capacity. If that happens, manually flush the buffer cache with:
> ```bash
> sudo sh -c 'sync; echo 3 > /proc/sys/vm/drop_caches'
> ```

For latest known issues, see the documentation linked under **Resources** for your hardware platform.