---
title: "Set Up Claude Code with Local Inference — Troubleshooting"
canonical: "https://build.nvidia.com/playbooks/local-coding-agent/troubleshooting.md"
---

| Symptom | Cause | Fix |
|---------|-------|-----|
| `ollama: command not found` | Ollama not installed or PATH not updated | Rerun `curl -fsSL https://ollama.com/install.sh \| sh` and open a new shell |
| `ollama launch` reports unknown command | `ollama launch` not available in this install | Install or reinstall Ollama from [ollama.com/download](https://ollama.com/download): `curl -fsSL https://ollama.com/install.sh \| sh` |
| Model load fails with version error | Installed Ollama does not accept the recommended model tag | Install or reinstall Ollama from [ollama.com/download](https://ollama.com/download): `curl -fsSL https://ollama.com/install.sh \| sh` |
| `model not found` in Claude Code | Model was not pulled | Run `ollama pull qwen3.6:27b` and retry with `ollama launch claude --model qwen3.6:27b` |
| `connection refused` to localhost:11434 | Ollama service not running | Start with `ollama serve` or `sudo systemctl start ollama` |
| Sharded GGUF model pull fails with HTTP 400 | The selected model tag is not supported by Ollama | Use the documented `qwen3.6:27b` model instead: `ollama pull qwen3.6:27b` |
| `CUDA error: context is destroyed` on a multi-GPU system | Ollama is initializing across a mixed-GPU topology | Pin Ollama to one GPU. For a foreground test, run `CUDA_VISIBLE_DEVICES=0 ollama serve`; for a system service, add `Environment="CUDA_VISIBLE_DEVICES=0"` (or another GPU index) to an Ollama systemd drop-in and restart Ollama |
| Claude Code edit task fails through a direct Ollama endpoint | Direct endpoint wiring can fail with some Ollama/model combinations | Launch Claude Code through Ollama instead: `ollama launch claude --model qwen3.6:27b` |
| `externally-managed-environment` or Python package install fails | System Python blocks direct package installs | Create and activate a virtual environment, then install pytest inside it: `python3 -m venv .venv`, `source .venv/bin/activate`, `python3 -m pip install -U pytest` |
| Slow responses or OOM | Insufficient GPU memory or fragmentation | Ensure no other heavy GPU workloads. If OOM persists, unload other models or set `OLLAMA_MAX_LOADED_MODELS=1` |
| `claude: command not found` after install | CLI not on PATH or install script did not complete | Restart the terminal or run `source ~/.bashrc` (or your shell profile). Check the install script output for the install path and add it to PATH |
| Claude Code install fails (Node.js / network) | Node.js missing or install script cannot download | Ensure Node.js is installed (`node --version`). Run the installer with Bash: `curl -fsSL https://claude.ai/install.sh \| bash`. If the install script fails with a network error, retry from a stable connection. See [Claude Code documentation](https://docs.claude.com/en/docs/claude-code/overview) for alternatives |

> [!NOTE]
> The recommended `qwen3.6:27b` workflow fits the Supported hardware platforms matrix defaults. Use `OLLAMA_MAX_LOADED_MODELS=1` if you hit memory limits with multiple models.

For latest known issues, see the documentation linked under **Resources** for your hardware platform.