---
title: "Set Up CLI Coding Agents with Local Inference"
publisher: "nvidia"
type: "playbook"
updated: "2026-09-22T17:55:05.308Z"
description: "Your chosen agent connected to a local model in one command"
canonical: "https://build.nvidia.com/playbooks/cli-coding-agent.md"
---

# Basic idea

Use [Ollama](https://ollama.com) on your **hardware platform** to run a local coding model and connect a CLI coding agent. This playbook supports three options: **[Claude Code](https://docs.claude.com/en/docs/claude-code)**, **[OpenCode](https://opencode.ai)**, and **[Codex CLI](https://github.com/openai/codex)**. Each agent is wired up with Ollama's built-in [launch method](https://ollama.com/blog/launch) (`ollama launch <agent>`), so you can work without environment variables, provider config files, or external cloud APIs.

# Choose your CLI agent

Pick the tab that matches the CLI agent you want to use:

- **Claude Code**: Fastest path to a working CLI agent with a local Ollama model.
- **OpenCode**: Open-source CLI installed locally, then configured and launched through Ollama.
- **Codex CLI**: OpenAI Codex CLI launched directly from Ollama against the local model.

# What you'll accomplish

You'll run a local coding model ([Qwen3.6](https://ollama.com/library/qwen3.6)) on your **hardware platform** with Ollama, launch your chosen CLI agent against it with a single command, and complete a small coding task end-to-end.

# What to know before starting

**Required:**

- Comfort with Linux command line basics
- Experience running terminal-based tools and editors
- Familiarity with Python for the short coding task

**Optional:**

- Experience browsing models on the [Ollama library](https://ollama.com/library) (for example, [Qwen3.6](https://ollama.com/library/qwen3.6))

# Supported hardware platforms

Use the matrix below to confirm your hardware platform, recommended default local settings, and whether multi-node applies.

| Hardware platform | OS | Memory | Recommended default local settings | Multi-node capable hardware |
| :---- | :---- | :---- | :---- | :---- |
| **DGX Spark** | DGX OS (Linux) | 128 GB Unified Memory | Ollama + `qwen3.6:35b-a3b-mtp-q4_K_M` (~23 GB); Claude Code / OpenCode / Codex via `ollama launch` | — |

# Prerequisites

**Hardware requirements**

- Supported hardware platform — see Supported hardware platforms matrix above
- Hardware platform powered on, networked, and reachable for local or SSH terminal access
- Sufficient memory for your chosen Qwen3.6 variant (about 23 GB for the default MTP Q4_K_M; about 39 GB for `q8_0`; about 71 GB for `bf16`)
- Enough free storage for model downloads

**Software requirements**

- A [current Ollama release](https://ollama.com/download) (required for [`ollama launch`](https://ollama.com/blog/launch) and the default MTP Q4_K_M model): `ollama --version`
- Internet access to download model weights
- For Codex CLI: Node.js / npm available to install `@openai/codex`
- Python 3 with `venv` support for the optional coding-task verification

# Find model recipes

Browse models you can pull and run with Ollama in the [Ollama library](https://ollama.com/library). Use the tags and sizes that fit your hardware platform’s memory and storage.

| Hardware platform | More recipes |
| ----------------- | ------------ |
| **DGX Spark** | [Ollama library](https://ollama.com/library) · [Qwen3.6](https://ollama.com/library/qwen3.6) |

Use the **Claude Code**, **OpenCode**, or **Codex CLI** tab for the base workflow.

# Time & risk

- **Estimated time:** 20 MIN (mostly model download time)
- **Risk level:** Low
- Large model downloads can fail if network connectivity is unstable
- Ollama must support [`ollama launch`](https://ollama.com/blog/launch) and the default MTP Q4_K_M model tag — verify with `ollama --version` and install from [ollama.com/download](https://ollama.com/download) if needed
- **Rollback:** Stop Ollama and delete the downloaded model from `~/.ollama/models` (see Cleanup in each agent tab). Cleanup is optional and removes downloaded model files.
- **Last Updated:** 08/03/2026
- Set up Claude Code, OpenCode, or Codex CLI against a local Qwen3.6 model with `ollama launch` on supported hardware platforms

## More

- [Claude Code](/playbooks/cli-coding-agent/claude-code.md)
- [OpenCode](/playbooks/cli-coding-agent/opencode.md)
- [Codex CLI](/playbooks/cli-coding-agent/codex.md)
- [Troubleshooting](/playbooks/cli-coding-agent/troubleshooting.md)