AI Coding Harness Architecture in 2026: Claude Code vs Codex CLI vs OpenCode and 5 More — Why the Loop Now Beats the Model
TL;DR: Frontier coding models have converged enough that the harness — the loop, sandbox, memory, and context rules wrapped around the model — now decides most of your day-to-day experience. Claude Code and Codex CLI ship the two strongest OS-level sandboxes but lock you to a vendor’s models. OpenCode (MIT) is the most architecturally open: client/server split, 75+ providers, and local models over any OpenAI-compatible endpoint.
| Claude Code | Codex CLI | OpenCode | |
|---|---|---|---|
| Best for | Anthropic-model users who want domain-gated network sandboxing | OpenAI-model users who want network off by default | Self-hosted stacks, remote/headless servers, model freedom |
| Sandbox | Seatbelt (macOS), bubblewrap + socat (Linux/WSL2) | Seatbelt sandbox-exec (macOS 12+), Landlock/seccomp (Linux) | Permission prompts per agent (plan denies edits), no kernel sandbox |
| The catch | No native Windows sandbox — WSL2 required | Network in workspace-write is opt-in via config | Isolation is policy, not kernel enforcement — bring your own container |
Honest take: If you run Anthropic or OpenAI models anyway, take the first-party harness — the sandbox engineering alone is worth it. For everything else, including local models on your own GPU, OpenCode’s architecture is the one the others are slowly converging toward.
The claim that “the harness matters more than the model” stopped being a hot take this year. Winder.AI’s comparison of AI agent harnesses (August 2026) put it plainly after operating Claude Code, Codex, OpenCode, Qwen Code, Gemini CLI, Goose, and Zed Agent (among others) side by side on the same infrastructure: the differentiating work has moved into the execution loop, the tool sandbox, and context management. (Disclosure worth knowing: Winder.AI sells Helix, an agent control room that runs these harnesses, so they profit from the “harness matters” thesis — but their per-tool observations match what the official docs say, and everything below is verified against first-party documentation as of September 22, 2026.)
This comparison covers eight harnesses on four architecture dimensions: sandboxing, instruction files, memory, and local model support.
What is an AI coding harness?
A coding harness is everything around the model: the loop that calls it, the tools it can execute, the sandbox those tools run in, the instructions and memory that survive between turns, and the rules deciding what reaches the context window. Two harnesses pointed at the same model produce very different agents — one asks permission for every shell command, another runs a kernel-enforced sandbox and never asks; one forgets everything on restart, another persists memory to disk.
That’s why the September 2026 question isn’t “which model is smartest” — for coding, the top models trade blows benchmark to benchmark — but “which loop do you want wrapped around it.” The four dimensions below are where the eight harnesses actually differ.
How do Claude Code and Codex CLI sandbox commands differently?
Both use kernel-level OS sandboxing, but they make opposite default choices about the network.
Codex CLI (openai/codex docs, Apache 2.0) sandboxes with Apple Seatbelt on macOS 12+ (sandbox-exec with a profile matching the --sandbox flag) and Landlock/seccomp on Linux. It has three modes — read-only, workspace-write (edits and commands allowed in the workspace and /tmp), and danger-full-access — and in workspace-write, network is disabled by default unless you opt in:
# ~/.codex/config.toml
[sandbox_workspace_write]
network_access = true
Approval policy is separate from the sandbox: untrusted, on-failure, on-request, or never. The default “Auto” behavior pairs workspace-write with on-failure, so the agent runs freely inside the box and only escalates to you when a sandboxed command fails.
Claude Code (sandboxing docs) also uses Seatbelt on macOS; on Linux and WSL2 it uses bubblewrap for filesystem isolation plus socat to route traffic through a sandbox proxy, with an optional seccomp filter (npm install -g @anthropic-ai/sandbox-runtime) that adds Unix-socket blocking. Instead of network on/off, it does domain-gated networking: the first time a command needs a new domain, you get an approval prompt, and the allowlist accumulates. Commands that can’t run sandboxed fall back to a regular permission prompt titled “Bash command (unsandboxed)” — and you can kill that escape hatch entirely:
claude --settings '{"sandbox": {"enabled": true, "allowUnsandboxedCommands": false}}'
One real-world trap from the official troubleshooting section: on Ubuntu 24.04 and later, the default AppArmor policy blocks bubblewrap from creating the user namespaces it needs, so sandboxed commands fail with Operation not permitted until you loosen that policy. And there is no native Windows sandbox at all — Anthropic’s docs say to run Claude Code inside WSL2.
The philosophical split, which Winder.AI’s testing also surfaced: Codex leans on the sandbox to contain what generated code can touch, while Claude Code leans on interactive permission prompts plus the sandbox proxy for reach. Codex’s “network off unless configured” is the safer default for untrusted repos; Claude Code’s per-domain allowlist is less friction once you’ve approved your registry and package hosts.
Nobody else in this group ships comparable kernel-level isolation. Qwen Code lists “Sandbox” and Git worktree isolation among its features, Gemini CLI documents sandboxing and per-folder trust levels, and OpenCode, Cline, Goose, and Zed rely on permission prompts and review flows — policy enforcement, not kernel enforcement. If you run one of those against untrusted code, put the whole session in a container.
Which coding agents read AGENTS.md in September 2026?
The instruction-file landscape consolidated hard around AGENTS.md this year, but each harness treats it differently:
| Harness | Native file | Reads AGENTS.md? | Notes (verified Sep 22, 2026) |
|---|---|---|---|
| Claude Code | CLAUDE.md | Fallback since v2.1.277 (Sep 18) | Only when no CLAUDE.md/CLAUDE.local.md exists — details |
| Codex CLI | AGENTS.md | Yes, native | AGENTS.md is its primary instructions file |
| OpenCode | AGENTS.md | Yes, native | Falls back to CLAUDE.md, then ~/.claude/CLAUDE.md; instructions config accepts globs like .cursor/rules/*.md and remote URLs |
| Qwen Code | Auto-Memory | — | README documents “Auto-Memory, Auto-Skills” rather than a single instructions file |
| Gemini CLI | GEMINI.md | No | Custom context files are GEMINI.md only |
| Goose | .goosehints | No | Every line of .goosehints is sent with every request |
| Cline | .clinerules | No | Same .clinerules picked up by CLI, VS Code extension, and JetBrains plugin |
| Zed Agent | .rules | Yes | Reads, in order: .rules, .cursorrules, .windsurfrules, .clinerules, .github/copilot-instructions.md, AGENT.md, AGENTS.md, CLAUDE.md, GEMINI.md |
Zed is the outlier worth noticing: it reads everyone’s rules files, which makes it the cheapest harness to trial on a repo already configured for another agent. OpenCode’s remote-URL support in instructions is the team-scale feature — one hosted rules file, every developer’s agent pulls it.
If you run three or more of these tools on one repo, write AGENTS.md as the shared base. Codex, OpenCode, and Zed read it natively; Claude Code reads it when you don’t have a CLAUDE.md; only Gemini CLI, Goose, and Cline still need their own files.
Which harnesses run local models?
This is where the open-source group runs away from the vendor harnesses. Verified support as of September 22, 2026:
| Harness | Local model path | Providers |
|---|---|---|
| OpenCode | Ollama, LM Studio, llama.cpp — any OpenAI-compatible baseURL | 75+ via Models.dev / AI SDK |
| Cline | ”Ollama / LM Studio” plus any OpenAI-compatible API | BYOK everything |
| Goose | Ollama listed among 15+ providers | 70+ MCP extensions |
| Qwen Code | ”any third-party provider or local model (Ollama / vLLM)”; speaks OpenAI, Anthropic, Gemini, and Qwen protocols | Multi-protocol |
| Zed Agent | Local models via LLM provider settings | Plus external agents via ACP |
| Claude Code | Not official — works via proxy/router setups (our guide) | Anthropic-first |
| Codex CLI | Not documented in current official docs | OpenAI-first |
| Gemini CLI | No local model support in the official README | Gemini only; free tier: 60 req/min, 1,000 req/day with OAuth |
Two practical notes. First, “supports Ollama” does not mean “runs well on your laptop” — an agentic loop burns through context fast, and a 4096-token default Ollama context will silently cripple tool calling long before the model itself fails. Our setup guides for Qwen Code + Ollama, Aider + Ollama, and Goose + Ollama cover the per-tool traps. Second, hardware: a used RTX 3090’s 24 GB remains the entry point for coding-capable local models — see runaihome.com’s local model VRAM guide for what fits where, and if you’d rather test the workload before buying a GPU, renting an RTX 3090 starts around $0.07/hr on Vast.ai (marketplace pricing, floats).
Qwen Code deserves a specific flag here: v0.22.0’s README reports a 77.33% average on SWE-bench tasks. That’s vendor-reported and tied to Qwen’s own models — treat it as a claim about the harness+model pair, not an independent benchmark.
What memory actually survives a restart?
Four distinct architectures hide behind the word “memory”:
Static preprompts. Goose’s .goosehints is the purest form — every line is sent with every request, even “what time is it?”. Cline’s .clinerules, Gemini’s GEMINI.md, and Claude Code’s CLAUDE.md work the same way. Cheap, predictable, and it taxes every single call; we covered how this pattern degrades local backends in our large-CLAUDE.md gotchas piece.
Retrieval memory. Goose’s Memory Extension (an MCP server storing to ~/.goose/memory) is the counter-design: context stored with tags and retrieved on demand instead of shipped with every request. As of September 2026, Goose is the only harness in this group shipping that as a first-party extension.
Checkpoints. Cline tracks every change with checkpoints you can roll back; Zed puts a “Restore Checkpoint” button on every message that performed an edit and lets you accept or reject each individual hunk. This is memory of what the agent did, and it’s the difference between “undo” and “git archaeology” after a bad run.
Session state elsewhere. OpenCode’s client/server split means the session lives in the server, not the terminal window: opencode serve runs a headless HTTP server (default port 4096) exposing an OpenAPI 3.1 endpoint, and any TUI can attach with --hostname/--port. Your laptop dying doesn’t kill the agent’s run on the box that matters. None of the other seven has an equivalent documented remote-attach story.
Where do Cursor and the IDE-bound agents fit?
Cursor’s agent is real, but its harness isn’t separable from its IDE and cloud — you can’t run it headless on a server, point it at your own endpoint fleet, or audit its loop, which is why it doesn’t fit an architecture comparison of harnesses you can host and inspect (our Cursor 3 review covers it as a product). The interesting 2026 development is that the editor question and the harness question have decoupled: Zed’s Agent Panel runs external agents over the Agent Client Protocol, and its ACP registry currently lists Claude, Codex, Gemini CLI, OpenCode, Copilot, Cursor, Pi, and Poolside — each owning its own auth and billing. You can pick your harness for its architecture and still get editor-native diffs and per-hunk review.
Goose made the other notable structural move: Block contributed it as a founding project of the Linux Foundation’s Agentic AI Foundation in December 2025, with the repo transferred in April 2026 — alongside MCP and the AGENTS.md spec. Vendor-neutral governance is an architecture feature too, if you’re a team betting years on a tool.
Which harness should you pick?
| Your situation | Pick | Why | Cost to run |
|---|---|---|---|
| You pay for Claude anyway | Claude Code | Best domain-gated sandbox; CLAUDE.md + AGENTS.md fallback | Subscription/API |
| You pay for OpenAI anyway | Codex CLI | Network-off-by-default sandbox; native AGENTS.md | Subscription/API |
| Local models, self-hosted, headless server | OpenCode | Client/server, 75+ providers, MIT, 209k GitHub stars | $0 + your GPU |
| VS Code / JetBrains, per-edit control | Cline | Approval per edit/command, checkpoints, Apache 2.0 | $0 + BYOK |
| Zero budget, cloud model | Gemini CLI | 1,000 free requests/day on OAuth | $0 |
| Automation beyond code (MCP-heavy) | Goose | 70+ MCP extensions, desktop + CLI + API, AAIF governance | $0 + BYOK |
| Qwen ecosystem, multi-protocol | Qwen Code | Ollama/vLLM native, git-worktree isolation | $0 + BYOK |
| You want one editor, any harness | Zed Agent | ACP runs Claude/Codex/Gemini/OpenCode inside Zed; reads every rules file | $0 + agent’s own billing |
Honest take: The winner by architecture is OpenCode — it’s the only harness here with a documented client/server split, the widest provider surface, and instruction-file interop with everyone else’s config. The winner by engineering depth is Claude Code’s sandbox, and Codex’s network-off default is the one I’d hand an intern. If you’re on neither vendor’s models, don’t overthink it: OpenCode for the terminal, Cline for the IDE.
FAQ
Does a stronger harness really beat a stronger model? For agentic coding, often yes — a model that can’t run tests safely (no sandbox), forgets project rules (no instruction file), or loses its session when your SSH drops (no server mode) wastes more time than a few benchmark points recover. The gap between top coding models is now smaller than the gap between the best and worst harness on this page.
Is Gemini CLI deprecated? No. Despite recurring community chatter, the official repository is actively maintained as of September 22, 2026, with preview/stable/nightly release channels and a public roadmap — and its free tier (60 req/min, 1,000 req/day via OAuth) is still the largest zero-cost allowance in this comparison.
Can I use one AGENTS.md for all eight tools? Not yet. Codex, OpenCode, and Zed read it natively; Claude Code v2.1.277+ reads it only when no CLAUDE.md exists anywhere above your working directory. Gemini CLI, Goose, and Cline still require GEMINI.md, .goosehints, and .clinerules respectively. Symlinks close most of the gap.
Which harness is safest for untrusted repositories? Codex CLI in default workspace-write — filesystem writes confined to the workspace and /tmp, network fully off unless you enable it in config. Claude Code with allowUnsandboxedCommands: false is the close second, with the advantage of granular domain allowlists when you do need network.
Sources
- A Comparison of AI Agent Harnesses in 2026 — Winder.AI
- Claude Code sandboxing documentation — Anthropic
- Codex CLI repository and sandbox docs — OpenAI
- OpenCode repository (MIT) — sst and OpenCode docs: rules, providers, server
- Qwen Code v0.22.0 README — QwenLM
- Gemini CLI README — Google
- Goose README — Block / AAIF and goose (AI agent) — Wikipedia
- Cline README — Cline Bot Inc
- Zed Agent Panel, rules, and external agents docs — Zed Industries
Last updated September 22, 2026. Sandbox behavior, instruction-file support, and provider lists change fast in this category; verify against the linked official docs before making a team-wide decision.
Was this article helpful?
Thanks for the feedback — it helps improve future articles.
Need hands-on help?
I offer 1-on-1 technical consulting for local AI setup, GPU selection, and AI coding tool configuration — same topics covered on this site.
Book a session — $49 / hour →Know which coding tool is worth paying for
Hands-on comparisons of AI coding assistants and what each one costs to run — including the local-model path. Sent only when something changes. Unsubscribe anytime.