AI Coding Harness Architecture in 2026: Claude Code vs Codex CLI vs OpenCode and 5 More — Why the Loop Now Beats the Model

claude-codecodexopencodeqwen-codegemini-cligooseclinezedcomparisonlocal-llmarchitecture

TL;DR: Frontier coding models have converged enough that the harness — the loop, sandbox, memory, and context rules wrapped around the model — now decides most of your day-to-day experience. Claude Code and Codex CLI ship the two strongest OS-level sandboxes but lock you to a vendor’s models. OpenCode (MIT) is the most architecturally open: client/server split, 75+ providers, and local models over any OpenAI-compatible endpoint.

Claude CodeCodex CLIOpenCode
Best forAnthropic-model users who want domain-gated network sandboxingOpenAI-model users who want network off by defaultSelf-hosted stacks, remote/headless servers, model freedom
SandboxSeatbelt (macOS), bubblewrap + socat (Linux/WSL2)Seatbelt sandbox-exec (macOS 12+), Landlock/seccomp (Linux)Permission prompts per agent (plan denies edits), no kernel sandbox
The catchNo native Windows sandbox — WSL2 requiredNetwork in workspace-write is opt-in via configIsolation is policy, not kernel enforcement — bring your own container

Honest take: If you run Anthropic or OpenAI models anyway, take the first-party harness — the sandbox engineering alone is worth it. For everything else, including local models on your own GPU, OpenCode’s architecture is the one the others are slowly converging toward.

The claim that “the harness matters more than the model” stopped being a hot take this year. Winder.AI’s comparison of AI agent harnesses (August 2026) put it plainly after operating Claude Code, Codex, OpenCode, Qwen Code, Gemini CLI, Goose, and Zed Agent (among others) side by side on the same infrastructure: the differentiating work has moved into the execution loop, the tool sandbox, and context management. (Disclosure worth knowing: Winder.AI sells Helix, an agent control room that runs these harnesses, so they profit from the “harness matters” thesis — but their per-tool observations match what the official docs say, and everything below is verified against first-party documentation as of September 22, 2026.)

This comparison covers eight harnesses on four architecture dimensions: sandboxing, instruction files, memory, and local model support.

What is an AI coding harness?

A coding harness is everything around the model: the loop that calls it, the tools it can execute, the sandbox those tools run in, the instructions and memory that survive between turns, and the rules deciding what reaches the context window. Two harnesses pointed at the same model produce very different agents — one asks permission for every shell command, another runs a kernel-enforced sandbox and never asks; one forgets everything on restart, another persists memory to disk.

That’s why the September 2026 question isn’t “which model is smartest” — for coding, the top models trade blows benchmark to benchmark — but “which loop do you want wrapped around it.” The four dimensions below are where the eight harnesses actually differ.

How do Claude Code and Codex CLI sandbox commands differently?

Both use kernel-level OS sandboxing, but they make opposite default choices about the network.

Codex CLI (openai/codex docs, Apache 2.0) sandboxes with Apple Seatbelt on macOS 12+ (sandbox-exec with a profile matching the --sandbox flag) and Landlock/seccomp on Linux. It has three modes — read-only, workspace-write (edits and commands allowed in the workspace and /tmp), and danger-full-access — and in workspace-write, network is disabled by default unless you opt in:

# ~/.codex/config.toml
[sandbox_workspace_write]
network_access = true

Approval policy is separate from the sandbox: untrusted, on-failure, on-request, or never. The default “Auto” behavior pairs workspace-write with on-failure, so the agent runs freely inside the box and only escalates to you when a sandboxed command fails.

Claude Code (sandboxing docs) also uses Seatbelt on macOS; on Linux and WSL2 it uses bubblewrap for filesystem isolation plus socat to route traffic through a sandbox proxy, with an optional seccomp filter (npm install -g @anthropic-ai/sandbox-runtime) that adds Unix-socket blocking. Instead of network on/off, it does domain-gated networking: the first time a command needs a new domain, you get an approval prompt, and the allowlist accumulates. Commands that can’t run sandboxed fall back to a regular permission prompt titled “Bash command (unsandboxed)” — and you can kill that escape hatch entirely:

claude --settings '{"sandbox": {"enabled": true, "allowUnsandboxedCommands": false}}'

One real-world trap from the official troubleshooting section: on Ubuntu 24.04 and later, the default AppArmor policy blocks bubblewrap from creating the user namespaces it needs, so sandboxed commands fail with Operation not permitted until you loosen that policy. And there is no native Windows sandbox at all — Anthropic’s docs say to run Claude Code inside WSL2.

The philosophical split, which Winder.AI’s testing also surfaced: Codex leans on the sandbox to contain what generated code can touch, while Claude Code leans on interactive permission prompts plus the sandbox proxy for reach. Codex’s “network off unless configured” is the safer default for untrusted repos; Claude Code’s per-domain allowlist is less friction once you’ve approved your registry and package hosts.

Nobody else in this group ships comparable kernel-level isolation. Qwen Code lists “Sandbox” and Git worktree isolation among its features, Gemini CLI documents sandboxing and per-folder trust levels, and OpenCode, Cline, Goose, and Zed rely on permission prompts and review flows — policy enforcement, not kernel enforcement. If you run one of those against untrusted code, put the whole session in a container.

Which coding agents read AGENTS.md in September 2026?

The instruction-file landscape consolidated hard around AGENTS.md this year, but each harness treats it differently:

HarnessNative fileReads AGENTS.md?Notes (verified Sep 22, 2026)
Claude CodeCLAUDE.mdFallback since v2.1.277 (Sep 18)Only when no CLAUDE.md/CLAUDE.local.md exists — details
Codex CLIAGENTS.mdYes, nativeAGENTS.md is its primary instructions file
OpenCodeAGENTS.mdYes, nativeFalls back to CLAUDE.md, then ~/.claude/CLAUDE.md; instructions config accepts globs like .cursor/rules/*.md and remote URLs
Qwen CodeAuto-Memory—README documents “Auto-Memory, Auto-Skills” rather than a single instructions file
Gemini CLIGEMINI.mdNoCustom context files are GEMINI.md only
Goose.goosehintsNoEvery line of .goosehints is sent with every request
Cline.clinerulesNoSame .clinerules picked up by CLI, VS Code extension, and JetBrains plugin
Zed Agent.rulesYesReads, in order: .rules, .cursorrules, .windsurfrules, .clinerules, .github/copilot-instructions.md, AGENT.md, AGENTS.md, CLAUDE.md, GEMINI.md

Zed is the outlier worth noticing: it reads everyone’s rules files, which makes it the cheapest harness to trial on a repo already configured for another agent. OpenCode’s remote-URL support in instructions is the team-scale feature — one hosted rules file, every developer’s agent pulls it.

If you run three or more of these tools on one repo, write AGENTS.md as the shared base. Codex, OpenCode, and Zed read it natively; Claude Code reads it when you don’t have a CLAUDE.md; only Gemini CLI, Goose, and Cline still need their own files.

Which harnesses run local models?

This is where the open-source group runs away from the vendor harnesses. Verified support as of September 22, 2026:

HarnessLocal model pathProviders
OpenCodeOllama, LM Studio, llama.cpp — any OpenAI-compatible baseURL75+ via Models.dev / AI SDK
Cline”Ollama / LM Studio” plus any OpenAI-compatible APIBYOK everything
GooseOllama listed among 15+ providers70+ MCP extensions
Qwen Code”any third-party provider or local model (Ollama / vLLM)”; speaks OpenAI, Anthropic, Gemini, and Qwen protocolsMulti-protocol
Zed AgentLocal models via LLM provider settingsPlus external agents via ACP
Claude CodeNot official — works via proxy/router setups (our guide)Anthropic-first
Codex CLINot documented in current official docsOpenAI-first
Gemini CLINo local model support in the official READMEGemini only; free tier: 60 req/min, 1,000 req/day with OAuth

Two practical notes. First, “supports Ollama” does not mean “runs well on your laptop” — an agentic loop burns through context fast, and a 4096-token default Ollama context will silently cripple tool calling long before the model itself fails. Our setup guides for Qwen Code + Ollama, Aider + Ollama, and Goose + Ollama cover the per-tool traps. Second, hardware: a used RTX 3090’s 24 GB remains the entry point for coding-capable local models — see runaihome.com’s local model VRAM guide for what fits where, and if you’d rather test the workload before buying a GPU, renting an RTX 3090 starts around $0.07/hr on Vast.ai (marketplace pricing, floats).

Qwen Code deserves a specific flag here: v0.22.0’s README reports a 77.33% average on SWE-bench tasks. That’s vendor-reported and tied to Qwen’s own models — treat it as a claim about the harness+model pair, not an independent benchmark.

What memory actually survives a restart?

Four distinct architectures hide behind the word “memory”:

Static preprompts. Goose’s .goosehints is the purest form — every line is sent with every request, even “what time is it?”. Cline’s .clinerules, Gemini’s GEMINI.md, and Claude Code’s CLAUDE.md work the same way. Cheap, predictable, and it taxes every single call; we covered how this pattern degrades local backends in our large-CLAUDE.md gotchas piece.

Retrieval memory. Goose’s Memory Extension (an MCP server storing to ~/.goose/memory) is the counter-design: context stored with tags and retrieved on demand instead of shipped with every request. As of September 2026, Goose is the only harness in this group shipping that as a first-party extension.

Checkpoints. Cline tracks every change with checkpoints you can roll back; Zed puts a “Restore Checkpoint” button on every message that performed an edit and lets you accept or reject each individual hunk. This is memory of what the agent did, and it’s the difference between “undo” and “git archaeology” after a bad run.

Session state elsewhere. OpenCode’s client/server split means the session lives in the server, not the terminal window: opencode serve runs a headless HTTP server (default port 4096) exposing an OpenAPI 3.1 endpoint, and any TUI can attach with --hostname/--port. Your laptop dying doesn’t kill the agent’s run on the box that matters. None of the other seven has an equivalent documented remote-attach story.

Where do Cursor and the IDE-bound agents fit?

Cursor’s agent is real, but its harness isn’t separable from its IDE and cloud — you can’t run it headless on a server, point it at your own endpoint fleet, or audit its loop, which is why it doesn’t fit an architecture comparison of harnesses you can host and inspect (our Cursor 3 review covers it as a product). The interesting 2026 development is that the editor question and the harness question have decoupled: Zed’s Agent Panel runs external agents over the Agent Client Protocol, and its ACP registry currently lists Claude, Codex, Gemini CLI, OpenCode, Copilot, Cursor, Pi, and Poolside — each owning its own auth and billing. You can pick your harness for its architecture and still get editor-native diffs and per-hunk review.

Goose made the other notable structural move: Block contributed it as a founding project of the Linux Foundation’s Agentic AI Foundation in December 2025, with the repo transferred in April 2026 — alongside MCP and the AGENTS.md spec. Vendor-neutral governance is an architecture feature too, if you’re a team betting years on a tool.

Which harness should you pick?

Your situationPickWhyCost to run
You pay for Claude anywayClaude CodeBest domain-gated sandbox; CLAUDE.md + AGENTS.md fallbackSubscription/API
You pay for OpenAI anywayCodex CLINetwork-off-by-default sandbox; native AGENTS.mdSubscription/API
Local models, self-hosted, headless serverOpenCodeClient/server, 75+ providers, MIT, 209k GitHub stars$0 + your GPU
VS Code / JetBrains, per-edit controlClineApproval per edit/command, checkpoints, Apache 2.0$0 + BYOK
Zero budget, cloud modelGemini CLI1,000 free requests/day on OAuth$0
Automation beyond code (MCP-heavy)Goose70+ MCP extensions, desktop + CLI + API, AAIF governance$0 + BYOK
Qwen ecosystem, multi-protocolQwen CodeOllama/vLLM native, git-worktree isolation$0 + BYOK
You want one editor, any harnessZed AgentACP runs Claude/Codex/Gemini/OpenCode inside Zed; reads every rules file$0 + agent’s own billing

Honest take: The winner by architecture is OpenCode — it’s the only harness here with a documented client/server split, the widest provider surface, and instruction-file interop with everyone else’s config. The winner by engineering depth is Claude Code’s sandbox, and Codex’s network-off default is the one I’d hand an intern. If you’re on neither vendor’s models, don’t overthink it: OpenCode for the terminal, Cline for the IDE.

FAQ

Does a stronger harness really beat a stronger model? For agentic coding, often yes — a model that can’t run tests safely (no sandbox), forgets project rules (no instruction file), or loses its session when your SSH drops (no server mode) wastes more time than a few benchmark points recover. The gap between top coding models is now smaller than the gap between the best and worst harness on this page.

Is Gemini CLI deprecated? No. Despite recurring community chatter, the official repository is actively maintained as of September 22, 2026, with preview/stable/nightly release channels and a public roadmap — and its free tier (60 req/min, 1,000 req/day via OAuth) is still the largest zero-cost allowance in this comparison.

Can I use one AGENTS.md for all eight tools? Not yet. Codex, OpenCode, and Zed read it natively; Claude Code v2.1.277+ reads it only when no CLAUDE.md exists anywhere above your working directory. Gemini CLI, Goose, and Cline still require GEMINI.md, .goosehints, and .clinerules respectively. Symlinks close most of the gap.

Which harness is safest for untrusted repositories? Codex CLI in default workspace-write — filesystem writes confined to the workspace and /tmp, network fully off unless you enable it in config. Claude Code with allowUnsandboxedCommands: false is the close second, with the advantage of granular domain allowlists when you do need network.

Sources

Last updated September 22, 2026. Sandbox behavior, instruction-file support, and provider lists change fast in this category; verify against the linked official docs before making a team-wide decision.

Was this article helpful?

Know which coding tool is worth paying for

Hands-on comparisons of AI coding assistants and what each one costs to run — including the local-model path. Sent only when something changes. Unsubscribe anytime.