AI Coding Is 2x, Not 10x — and Sometimes Minus 19%: What Four Controlled Studies Say About Your $20/Month

productivitycost-analysiscursorcopilotclaude-coderesearchroi

TL;DR: Every controlled study of AI coding productivity lands between -19% and +56% — nowhere near the 10x in vendor marketing, and the widely shared “2x, not 10x” essay (Hacker News, 300+ points) is the honest ceiling for the coding portion of the job. Which end of that range you land on is predicted almost entirely by task type: greenfield boilerplate gains the most, expert work in a mature codebase can go negative. At $10–$20/month, the subscription still pays for itself with 12–40 minutes saved per month — the real risk isn’t the fee, it’s trusting your own perception of speedup, which the data shows is off by roughly 40 percentage points.

Peng et al. lab experimentMicrosoft/Accenture RCTsGoogle enterprise RCTMETR expert RCT
Result+55.8% faster+26% more tasks~21% faster19% slower
Who95 freelance devs4,867 working devs96 Google engineers16 expert OSS maintainers
TaskHTTP server from scratch (JavaScript)Everyday work over monthsComplex internal task (10 files, 474 lines)246 real issues in their own large repos
Tooling, eraGitHub Copilot (GPT-3.5), 2023GitHub Copilot (GPT-3.5), 2022–23Google-internal AI in VS Code, 2024Cursor Pro + Claude 3.5/3.7 Sonnet, early 2025

Honest take: Pay the $10–$20. The break-even math is trivial for anyone billing real money. But stop deciding which tool and which tier based on a felt sense of speed — the METR developers felt 20% faster while measuring 19% slower. Route your task mix instead: hand the agent greenfield, boilerplate, and test scaffolding, keep the mature-codebase surgery for yourself, and measure your own numbers before rolling anything out to a team.

What does the “2x, not 10x” claim actually say?

The essay behind the Hacker News discussion “2x, not 10x: The Real Impact of Coding with LLMs in 2026” (HN item 49047839) argues that a realistic productivity gain from LLM coding tools is about 2x on the coding work itself — not the 10x that vendor keynotes and social media threads promise. Three specifics from the essay are worth keeping:

  1. LLMs shine inside tight, automated feedback loops — code that can be run, tested, and lint-checked immediately. They are far weaker on maintainability, documentation, and the structural decisions that make code cheap to change later.
  2. A working implementation used to mean a task was 80% done. Now it means about 20% done. The model gets you to “it runs” fast; review, integration, hardening, and documentation — the part that was always the hidden majority — is still yours.
  3. Future gains will come from retooling workflows around what models can already do, not from waiting for a model release to fix the fundamentals.

Simon Willison, who has published some of the most-read practitioner notes on LLM coding, put the same idea in one sentence: LLMs make him roughly 2–5x more productive on the parts of the job that involve typing code into a computer — which is a small portion of what a software engineer actually does. That framing turns out to reconcile every study result below.

What do the controlled studies actually measure?

Four studies with real control groups exist as of September 2026, and they disagree with each other in a way that is informative rather than embarrassing — each one measured a different slice of software work.

Peng et al. (arXiv 2302.06590), the number vendors quote. 95 freelance developers were asked to implement an HTTP server in JavaScript as fast as possible. The Copilot group finished in an average of 71 minutes versus 161 minutes for the control group — 55.8% faster. This is the ceiling case: a self-contained greenfield task with zero legacy context, exactly what LLMs are best at. It says almost nothing about week-to-week work in an existing codebase.

The Microsoft/Accenture field RCTs, the best large-scale number. Researchers from Microsoft, MIT, Princeton, and Wharton ran three randomized trials at Microsoft, Accenture, and an unnamed electronics company — 4,867 developers doing their normal jobs. Copilot access raised completed tasks by 26.08%, commits by 13.55%, and builds by 38.38%. Two details matter: less-experienced developers gained noticeably more than senior ones, and the trials ran in 2022–23 on a GPT-3.5-based Copilot — several model generations ago.

The Google enterprise RCT (arXiv 2410.12944). 96 full-time Google engineers completed a realistic complex task — editing 10 files and 474 lines to build a logging service on internal infrastructure — with or without AI completion, Smart Paste, and natural-language-to-code. With AI: 96 minutes on average. Without: 114. About 21% faster, with a wide confidence interval the authors flag themselves. Engineers who spent more hours per day on code got more benefit.

The METR expert RCT, the uncomfortable one. In early 2025, METR recruited 16 experienced open-source maintainers and randomized 246 real issues from their own large repositories to AI-allowed or AI-forbidden conditions. The AI-allowed condition — mostly Cursor Pro with Claude 3.5/3.7 Sonnet — took 19% longer. Not less gain. Slower.

Plotted on one axis, the four results tell a single story: the newer the codebase and the less context required, the bigger the win. The more expertise and repo familiarity the human already has, the smaller the win — until it crosses zero.

Why did experienced developers get 19% slower in the METR trial?

Because the conditions inverted every advantage the tools have. The METR developers were working in large repositories they had contributed to for years. Their unassisted baseline was already fast: they knew where everything lived, what would break, and what the maintainer would reject. Against that baseline, the AI condition added prompt-writing time, waiting time, and — the biggest cost — review-and-rewrite time on generated code that didn’t meet the repo’s standards. The models also navigated the large codebases poorly, missing conventions a maintainer applies without thinking.

The perception gap is the most useful finding for anyone holding a company card. Before the study, the developers forecast AI would make them 24% faster. Afterward — having actually been measured at 19% slower — they still believed it had made them about 20% faster. That is a roughly 40-point gap between felt speed and measured speed, in the direction that sells subscriptions. AI-assisted coding feels fast because the keystrokes-to-running-code interval shrinks, while the added review and correction time doesn’t register as “the tool being slow.”

METR’s own caveats apply in both directions: 16 developers is small, the setup (deep expert, mature high-standards repo) is deliberately the hardest case for AI, and tooling has improved since early 2025. The study does not claim AI slows everyone down — it claims your intuition about the speedup is not evidence.

How can “2x” and “+26%” both be true?

Simple arithmetic, and it’s worth doing explicitly because it collapses most online arguments about this topic.

The RCTs measure whole-task or whole-job throughput. The 2x claims measure the coding slice — the time spent actually producing code. Suppose coding is 40% of a developer’s week (the rest: review, design, meetings, debugging in production, reading). Double the speed of that 40% slice and total throughput becomes:

1 / (0.60 + 0.40/2) = 1 / 0.80 = 1.25   → +25% overall

A genuine 2x on the coding slice yields +25% overall — almost exactly what the Microsoft (+26%) and Google (+21%) trials measured. The vendor lab number (+56%) is what you get when the task is 100% coding slice. And the ceiling is instructive: with coding at 40% of the job, even an infinitely fast code generator caps out at 1.67x overall. There is no 10x available at the whole-job level without also automating review, design, and coordination — which is precisely the pitch behind autonomous agents, and precisely where the “80% done is now 20% done” inversion bites hardest.

So “2x, not 10x” and the RCTs are the same finding at different zoom levels. Neither supports the 10x story.

Is $20/month still worth it if the real gain is 1.2x?

Yes, and it isn’t close — the subscription fee is the wrong thing to scrutinize. Prices verified September 30, 2026 against the official plan pages (Cursor’s site blocks our fetcher; its prices are cross-checked against its published plan lineup and were last verified directly September 21, 2026):

PlanPrice/monthWhat you get
GitHub Copilot Free$02,000 completions + limited chat/agent per month
GitHub Copilot Pro$10Unlimited completions + $15 in AI credits
Claude Pro (includes Claude Code)$20 ($17 annual)Claude Code CLI + chat
Cursor Pro$20$20 agent-credit pool + Tab completions
GitHub Copilot Pro+$39$70 in AI credits
Copilot Max / Claude Max / Cursor Ultra$100 / $100+ / $200Heavy agentic use

Break-even on a $20 plan, by what your working hour is worth: at $30/hour it needs to save 40 minutes a month; at $60/hour, 20 minutes; at $100/hour, 12 minutes. A 21–26% throughput gain on even a few hours of coding per week clears that by two orders of magnitude. This is why the ROI argument survives the death of the 10x story completely intact — at $150/hour, a measured 1.2x on ten coding hours a month returns about $300 against a $20 fee.

The decision the data actually complicates is the upgrade ladder and the team rollout:

  • The jump from $10–$20 to $100–$200 tiers buys usage volume, not a bigger multiplier. If your measured gain is 1.2x, five times more agent credits doesn’t make it 2x — it makes the same 1.2x available for more hours. That’s worth it only if you’re hitting caps on work that’s already paying off.
  • A team-wide mandate multiplies the METR risk. Your strongest engineers in your oldest codebase are exactly the profile that measured negative. A blanket “everyone uses the agent for everything” policy can burn senior time to subsidize junior gains — the Microsoft RCTs found juniors benefit most, so target the rollout there first.

For a full plan-by-plan cost breakdown, see our AI code editor cost comparison and what a multi-tool stack really costs.

How do you measure your own multiplier instead of guessing?

Run the METR protocol on yourself — it takes two weeks and a git log. The problem it solves is the one the study exposed: developers misjudged their own speedup by ~40 points, so any decision based on “it feels faster” is built on the least reliable instrument available.

  1. Forecast first. Write down how much faster you think the tool makes you. Sealed prediction, before measuring.
  2. Split real tasks. For two weeks, alternate comparable tasks: AI-assisted and unassisted. Don’t cherry-pick — boring tickets count.
  3. Measure completion, not vibes. Track wall-clock per task, plus a throughput proxy from your own history:
$ git log --author="$(git config user.email)" --since="14 days ago" --oneline | wc -l
42

Compare merged work, not generated lines — line counts reward exactly the verbose, rework-prone output the essay warns about. Add a quality check: how many of the AI-assisted changes needed a follow-up fix commit within the window?

  1. Compare against the sealed forecast. If your measured number is far below your forecast, you’ve reproduced the METR result at n=1 and saved yourself a $200/month tier — or a team rollout that quietly taxes your best people.

What actually moves you toward the 2x end of the range?

Task routing and feedback loops — the two levers every study and the essay agree on.

Give the model a loop it can close. The essay’s core observation is that LLMs excel when output can be executed and checked immediately. An agent with a test suite, a type checker, and a linter it can run converges on working code; the same agent free-styling into an untested module produces the “20% done” artifact that eats your afternoon in review. Before blaming the model or upgrading tiers, invest in the boring substrate: fast tests, strict types, pre-commit checks. That investment also pays when the human is typing.

Route by task, not by loyalty. The study spread is a routing table. Greenfield scaffolding, test generation, one-off scripts, unfamiliar-framework boilerplate: agent, expect Peng-style gains. Deep changes in a mature codebase you know cold, security-sensitive paths, anything where the repo’s conventions live in your head: keep the keyboard, expect METR-style losses if you delegate. Language and experience level shift the table too — our beginners vs. experienced developers breakdown covers why the same tool reads as magic to one dev and friction to another.

Budget for the 80/20 inversion. If “it runs” is now the 20% mark, schedule the remaining 80% — review, integration, docs — instead of discovering it after the sprint commitment. Teams that count an agent’s first working draft as “nearly done” are the ones that later report AI “didn’t help” — the time didn’t disappear, it moved downstream where nobody logged it.

None of this requires the expensive tiers, and none of it requires cloud models at all. The measured gains in every study came from GPT-3.5-era or early-2025 models — well within reach of current open-weight models running locally. A Cline + local model setup captures the boilerplate-and-tests lift at $0/month in subscriptions if privacy or budget rules out cloud tools; see runaihome.com’s guide to the best local models by VRAM for the hardware side, or rent a used-class GPU from $0.07/hour on Vast.ai to test the workload before buying anything.

FAQ

Is AI coding actually making developers slower? Only in one measured setting so far: experienced maintainers working in large repositories they know deeply (METR, early 2025 — 19% slower with Cursor Pro + Claude 3.5/3.7 Sonnet). Every broader trial measured gains of 21–56%. The honest summary is a range, not a verdict: your task mix decides your multiplier.

Do beginners or experts gain more from AI coding tools? Beginners, consistently. The 4,867-developer Microsoft/Accenture RCTs found less-experienced developers gained noticeably more than senior ones, and the only negative result on record involved deep experts. This has a direct policy implication: mandate-from-the-top rollouts help the people who need it least urgently and can tax your strongest engineers.

Aren’t these studies outdated now that models are much better? Partly. The +26% field result ran on GPT-3.5-era Copilot in 2022–23, and the -19% result used early-2025 tooling — current models are stronger on both counts. But no 2026 study of comparable rigor exists yet, so treat today’s tools as “at least this good” rather than assuming the gap to 10x has closed. The perception-gap finding, meanwhile, has no reason to have expired.

Does a 1.2x real gain justify $200/month tiers? Only if you’re hitting usage caps on work that’s already measurably paying off. The expensive tiers sell volume, not a bigger multiplier. Measure first (two-week protocol above), then buy the credits your measured workflow consumes — our cost comparison has the tier-by-tier math.

Sources

Last updated September 30, 2026. Study results are fixed, but tool pricing changes frequently — verify current plan prices before purchasing.

Was this article helpful?

Know which coding tool is worth paying for

Hands-on comparisons of AI coding assistants and what each one costs to run — including the local-model path. Sent only when something changes. Unsubscribe anytime.