The AI Margin Collapse, Re-Checked: Coding Model Prices Didn't Race to Zero — They Converged on the Middle

pricingcost-analysiscursorclinedeepseekglmclaudebyoklocal-llm

TL;DR: Martin Alderson’s July 6, 2026 thesis — GLM 5.2 at ~$4.40/M output would collapse frontier labs’ ~90% inference margins — got a real-world test within eleven weeks. The result isn’t a race to zero: DeepSeek raised output prices 114–329% on September 10, Z.ai moved its flagship behind $18–160/month plans, and Anthropic cut Opus to $4/$20. Prices are converging on a $1–5/M middle band from both directions.

Budget API tier (DeepSeek, GLM routers)Frontier API tier (Anthropic)Flat-rate plans (Cursor, Copilot, GLM Coding Plan)
Direction since JulyUp — V4.1-Flash output $1.20/M peak vs V4-Flash’s $0.28Down — Opus 5.5 at $4/$20, 20% below Opus 5; Sonnet 5’s Sep 1 hike canceledMetered — Cursor pools are now dollar-denominated; Copilot moved to AI Credits June 1
What it signalsLoss-leader pricing is endingMiddle tier is contested groundThe all-you-can-eat subsidy is what actually collapsed
The catchCache and off-peak rules decide your real billCache reads ($0.20/M) are the real moatCredit pool = plan price; overage is BYOK math in disguise

Honest take: The margin collapse was real but hit the wrong victim first — flat-rate subscriptions, not frontier APIs. Route 70–80% of routine agent work to the $0.60–4.40/M band (V4.1-Flash off-peak, GLM via routers), keep Opus 5.5 or Sonnet 5 for the hardest multi-file work, and stop expecting 2025-style $0.14/M pricing to come back.

What did the July 2026 margin-collapse thesis actually claim?

The claim, from Martin Alderson’s “GLM 5.2 and the coming AI margin collapse” (July 6, 2026; 100+ points on Hacker News), was specific: when Anthropic charges $25/M output tokens, that’s roughly a 90% gross margin over compute rack rates — and GLM 5.2, the first open-weight model Alderson considered a genuine Opus-class competitor for agentic coding, was serving the same work at $1.40/$4.40 per million tokens on Z.ai’s API, with third-party hosts undercutting even that at roughly $0.73–1.05 input and $2.29–3.30 output. At 15–20% of frontier rack rate for comparable agentic quality, the 90% margin was the arbitrage, and the prediction was that frontier per-token prices had to fall.

The sharpest counterargument arrived in the same HN thread: coding workloads are input-dominated and cache-dominated. Anthropic sells cache reads at a ~95% discount, so a Cline session that keeps resending a growing context pays cache-read rates on most of its tokens. On that math, one commenter calculated GLM would need to be roughly 1/36th of Claude’s price to win a cache-heavy coding workload on cost — and it was only 1/4th to 1/8th depending on host. Keep that number in mind; it explains most of what happened next.

What actually happened to coding-model prices between July and September 2026?

Three moves, none of them “race to zero”:

EventDatePrice change
DeepSeek retires V4-Flash, launches V4.1-FlashSep 10, 2026Output $0.28 → $1.20/M peak (+329%), $0.60 off-peak (+114%); input $0.14 → $0.30 peak; cache hits cut to $0.003–0.006/M
Z.ai restructures GLM accessJul 30–Aug 26, 2026Coding Plan moves to token-metered credits; GLM-5.3 ships Aug 14 as plan-first (Lite $18 / Pro $72 / Max $160 per month); general per-token API stays $1.40/$4.40
Anthropic cuts frontier pricingSep 22, 2026Opus 5.5 replaces Opus 5 at $4/$20 (was $5/$25, −20%); cache reads −60% to $0.20/M; Sonnet 5’s planned Sep 1 increase to $3/$15 canceled, stays $2/$10

All three verified September 25, 2026 against claude.com/pricing, DeepSeek’s rate card as covered in our September 17 breakdown, and Z.ai plan trackers. The cheapest tier got more expensive, the most expensive tier got cheaper, and the open-weight champion started selling subscriptions. That’s not a collapse toward zero — it’s compression toward a $1–5/M output band where DeepSeek’s peak rate ($1.20), GLM’s API rate ($4.40), router-hosted GLM ($2.29–3.30), and Opus 5.5 ($20, but $0.20 cached input) all now compete on workload shape rather than sticker price.

Why did DeepSeek — the cheapest provider — raise prices if margins are collapsing?

Because the collapse thesis assumed budget pricing was sustainable, and DeepSeek’s September repricing is the strongest evidence it wasn’t. V4-Flash at $0.14/$0.28 per million tokens was the number every “open weights will eat the frontier” argument leaned on, including Alderson’s. Eleven weeks later that price no longer exists: the same deepseek-flash endpoint silently serves V4.1-Flash at up to 4.3× the output rate.

Look at where DeepSeek did cut, though: cache hits dropped to $0.003–0.006/M — 46× cheaper than V4-Flash’s old cache-miss input rate — and off-peak hours (everything outside Mon–Fri 01:00–04:00 and 06:00–10:00 UTC) stay at half price. That’s DeepSeek adopting Anthropic’s playbook: charge for fresh compute, nearly give away cached context, and shape demand around cluster utilization. The cache-economics counterargument from the July HN thread didn’t just survive — both sides of the market converged on it. For what the change costs a real Cline user (between a wash and ~30% more), see our bill-math breakdown.

Did the margin collapse hit anyone? Yes: flat-rate subscriptions

The casualty the July thesis didn’t predict: all-you-can-eat pricing. Between June and September 2026, every major flat-rate coding plan became a metered one.

GitHub Copilot moved to AI Credits on June 1, 2026 — the change that produced 10–50× bill spikes for agentic power users. Cursor’s paid tiers now carry a dollar-denominated credit pool equal to the plan price — $20 of usage on Pro ($20/month), $60 on Pro+, $200 on Ultra, verified September 25, 2026 via cursor.com/pricing coverage. Even Z.ai, the open-weight vendor, replaced its Coding Plan’s prompt quotas with token-metered credits on July 30.

The economics are straightforward: agent sessions run for hours and burn tokens at rates no $20 flat fee can absorb when the marginal serving cost is real. Alderson’s 90%-margin arithmetic was right about where the cushion was — but the cushion was spent subsidizing flat-rate plans, and vendors chose to end the subsidy rather than cut API rack rates. If your bill went up this summer, this — not open-weight competition failing — is why.

Where does the $0 floor stand — can a local model replace the API tier?

Stronger than in July, with the same ceiling. Qwen3.8-27B shipped August 14, 2026 under Apache 2.0: a 27.78B dense model with a Qwen-reported 61.7 on SWE-bench Pro — within 1.5 points of Sonnet 5 — whose 18GB Ollama build runs on a single 24GB card. GLM-5.3’s weights followed in late August. The open-weight floor under the whole market is lower and higher-quality than it was when the thesis was written.

What hasn’t changed is what the floor can’t reach. A used RTX 3090 runs $1,150–1,350 (September 2026 street price — up from ~$1,010 in March, since the DRAM crisis repriced GPUs too), and at DeepSeek’s off-peak cache-hit rates a typical solo Cline bill is $1.20–2.40/month. The API has to get much more expensive before hardware pays for itself on cost alone; the local case remains privacy, rate-limit immunity, and offline work. If that’s your case, rent a 24GB card from about $0.07/hr on Vast.ai to test your workload before buying a used RTX 3090 — and see the sister-site guide to local AI models by VRAM for which GPU tier buys which model class.

How should a Cursor or Cline user route models in the convergence era?

Route by workload shape, not by sticker price — the September price moves made session shape (cache ratio, time of day, task difficulty) worth more than provider loyalty.

A concrete comparison, for one uncached agent session of 200K input / 25K output tokens (the session shape from our Opus cost analyses), at prices verified September 25, 2026:

session = {"in": 0.2, "out": 0.025}  # millions of tokens
rates = {  # (input $/M, output $/M)
    "V4.1-Flash off-peak": (0.15, 0.60),
    "V4.1-Flash peak":     (0.30, 1.20),
    "GLM-5.3 via router":  (1.40, 4.40),
    "Sonnet 5":            (2.00, 10.00),
    "Opus 5.5":            (4.00, 20.00),
}
for m, (i, o) in rates.items():
    print(f"{m:22s} ${session['in']*i + session['out']*o:.2f}")

Output:

V4.1-Flash off-peak    $0.04
V4.1-Flash peak        $0.09
GLM-5.3 via router     $0.39
Sonnet 5               $0.65
Opus 5.5               $1.30

A nearly 30× spread, uncached. With warm caches the spread narrows hard: Opus 5.5’s $0.20/M cache reads mean a read-dominated session costs a fraction of the naive number, and DeepSeek’s $0.003 cache hits do the same at the budget end. Two practical consequences:

  1. Escalation beats loyalty. Run routine single-file edits, test generation, and boilerplate on V4.1-Flash or router-hosted GLM; escalate multi-file refactors and architecture work to Sonnet 5 or Opus 5.5. At the spread above, even escalating 20% of sessions leaves you at roughly a third of an all-Opus bill.
  2. Watch your harness’s default. One real trap we hit in practice: Cline’s Anthropic provider default has resolved to Fable 5.1 ($10/$50) since v4.1.17 — 2.5× Opus 5.5’s price for a model Anthropic’s own September 22 launch table scores lower on coding. The fix takes a minute: open Cline’s provider settings and pin a model explicitly (Sonnet 5 at $2/$10, or Opus 5.5 once your catalog update lands). In a converging market, unpinned defaults are where the margin hides.

This is the same conclusion our multi-tool cost analysis reached from the subscription side; the full price landscape lives in the AI code editor cost comparison.

What still doesn’t commoditize?

Four things kept their pricing power through the convergence, and they’re where the next margin lives:

  • Cache infrastructure. Both the cheapest and the most expensive vendors now make their real money on the cache-miss/cache-hit spread. Whoever holds your warm context holds you.
  • The hardest 5% of tasks. Opus 5.5 at $4/$20 and Fable 5.1 at $10/$50 exist because no open-weight model has independently matched them on the multi-file agentic benchmarks that decide whether a refactor lands or burns an afternoon.
  • Harness integration. Cursor’s indexing and tab model, Claude Code’s agent loop — the harness-architecture comparison covers why the loop now matters more than the model. Switching costs live in the tool, not the tokens.
  • Reliability at long context. Router-hosted open models quantize, rate-limit, and vary between hosts; a September price that undercuts Z.ai by 30% can mean an fp4 quant that fails tool calls. This doesn’t apply to first-party endpoints, which is partly what their premium buys.

What to actually do

Prices as of September 25, 2026, all verified in the sections above:

Your situationThe movePriceWhere
Routine agent work, US working hoursDeepSeek V4.1-Flash off-peak as your Cline/BYOK default$0.15/$0.60 per M, $0.003 cache hitsDeepSeek API
All-day agentic coding, want flat costGLM Coding Plan Lite on GLM-5.3$18/monthZ.ai
Hardest 20% of tasks, escalation targetOpus 5.5 ($4/$20) or Sonnet 5 ($2/$10) via API key~$0.65–1.30 per big uncached sessionclaude.com/pricing
Code that can’t leave the machineQwen3.8-27B local on a 24GB card — test the workload rented firstfrom $0.07/hr rented on Vast.ai; used RTX 3090 ~$1,150–1,350 to ownOllama, weights free

FAQ

Was the margin-collapse thesis wrong? Half wrong. Frontier prices did fall (Opus −20% in September, Sonnet’s hike canceled), which the thesis predicted. But the budget tier rose toward the middle instead of dragging everything to zero, and the clearest casualty was flat-rate subscription pricing — a mechanism the thesis didn’t address.

Is GLM still 15–20% of frontier price? On sticker, roughly yes: $4.40/M output vs Opus 5.5’s $20. In practice the gap is narrower, because Anthropic’s $0.20/M cache reads dominate real coding sessions and Z.ai’s flagship access is now plan-first. For cache-heavy workloads, the effective multiple is closer to the 1/4–1/8 the July HN counterargument computed — which is why frontier APIs didn’t have to panic-cut.

Should I buy a GPU now that API prices are rising? Not on cost grounds. Even after DeepSeek’s 114–329% increase, a typical solo agentic workload runs single-digit dollars a month off-peak, against $1,150–1,350 for a used RTX 3090 whose price also rose this year. Buy local for privacy or offline needs; rent first to size the workload.

Does this affect Cursor Pro users who never touch BYOK? Yes, through the credit pool. Cursor Pro’s $20/month now maps to $20 of model usage at rates close to API list price, so every provider repricing flows through to how far your pool stretches. The September Opus cut means Opus-heavy Cursor sessions burn ~20% fewer credits than they did on Opus 5.

Sources

Last verified September 25, 2026. Every provider named here repriced at least once between June and September 2026; check the official rate card before wiring a bill-sensitive workflow to any of these numbers.

Was this article helpful?

Know which coding tool is worth paying for

Hands-on comparisons of AI coding assistants and what each one costs to run — including the local-model path. Sent only when something changes. Unsubscribe anytime.