Claude Haiku 5.5 for Cursor, Cline, and Claude Code: $0.10/M Input, the 100K Pricing Cliff, and When the Cheap Model Is Enough

claudeanthropiccursorclineclaude-codepricingcost-analysisapibyok

TL;DR: Claude Haiku 5.5 (released October 7, 2026) costs $0.10/$0.50 per million tokens on prompts up to 100K tokens — 20x cheaper than Sonnet 5.5 on both sides — and scores 39.2% on Terminal-Bench 4.0, a benchmark where its predecessor scored 0.0%. Above 100K prompt tokens the price jumps 5x to $0.50/$2.50. Cursor added it to the model picker on launch day; Claude Code v2.1.293 resolves the haiku alias to it; Cline’s built-in catalog is unconfirmed, but the OpenRouter route works today.

Claude Haiku 5.5Claude Sonnet 5.5Claude Opus 5.5
Input / output per MTok$0.10 / $0.50 (≤100K prompt), $0.50 / $2.50 above$2 / $10$4 / $20
Cache read per MTok$0.01 (≤100K), $0.05 above$0.10 (cut from $0.20 on Oct 7)$0.20
Terminal-Bench 4.0 (vendor-run)39.2%70.6%66.4%
FrontierCode 1.1 Main (vendor-run)46.4%52.1% (at xhigh)54.4%
Comparative latencyFastestFastModerate
The catchFinishes barely over half of what Sonnet does on terminal-agent tasks20x the token price for short prompts40x the token price; slower

Honest take: Haiku 5.5 is the first Haiku you can hand real coding work to — but not your main agent loop. Route it the subagent, lint-fix, test-loop, and summarization traffic where a failure is cheap to catch, keep Sonnet 5.5 as the driver, and your Anthropic bill drops hard without your merge quality following it. If you were paying Haiku 4.5 rates for background tasks, switching is free money: 10x cheaper on short requests and a model that actually completes agentic tasks.

How much does Claude Haiku 5.5 cost, exactly?

$0.10 per million input tokens and $0.50 per million output tokens — as long as the prompt stays at or under 100K tokens. Past that threshold, input is $0.50/M and output $2.50/M. All numbers below were verified against Anthropic’s announcement page and the Claude Platform model docs on October 10, 2026.

Per million tokens≤100K prompt>100K promptHaiku 4.5 (old)
Input$0.10$0.50$1.00
Output$0.50$2.50$5.00
Cache read$0.01$0.05$0.10
Cache write$0.125$0.625$1.25

Anthropic frames the tiering as a 75% average cost cut versus Haiku 4.5 — 90% lower for requests under 100K tokens (which it says covered about 90% of Haiku 4.5 traffic) and 50% lower above. Batch API requests take a further 50% off. One footnote worth knowing before you compare invoices: Haiku 5.5 uses a new tokenizer that consumes slightly more tokens for the same text, so your per-task savings land a little below the per-token math.

The spec sheet is no longer the small-model compromise it used to be. Haiku 5.5 ships with a 1M-token context window, 128K max output (300K on the Batch API with the output-300k-2026-03-24 beta header), a June 2026 knowledge cutoff, tool use, vision, and adaptive thinking with five effort levels (low through max, default medium) — the first Haiku-class model with an effort dial. API model ID: claude-haiku-5-5, a pinned dateless snapshot with no separate alias. Retirement is committed no sooner than October 7, 2027.

Same-day pricing news that affects the comparison column: Anthropic cut Sonnet 5.5 cache reads 50%, from $0.20 to $0.10 per million. If your harness leans hard on prompt caching, Sonnet got cheaper this week too.

What does a month of agentic coding cost on Haiku 5.5 vs Sonnet 5.5?

Roughly $2.40 versus $48 for the same short-prompt workload. Take a concrete day: 40 agent requests averaging 20K input and 1.5K output tokens each — a normal day of scoped edits, test loops, and commit messages, with every prompt under the 100K threshold. That is 800K input + 60K output daily.

Monthly (22 workdays)Haiku 5.5Sonnet 5.5Opus 5.5
Input (17.6M tokens)$1.76$35.20$70.40
Output (1.32M tokens)$0.66$13.20$26.40
Total$2.42$48.40$96.80

The 20x ratio holds on both input and output as long as prompts stay short. That changes what “worth it” means: at $2.42/month, Haiku 5.5 costs less than a tenth of a Cursor Pro seat, and the question stops being “can I afford to run an agent” and becomes “which tasks can this model actually finish.”

One more input for the math if you subscribe to Claude: alongside the Haiku 5.5 launch, Anthropic announced monthly API credits for subscribers — $100 for Max 5x, $200 for Max 20x, and up to $500 pooled for Team. A Max subscriber’s entire Haiku experimentation budget now fits inside credits they already get.

The 100K-token cliff: the trap in every long Cline session

Here is the problem we’d flag before anything else: agentic coding sessions grow their context, and Haiku 5.5’s pricing punishes that growth 5x. A Cline or Claude Code session that accumulates 150K tokens of files, diffs, and tool results re-sends that entire prompt every turn — and every one of those turns now bills input at $0.50/M instead of $0.10/M. A single 150K-token turn costs $0.075 in input instead of the $0.015 you’d pay if the same tokens sat under the threshold. Run 30 turns like that and the “cheap model” quietly costs 5x its sticker price.

Two fixes, both verifiable in the docs:

  1. Prompt caching absorbs most of it. Cache reads bill at $0.05/M even on the long-prompt tier, so a cached 150K-token prefix costs $0.0075 per turn, not $0.075. Claude Code and Cline both cache aggressively by default; if you’re wiring Haiku 5.5 into a custom harness, caching is not optional at this price structure.
  2. Compact before the cliff, not at the context limit. Claude Code auto-compacts around 967K tokens by default — which, on Haiku 5.5, means your session can spend 850K+ tokens in the expensive tier before compaction triggers. Lowering the threshold with /autocompact (or autoCompactWindow in settings) to keep working context under 100K preserves the 10-cent tier for the whole session.

This is the single biggest difference between Haiku 5.5 benchmarks and Haiku 5.5 invoices. The model is priced for short, frequent, well-scoped requests; it is merely tolerable for marathon sessions.

Is Haiku 5.5 actually good enough for coding work?

For scoped tasks, yes — and that’s new for a Haiku. On Anthropic’s own launch numbers (vendor-run, flag accordingly), Haiku 5.5 scores 39.2% on Terminal-Bench 4.0. Haiku 4.5 scored 0.0% on the same benchmark — it could not complete terminal-agent tasks at all — and OpenAI’s GPT-6 Luna, a model in a similar price-speed class, sits at 16.4%. On FrontierCode 1.1 Main, Haiku 5.5’s 46.4% actually edges out GPT-6 Luna’s 42.4% and lands within six points of Sonnet 5.5’s 52.1%.

The ceiling is just as clear. Sonnet 5.5 scores 70.6% on Terminal-Bench 4.0 — Haiku completes barely over half of what Sonnet does when the task is “operate a terminal until the job is done.” Computer use shows the same shape: 72.4% on OSWorld 2.1 (offline subset) against Sonnet’s 83.9%, a massive jump from Haiku 4.5’s 15.7% but still clearly a tier down.

Independent measurement backs the general picture. Artificial Analysis scores Haiku 5.5 between 29 and 43 on its Intelligence Index depending on effort level (low through max), with output speeds of roughly 137–243 tokens/second on the Anthropic API — against a median of about 111 tok/s for reasoning models in its price band. OpenRouter separately reports measuring 100+ tokens/second. No independent SWE-bench Verified score had been published as of October 10; until one lands, treat the coding benchmarks above as Anthropic’s numbers.

What that means in practice: give Haiku 5.5 the work where a failure is cheap to detect — a lint fix the linter re-checks, a test loop the test suite judges, a single-file refactor you review in ten seconds, a commit message, a codebase summary. Keep Sonnet 5.5 on multi-file changes where a subtly wrong edit costs you a review cycle. If Haiku fails a task and you re-run it on Sonnet, you paid for both — which is why routing by task type beats routing by price.

Which effort level should you run — and the xhigh latency trap

Stay at medium or high for interactive coding. The effort dial is Haiku 5.5’s headline feature, and it hides a trap: Artificial Analysis measured time-to-first-answer-token at roughly 13 seconds at medium but around 87 seconds at xhigh on its 10K-input-token workload, because the model spends the gap thinking. An effort level that adds over a minute of dead air defeats the entire reason you picked the fastest model in the lineup. Reserve xhigh/max (Intelligence Index 41–43) for batch jobs and background subagent work where nobody is watching a spinner; the default medium is where the latency story — Anthropic’s “fastest model to date,” Asana’s reported 2.5x faster agent turns, Box’s reported halved latency versus Haiku 4.5 — actually holds.

How do you turn on Haiku 5.5 in Cursor?

Settings → Models → toggle on Claude Haiku 5.5. Cursor shipped it on launch day (October 7) and its docs list the same tiered pricing as Anthropic’s: $0.10/$0.50 under 100K input tokens, $0.50/$2.50 above. Cursor’s own framing: on shorter requests it costs 10x less than Claude Haiku 4.5. The effort variants surface as separate model entries rather than a dial.

Where it fits in a Cursor workflow: Tab completion stays on Cursor’s proprietary model regardless, so Haiku 5.5 competes for your chat and agent traffic. Pointing quick, scoped agent asks at Haiku 5.5 and keeping Sonnet 5.5 or Opus 5.5 for Composer-scale multi-file work mirrors the routing logic above — and on usage-based pricing, every short request you move to Haiku is a 95% discount against Sonnet.

How do you use Haiku 5.5 in Claude Code?

Update first — Haiku 5.5 requires Claude Code v2.1.293 or later (claude update). On the Anthropic API, the haiku alias now resolves to Haiku 5.5, so selection is one command:

/model haiku

Or pin the exact ID in settings.json so an alias change never surprises you:

{
  "model": "claude-haiku-5-5"
}

Watch the platform asterisk: the haiku alias resolves to Haiku 5.5 only on the Anthropic API. On Amazon Bedrock, Google Cloud, and Microsoft Foundry it still resolves to Haiku 4.5 as of October 10 — if you run Claude Code through Bedrock, /model haiku does not get you the new model or the new pricing. Use the full ID where your platform offers it, or pin via ANTHROPIC_DEFAULT_HAIKU_MODEL.

The highest-leverage setup isn’t making Haiku your main model — it’s making it your subagent model. Claude Code lets you set a per-agent model in the subagent’s frontmatter:

---
name: test-runner
model: haiku
---

or set one default for every subagent, teammate, and workflow agent that doesn’t specify its own:

export CLAUDE_CODE_SUBAGENT_MODEL=claude-haiku-5-5

Main loop on Sonnet 5.5 or Opus 5.5, fan-out work on Haiku 5.5 at a twentieth of Sonnet’s token price — this is exactly the architecture Anthropic itself pitches the model for (subagents, classification, extraction, summarization), and it’s the pattern we costed out in our subagent model routing guide. Haiku 5.5 also uses the full 1M-token context window on the API by default, no [1m] alias suffix needed.

Can Cline use Haiku 5.5 yet?

Through OpenRouter, yes; through the built-in Anthropic catalog, we could not confirm it as of October 10, 2026. Cline’s docs were unreachable during this check and no changelog entry confirms a native listing — the same lag we documented when Sonnet 5.5 shipped and Cline’s built-in catalog trailed the release. The route that works today: select OpenRouter as the provider and use model ID anthropic/claude-haiku-5.5 (note the dot — OpenRouter’s ID format differs from Anthropic’s dashed claude-haiku-5-5). OpenRouter listed the model on launch day. If you use Cline’s native Anthropic provider, check whether the model dropdown has updated before assuming; a mistyped manual ID fails loudly, which beats silently running an older Haiku.

Where Haiku 5.5 is the wrong choice

Don’t route these at it:

  • Long-horizon multi-file agent runs. The 39.2% vs 70.6% Terminal-Bench gap is the difference between an agent that usually finishes and one that usually doesn’t. A failed 40-minute agent run costs more than Sonnet’s tokens.
  • Sessions that live above 100K context. The 5x price cliff plus weaker long-horizon reliability stack against you. Compact aggressively or use Sonnet.
  • xhigh/max effort for interactive use. ~87 seconds to first token at xhigh per Artificial Analysis. Batch work only.
  • Anything where you’d re-do the failure on a bigger model anyway. Paying twice erases a 20x price edge faster than you’d think: if more than roughly 1 task in 20 falls back to Sonnet and the retry burns similar tokens, you’re still ahead — but the review time you spend catching the failures is the real cost, and it isn’t on the invoice.

And the zero-dollar alternative still exists: a local Qwen- or Devstral-class model behind an OpenAI-compatible endpoint costs nothing per token if you own the hardware. At $0.10/M, though, Haiku 5.5 makes the “is a local GPU worth it for cheap-tier tasks” math harder than it’s ever been — see the best local models by VRAM budget for what your hardware can actually serve, and the FOSS agent comparison if you want the no-cloud stack end to end.

FAQ

What is the exact model ID for Claude Haiku 5.5? claude-haiku-5-5 on the Anthropic API, Google Cloud, Microsoft Foundry, and Claude Platform on AWS; anthropic.claude-haiku-5-5 on Amazon Bedrock; anthropic/claude-haiku-5.5 on OpenRouter. It’s a pinned dateless snapshot — there is no dated variant to chase.

Is the context window 200K or 1M? 1M tokens, with 128K max output, per Anthropic’s platform docs (verified October 10, 2026). At least one third-party article circulated a 200K/February-2025-cutoff spec at launch; it’s wrong. Knowledge cutoff is June 2026.

Does Haiku 5.5 support tool use and vision? Yes. All current Claude models — Haiku 5.5 included — support tool use, vision/image input, and multilingual text, per the models overview. Tool use is the one that matters for Cline and Claude Code agent loops, and it’s there.

Why did my Haiku 5.5 bill spike mid-session? Almost certainly the 100K cliff: once a request’s prompt exceeds 100K tokens, input bills at $0.50/M and output at $2.50/M — 5x the short-prompt rate. Prompt caching ($0.05/M cache reads on the long tier) and earlier compaction are the fixes; see the cliff section above.

Should I replace Sonnet 5.5 with Haiku 5.5 to save money? No — split the work instead. Haiku 5.5 completes just over half of Sonnet 5.5’s Terminal-Bench tasks (39.2% vs 70.6%, vendor-run). Route short, verifiable tasks and subagents to Haiku and keep Sonnet as the main agent; that captures most of the savings with none of the reliability downgrade on the work that matters. Full Sonnet 5.5 numbers are in our Sonnet 5.5 breakdown, and the top-tier escalation case in the Opus 5.5 backend review.

Sources

Last updated October 10, 2026. Pricing and features change frequently; verify current state on the official pages before committing a team to a model.

Was this article helpful?

Know which coding tool is worth paying for

Hands-on comparisons of AI coding assistants and what each one costs to run — including the local-model path. Sent only when something changes. Unsubscribe anytime.