MiniMax M3 as a Cursor and Cline Backend in 2026: $0.30/M Pricing, the Interleaved-Thinking Trap, and Why H3 Is Not the Model You Want
TL;DR: MiniMax M3 — the 428B-parameter open-weight MoE released June 1, 2026 — wires into Cline natively and into Cursor via BYOK, at a promo rate of $0.30/$1.20 per million tokens. It claims 80.5% on SWE-bench Verified, but the number is vendor-run. And the “MiniMax H3” everyone saw on r/LocalLLaMA this week is a video model, not a coding model.
| MiniMax M3 | DeepSeek V4-Flash | Kimi K3 | |
|---|---|---|---|
| Best for | Long-context agent loops, 1M-token codebase jobs | Cheapest competent daily driver | Frontend/UI generation |
| Input / Output per 1M | $0.30 / $1.20 promo (list $0.60 / $2.40), $0.06 cache read | $0.14 / $0.28, $0.0028 cache hit | $3.00 / $15.00, $0.30 cache hit |
| SWE-bench Verified | 80.5% (vendor-run) | 79% | 76.8% (vendor) |
| Context window | 1M tokens (input >512K bills 2×) | 1M tokens | 1M tokens |
| The catch | Strip its <think> blocks and tool use falls apart | DeepSeek announced a “significant” price increase Aug 6 | ~2× verbosity at Sonnet prices |
Honest take: DeepSeek V4-Flash is still the better default at half the price — but its price is about to go up, and M3 is the strongest hedge: same 1M context, credible coding scores, and a subscription plan V4-Flash doesn’t offer. Set it up in Cline now so switching is a dropdown click, not a research project.
First, the name problem: H3 is not a coding model
If you landed here after the August 3 Hugging Face drop that lit up r/LocalLLaMA, here’s the correction nobody’s headline made room for: MiniMax H3 is Hailuo 3.0, a 33B omni-modal video generation model. It produces 4–15 second clips at up to 2K/24fps with native stereo audio. It has no chat completions endpoint, no tool calling, and no business being anywhere near your editor. Its community license also excludes the US, EU, UK, and South Korea from the “Applicable Territory” for local deployment, per TechTimes’ August 4 coverage — so for most readers of this site, H3 isn’t even legally self-hostable.
The MiniMax model you point Cursor or Cline at is M3, the M-series text flagship released June 1, 2026. Everything below is about M3, verified against MiniMax’s platform docs and current pricing pages on August 8, 2026.
What M3 actually is
M3 is a 428-billion-parameter mixture-of-experts model with roughly 23B parameters active per token, a 1M-token context window, and native image and video input (it reads screenshots; it doesn’t generate video — that’s H3’s job). The long-context trick is MiniMax Sparse Attention (MSA): a block-sparse scheme where each attention layer scores the preceding context in 128-token blocks and keeps only the 16 most relevant, cutting per-token attention compute to roughly a twentieth of dense attention at the full window.
The benchmark table MiniMax published is aggressive: 80.5% on SWE-bench Verified (a hair under DeepSeek V4’s 80.6%), 59.0 on SWE-bench Pro, and 66.0 on Terminal-Bench 2.1 — the last one roughly matching Claude Opus 4.7’s 66.1 baseline while trailing Opus 4.8’s 74.6. Two caveats before you reprice your stack around those numbers. First, they’re vendor-run on MiniMax’s own harness; TechTimes’ launch-day analysis flatly called them “frontier claims, unverified benchmarks,” and independent leaderboard placements have been slower to appear. Second, VentureBeat’s framing — matching GPT-5.5-class performance “for 5–10% of the cost” — is a pricing argument, not a capability one. The honest read: M3 is in the top open-weight tier for agentic coding, close to DeepSeek V4 and ahead of Kimi K3 on backend work, and nobody outside MiniMax has proven more than that.
The weights are live at MiniMaxAI/MiniMax-M3 on Hugging Face under the MiniMax Community License: commercial use requires displaying “Built with MiniMax M3” and a one-time notice to MiniMax, escalating to written authorization only past $20M in annual revenue. That’s Kimi-style attribution licensing, not Apache — fine for a solo developer’s BYOK use, worth a legal read if you’re shipping M3 inside a product.
Can you run those weights at home? At 428B parameters, a Q4 quant is still north of 200GB — this is Mac Studio 512GB or multi-GPU server territory, not RTX 3090 territory. Our sister site has the full VRAM math in its MiniMax M3 local hardware guide; the one-line version is that almost everyone should use the API and treat “open weights” as an insurance policy, the same conclusion we reached for Kimi K3.
Pricing: the promo rate, the Token Plan, and the DeepSeek wrinkle
MiniMax lists M3 at $0.60/$2.40 per million input/output tokens, with a “Permanent 50% off” banner bringing the billed rate to $0.30 input / $1.20 output, plus $0.06 per million on prompt-cache reads — for requests up to 512K input tokens. Cross that line and the rate doubles to the list price. OpenRouter carries minimax/minimax-m3 at the same $0.30/$1.20 if you’d rather not open another account.
There’s also a subscription route the pure-API rivals don’t offer: the MiniMax Token Plan at $20 (Plus), $50 (Max), and $120 (Ultra) per month, with monthly M3 quotas of roughly 1.6B, 5.1B, and 9.8B tokens respectively, shared across MiniMax’s text, image, and speech models through a subscription key. At 1.6B tokens for $20, the Plus tier is dramatically cheaper per token than pay-as-you-go — but only if you actually burn that volume, which means multiple concurrent agents, not one Cline session.
For a single developer, the pay-as-you-go math at a heavy 500K tokens/day (400K in, 100K out) works out to:
| Backend | Per day | Per month (30 days) |
|---|---|---|
| DeepSeek V4-Flash | $0.084 | $2.52 |
| MiniMax M3 (promo) | $0.24 | $7.20 |
| MiniMax M3 (list, if promo ends) | $0.48 | $14.40 |
| Kimi K3 | $2.70 | $81.00 |
V4-Flash wins on today’s sticker. But on August 6, Bloomberg reported DeepSeek plans a “significant” increase to its API pricing in the near future. Nobody knows the new rates yet; what the table tells you is that M3 at promo pricing is the closest fallback that doesn’t give up the 1M context window — and that’s exactly why it’s worth having configured before you need it.
Cline setup: use the native provider, not the generic OpenAI one
Cline ships a first-party MiniMax provider, and MiniMax’s own docs document the flow:
- Create an API key at
platform.minimax.io(international) — China-mainland accounts live onplatform.minimaxi.com, and keys are not interchangeable between the two. - In Cline’s settings, set API Provider → MiniMax, set MiniMax Entrypoint to
api.minimax.io(orapi.minimaxi.comfor China), and paste the key. - Pick
MiniMax-M3as the model and save.
Sanity-check the key from a terminal before debugging anything inside an editor:
curl https://api.minimax.io/v1/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $MINIMAX_API_KEY" \
-d '{"model": "MiniMax-M3", "messages": [{"role": "user", "content": "Say ready."}]}'
A working key returns HTTP 200 and a standard chat.completion JSON object whose message.content ends in “ready” — often preceded by a <think>...</think> block, which brings us to the part that actually breaks people’s setups.
The interleaved-thinking trap (and the fix)
M3 reasons between tool calls: it emits <think> blocks mid-loop — plan, call a tool, read the result, think again, call the next tool. MiniMax’s function-calling docs are explicit that the reasoning content must be fed back to the model on every subsequent turn. Strip it from the conversation history and multi-turn tool use degrades in a very recognizable way: the agent re-reads files it just read, forgets the plan it announced two steps ago, and loops on the same edit. MiniMax documented the same requirement for M2 (“Interleaved Thinking Unlocks Reliable Agentic Capability”), and M3 inherits it wholesale.
The fix depends on which surface you’re using:
- Cline’s native MiniMax provider handles it for you — the maintainers landed adaptive-thinking support for M3 specifically. This is the single best reason to select “MiniMax” in the provider dropdown instead of wiring M3 through Cline’s generic OpenAI-compatible provider, which treats reasoning as disposable.
- On the Anthropic-compatible endpoint (
https://api.minimax.io/anthropic, standard/v1/messagessemantics, works with Claude Code–style CLIs), append the model’s entire response content list — thinking, text, and tool_use blocks — to your message history, exactly as you would for Claude’s extended thinking. - On the raw OpenAI-compatible endpoint, echo assistant messages back with their
<think>content intact rather than dropping it.
If your agent loops even with the native provider, the failure mode overlaps with the truncated-context problems we covered in the Cline/Ollama tool-use loop fix — check context limits before blaming the model.
Cursor setup: works for chat, second-class for agents
Cursor doesn’t list M3 in its native picker, so this is the standard BYOK route — with three sharp edges verified against Cursor’s community forum this week:
- Settings → Models → API Keys, paste your MiniMax key in the OpenAI API Key field, enable Override OpenAI Base URL, and set it to
https://api.minimax.io/v1. AddMiniMax-M3as a custom model name. - The override field is hidden on the Hobby tier — you need Cursor Pro ($20/month) or above, which undercuts the “cheap backend” logic unless you’re on Pro anyway.
- The override is global: while it’s on, Cursor’s built-in models route through it too, so flip it off when you switch back. And if an old Anthropic key sits in the adjacent field, delete it — users report Cursor misrouting custom-model requests and returning 401s when both are populated. Custom models are also silently ignored inside Cursor’s subagents, per an open bug report.
Cursor’s Tab autocomplete keeps running on Cursor’s own models regardless — same as every BYOK backend we’ve tested, from DeepSeek V4-Flash to GLM-5.2. One more honest caveat: Cursor owns the conversation history on the OpenAI-compat path, and there’s no public documentation that it echoes M3’s reasoning blocks between tool calls. For chat and single-shot edits it works fine; for long agent chains, Cline’s native provider is the setup that matches how this model was trained. If you want a fully local fallback in the same editors instead, that’s what Inkling Small is for.
Verdict
Pick DeepSeek V4-Flash today if you want the cheapest competent agent backend — $2.52/month at heavy solo usage is still untouchable. Pick MiniMax M3 if you’re betting DeepSeek’s announced price hike lands hard, if you regularly stuff 500K+ tokens of codebase into context, or if you’re running enough parallel agents that the $20 Token Plan’s 1.6B monthly tokens beat any per-token rate on this page. Skip Kimi K3 for backend work entirely; its crown is frontend. And whatever you do, don’t ollama pull anything with “H3” in the name expecting it to write code — it will render you a very nice video of code instead.
FAQ
Is MiniMax M3 the same as MiniMax H3? No. M3 (June 1, 2026) is the 428B text/coding model with OpenAI- and Anthropic-compatible APIs. H3 (weights released August 3, 2026) is Hailuo 3.0, a 33B video generation model whose license excludes US/EU/UK/Korea local deployment.
Does M3 support tool calling in Cline and Cursor?
Yes — M3 was trained for long-horizon tool loops, and Cline’s native MiniMax provider supports it including the interleaved-thinking requirement. On generic OpenAI-compatible integrations, you must preserve <think> content across turns or tool use degrades.
What does M3 cost compared to DeepSeek V4-Flash? M3’s billed promo rate is $0.30/$1.20 per million input/output tokens ($0.06 cache reads); V4-Flash is $0.14/$0.28 ($0.0028 cache hits). V4-Flash is roughly half the price today, but DeepSeek announced a significant price increase on August 6, 2026.
Can I run MiniMax M3 locally?
The weights are on Hugging Face (MiniMaxAI/MiniMax-M3, MiniMax Community License), but at 428B parameters a Q4 quant exceeds 200GB — server-class hardware. See the runaihome.com hardware guide for the VRAM math.
Which endpoint should Claude Code users pick?
The Anthropic-compatible surface at https://api.minimax.io/anthropic (standard /v1/messages), with model ID MiniMax-M3. Append full response content blocks to history to keep interleaved thinking working.
Sources
- MiniMax M3: Frontier Coding, 1M Context, Native Multimodality — MiniMax Research — release announcement, specs, vendor benchmarks. Verified Aug 8 2026.
- Tool Use & Interleaved Thinking — MiniMax API Docs — reasoning-preservation requirement for M3 function calling.
- Anthropic SDK — MiniMax API Docs — Anthropic-compatible endpoint and model ID.
- Cline — MiniMax API Docs — official Cline provider setup steps and entrypoints.
- Cursor — MiniMax API Docs — official Cursor BYOK configuration.
- MiniMaxAI/MiniMax-M3 — Hugging Face — open weights and MiniMax Community License terms.
- MiniMax M3 — API Pricing & Benchmarks — OpenRouter — $0.30/$1.20 rates and permanent 50% discount listing.
- MiniMax Pricing 2026: API Costs, Token Plans & Credits — Token Plan tiers ($20/$50/$120) and monthly quotas.
- MiniMax M3 Open-Weight Coding Model: Frontier Claims, Unverified Benchmarks — TechTimes — June 1 launch analysis, vendor-run benchmark caveats.
- MiniMax H3 Open Weights Exclude US, EU, UK, and Korea From Local Deployment — TechTimes — H3 video model release and license territories.
- MiniMax-M3 debuts, eclipsing GPT-5.5 and Gemini 3.1 Pro on key benchmark performance for just 5-10% of the cost — VentureBeat — launch coverage and cost framing.
- DeepSeek Plans ‘Significant’ Price Increase for AI Services — Bloomberg — Aug 6 pricing news; current V4-Flash $0.14/$0.28 rates.
- Kimi API Pricing (August 2026): Kimi K3 at $3/$15 — BenchLM.ai — K3 comparison rates.
- Cline v3.35: Native Tool Calling, Auto-Approve Menu, and Free MiniMax M2 — Cline Blog — Cline native tool calling and MiniMax provider background.
- The custom override of the OpenAI base URL is unusable — Cursor Community Forum — BYOK override behavior, Pro-tier gating, and known conflicts.
Last verified: August 8, 2026. Pricing and features change fast — check the official pages before committing.
Was this article helpful?
Thanks for the feedback — it helps improve future articles.