Grok 4.7 for Cursor and Cline: Same $2/$6 Sticker, a 47% Verbosity Tax, and When It Actually Beats Claude Fable 5.1
TL;DR: Grok 4.7 (released September 21, 2026) is SpaceXAI’s best coding model — 4th on the Artificial Analysis Coding Agent Index — at the same $2/$6 per million tokens as Grok 4.6. The catch: it writes roughly 2.2x more output tokens per task, so measured cost per completed task rose 47%, from $1.86 to $2.73. It’s still cheaper per task than Claude Fable 5.1 ($3.76), but the gap is 27%, not the 5–8x the sticker prices suggest.
| Grok 4.7 | Claude Opus 5.5 | Claude Fable 5.1 | |
|---|---|---|---|
| Best for | Budget agentic sessions, effort capped at high | The price/accuracy middle ground | Accuracy-critical refactors, cache-heavy loops |
| Price per MTok (in/out) | $2 / $6 (2x above 200K input) | $4 / $20 | $10 / $50 |
| Measured cost per AA Intelligence Index task | $2.73 (high) | — | $3.76 (max) |
| The catch | ~81K output tokens per task at xhigh — verbosity eats the discount | Not selectable in Cline’s catalog yet | Cache reads are cheap ($0.25/M) but everything else isn’t |
Honest take: If you’re already on Grok 4.6 in Cursor or Cline, upgrade — same price, better model, just cap reasoning effort at high. If you’re choosing fresh and your work is accuracy-critical, Fable 5.1 or Opus 5.5 is still the right call; Grok 4.7’s per-task discount is real but far smaller than the price list implies.
What did SpaceXAI actually ship in Grok 4.7?
Grok 4.7 is SpaceXAI’s flagship coding and knowledge-work model, released September 21, 2026, with the API model ID grok-4.7 (period, not hyphen) at https://api.x.ai/v1. Pricing is unchanged from Grok 4.6: $2 per million input tokens, $6 per million output, with cached prompt prefixes at $0.50 per million — but input pricing doubles above 200K tokens of context, a detail that matters for large-repo agent sessions (more on that below).
The verified spec sheet, per SpaceXAI’s model docs and launch coverage (all checked September 26, 2026):
| Spec | Grok 4.7 |
|---|---|
| Release date | September 21, 2026 |
| API model ID | grok-4.7 |
| Context window | 500K tokens (input price doubles past 200K) |
| Input / output per MTok | $2.00 / $6.00 |
| Cached input per MTok | $0.50 |
| Reasoning levels | 4 (up to xhigh) |
| Modality | Text + image input; function calling, web search, code execution |
| Where it runs | xAI API, Grok Build, Cursor |
One asterisk: Grok 4.7 Fast, the speed-optimized variant, is currently available only inside Cursor and Grok Build — not through the public xAI API. If your workflow is Cline with an xAI key, you get the standard model only.
How good is Grok 4.7 at coding, really?
Grok 4.7 is the 4th-best coding agent stack money can buy right now, behind Claude Fable 5.1, GPT-6 Astra, and Claude Opus 5 — per Artificial Analysis’s independent September 2026 evaluation of models in their native harnesses. That’s a genuine jump: Grok Build with Grok 4.7 at xhigh effort scores 56 on the Coding Agent Index, up from 47 with Grok 4.6.
The component scores (Artificial Analysis, September 2026, Grok Build harness at xhigh):
| Benchmark | Grok 4.6 | Grok 4.7 |
|---|---|---|
| DeepSWE v1.1 | 65% | 73% |
| Terminal-Bench 4.0 | 18% | 33% |
| SWE-Atlas-QnA | 58% | 63% |
| Coding Agent Index | 47 | 56 |
Standalone (outside the Grok Build harness), the model posts 71.0% on DeepSWE v1.1 at high effort and 46.3% on CursorBench 4.0. On the broader Intelligence Index it scores 46 — up 2 points from Grok 4.6, enough to put SpaceXAI in Artificial Analysis’s top 4 AI labs for the first time.
Where it wins and loses is unusually legible, because xAI published a head-to-head table: Grok 4.7 beats Grok 4.6 on all seven benchmarks shown and leads Claude Fable 5.1 on EEBench (enterprise engineering) and the Harvey Legal Agent benchmark. Fable 5.1 at max effort leads Grok 4.7 on CursorBench, Terminal-Bench, GDPval, and HealthBench. Translation for daily coding: on hard, accuracy-critical software-engineering tasks, Fable 5.1 is still the stronger model; Grok 4.7’s wins are in enterprise document-plus-code and legal-agent territory.
What does Grok 4.7 actually cost per task? (The verbosity tax)
The number that matters: $2.73 per completed Intelligence Index task at high effort, up 47% from Grok 4.6’s $1.86 — at identical per-token prices. The model didn’t get more expensive; it got more talkative. Artificial Analysis measured Grok 4.7 at xhigh consuming roughly 81,000 output tokens per Intelligence Index task, versus 36,000 for Grok 4.6 at high and 27,000 for GPT-6 Astra at max. Running the full Intelligence Index took about 200M tokens against a median of 88M across models.
This is the reversal of the Grok 4.5 story. In mid-2026, Grok’s pitch was token efficiency — fewer output tokens per task than Anthropic’s models, multiplying the sticker discount. Grok 4.7 spends tokens to buy benchmark points. The discount survives, but shrunken:
| Cost view | Grok 4.7 | Claude Fable 5.1 | Sticker suggests |
|---|---|---|---|
| Input per MTok | $2 | $10 | Grok 5x cheaper |
| Output per MTok | $6 | $50 | Grok 8.3x cheaper |
| Measured cost per AA Intelligence Index task | $2.73 (high) | $3.76 (max) | Grok 27% cheaper |
For a typical uncached Cursor/Cline session — 200K input, 25K output, the shape we’ve used across our backend cost series — Grok 4.7 runs $0.55, versus $1.30 on Opus 5.5 and $3.25 on Fable 5.1. But if Grok emits 2–3x the output tokens to finish the same task, that 25K becomes 50–75K and the session lands at $0.70–$0.85 — still the cheapest, just not by the margin the rate card promises. And more tokens means more latency and more agent-loop turns, which no rate card shows.
Cache-heavy sessions flip the table. Grok 4.7 charges $0.50 per million cached input tokens; Fable 5.1’s cache reads cost $0.25 per million. A long Cline or Claude Code loop that accumulates 6M cache reads pays $3.00 in cache on Grok versus $1.50 on Fable 5.1 — the budget model is 2x more expensive on the one line item that dominates marathon agentic sessions. (Anthropic separately charges cache writes at $12.50/M for 5-minute TTL; xAI’s docs list cached-input pricing only, so a full apples-to-apples cache comparison isn’t possible from published numbers.) The pattern across vendors — cheap stickers, expensive tasks — is the same one we documented in the margin-collapse re-check: pricing is converging at the task level even when rate cards look far apart.
How do you set up Grok 4.7 in Cursor?
Grok 4.7 is in Cursor’s native model list — no BYOK required. Open Cursor Settings → Models, enable Grok 4.7 (and Grok 4.7 Fast, the Cursor-only speed variant), and select it from the model picker in the agent pane. Usage bills against your Cursor plan’s model quota like any other frontier model.
If you’d rather pay xAI directly and keep Cursor’s flat plan for other models, add an xAI key under Settings → Models → API Keys using the OpenAI-compatible base URL https://api.x.ai/v1 and model grok-4.7. BYOK makes sense once your Grok usage would otherwise burn through Cursor’s included quota — at $2/$6 with Grok’s verbosity, budget roughly $2.70–$3 per heavy agentic task, not the $1–$2 you’d extrapolate from Grok 4.6 habits.
How do you set up Grok 4.7 in Cline?
Cline ships a native xAI provider, so setup is three steps: generate an API key at x.ai, open Cline’s settings, choose xAI (Grok) as the provider, and paste the key (full walkthrough at docs.cline.bot/provider-config/xai-grok). Set the model to grok-4.7. Tool calling — which Cline’s entire edit loop depends on — is supported; the model ships function calling, web search, and code execution natively.
Sanity-check the key and model ID from a terminal before pointing Cline at it:
$ curl https://api.x.ai/v1/chat/completions \
-H "Authorization: Bearer $XAI_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model": "grok-4.7", "messages": [{"role": "user", "content": "Reply with OK"}]}'
A working key returns a standard OpenAI-shaped completion:
{"id":"...","object":"chat.completion","model":"grok-4.7",
"choices":[{"message":{"role":"assistant","content":"OK"}, ...}],
"usage":{"prompt_tokens":11,"completion_tokens":1,...}}
A 404 with “model not found” here almost always means the hyphenated grok-4-7 spelling — the ID uses a period.
One problem you will hit on large repos: the silent 2x price cliff at 200K context. Cline sessions that index a big codebase, or that run long enough to accumulate a large conversation, cross 200K input tokens without any visible warning — and every input token past that line bills at $4/M instead of $2/M. The fix is boring but effective: set Cline’s context-window limit at or below 200K for this provider, start fresh tasks instead of continuing marathon conversations, and keep repo context scoped with .clinerules file-access rules rather than letting the agent read everything. The 500K window is real, but on this rate card it’s a premium feature, not a default.
When should you pick Grok 4.7 — and when not?
Pick Grok 4.7 if you’re cost-first and effort-disciplined; skip it if your work is accuracy-critical or cache-heavy. Specifically:
Choose Grok 4.7 when:
- You’re already on Grok 4.6 — same price, strictly better scores. Cap effort at high; xhigh’s 81K output tokens per task is where the verbosity tax gets ugly.
- Your sessions are short-to-medium, low cache reuse, mid-difficulty — the profile where $0.55–$0.85 per session versus $3.25 on Fable 5.1 compounds fast.
- Your tasks look like EEBench or legal-agent work: enterprise codebases entangled with documents, contracts, compliance text. That’s where it beats Fable 5.1 outright.
Stay on Claude (or Astra) when:
- The task is a hard, multi-file, accuracy-critical refactor. Fable 5.1 leads on CursorBench and Terminal-Bench; a failed $0.85 attempt plus your review time costs more than a $3.25 success. Opus 5.5 at $4/$20 is the price/accuracy middle ground.
- Your sessions are marathon agentic loops dominated by cache reads. At $0.50/M cached versus Fable’s $0.25/M, Grok’s advantage inverts exactly where agentic spend concentrates.
- You need the frontier at any price — GPT-6 Astra runs $7.09 per Coding Agent Index task, about 40% less than Fable 5.1 in Claude Code for the same score, and both outrank Grok Build. Our Astra cost breakdown has that math.
And if $2.73 per task still reads as rent-seeking: open-weight models at $0/token on your own GPU remain the floor. A used 24GB card runs surprisingly capable coding models locally — see runaihome.com’s guide to the best local models by VRAM tier for the hardware math, and aifoss.dev’s Ollama review for the serving stack. For quick agentic work where speed matters more than depth, Grok Code Fast 1 at $0.20/$1.50 is still xAI’s actual budget lane.
FAQ
Is Grok 4.7 better than Claude Fable 5.1 for coding? No, on balance. Fable 5.1 at max effort leads Grok 4.7 on CursorBench, Terminal-Bench, GDPval, and HealthBench; Grok 4.7 leads on EEBench and Harvey Legal Agent. On the Artificial Analysis Coding Agent Index, Grok Build + Grok 4.7 ranks 4th behind Fable 5.1, GPT-6 Astra, and Claude Opus 5. Grok’s case is cost per task ($2.73 vs $3.76), not quality.
Did Grok 4.7’s price change from Grok 4.6? No — $2/$6 per million tokens in/out, $0.50/M cached input, unchanged. But measured cost per completed task rose 47% because the model generates roughly 2.2x more output tokens. Same rates, bigger meter.
What’s the exact model ID for the API?
grok-4.7, with a period — grok-4-7 returns a model-not-found error. Base URL https://api.x.ai/v1, OpenAI-compatible chat completions and responses endpoints.
Can I use Grok 4.7 Fast in Cline?
Not as of September 26, 2026. Grok 4.7 Fast is only available inside Cursor and Grok Build; the public xAI API — which Cline’s xAI provider uses — serves the standard grok-4.7 only.
Does the 500K context window cost extra? Yes, effectively. Input tokens beyond 200K of context bill at double rate ($4/M instead of $2/M). Keep Cline/Cursor sessions under 200K context unless the task genuinely needs more.
Sources
- Grok 4.7 model page — SpaceXAI Docs
- SpaceXAI API pricing and models — x.ai
- Benchmarking Grok 4.7 — Artificial Analysis
- Grok 4.7 pairs coding gains with the same affordable pricing — but high token consumption threatens real-world ROI — VentureBeat
- SpaceXAI Launches Grok 4.7: Low Prices, Heavy Token Use — TechRepublic
- Grok 4.7 Brings Big Coding Upgrades to Challenge Claude AI at Unchanged Pricing — Android Headlines
- Claude Fable 5.1 tops the Artificial Analysis Intelligence Index — Artificial Analysis
- Pricing — Claude Platform Docs (Anthropic)
- xAI Grok provider setup — Cline Docs
Last updated September 26, 2026. Pricing and features change frequently; verify current state on the official pages before committing to a provider.
Was this article helpful?
Thanks for the feedback — it helps improve future articles.
Need hands-on help?
I offer 1-on-1 technical consulting for local AI setup, GPU selection, and AI coding tool configuration — same topics covered on this site.
Book a session — $49 / hour →Know which coding tool is worth paying for
Hands-on comparisons of AI coding assistants and what each one costs to run — including the local-model path. Sent only when something changes. Unsubscribe anytime.