Grok 4.7: The Benchmarks, the Price, and the Gap It Still Has to Close
xAI released Grok 4.7 on Monday, after a launch window that slipped at least five times since late July. It is the company’s strongest model yet: 2.1 trillion parameters, roughly 40% up on Grok 4.6, trained with a longer reinforcement-learning run and supplemented with engineering data from SpaceX — Starlink telemetry, manufacturing records, failure logs.
It is also, on paper, one of the cheapest ways to buy frontier-adjacent intelligence. So the interesting question is not whether Grok 4.7 is good. It is whether good enough at this price is the trade most teams should now be making.
What Grok 4.7 actually is
- 500,000-token context window with a May 2026 knowledge cutoff.
- Text and image input, text output.
- Four reasoning levels — low, medium, high and xhigh — defaulting to high.
- Native tool use: function calling, web search, X search and code execution.
- Available immediately in the Grok app, Cursor, Grok Build and the xAI API — no waitlist.
The headline engineering change is not the parameter count. xAI says the longer RL run targets tasks that take hours rather than seconds, and the model’s ability to check its own work before handing it back. That is the right thing to be working on: for business use, an agent you have to supervise closely is not much cheaper than doing the task yourself.
The benchmarks: real gains, still second
xAI’s own published numbers show large step-ups over Grok 4.6:
- CursorBench 4.0: 46.3% (from 40.4%)
- DeepSWE v1.1: 71.0% (from 65.2%)
- Terminal-Bench 4.0: 38.0% (from 20.3%) — vendor harness
- EEBench (electrical engineering): 64.0%
- Harvey Legal Agent Benchmark: 19.6% (from 15.8%)
- HealthBench Professional: 56.7% (from 48.5%)
Those are vendor-run evaluations, so treat them as direction rather than gospel. Independent measurement is more sober. On Artificial Analysis’s Intelligence Index (v4.3.2), a composite of ten benchmarks, Grok 4.7 scores 46 and sits mid-pack: Claude Fable 5.1 and GPT-6 both score 53. On GDPval, which rates economically valuable knowledge work on an Elo scale, Grok 4.7 posts 1,695 against Fable 5.1’s 1,735. On AA-Briefcase, which strings a multi-hour office task together end to end, it manages 1,657 to Fable 5.1’s 1,678.
The gap widens where it matters most for production: agentic coding. Independent runs of Terminal-Bench 4.0 have placed Grok 4.7 well behind GPT-6 Astra and Fable 5.1 — and behind at least one cheap challenger model. Vendor harnesses flatter; third-party harnesses don’t.
The pricing is the actual story
Grok 4.7 costs $2 per million input tokens and $6 per million output tokens for prompts under 200,000 tokens, with cached input at $0.50 and every rate doubling above that threshold. That is identical to Grok 4.6 — a rare thing in a market where upgrades have mostly arrived with a price rise.
Set against the current Western frontier, the difference is stark. GPT-5.6 Sol lists at $4 input / $20 output; Claude Fable 5.1 at $10 input / $50 output. On list price, Grok 4.7 undercuts them by roughly 50% to 80% — sitting closer to the pricing of Chinese open-weight models than to its American rivals.
The honest caveat: cheaper per token only matters if fewer tokens finish the job. Cursor’s benchmark plots accuracy against total cost per task, not token price — and there Grok 4.7 lands mid-pack. Cheaper per task than several rivals, but still short of Fable 5.1, which wins across the price curve.
Cheaper per token only matters if fewer tokens finish the job.
The competition, and what catching up looks like
Three competitors define the field right now. Claude Fable 5.1 leads the knowledge-work evaluations and the cost-efficiency curve. GPT-6 Astra leads agentic coding and electrical engineering. And underneath them, cheaper models — DeepSeek’s Flash tier, the open-weight Kimi K3 — are closing on the same benchmarks at lower prices. That lower-middle is exactly the squeeze xAI is trying to occupy.
For Grok to move from “credible and cheap” to “default choice”, three things have to happen:
- Close the agentic-coding gap. The multi-hour autonomy tests are where enterprises actually deploy, and that is where the deficit is widest.
- Fix multimodality. Musk has said publicly that it still needs work — a real constraint for document and video-heavy workflows.
- Make the release cadence boring. Five delays on a flagship is a signal procurement teams notice. Roadmap reliability is part of the product.
Credit where it is due
It is worth saying how far this line has come. Grok 4.5 launched in July and finished third behind the frontier pair. Two months later, Grok 4.7 is second on several professional evaluations, has grown its base model by 40%, has held its price flat, and ships with a safety stack and native support for xAI’s agent harness. Two meaningful releases in two months, with a stated path onward, is genuine momentum — and the company’s own founder, notably, is now framing the gap more modestly than the launch marketing does.
What it means if you are choosing a model
Practical read. If your workload is document-heavy knowledge work or cost-sensitive batch processing, Grok 4.7 has earned a place in your next bake-off — at $2/$6 the arithmetic genuinely changes. If your workload is autonomous coding or long multi-step agents, the leaders still earn their premium. Either way, list price is not the number that matters: measure cost per completed task on your own data, including retries and human correction. That number, not a benchmark table, is what decides.
Sources and caveats: figures are as published on 22 September 2026, from xAI’s model documentation and independent trackers including Artificial Analysis. Vendor-run results are marked as such. This market moves weekly — check current pricing and benchmarks before committing.
Work with KREO Studio
AI engineering, data science and design architecture, from Plymouth to the wider UK.
