xAI shipped Grok 4.7 on September 21 and pushed it to GitHub Copilot across every subscription plan the same day. The 2.1 trillion parameter model comes with a 500K context window, improved self-verification, and extended reinforcement learning aimed at multi-hour agentic tasks. The marketing frames it as xAI’s “most capable coding model yet.” Independent benchmarks tell a more complicated story.
Where Grok 4.7 Actually Stands
On the Artificial Analysis Intelligence Index (v4.3.2), which aggregates ten benchmarks, Grok 4.7 scores 46. Claude Fable 5.1 and GPT-6 Astra both score 53 — a 13% gap. That gap narrows for specific workloads, but it does not disappear. On Terminal-Bench 4.0, the benchmark most directly relevant to agentic coding — xAI’s stated focus — Grok 4.7 lands at roughly 26–38% depending on configuration. GPT-6 Astra reaches 60%. Claude Fable 5.1 reaches 55%. DeepSeek V4.1 Flash, at a fraction of the cost, competes directly with Grok 4.7 on this same benchmark.
CursorBench 4.0 shows a similar picture: Grok 4.7 at 46.3% versus Claude Fable 5.1 at 51.8%. According to The Decoder’s benchmark analysis, the gap is consistent and most pronounced in agentic coding — precisely the use case xAI is targeting. Calling Grok 4.7 the “most capable” anything requires a very selective reading of the data.
Related: GPT-6 Astra: Async Tools, Mid-Turn Steering, and 1M Context
The Hidden Cost Trap in Grok 4.7
xAI kept pricing unchanged from Grok 4.6: $2 per million input tokens and $6 per million output tokens. Compared to Claude Fable 5.1, that is anywhere from one-fifth to one-eighth of the cost on a per-token basis. However, per-token pricing is not the same as per-task cost — and Grok 4.7 burns through tokens at a rate that undercuts the pricing advantage at scale.
Independent testing by Artificial Analysis found Grok 4.7 xHigh using roughly 81,000 output tokens per task. GPT-6 Astra averages around 27,000 — meaning Grok 4.7 generates nearly three times the output for equivalent work. Even compared to its predecessor, Grok 4.7 uses 125% more output tokens than Grok 4.6. VentureBeat’s analysis is direct: “A model charging less per token can still be more expensive on a finished workload.” Before routing production traffic to Grok 4.7, benchmark it on representative internal jobs and track cost per successful completion — not API rate sheets.
What Grok 4.7 Genuinely Improved
The benchmark gap versus Claude and GPT-6 should not obscure the real progress over Grok 4.6. On DeepSWE v1.1 at high effort, Grok 4.7 scores 71–73%, up from 65.2%. On CursorBench 4.0, it moved from 40.4% to 46.3%. On AA-Briefcase, the agentic knowledge work benchmark, it posts 1657 Elo — an improvement of 111 points over its predecessor, placing it just behind Claude Opus 5 and Fable 5.1. Hallucination rate dropped from 34% to 29.3%, and the 500K context window is a meaningful expansion.
For teams already running Grok 4.6, the upgrade is clearly worth taking. The question is not whether Grok 4.7 is better than 4.6 — it is. The question is whether it competes with Claude Fable 5.1 or GPT-6 Astra for demanding agentic workloads, and the answer there is more nuanced.
GitHub Copilot: Grok 4.7 Rolls Out Across All Plans
According to the GitHub Changelog, Grok 4.7 landed in GitHub Copilot on September 21 — the same day as the model launch. It is available on Pro, Pro+, Max, Business, and Enterprise plans via VS Code, Visual Studio, JetBrains, Xcode, Eclipse, the CLI, and the Copilot mobile app. Billing runs at provider list pricing under usage-based billing, and the rollout is gradual, so not every user will see it immediately.
Enterprise and Business administrators can disable access through model policies in Copilot settings — new models are enabled by default. The broad rollout puts Grok 4.7 in front of millions of developers without additional setup, but the usage-based billing means organizations should monitor token consumption before treating it as a cost-saving default.
Related: Claude Managed Agents: Budget Caps, Advisor, and Geo-Pinned Inference
Key Takeaways
- Grok 4.7 scores 46 on the Artificial Analysis Intelligence Index against 53 for both Claude Fable 5.1 and GPT-6 Astra — the gap is consistent and most significant on Terminal-Bench 4.0 agentic coding tasks.
- Per-token pricing does not equal per-task cost: Grok 4.7 xHigh uses ~81K output tokens per task versus ~27K for GPT-6 Astra, eroding the cost advantage at production scale.
- Real improvements over Grok 4.6 are genuine — DeepSWE up to 71–73%, AA-Briefcase +111 Elo, hallucination rate down to 29.3% — the comparison baseline matters.
- GitHub Copilot integration is live across all plans and 7 IDEs; enterprise admins should review model policies and track token consumption before defaulting to Grok 4.7.
- Grok 4.7 is a legitimate option for cost-sensitive, lower-complexity workloads — but benchmark real-world task cost before committing production agentic workflows.













