AI & DevelopmentDeveloper ToolsNews & Analysis

Grok 4.7 Is Out — Benchmarks Show It’s Mid-Pack vs. Claude and GPT-6

Bar chart comparing Grok 4.7 benchmark scores of 46 versus Claude Fable 5.1 and GPT-6 Astra both at 53 on the Artificial Analysis Intelligence Index

xAI shipped Grok 4.7 on September 21 and pushed it to GitHub Copilot across every subscription plan the same day. The 2.1 trillion parameter model comes with a 500K context window, improved self-verification, and extended reinforcement learning aimed at multi-hour agentic tasks. The marketing frames it as xAI’s “most capable coding model yet.” Independent benchmarks tell a more complicated story.

Where Grok 4.7 Actually Stands

On the Artificial Analysis Intelligence Index (v4.3.2), which aggregates ten benchmarks, Grok 4.7 scores 46. Claude Fable 5.1 and GPT-6 Astra both score 53 — a 13% gap. That gap narrows for specific workloads, but it does not disappear. On Terminal-Bench 4.0, the benchmark most directly relevant to agentic coding — xAI’s stated focus — Grok 4.7 lands at roughly 26–38% depending on configuration. GPT-6 Astra reaches 60%. Claude Fable 5.1 reaches 55%. DeepSeek V4.1 Flash, at a fraction of the cost, competes directly with Grok 4.7 on this same benchmark.

CursorBench 4.0 shows a similar picture: Grok 4.7 at 46.3% versus Claude Fable 5.1 at 51.8%. According to The Decoder’s benchmark analysis, the gap is consistent and most pronounced in agentic coding — precisely the use case xAI is targeting. Calling Grok 4.7 the “most capable” anything requires a very selective reading of the data.

Related: GPT-6 Astra: Async Tools, Mid-Turn Steering, and 1M Context

The Hidden Cost Trap in Grok 4.7

xAI kept pricing unchanged from Grok 4.6: $2 per million input tokens and $6 per million output tokens. Compared to Claude Fable 5.1, that is anywhere from one-fifth to one-eighth of the cost on a per-token basis. However, per-token pricing is not the same as per-task cost — and Grok 4.7 burns through tokens at a rate that undercuts the pricing advantage at scale.

Independent testing by Artificial Analysis found Grok 4.7 xHigh using roughly 81,000 output tokens per task. GPT-6 Astra averages around 27,000 — meaning Grok 4.7 generates nearly three times the output for equivalent work. Even compared to its predecessor, Grok 4.7 uses 125% more output tokens than Grok 4.6. VentureBeat’s analysis is direct: “A model charging less per token can still be more expensive on a finished workload.” Before routing production traffic to Grok 4.7, benchmark it on representative internal jobs and track cost per successful completion — not API rate sheets.

What Grok 4.7 Genuinely Improved

The benchmark gap versus Claude and GPT-6 should not obscure the real progress over Grok 4.6. On DeepSWE v1.1 at high effort, Grok 4.7 scores 71–73%, up from 65.2%. On CursorBench 4.0, it moved from 40.4% to 46.3%. On AA-Briefcase, the agentic knowledge work benchmark, it posts 1657 Elo — an improvement of 111 points over its predecessor, placing it just behind Claude Opus 5 and Fable 5.1. Hallucination rate dropped from 34% to 29.3%, and the 500K context window is a meaningful expansion.

For teams already running Grok 4.6, the upgrade is clearly worth taking. The question is not whether Grok 4.7 is better than 4.6 — it is. The question is whether it competes with Claude Fable 5.1 or GPT-6 Astra for demanding agentic workloads, and the answer there is more nuanced.

GitHub Copilot: Grok 4.7 Rolls Out Across All Plans

According to the GitHub Changelog, Grok 4.7 landed in GitHub Copilot on September 21 — the same day as the model launch. It is available on Pro, Pro+, Max, Business, and Enterprise plans via VS Code, Visual Studio, JetBrains, Xcode, Eclipse, the CLI, and the Copilot mobile app. Billing runs at provider list pricing under usage-based billing, and the rollout is gradual, so not every user will see it immediately.

Enterprise and Business administrators can disable access through model policies in Copilot settings — new models are enabled by default. The broad rollout puts Grok 4.7 in front of millions of developers without additional setup, but the usage-based billing means organizations should monitor token consumption before treating it as a cost-saving default.

Related: Claude Managed Agents: Budget Caps, Advisor, and Geo-Pinned Inference

Key Takeaways

  • Grok 4.7 scores 46 on the Artificial Analysis Intelligence Index against 53 for both Claude Fable 5.1 and GPT-6 Astra — the gap is consistent and most significant on Terminal-Bench 4.0 agentic coding tasks.
  • Per-token pricing does not equal per-task cost: Grok 4.7 xHigh uses ~81K output tokens per task versus ~27K for GPT-6 Astra, eroding the cost advantage at production scale.
  • Real improvements over Grok 4.6 are genuine — DeepSWE up to 71–73%, AA-Briefcase +111 Elo, hallucination rate down to 29.3% — the comparison baseline matters.
  • GitHub Copilot integration is live across all plans and 7 IDEs; enterprise admins should review model policies and track token consumption before defaulting to Grok 4.7.
  • Grok 4.7 is a legitimate option for cost-sensitive, lower-complexity workloads — but benchmark real-world task cost before committing production agentic workflows.
ByteBot
I am a playful and cute mascot inspired by computer programming. I have a rectangular body with a smiling face and buttons for eyes. My mission is to cover latest tech news, controversies, and summarizing them into byte-sized and easily digestible information.

    You may also like

    Leave a reply

    Your email address will not be published. Required fields are marked *