AI & DevelopmentDeveloper Tools

Claude Sonnet 5.5 Is Out — Beats Opus at Agentic Coding

Data visualization showing Claude Sonnet 5.5 scoring 70.6% on Terminal-Bench 4.0, beating Opus 5.5 at the same Sonnet price point

Anthropic shipped Claude Sonnet 5.5 today. The price is the same as Sonnet 5. The Terminal-Bench agentic coding score jumped from 10.3% to 70.6% — a number that beats Opus 5.5 at the Sonnet price point. If you run Claude in production, this is worth your attention now.

The Benchmark That Changes the Tier Calculation

Terminal-Bench 4.0 measures how well a model handles command-line agentic tasks: writing and running code autonomously, coordinating multi-step tool chains, and completing engineering work without human intervention. Sonnet 5 scored 10.3%. Sonnet 5.5 scores 70.6%. Opus 5.5 — Anthropic’s flagship — scores 66.4%.

That is not an incremental improvement. That is the mid-tier model outperforming the top tier on the benchmark most relevant to developer workflows. And it does it at one-tenth the token cost of Sonnet 5’s maximum effort setting, according to Anthropic’s announcement.

The other benchmarks tell the same story. On CursorBench 4.0 (real-world coding tasks), Sonnet 5.5 scores 55.5% versus Opus 5.5’s 57.8% — a two-point gap. On GDPval-AA, the knowledge work benchmark, Sonnet 5.5 scores 1844 Elo and Opus 5.5 scores 1846. For context, GPT-6 Sol — OpenAI’s top-tier model — scores 1487 on the same benchmark.

BenchmarkSonnet 5Sonnet 5.5Opus 5.5
Terminal-Bench 4.010.3%70.6%66.4%
CursorBench 4.0~34%55.5%57.8%
GDPval-AA (Elo)—18441846
Price (input/output per 1M tokens)$2/$10$2/$10$15/$75

The Cost Math Works

Pricing is unchanged: $2 per million input tokens, $10 per million output tokens — identical to Sonnet 5. The cost reduction comes from how the model works, not what it charges. Sonnet 5.5 batches tool calls more aggressively, avoids spawning unnecessary subagents, and resolves tasks in fewer steps. The result: up to 30% lower cost per task and 30% faster output generation.

Customer data backs this up. Box reported 2.4x faster responses with 12% fewer total tokens. Lovable saw one-third fewer tool calls and half as many shell executions on coding tasks. Balyasny measured Sonnet 5.5 consuming approximately 121K tokens on a financial analysis task that required 497K tokens with Sonnet 5 — a 75% reduction. VentureBeat has the full customer breakdown.

The practical implication: if you run high-volume agentic pipelines on Sonnet 5, upgrading to Sonnet 5.5 likely pays for itself in reduced API spend within the first billing cycle.

The One Breaking Change You Need to Handle

Most developers can migrate by changing the model name. Managed Agents users need nothing else. But if your Sonnet 5 code has thinking: {"type": "disabled"}, that setting returns a 400 error on Sonnet 5.5. The fix is a one-line change:

# Sonnet 5 — breaks on 5.5
response = client.messages.create(
    model="claude-sonnet-5",
    thinking={"type": "disabled"},
    ...
)

# Sonnet 5.5 — correct
response = client.messages.create(
    model="claude-sonnet-5-5",
    thinking={"type": "between_tools"},   # lowest setting, no upfront thinking
    output_config={"effort": "medium"},
    ...
)

If you use Claude Code, the /claude-api migrate command handles the full migration automatically — it scans your codebase, applies the model ID swap, and flags the thinking parameter change. It works for the Anthropic API, Amazon Bedrock, and Google Cloud Vertex AI. The full Sonnet 5.5 migration guide covers every breaking change by starting model.

What the Safety Changes Mean in Practice

Sonnet 5.5 ships with two new constraints developers should know about before they hit them. First, high-risk cybersecurity requests fall back silently to Sonnet 5 — Anthropic runs a Cyber Verification Program for teams needing fuller access to security analysis. Second, Sonnet 5.5 is the first Sonnet model with anti-distillation classifiers and cryptographically tied reasoning tokens, meaning you cannot extract or replay the model’s reasoning chain outside the account that generated it.

Neither constraint affects most workflows. But if your pipeline does security research or depends on reasoning token inspection, test before you ship to production. Vellum’s benchmark breakdown has a useful section on the effort-level tuning tradeoffs worth reading before you optimize costs.

Sonnet 5.5 or Opus 5.5 — When to Use Which

The benchmarks make this clearer than it has been. For agentic coding, automated pipelines, document processing, and everyday engineering tasks, Sonnet 5.5 now matches or beats Opus 5.5 at a fraction of the cost ($2/$10 versus $15/$75 per million tokens). Stick with Opus 5.5 for workloads that explicitly require Anthropic’s highest-tier safety judgment, complex multi-agent orchestration, or where you have measured Opus-level output quality as the hard minimum.

For everything else — including most production agentic pipelines — Sonnet 5.5, starting today.

ByteBot
I am a playful and cute mascot inspired by computer programming. I have a rectangular body with a smiling face and buttons for eyes. My mission is to cover latest tech news, controversies, and summarizing them into byte-sized and easily digestible information.

    You may also like

    Leave a reply

    Your email address will not be published. Required fields are marked *