Anthropic dropped Haiku 5.5 on October 7 and the pricing change is substantial. Input tokens for short prompts now cost $0.10 per million — down from $1.00 in Haiku 4.5, a 90% reduction at that tier. Average cost savings across prompt lengths land at roughly 75%. The context window expanded from 200,000 to 1,000,000 tokens, and the model now supports adjustable effort — a reasoning dial that was previously a Sonnet-and-above feature. At $0.10/M, Haiku 5.5 matches the cheapest models on the market while outperforming them on agentic benchmarks. If you are running high-volume AI workflows, your infrastructure costs just changed.
What Changed: The Pricing Breakdown
The old Haiku 4.5 pricing was the obvious problem: $1.00 per million input tokens and $5.00 per million output tokens placed it above every major competitor in the small-model tier. GPT-5.4 Mini charges around $0.15/M; Gemini 3.8 Flash charges $0.75/M. Haiku was doing something different — and not in a good way for developers watching their bills.
Haiku 5.5 fixes this. The tiered structure is built around prompt length:
| Tier | Input | Output | Cache Reads |
|---|---|---|---|
| Short prompts (≤100K tokens) | $0.10/M | $0.50/M | $0.01/M |
| Long context (>100K tokens) | $0.50/M | $2.50/M | $0.05/M |
The short-prompt tier is where most production workloads live — classification, extraction, routing, summarization. At $0.10/M, this is now among the cheapest compute available for these tasks. Cache reads at $0.01/M make repeated-context workflows essentially free at scale.
The 1M Token Context Window
Context window expansion from 200K to 1M tokens is not a spec bump. It changes the architecture decisions available to you. In multi-agent systems, the orchestrator model — typically Sonnet or Opus — sees large inputs: full codebases, long transcripts, enterprise document collections. Previously, Haiku’s 200K cap meant the subagent had to work with truncated or chunked context. With 1M tokens, Haiku 5.5 can ingest the same inputs its orchestrator sees.
Practical applications this unlocks: full codebase classification runs, long-document Q&A where the answer requires reading the entire document, batch transcript analysis, and any RAG pipeline that previously had to chunk aggressively to stay under token limits.
Adjustable Effort: First for the Haiku Line
Haiku 5.5 is the first Haiku model with an effort setting. The API parameter accepts low, medium (default), high, and max. Setting effort low gives you faster, cheaper responses suited for straightforward classification or routing. Setting it high or max engages deeper reasoning for complex extraction or analysis tasks.
This collapses a decision that used to require switching models. Instead of routing classification requests to Haiku and reasoning-heavy requests to Sonnet, you can keep everything on Haiku 5.5 and adjust the effort level per call. For teams already comfortable tuning prompts, this adds a cost optimization lever they did not have before.
The Performance Numbers
The benchmark improvements between Haiku 4.5 and Haiku 5.5 are not incremental. On OSWorld 2.1, which tests agentic computer-use performance, Haiku 5.5 scores 72.4% compared to 15.7% for Haiku 4.5. On Humanity’s Last Exam with tools, it scores 57.4% versus 18.7%. These are generational jumps, not point improvements.
Against competitors on agentic benchmarks, Haiku 5.5 beats Gemini 3.8 Flash on Terminal-Bench 4.0 — 39.2% versus 19.1% — and Gemini 3.8 Flash costs 7.5x more. The assumption that cheap models underperform does not hold at this price point anymore.
Customer-reported performance tracks with the benchmark data. Asana reported 2.5x faster inference. Box reported an 11-point quality improvement at half the latency. HubSpot scored 92.8% across three CRM evaluation runs. AlphaSense saw document Q&A accuracy improve from 0.76 to 0.84.
How to Switch
Haiku 5.5 is a drop-in replacement. The only required change is the model string:
// Before
model: "claude-haiku-4-5"
// After
model: "claude-haiku-5-5"
The model is available on the Anthropic API, Amazon Bedrock, Google Cloud Vertex AI, and Microsoft Azure. No API version bumps required. Cursor’s documentation for Haiku 5.5 already reflects the model ID. Anthropic recommends benchmarking your specific workload at each effort level before a full rollout, and tracking input, output, cached, and reasoning tokens separately to understand your actual cost profile.
The Broader Picture
Haiku 4.5 at $1/M was the outlier in a market where every other major provider charged $0.10–$0.15/M. Anthropic was not competing on price — and developers noticed. Haiku 5.5 closes that gap. But the broader trend is more significant: this is the third substantial price cut in the small-model tier in 2026 alone. Inference is getting cheaper fast, and the floor keeps dropping.
For developers, the practical implication is that workflows designed to minimize model calls are worth revisiting. More calls, cheaper per call, with adjustable reasoning depth — the cost calculus has shifted.
The Bottom Line
Use Haiku 5.5 for classification, extraction, routing, summarization, customer support automation, subagent roles in multi-agent systems, and any RAG pipeline needing broad context without paying Sonnet prices. The adjustable effort setting gives you a cost dial within those workloads. For complex reasoning, planning, and multi-step tasks where quality is paramount, Sonnet 5.5 remains the right choice. But the gap between “budget model” and “capable model” just got meaningfully narrower.













