Cognition shipped SWE-2 on September 10, 2026 — yesterday — claiming its new coding model scores within one benchmark point of Anthropic’s Fable 5.1 on FrontierCode 1.1 Main while costing 64% less to run. To sweeten the launch, Cognition is offering SWE-2 free for one month across every Devin Pro ($20), Max, and Teams subscription. The cost story is real. The “matches frontier” claim is selective. Both things are true.
The Benchmark Numbers Worth Knowing
On FrontierCode 1.1 Main — the benchmark Cognition leads with — SWE-2 scores 50.0% against Fable 5.1’s 50.9% and GPT-6 Astra’s 53.3%. For everyday coding tasks, that gap is noise. The model also gets to its first real code edit in a median of 18 steps, down from 48 steps in SWE-1.7, and uses 58% fewer total turns to complete tasks at the medium effort level. On Terminal-Bench 2.1 — shorter agentic tasks — SWE-2 actually outperforms both at 92.8% vs 91.4% for Fable 5.1 and 89.9% for GPT-6 Astra.
The cost math is straightforward. Cognition trains its own model rather than routing through Anthropic or OpenAI, so the $20/month Devin Pro plan can bundle near-frontier coding intelligence without paying retail per-token rates. At 64% lower compute cost than Fable 5.1, and roughly a quarter of GPT-6 Astra’s cost, SWE-2 is a credible offer for developers doing high-volume coding work on a budget.
Related: Cognition Raises $2B at $48B Valuation: What It Means for Developers
The SWE-2 Gap Cognition Doesn’t Headline
Terminal-Bench 4 tests long-horizon agentic work — multi-file changes, extended planning sessions, the kind of tasks that run for hours. On that benchmark, SWE-2 scores 27.3%. Fable 5.1 scores 55.8%. GPT-6 Astra scores 57.9%. Even DeepSeek v4.1 Flash, released the same day, posts 31.2% — outperforming SWE-2 on the harder metric. This is what the Hacker News thread at 402 upvotes is primarily arguing about.
The contrast is stark enough to matter. A model that scores 92.8% on Terminal-Bench 2.1 and 27.3% on Terminal-Bench 4 is very good at short, well-scoped tasks and significantly weaker at the complex agentic runs that demanding engineering workflows require. Cognition does not dispute this — the company simply doesn’t feature it in the announcement. Developers doing routine PR reviews, small feature implementations, and bug fixes will find SWE-2 competitive. Those running overnight agentic sessions on large codebases will want Fable 5.1.
How Cognition Built SWE-2
SWE-2 is post-trained from Kimi K3, a 2.8 trillion-parameter model from Moonshot AI that was pre-trained on agentic coding tasks. The key innovation is in how Cognition applied reinforcement learning: instead of tuning each effort level (medium, high, max) separately, they used a linear cost penalty formula that trains all three levels in a single run while optimizing the full cost-performance curve. According to their benchmark analysis at BenchLM, this approach is mathematically distinct from simply capping a high-capability model at lower cost tiers.
The result is a 1M-context model with FP8/FP4 inference optimizations that achieves similar throughput to SWE-1.7 despite running on a base model three times larger. That is the infrastructure bet paying off — not raw parameter count, but inference efficiency built for the cost floor Cognition is targeting.
The Catch Most Developers Will Hit
SWE-2 has no standalone API, no open weights, and no access outside Devin’s product suite. If your workflow involves pulling a model through OpenRouter, building a custom coding agent harness, or integrating into CI/CD pipelines directly, SWE-2 is off the table — at least for now. The top community complaint on the launch thread was blunt: “I don’t want to use your CLI. I already have my own harnesses.” Additionally, Cognition has not announced a standalone API, and no open weights are planned.
Cognition’s bet is that the $20 plan’s value proposition is strong enough that developers shift to Devin’s ecosystem rather than demanding an API. That’s a reasonable bet for individual developers and small teams. However, it is a harder sell for engineering teams with established tooling investments — and precisely the gap that open-weight alternatives continue to exploit.
Key Takeaways
- SWE-2 matches Fable 5.1 within one FrontierCode point at 64% lower compute cost — the cost advantage is real for short-to-medium coding tasks
- Terminal-Bench 4 scores tell a different story: 27.3% for SWE-2 vs 55.8% for Fable 5.1 — a significant gap for complex, long-horizon engineering workflows
- No standalone API and no open weights mean the cost savings only apply if you’re running Devin’s managed environment
- The Pareto RL training approach is a genuine technical contribution — not just a capability-capped version of a larger model
- Free for one month on the $20 Devin Pro plan — worth testing for routine coding work if you have not committed to Copilot or Claude Code













