DigestAI news desk
OpenAI board member warns company is not on track to prevent catastrophic AI loss of control OpenAI Unveils GPT‑6 Astra: Record‑Breaking 3D Rendering, Loop‑Transformer Architecture OpenAI launches Agents API beta for long-running cloud agents OpenAI solves Navier-Stokes problem, sparking academic controversy over data use OpenAI Introduces ChatGPT for Financial Services The Waymo effect: AI making research less collaborative RTK Token Savings Debunked: Cost Benchmarks Disagree Ypsilanti Township Residents Protest Nuclear AI Data Center Proposal
Generative AI & Models updated 2 min read

Cognition's SWE-2 Breaks Terminal-Bench 2.1 with 92.8%

Cognition's SWE-2 model has achieved a remarkable 92.8% on the Terminal-Bench 2.1 benchmark, a significant improvement over the 73.0% of its predecessor, DeepSWE 1.1. The model, which has 2.8 trillion total parameters and is a post-training upgrade of the Kimi K3 base, has shown substantial gains in efficiency and performance. Cognition's approach includes using a mix of Multi-Head Attention…

1 source HN 61

Key points

  • Cognition's SWE-2 achieves 92.8% on Terminal-Bench 2.1
  • Model uses 2.8 trillion parameters and post-training on Kimi K3 base
  • Improves efficiency by 15% and boosts throughput by 10-20%
Full story from tokenstead.ai · via Hacker News Open source ↗

Cognition's SWE-2 achieves 92.8 on Terminal-Bench 2.1

tokenstead.ai · 10 September 2026

2.8T total params, 104B active per token (MoE) - the Kimi K3 base with Cognition’s post-training on top, and the first time Cognition has scaled RL into the multi-trillion-parameter regime. The base had already been RL-heavy for agentic coding; Cognition’s pass added another 5 to 6 points on most benchmarks.

-

Serving stack: MoE inference on NVFP4 and FP8 kernels with quantization-aware training; FP8 carries K, Q, V, and score computations in the MLA layers. A draft model retrained with SpecForge gives 15% longer accept lengths, and a prefill delayer lifts TPM per GPU and tokens/sec per request by 10 to 20% (TTFT takes the hit).

Effort levels: mean steps per run 53 (medium), 80 (high), 98 (max), against 127 for SWE-1.7. Medium posts a higher FrontierCode score than SWE-1.7 with 58% fewer turns and 81% lower average cost, and lands its first real edit at a median of step 18 (SWE-1.7: 48).

Benchmarks (Cognition self-reported): FrontierCode 1.1 Main 50.0, DeepSWE 1.1 73.0, Terminal-Bench 2.1 92.8, Terminal-Bench 4.0 27.3. The headline: 50.0 on FrontierCode is one point behind Claude Fable 5.1 (50.9) and 3.3 behind GPT-6 Astra (53.3) - at a claimed 64% lower cost than Fable 5.1 and a quarter of Astra’s. Terminal-Bench 2.1 is the highest number in the published table. The soft spot is Terminal-Bench 4.0, where SWE-2’s 27.3 trails Fable 5.1 (55.8) and GPT-6 Astra (57.9) by a wide margin - long-horizon agentic work is where the gap to the frontier still lives.

Proprietary weights, no local run. Cognition has not published SWE-2 weights, so there is nothing to download and no quant ladder to wait for. It is available today in Devin Desktop and CLI, with rollout on Devin Web and Fusion. Cognition publishes no per-token API for SWE-2, so the cost-per-task comparisons (64% cheaper than Fable 5.1 at FrontierCode parity) are the pricing surface, not a $/1M rate card. Every figure here is Cognition’s own number, pending independent replication.

  • 2800.0B
  • proprietary
  • 🇺🇸 USA
  • Sep 2026

What people are building with SWE-2

Real demos from X

Benchmark scores

Vendor-reported - from the developer's own model card / tech report

Ran this model on your own hardware? Join free and add your measured tok/s to the community numbers.

Or run it in the cloud

Inference cost over time

Data accumulates from the first daily sync - longer ranges populate over time. Prices come from OpenRouter snapshots, not a historical API.

This text was published by tokenstead.ai . It is reproduced here with attribution so you can read it in full; the rights remain with the publisher. Read it at the source ↗

Coverage and discussion

1 source
Topics · follow one to build your own front page

The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us.

Comments

via GitHub Discussions

Related stories