Cognition, the startup behind the Devin AI software engineer, launched SWE-2 on Sept. 10, a coding model built on Moonshot AI's Kimi K3 base model, according to the company's announcement. It is available now through Devin's desktop, command line, web and Fusion interfaces.
On the FrontierCode 1.1 Main benchmark, SWE-2 scored 50.0%, within one point of Fable 5.1's 50.9% while costing 64% less per task, and within a few points of GPT-6 Astra's 53.3% at about a quarter of the cost, Cognition said.
Compared with Cognition's earlier SWE-1.7 model, SWE-2 reached its first code edit after a median of 18 steps, versus 48 for the prior model, and used 58% fewer turns while costing 81% less to run, the company said. Cognition described the release as the first scaling of reinforcement learning to what it called the multi-trillion-parameter regime.
SWE-2 runs on Kimi K3, built by Moonshot AI, the same company Anthropic accused hours later of routing a mass Claude-distillation campaign through the Chinese military, in a report published the same day.
The real fight in coding agents right now is cost per completed task, not raw capability, and Cognition is explicitly pricing SWE-2 as the budget option against the two labs it benchmarks itself against. Teams picking a coding agent should run their own cost-per-merged-pull-request numbers rather than trust a single benchmark table.