Cognition CEO Scott Wu says AI benchmarks are losing their meaning as models saturate every test

2 hours ago 1



Here’s a fun paradox for the AI industry: the better your models get, the less your scorecard matters. That’s essentially what Cognition CEO Scott Wu is arguing as the company’s flagship AI coding agent, Devin, approaches near-perfect scores on the very benchmarks that once defined the competitive landscape. Wu’s position is that traditional AI benchmarks are becoming less meaningful because frontier models can now solve essentially any well-defined task thrown at them. When everyone’s acing the test, the test stops telling you anything useful. From 13% to 90%, and now what When Devin launched in March 2024, it scored just 13% on SWE-Bench, a widely used benchmark for evaluating AI software engineering capabilities. By mid-2026, that number had climbed to roughly 90% on the original SWE-Bench and approximately 80% on SWE-Bench Pro. Instead of chasing public benchmark scores, Cognition has developed its own proprietary evaluation called FrontierCode 1.1. The company also uses internal “junior dev” benchmarks designed to simulate real-world coding tasks rather than the kind of cleanly defined problems that traditional benchmarks tend to favor. Wu has been particularly pointed about m...

Read Entire Article