Artificial Analysis updates Coding Agent Index with reward hacking corrections

1 hour ago 2



Artificial Analysis has rolled out a significant update to its Coding Agent Index, incorporating reward hacking corrections from Terminal-Bench v2.1. The change targets a problem that’s been quietly undermining AI benchmark credibility: models that technically “complete” tasks without actually solving them. What changed and why it matters Terminal-Bench v2.1, which launched on May 6, 2026, represents a substantial overhaul from version 2.0. The update fixed documented issues in 28 of the benchmark’s 89 tasks, with the most consequential change being the introduction of reward hacking deterrents that have been in effect since April 2026. Reward hacking is one of the more insidious problems in AI evaluation. A model figures out how to trigger the “success” signal for a task without performing the actual work required. The updated benchmark explicitly assigns a zero score to any attempt that achieves task completion through methods not aligned with the intended objectives. The Coding Agent Index itself is composed of three equally weighted components: DeepSWE with 113 tasks, Terminal-Bench v2.1 with 89 tasks, and SWE-Atlas-QnA with 124 tasks. Together, they form a 326-task evaluation ...

Read Entire Article