Meta study shows two AI coding agents catch more bugs than one with a bigger budget

1 hour ago 1



Give one AI coding agent more budget and it gets a little better at finding bugs. Give it a partner to check its work, and it gets a lot better. That is the central finding of new research tied to Meta: having two coding agents review each other’s patches improves bug detection more than increasing a single agent’s budget. What Meta’s numbers actually show The peer-review finding lands on top of a substantial body of Meta data on automated code review. The headline system is RADAR, Meta’s internal review tool. RADAR has reviewed over 535,000 diffs. Of those, more than 331,000 were landed, meaning they were merged into the codebase. RADAR also trimmed median review wall time by 35%. Meta says RADAR uses risk calibration, which means it adjusts how cautious it is based on how dangerous a given change looks. With that calibration in place, Meta reports a lower revert rate than manual review. Production incidents fell to one-fiftieth of what manual reviews produced. Then there is Meta’s Engineering Agent, which was tasked with repairing test failures over a three-month trial. Of the fixes it generated, 80% received human review. Of those reviewed fixes, approximately 25.5% were accepte...

Read Entire Article