MIT and Sakana AI’s SIFT framework cuts the cost of judging self-improving coding agents

1 hour ago 1



Coding agents that rewrite their own code have a quiet problem. Every time they tweak themselves, someone has to check whether the tweak helped. Researchers from MIT and Sakana AI think they have found a cheaper way to do that checking. Their framework, called SIFT, short for Self-Improvement via Fast Tree-search, posted a full score of 35.1% on the Polyglot coding benchmark after just 30 expansion steps. The comparison point matters. The earlier Darwin Gödel Machine (DGM) approach reached 30.7% on the same benchmark, but only after 80 nodes. How SIFT picks winners without running the full gauntlet Recursive self-improving coding agents work like a writer revising their own drafts. The agent proposes a change to its own source code, hoping the new version performs better on real tasks. The catch is verification. Testing every proposed patch against a full benchmark burns serious compute, and the bill compounds quickly when an agent generates many candidates. SIFT sidesteps much of that cost with a referee. Instead of benchmarking every change, it asks a large language model to compare two candidate modifications and say which one looks better. Those head-to-head verdicts are then c...

Read Entire Article