Vals AI built a test to see if AI models can create their own successors, and the results are fascinating

1 week ago 5



If you’ve ever wondered how close we are to AI systems that can design better versions of themselves, Vals AI just built the scoreboard. The company’s Recursive Self-Improvement Index, or RSI Index, attempts to quantify something that until now has lived mostly in the realm of theoretical worry: how capable are today’s frontier models at conducting the research and development needed to build their successors? The top-performing model on the index, Claude Fable 5.1, scores 35.03%. Which sounds low until you understand the scale. How the RSI Index actually works Vals AI designed the index around five specific tasks that mirror the real workflow of AI development. Models are evaluated under fixed compute and time constraints, meaning they can’t just brute-force their way to a good score by burning through unlimited resources. The scoring system is anchored to reference points rather than arbitrary grades. A score of 0 represents a baseline, 0.5 maps to performance levels already documented in published research, and a theoretical optimum sits at 1. Claude Fable 5.1’s score of 35.03% means it’s performing meaningfully but still falls short of replicating techniques already known to re...

Read Entire Article