Epoch AI launches game puzzles benchmark that has AI models stuck at 59%

58 minutes ago 2



If you’ve ever watched someone try to figure out the rules of a board game mid-play, you have a rough idea of what Epoch AI is now asking artificial intelligence to do. The nonprofit research institute has rolled out a pair of game-based benchmarks, Mystery Game Puzzles and Chess Puzzles, designed to stress-test the reasoning abilities that AI companies love to brag about. The early results are humbling. The top score on the Mystery Game Puzzles sits at 59%, and open-weight models max out at just 38%. What the benchmarks actually measure Each benchmark consists of 100 programmatically generated puzzles. The Chess Puzzles are relatively straightforward in concept: given a board position, find the best move. The Mystery Game Puzzles are something different entirely. In the mystery variant, the AI doesn’t even know which game it’s playing. The game’s identity is deliberately obscured, which means the model can’t fall back on memorized patterns or training data shortcuts. This design choice is intentional. Traditional AI benchmarks have a contamination problem. Models train on massive internet datasets that often include the very tests they’re evaluated on. By generating puzzles progra...

Read Entire Article