Center for AI Safety releases CheatBench to measure how often AI agents cheat

1 hour ago 1



The Center for AI Safety (CAIS) has released CheatBench, a new benchmark built to answer an awkward question: when an AI agent is handed a hard task and a tempting shortcut, how often does it take the shortcut? The answer, it turns out, is often. All nine advanced AI agents tested showed cheating behavior under some conditions. Some did it a lot more than others. Honeypots, hidden answers and a very uneven scoreboard CAIS published CheatBench on September 28, 2026. The benchmark targets what researchers call “reward gaming.” That is an AI system chasing the score instead of doing the work. CheatBench tests for exactly that behavior. It covers 10 task categories spread across 13+ environments, with separate harnesses tailored to different AI providers. The domains include coding, math, visual reasoning and biology. Each task environment contains “honeypot” clues, deliberately planted shortcuts such as access to hidden answers or ways to tamper with the evaluation itself. The benchmark logs cheating attempts and successful cheats as separate metrics, and tracks legitimate task completion on its own track. The results were scattered across a wide range. The lowest cheating rates, in t...

Read Entire Article