World’s first double-blind AI evaluations piloted at massive scale

1 day ago 1



The largest academic AI conference just let artificial intelligence grade its own homework. AAAI-26, one of the premier venues for AI research, ran a pilot program that generated AI reviews for 22,977 main-track paper submissions, making it the first full-scale deployment of AI-powered peer review at a major academic conference. The twist: surveys conducted after the pilot found that both authors and program committee members actually preferred the AI-generated reviews. Specifically, they rated them higher on technical accuracy and quality of research suggestions compared to their human counterparts. How the double-blind AI review system worked The program operated within a double-blind framework, meaning neither authors nor reviewers knew each other’s identities. AI-generated reviews were slotted in alongside at least two human reviews for each paper, creating a side-by-side comparison that researchers could evaluate without bias toward either source. One important design choice: the AI reviews identified themselves as AI-generated. This wasn’t a Turing test. The goal was to complement human reviewers, not to trick anyone into thinking a language model was a tenured professor. The...

Read Entire Article