Newer AI models missed more payment fraud in Coinbase’s benchmark

1 hour ago 1



Coinbase reported Oct. 7 that newer versions of three major AI model families caught fewer fraudulent payments and a smaller share of fraud value in a historical test of payment screening for its Onramp service, despite an unchanged decision policy. The findings challenge the assumption that upgrading a model improves an existing payment screener.The company’s evaluation replayed 16,140 transactions across 7,293 users, including 813 confirmed fraudulent transactions. The cohort covered nine weeks before its risk agent rolled out, retaining all matured fraud cases while sampling legitimate traffic.Each candidate reviewed recent transaction behavior under fixed guidance and the same policy for turning risk classifications into decisions. This isolated the decision model’s behavior within that setup, rather than comparing redesigned screening systems.Results from a fixed historical replayCoinbase compared Opus 4.5 with Opus 5, Sonnet 4.6 with Sonnet 5, and GPT-5.4 with GPT-5.6 (sol). Every newer version had lower recall, a lower combined precision-and-recall score called F1, and lower dollar-weighted recall. Recall measures the share of fraud cases a model catches; dollar-weighted rec...

Read Entire Article