Chutes AI trains an 8B model on distributed gaming GPUs, runs it on a phone CPU

1 hour ago 1



Training an AI model usually conjures images of warehouse-sized data centers humming with specialized chips. Chutes AI just did it with the same graphics cards people buy to play video games. At the Exploit Summit in Montreal, held September 28-29, 2026, Jon Durbin presented an 8 billion parameter model trained on distributed consumer GPUs. The finished model runs on a phone’s CPU. What Chutes AI actually built The system is called Parallax, and it is Chutes AI’s approach to decentralized model training. Instead of renting one giant cluster, Parallax stitches together hardware scattered across the globe. For this run, the setup used 240 RTX 5090 GPUs spread across 30 hosts in 13 countries. The price tag is the headline number. Chutes AI put the training cost at approximately $6,500, or about $11 per billion tokens processed. The 8B model reportedly hits approximately 59.6 tokens per second on mobile CPUs, generating text on a phone processor with no cloud connection required. Chutes AI also showed a larger variant with 40.75 billion parameters. That version ran at 26.9 tokens per second while using 12GB of peak memory. The privacy and resilience pitch Local inference was a central ...

Read Entire Article