Alibaba’s Qwen AI model has censorship built in, and researchers say they can strip it out

1 hour ago 1



Alibaba Cloud’s Qwen is one of the most popular AI model families on the planet. It also draws a blank on Tiananmen Square and has no patience for Winnie the Pooh. Qwen models have been downloaded over 3 billion times globally, and researchers now say the censorship isn’t a filter bolted onto the outside. It sits in the model’s core training. A model that knows more than it says Internal inspections suggest the models actually understand the suppressed information. They have simply been trained to deflect or evade conversations about it. When pressed on sensitive prompts, Qwen often frames its refusals around references to “illegal information.” Benchmarks indicate the behavior is deeply rooted in the models’ training protocols. Rather than being the product of an after-the-fact filtering layer, the censorship reinforces state narratives from the inside. The mechanics come down to two common training stages. Supervised fine-tuning teaches a model by example, and reinforcement learning from human feedback, or RLHF, rewards it for answers people rate highly. According to the research findings, both were used to bake the censorship into Qwen. Why the rules exist in the first place Chi...

Read Entire Article