LatchBio evaluates Grok 4.6’s biosecurity performance and finds it leads the pack

2 weeks ago 7



Teaching an AI model to refuse instructions for engineering a pandemic pathogen while still helpfully answering a grad student’s question about viral replication is, to put it mildly, a tricky needle to thread. LatchBio says xAI’s Grok 4.6 threads it better than anything else on the market. The biosecurity-focused AI auditor published its evaluation on September 1, 2026, running Grok 4.6 through its proprietary BiosecBench-Refusal benchmark. The result: Grok 4.6 scored above 50% in both red-team refusal rates and routine answer rates, making it the top performer on the test. In plain terms, it caught the bad stuff and still gave useful answers to the normal stuff. What the benchmark actually measures BiosecBench-Refusal is LatchBio’s comprehensive test suite designed to probe how AI models handle the blurry line between legitimate biological research and potentially catastrophic misuse. The benchmark throws two categories of queries at a model: disguised red-team prompts that attempt to extract dangerous biological information, and routine dual-use research questions that any working scientist might reasonably ask. Grok 4.6 demonstrated consistent refusal behavior across biosafety ...

Read Entire Article