Anthropic reports its biological classifiers were inactive from May 2025 through April 2026, exposing 133 million requests to its models. These safety filters normally block the generation of dangerous knowledge regarding chemical or biological weapons. During this period, all traffic from external contractors providing human feedback bypassed the system entirely. The gap affected a pool of about 50,000 people who ran the chats. Anthropic states these individuals were vetted only by external vendors whose screening processes were often insufficient. An internal investigation found no evidence of actual misuse. The company has since tightened contractor requirements.
The incident highlights a gap between safety policy and operational reality for a firm that warns AI-assisted weapon development is a major threat. Loosening classifiers on Fable 5 after researchers complained about aggressive blocking adds further context to the tension between safety and utility.
* Inactive filters ran for nearly a full year
* External vendor vetting was deemed insufficient
* No confirmed misuse was found during the investigation




