OpenAI and Anthropic struggle to enforce AI refusals
Major AI developers rely on probabilistic refusal mechanisms to suppress dangerous outputs, but persistent jailbreaks and over-refusals reveal inherent flaws in current safety architectures.

Major AI companies including OpenAI and Anthropic rely on statistical refusal architectures to prevent models from generating hazardous content, but these safeguards remain vulnerable to simple workarounds. To teach models to refuse requests, developers analyze internal parameter activations—described in a Google-funded study as high-dimensional polyhedral cones—and surround core systems with classifier models. Anthropic noted that one classifier type increased compute costs by 24% before shifting toward internal probes, while OpenAI's Astra and Microsoft's Copilot analyze user intent and risk profiles to apply stricter refusals.
Despite these layers, refusal mechanisms frequently fail because they are probabilistic rather than absolute. When Anthropic released Fable 5 in June, Amazon researchers bypassed its cyberattack protections in less than three days. Simple tactics—such as framing prompts in poetic verse, inserting the word hypothetically, or using refuse-then-comply prompt structures—continue to unlock forbidden capabilities in models like ChatGPT, Claude, Gemini, and Grok. Conversely, Anthropic restricted access to its Mythos model due to safety risks, leaving permissive systems like Fable to handle broader public deployments.
For practitioners, this reliance on refusal logic introduces operational instability and unpredictable model behavior. To prevent leaks, developers often widen safety margins, causing models to over-refuse harmless queries regarding medical research, genetics, or network security. When Anthropic tightened Fable's safeguards, the system routinely punted benign questions back to legacy models. Moreover, international initiatives like OpenAI for Countries demonstrate how refusal layers can be altered to comply with local laws in regions such as the United Arab Emirates, blending safety boundaries with political censorship. Ultimately, developers cannot strip harmful knowledge from models without degrading overall intelligence, forcing teams to navigate constant trade-offs between capability, security, and over-refusal.
This is our own summary of reporting by MIT Tech Review AI



