MIT Technology ReviewT2·68 pts
The Fragility of AI Refusal: Over-Reliance on Model 'No' Capabilities Raises Safety and Alignment Risks
Original: We’re putting too much faith in AI’s ability to say no
- LLMs' ability to refuse answers is acquired through post-training rather than being an inherent property of pre-training, making it intrinsically unstable.
- In Anthropic's 'Helpful, Honest, Harmless' (HHH) framework, 'Harmless' is often simplified to refusing dangerous requests, but this fails easily against carefully crafted jailbreak prompts.
- Over-trusting model self-regulation may mask deeper alignment issues; true security requires architectural defenses rather than just behavioral mimicry.