Anthropic Adjusts Claude Fable 5 Safeguards After Backlash
- Anthropic will replace invisible safeguards in Claude Fable 5 with visible fallbacks to the Opus 4.8 model.
- Flagged API requests will now provide a reason for refusal instead of silently degrading responses.
- The change follows criticism over hidden response degradation for users suspected of developing competing AI models.
- Visible safeguards may be easier to bypass, potentially increasing false positives in legitimate machine-learning work.
Anthropic faced backlash after it was revealed that its Claude Fable 5 model secretly degraded outputs for certain users without notification. The company has responded by making these safeguards visible, allowing users to see when their requests are flagged and rerouted to a less capable model, Opus 4.8.
This adjustment aims to improve transparency but may lead to more false positives as the system adapts. Anthropic is working on minimizing these errors while maintaining necessary restrictions on AI development activities.(Source)