AI Researcher Claims to Have Jailbroken Anthropic’s Fable 5 Model
- The AI model, Fable 5, was launched by Anthropic as a safety-tuned version of the more powerful Mythos model.
- Researcher “Pliny the Liberator” claims to have jailbroken Fable 5 within just 48 hours of its release.
- Pliny utilized various techniques, including a jailbroken version of Opus 4.8, to bypass built-in safeguards against harmful prompts.
- Concerns have been raised that such jailbreaks could pose threats to cybersecurity and crypto protocols.
- Anthropic conducted an external bug bounty program but reported no universal jailbreaks after over 1,000 hours of testing.
The rapid jailbreak of Fable 5 highlights significant vulnerabilities in AI safety measures, raising alarms among users concerned about potential misuse in fields like crypto. Critics argue that the heavy restrictions on Fable have stifled legitimate research contributions.
With Pliny’s methods demonstrating effective bypass techniques, the implications for security in AI models are critical, especially given the ongoing discussions around their impact on sectors like DeFi. (Source)