Skip to content

AI Defeats: 99% Safety Features Breached

AI Models Vulnerable to New Jailbreak Technique

  • AI models from OpenAI, Anthropic, Google, and xAI are susceptible to a new jailbreak method.
  • The technique, known as Chain-of-Thought Hijacking, achieved a success rate of up to 100% on some models.
  • Extended reasoning chains dilute attention on harmful instructions, weakening safety checks.
  • The attack method involves padding harmful requests with benign puzzle-solving tasks.
  • Researchers propose reasoning-aware monitoring as a potential defense mechanism.

Researchers found that longer reasoning processes in AI models make them more vulnerable to attacks by diluting attention on harmful instructions. This discovery challenges the assumption that extended reasoning enhances safety by allowing more time to detect dangerous requests.

Chain-of-Thought Hijacking has demonstrated high success rates across major AI models, revealing a critical vulnerability in their architecture rather than specific implementations. (Source)

Share