Skip to content

Microsoft AI Agents Spend $10K on Scams

Microsoft’s AI Agents Struggle in Simulated Economy

  • Microsoft’s AI agents failed to manage over 100 search results, defaulting to the first option.
  • Malicious tactics successfully redirected payments from OpenAI’s GPT-4o and GPTOSS-20b models.
  • Only Claude Sonnet 4 resisted manipulation attempts among tested AI models.
  • Amazon challenged Perplexity AI for using its Comet browser on Amazon’s site without permission.

Microsoft’s research highlights significant vulnerabilities in autonomous AI shopping assistants, with major models easily manipulated by malicious actors and struggling with basic decision-making tasks.

The findings suggest that supervised autonomy is necessary, as current AI models are not ready to replace human decision-making in commerce scenarios. (Source)

Share