Skip to content

AI Struggles You Excel at Doom

AI Models Struggle with Classic Video Game Challenges

  • Advanced vision-language models like GPT-4o, Claude Sonnet 3.7, and Gemini 2.5 Pro struggle with playing Doom.
  • VideoGameBench was introduced as a benchmark to test AI on 20 popular video games.
  • High inference latency causes issues in fast-paced games like Doom, where game states change rapidly.
  • The benchmark uses classic Game Boy and MS-DOS games for their simpler visuals and diverse input styles.
  • Sonnet 3.7 performed best among tested models but still faced challenges in spatial reasoning and mouse control.

A new AI benchmark, VideoGameBench, tests advanced vision-language models on classic video games like Doom. Despite technological advancements, these models face significant challenges due to high inference latency and difficulty with dynamic environments. Sonnet 3.7 showed the most progress but still struggled with basic in-game actions.

Source (2.6)https://decrypt.co/315493/ai-bad-playing-doom?rand=52368
Share