Skip to content

AI Benchmark Reveals AGI Still Distant

ARC-AGI-3 Benchmark Reveals AI’s Struggle with Generalization

  • The ARC-AGI-3 benchmark shows top AI models scoring below 1%, while humans achieve perfect performance.
  • Google’s Gemini 3.1 Pro leads with a score of just 0.37%, followed by OpenAI’s GPT-5.4 at 0.26%.
  • Nvidia CEO Jensen Huang claimed AGI achievement, yet the benchmark results contradict this assertion.
  • ARC Prize Foundation’s test uses Relative Human Action Efficiency (RHAE) to measure AI performance against human efficiency.

The ARC-AGI-3 benchmark highlights a significant gap between current AI capabilities and true artificial general intelligence (AGI). Despite industry claims, top AI models struggle to match human-level reasoning and adaptability in unfamiliar environments.

With Google’s Gemini leading at only 0.37%, the results suggest that current AI systems are far from achieving AGI, emphasizing the need for further development in reasoning and generalization capabilities. (Source)

Share