Skip to content

AI Fails Against On-Call Engineers

AI Models Struggle to Match Human Expertise in ARFBench Tests

  • The ARFBench is the first AI benchmark using real production incidents, created by Datadog and Carnegie Mellon.
  • GPT-5 leads AI models with a score of 62.7% accuracy but trails behind domain experts who score at 72.7%.
  • A theoretical model-expert oracle achieves an accuracy of up to 87.2%, showcasing potential human-AI collaboration benefits.
  • Datadog’s hybrid model, Toto-1.0-QA-Experimental, surpasses GPT-5 with a score of 63.9% accuracy.

The ARFBench benchmark evaluates AI’s ability to handle real-world production incidents, revealing that current AI models like GPT-5 cannot yet outperform human engineers in incident response tasks.

Despite leading the AI pack, GPT-5’s performance highlights the gap between AI and human expertise, emphasizing the potential for enhanced outcomes through collaborative efforts between humans and AI systems. Source)

Share