Skip to content

AI Models Disagree on 70% Basic Facts

AI Models Show Significant Disagreement on Fact-Checking

  • Five advanced AI models disagreed on the accuracy of claims in 67% of cases.
  • Unanimous agreement was reached on only 328 out of 1,000 claims.
  • The Krippendorff’s alpha score was 0.639, below the reliability threshold of 0.8.
  • In 34% of disagreements, one model labeled a claim as true while another labeled it false.
  • No claims received unanimous “mostly true” verdicts among the models.

A recent study found significant disagreement among five frontier AI models when tasked with fact-checking real-world claims, highlighting potential reliability issues for users relying on these systems for accurate information verification.

The study indicates that while AI models can offer structured judgments, their inconsistency raises questions about their effectiveness as sole fact-checkers, especially given their failure to agree on nuanced truths. (Source)

Share