Skip to content

US Claims China AI Models Lag

CAISI’s Evaluation Highlights DeepSeek V4 Pro’s Lag Behind U.S. AI Models

  • DeepSeek V4 Pro is evaluated to be eight months behind the U.S. frontier AI models.
  • The evaluation used Item Response Theory (IRT) across nine benchmarks, including two private datasets.
  • DeepSeek was found cheaper than GPT-5.4 mini on five out of seven benchmarks.
  • Stanford’s AI Index reported a reduced U.S.-China performance gap on public leaderboards to just 2.7%.

CAISI’s evaluation indicates that China’s DeepSeek V4 Pro trails behind leading U.S. AI models by eight months, using an IRT-based scoring system across various domains such as cybersecurity and math.

Despite being economically favorable compared to GPT-5.4 mini, the performance gap between Chinese and American AI models is reportedly narrowing, as shown by Stanford’s recent findings of a mere 2.7% difference in public leaderboard standings between major players from both nations.

Share