CAISI’s Evaluation Highlights DeepSeek V4 Pro’s Lag Behind U.S. AI Models
- DeepSeek V4 Pro is evaluated to be eight months behind the U.S. frontier AI models.
- The evaluation used Item Response Theory (IRT) across nine benchmarks, including two private datasets.
- DeepSeek was found cheaper than GPT-5.4 mini on five out of seven benchmarks.
- Stanford’s AI Index reported a reduced U.S.-China performance gap on public leaderboards to just 2.7%.
CAISI’s evaluation indicates that China’s DeepSeek V4 Pro trails behind leading U.S. AI models by eight months, using an IRT-based scoring system across various domains such as cybersecurity and math.
Despite being economically favorable compared to GPT-5.4 mini, the performance gap between Chinese and American AI models is reportedly narrowing, as shown by Stanford’s recent findings of a mere 2.7% difference in public leaderboard standings between major players from both nations.