AI Prediction Index

Center for AI Safety

Co-creator of the Humanity's Last Exam benchmark with Scale AI

0%
success rate · #142 of 144
0correct0half right1missed0on the clockFull leaderboard
CapabilitiesIncorrect
“While current LLMs achieve very low accuracy on HLE, recent history shows benchmarks are quickly saturated. Given the rapid pace of AI development, it is plausible that models could exceed 50% accuracy on HLE [Humanity's Last Exam] by the end of 2025.”

Co-creator of the Humanity's Last Exam benchmark with Scale AI

Said Jan 23, 2025Deadline Dec 31, 2025Humanity's Last Exam, arXiv:2501.14249, January 23, 2025

No model exceeded 50% on HLE by the end of 2025. The best published results were Gemini 3 Pro at 37.5% and Gemini 3 Deep Think at 41.0% without tools (roughly 45% with search and code execution); scores above 50% only appeared in 2026.