CapabilitiesIncorrect
“While current LLMs achieve very low accuracy on HLE, recent history shows benchmarks are quickly saturated. Given the rapid pace of AI development, it is plausible that models could exceed 50% accuracy on HLE [Humanity's Last Exam] by the end of 2025.”
No model exceeded 50% on HLE by the end of 2025. The best published results were Gemini 3 Pro at 37.5% and Gemini 3 Deep Think at 41.0% without tools (roughly 45% with search and code execution); scores above 50% only appeared in 2026.