AI Prediction Index

One prediction, on the record

“Early data points suggest that the upcoming ARC-AGI-2 benchmark will still pose a significant challenge to o3, potentially reducing its score to under 30% even at high compute (while a smart human would still be able to score over 95% with no training).”

Creator of Keras; co-founder of the ARC Prize Foundation and Ndea (formerly Google senior staff engineer)

When ARC-AGI-2 launched in March 2025, ARC Prize estimated o3-preview-low would score ~4% and o3-high only 15–20% even at very high compute; subsequent official testing put o3 (medium) at roughly 3%. Every measured or estimated o3 configuration fell far below the 30% ceiling Chollet predicted.