CapabilitiesCorrect
“Early data points suggest that the upcoming ARC-AGI-2 benchmark will still pose a significant challenge to o3, potentially reducing its score to under 30% even at high compute (while a smart human would still be able to score over 95% with no training).”
Creator of Keras; co-founder of the ARC Prize Foundation and Ndea (formerly Google senior staff engineer)
Said Dec 20, 2024Deadline Dec 31, 2025ARC Prize blog, 'OpenAI o3 Breakthrough High Score on ARC-AGI-Pub', December 20, 2024
When ARC-AGI-2 launched in March 2025, ARC Prize estimated o3-preview-low would score ~4% and o3-high only 15–20% even at very high compute; subsequent official testing put o3 (medium) at roughly 3%. Every measured or estimated o3 configuration fell far below the 30% ceiling Chollet predicted.