AI Prediction Index

One prediction, on the record

CapabilitiesIncorrect
“Our projections predict that we will have exhausted the stock of low-quality language data by 2030 to 2050, high-quality language data before 2026, and vision data by 2030 to 2060. This might slow down ML progress.”

Researcher, Epoch AI; lead author of 'Will we run out of ML data? Evidence from projecting dataset size trends'

High-quality language data was not exhausted before 2026: Epoch's own June 2024 update explicitly revised the estimate ('our 2022 paper predicted that high-quality text data would be fully used by 2024, whereas our new results indicate that might not happen until 2028') and put full use of the ~300-trillion-token stock of public human text at 2026–2032, while frontier LLM training data kept scaling through 2025 without hitting the projected wall. The low-quality-text (2030–2050) and vision (2030–2060) components are not yet due.