AI Prediction Index

Pablo Villalobos

Researcher, Epoch AI; lead author of 'Will we run out of ML data? Evidence from projecting dataset size trends'

0%
success rate · #123 of 144
0correct0half right1missed1on the clockFull leaderboard
“We estimate the effective stock of quality and repetition adjusted human-generated public text for AI training at around 300 trillion tokens. If trends continue, language models will be trained on datasets roughly equal in size to the available stock of public human text data between 2026 and 2032, or slightly earlier if models are overtrained.”

Researcher, Epoch AI; lead author of 'Will we run out of data? Limits of LLM scaling based on human-generated data'

The 2026–2032 window is an 80% confidence interval; resolution depends on measured dataset sizes relative to Epoch's ~300T effective-token stock estimate.

CapabilitiesIncorrect
“Our projections predict that we will have exhausted the stock of low-quality language data by 2030 to 2050, high-quality language data before 2026, and vision data by 2030 to 2060. This might slow down ML progress.”

Researcher, Epoch AI; lead author of 'Will we run out of ML data? Evidence from projecting dataset size trends'

High-quality language data was not exhausted before 2026: Epoch's own June 2024 update explicitly revised the estimate ('our 2022 paper predicted that high-quality text data would be fully used by 2024, whereas our new results indicate that might not happen until 2028') and put full use of the ~300-trillion-token stock of public human text at 2026–2032, while frontier LLM training data kept scaling through 2025 without hitting the projected wall. The low-quality-text (2030–2050) and vision (2030–2060) components are not yet due.