“We estimate the effective stock of quality and repetition adjusted human-generated public text for AI training at around 300 trillion tokens. If trends continue, language models will be trained on datasets roughly equal in size to the available stock of public human text data between 2026 and 2032, or slightly earlier if models are overtrained.”
Researcher, Epoch AI; lead author of 'Will we run out of data? Limits of LLM scaling based on human-generated data'
The 2026–2032 window is an 80% confidence interval; resolution depends on measured dataset sizes relative to Epoch's ~300T effective-token stock estimate.