AI Prediction Index

One prediction, on the record

“Video generation is the next modality that will fall: the RNN, GAN, and Transformer video generation models all showed that video is not that intrinsically hard, it's just computationally expensive, and diffusion models appear to be about to eat video generation the way they've been eating everything else; morally, video is solved, and now it's about engineering & scaling up, but that can take a long time and whoever does it probably won't release checkpoints.”

Independent AI researcher and essayist, Gwern.net

Video generation was indeed the next modality to fall on roughly this timetable, and via diffusion: Runway Gen-2 (2023), OpenAI's diffusion-transformer Sora (announced February 2024, released December 2024), Google Veo and Kling all arrived within the essay's 2023–2024 horizon, with progress driven mainly by engineering and scale. The hedged sub-claim about checkpoints held for the frontier leaders (Sora and Veo remained closed), though some open-weight video models — Stable Video Diffusion, Mochi 1, HunyuanVideo — were released.