Energy & Data CentersIncorrect
“The amount of compute needed to train Llama 4 will likely be almost 10x more than what we used to train Llama 3, and future models will continue to grow beyond that.”
Llama 4 as actually trained fell far short of a 10x compute scale-up over Llama 3.1 405B (~3.8e25 FLOP): Epoch AI estimates Llama 4 Maverick at roughly 2.2e24 FLOP, and the largest model, Behemoth — reported as trained on about 32,000 H100s — at roughly 5e25 FLOP, only around 1.3x Llama 3.1 405B. Behemoth was delayed through 2025 and never publicly released.