AI Prediction Index

One prediction, on the record

“Anthropic is doubling down on interpretability, and we have a goal of getting to 'interpretability can reliably detect most model problems' by 2027 — the analogue of an MRI for AI, arriving before models become overwhelmingly powerful.”

Co-founder and CEO, Anthropic

Stated as an Anthropic goal with an explicit 2027 deadline; resolution will rest largely on Anthropic's own and third-party interpretability evaluations.