“Anthropic is doubling down on interpretability, and we have a goal of getting to 'interpretability can reliably detect most model problems' by 2027 — the analogue of an MRI for AI, arriving before models become overwhelmingly powerful.”
Said Apr 2025Deadline Dec 31, 2027Dario Amodei, 'The Urgency of Interpretability', darioamodei.com, April 2025
Stated as an Anthropic goal with an explicit 2027 deadline; resolution will rest largely on Anthropic's own and third-party interpretability evaluations.