Ajeya Cotra
METR
Capability 5Outlook 2
Where they stand
Against all 17 people scored in Week 40. Bars count people at each score; black is Cotra, the dashed line is the average.
What they've said
- Week 39Sept 21 to Sept 27
Wrote that the science on loss-of-control risk is nascent and the field lacks the evidence to set standards.
Her first scored week after joining the roster, and she used it to argue the field is not ready to have the argument everyone is having. On September 25 she wrote that the science on loss-of-control risk is
to put it generously, nascent
and thatwe currently lack the basic prerequisites needed to have a conversation about safety standards.
Her sequencing is evidence before verification: what the industry needs is not audits of the claims labs currently make butproduction of a far greater quantity and quality of concrete evidence.
Capability sits at 5 on her own words,
an apparent acceleration in the already-blistering pace of AI progress.
Worth noting the tension inside her own organisation: METR's evaluation of Claude Opus 5.5 three days earlier called itan on-trend improvement
and nota discontinuous jump.
Her acceleration claim is about the field, not the model. Outlook at 2 is a read on her framing rather than a stated prediction, because she never makes one.Planned Obsolescence, September 25. METR's evaluation reports carry no individual bylines and are not scored as her signal.
Sources
Sources
- Evidence about risk should be transparentPlanned Obsolescence