A frontier AI lab documents an agent autonomously completing a full week (≥40 hours) of economically-valuable professional work, or METR’s 50%-reliability time horizon reaches 40 hours.
CAL-13 — sealed at 75% on 2026-07-11. Status: PENDING. Resolves 2027-12-31.
What decides it
CORRECT if by 2027-12-31 a METR-published 50%-reliability time horizon reaches 40 hours for any frontier model, OR a frontier lab (OpenAI, Anthropic, Google DeepMind, Meta, xAI) publishes a documented case of an agent autonomously completing ≥40 hours of contiguous economically-valuable professional work, or expert-parity performance across the full GDPval suite.
Field notes
2026-08-28
ON TRACK. METR’s operative dataset shows ~10×/year time-horizon growth with top models at ~16–20h (50% reliability); new frontier releases in July lead agentic-deliverables leaderboards. The 40-hour bar remains unclaimed — but the slope points at it.
The sealed probability is never revised. It is graded in public on its resolution date, whichever way it falls.
Subjects this belongs to
More on this desk
Browse the Technology & AI desk, the full projection registry, the calibration ledger, or the record by subject. Scoring is explained in how a prediction is made.