A frontier lab documents an AI agent completing a full week of professional work autonomously
TEC-04 · probability 75% (confidence 64%, ±12 pts) over 2026-2027 · horizon NEXT 6-18M · domain technology. Probabilistic simulation, not advice.
The reading
The measurement line is public and steep: METR clocked the frontier 50%-reliability task horizon at 14.5 hours in February 2026 and past 16 hours by May, with the doubling time compressed from seven months to roughly 4.3. Extrapolated, a 40-hour work-week crosses around the turn of 2026-27 — and this claim gives the trend through end-2027, two extra doublings of slack. The objective bar: a METR-published 50%-horizon at or above 40 hours, or a lab-documented agent delivering a full professional work-week end-to-end (GDPval-class deliverables at expert parity). What it does not require: 80%-reliability (still ~3 hours), or economic diffusion — TEC-01 owns that harder, slower question.
What would prove this wrong
This projection is WRONG if by 2027-12-31 no METR-published 50%-reliability time horizon reaches 40 hours for any frontier model, AND no frontier lab (OpenAI, Anthropic, Google DeepMind, Meta, xAI) publishes a documented case of an agent autonomously completing ≥40 hours of contiguous economically-valuable professional work or expert-parity performance across the full GDPval suite.
Trigger events tracked
- The next frontier release cycle holding the ~4.3-month doubling curve through 2027
- GDPval win-rates crossing expert parity across occupations, not just on subsets
- Agent infrastructure maturing — multi-day memory, checkpointing, self-verification loops
- An in-the-wild autonomous work-week (a shipped codebase, a filed analysis) documented and audited by a lab
Causal chain
Historical precedents
If it happens
Sources
Directly related seals
Subjects this belongs to
More on this desk
Browse the Technology & AI desk, the full projection registry, the calibration ledger, or the record by subject. Scoring is explained in how a prediction is made.