AI & Automation of Work
The question is not whether AI changes work but by how much and by when — the two things most commentary omits. Each entry names a threshold precise enough for a stranger to adjudicate, and a date by which it must be met.
The ledger — hash-sealed, scored in public
CORRECT if Pew Research or an equivalent probability-sampled national survey (Gallup, AP-NORC, Reuters Institute) published by 2029-06-30 finds >60% of US adults using generative-AI tools weekly or more often. Baseline at sealing: Pew (June 2025) found 34% of US adults had ever used ChatGPT.
Resolves by 2029-06-30. Status: PENDING.
CORRECT if by 2027-12-31 a METR-published 50%-reliability time horizon reaches 40 hours for any frontier model, OR a frontier lab (OpenAI, Anthropic, Google DeepMind, Meta, xAI) publishes a documented case of an agent autonomously completing ≥40 hours of contiguous economically-valuable professional work, or expert-parity performance across the full GDPval suite.
Resolves by 2027-12-31. Status: PENDING.
Live projections
This is no longer a forecast about capability — agents' reliable task-horizon has been doubling roughly every seven months (METR) — it is a forecast about diffusion. The first casualty is measurable: early-career employment in AI-exposed occupations fell ~13% relative to less-exposed peers (Stanford, 2025), and every prior automation wave says the transition gap, not the endpoint, is where societies break.
Wrong if: by 2030, early-career employment in AI-exposed occupations recovers to its 2022 trend AND aggregate white-collar task automation measured by major labor studies stays under 15%.
Window: 2026-2030.
The measurement line is public and steep: METR clocked the frontier 50%-reliability task horizon at 14.5 hours in February 2026 and past 16 hours by May, with the doubling time compressed from seven months to roughly 4.3. Extrapolated, a 40-hour work-week crosses around the turn of 2026-27 — and this claim gives the trend through end-2027, two extra doublings of slack. The objective bar: a METR-published 50%-horizon at or above 40 hours, or a lab-documented agent delivering a full professional work-week end-to-end (GDPval-class deliverables at expert parity). What it does not require: 80%-reliability (still ~3 hours), or economic diffusion — TEC-01 owns that harder, slower question.
Wrong if: by 2027-12-31 no METR-published 50%-reliability time horizon reaches 40 hours for any frontier model, AND no frontier lab (OpenAI, Anthropic, Google DeepMind, Meta, xAI) publishes a documented case of an agent autonomously completing ≥40 hours of contiguous economically-valuable professional work or expert-parity performance across the full GDPval suite.
Window: 2026-2027.
AlphaFold took the 2024 Nobel in Chemistry, and Insilico's rentosertib — target and molecule both machine-discovered — has already returned positive Phase 2a data. The binding constraint is no longer discovery but trial biology; once one AI-native molecule clears the FDA, the $2.6B-per-drug cost curve starts its terminal decline.
Wrong if: no drug whose target discovery AND molecular design were primarily AI-driven receives FDA approval by end-2030, or AI-discovered candidates fail Phase 3 at rates worse than the historical ~50% baseline.
Window: 2026-2030.
Tail scenarios
If an AI system reaches the point where it can meaningfully improve its own successor, capability could compound faster than institutions, alignment techniques, or human oversight can track. The danger is not cartoon malice but discontinuity: a fast jump that leaves verification behind while objectives remain imperfectly specified. Frontier labs and states are racing, which removes the option to simply slow down.
More than 70% of trades are already algorithmic, and AI increasingly governs power dispatch, logistics, air traffic and defensive triage. Each system is tested alone; the interactions between them are not. A shared dependency, an adversarial input, or an emergent feedback loop can propagate across domains faster than human operators can intervene. The 2010 Flash Crash — a trillion dollars gone in six minutes (SEC/CFTC, 2010) — was the small preview.
Other subjects on the record
Every claim above is also listed in the full projection registry and the calibration ledger. Scoring is explained in how a prediction is made.