Superhuman performance across all digital work by the end of 2027
AI will be able to do anything digital (that doesn’t require shaping atoms) at a superhuman level by the end of next year
Elon Musk, 2026-08-31 — X
Original post
Sealed at 4%
Locked 2026-08-31. Confidence 68%. Uncertainty ±4 pts. Resolves 2027-12-31. Status: PENDING.
How the engine got there
The gap between this statement and our own published position is one of category, not degree. We already carry TEC-04 / CAL-13 at 75% — that a frontier lab documents an agent autonomously completing a full 40-hour week of economically valuable work by end-2027 — which places this house near the optimistic edge of the field on agent capability, and that number was sealed in July 2026, well before this sentence existed. But "completes a specified 40-hour task" and "superhuman at anything digital" are not neighbouring points on one axis. The first is a duration milestone on scoped work. The second is unbounded generality plus a bar above the best human, across every digital domain at once. Our TEC-01 seal expects 25% of white-collar task-hours automated by 2030 — a decade-defining shift, and still an order of magnitude short of the proposition here. Decomposing: P(agent workweek by end-2027) = 0.75, already sealed; P(broad superhuman generality across all digital work | that milestone) we put near 0.05, because the residual gap is precisely the cluster — long-horizon coherence, genuine novelty, adversarial robustness — that has proven least responsive to scale. The product lands at 4%. This is the one statement on the sheet where the disagreement is not only about the date: sixteen months is not the objection, the word "anything" is.
What counts as correct
CORRECT if, on or before 2027-12-31, a credible independent evaluation — METR, a major multi-lab benchmark consortium, or concurring published evaluations from at least three of OpenAI, Anthropic, Google DeepMind and xAI — documents frontier AI exceeding expert-human performance on a broad suite of purely digital professional tasks spanning at least five distinct domains, with no evaluated domain in which median domain experts still outperform. Domain-specific superhuman performance does not resolve this CORRECT: "anything digital" is read as broad generality, which is the plain meaning of the sentence.
We are wrong if
any substantial category of digital professional work — novel research, long-horizon software architecture, legal drafting under adversarial review, original strategic analysis — still shows median human experts outperforming the best publicly available frontier system on published evaluations at 2027-12-31.
Causal chain
Agent task-horizons have been lengthening on a steep, measurable curve
METR-style evaluations show the duration of work an agent completes reliably roughly doubling on a months-scale cadence, which is what drives our own 75% seal on the workweek milestone.
Scoped, specifiable digital work becomes automatable far faster than the 2023 consensus expected.
Generality has not tracked duration
Gains concentrate in domains with dense training signal and checkable outputs. Work that is novel, adversarial or lacking a verifier improves far more slowly, and no scaling result yet demonstrates that this closes on the same curve.
A system can pass the workweek milestone while remaining below expert humans across whole categories of digital work.
"Superhuman at anything" requires the slowest component to finish
A conjunctive claim resolves on its weakest term. Broad superhuman generality cannot be reached by excelling in most domains; it requires no remaining domain where experts win.
The probability collapses toward the hardest residual category rather than the average of capabilities.
Historical precedent
1965-1975 — The first AI timeline compression
Leading researchers forecast machines doing "any work a man can do" within twenty years, from a base of genuine and rapid early progress. The progress was real; the generality claim was not.
2012-2020 — Deep learning vision-to-generality expectation
Superhuman narrow performance (ImageNet, Go) was widely read as imminent generality. Narrow superhuman results arrived on schedule; broad transfer took another decade and is still partial.
2023-2026 — Agent horizons lengthening
The current, and genuinely stronger, case — the measured doubling curve is the reason our own workweek seal sits at 75% rather than at coin-flip.
Directly related seals
More on this desk
Browse the Technology & AI desk, the full projection registry, the calibration ledger, or the record by subject. Scoring is explained in how a prediction is made.