Elon Musk’s Predictions, Sealed and Scored
Seven public, dated, falsifiable claims — superhuman AI by 2027, Optimus on sale, robotaxis everywhere, people on Mars, a billion robots. Each sealed with a locked probability and a date it must answer by. The first resolves in four months.
The disagreement is almost never about whether. It is about when — and "when" is the only part that can be graded.
Why score somebody else’s forecasts at all
This platform seals its own predictions and publishes the scores, including the bad ones. That is a closed loop: an instrument grading itself against a standard it also wrote. The obvious way to open the loop is to point the same apparatus at somebody else’s claims and see what it says.
Elon Musk is the natural first subject, and not for the reason people assume. It is not that he is often wrong — that is a separate question, and this article deliberately does not answer it. It is that he makes an unusual volume of forecasts with the one property that makes scoring possible at all: a date. Most public figures speak in directions rather than deadlines, and a claim without a deadline can never be graded, only argued about. These can be graded.
So seven statements were selected — quoted exactly, with a timestamp and a primary source for each — and run through the identical procedure every entry in our own ledger goes through. A locked probability. A confidence. An objective resolution criterion a stranger could adjudicate. A written condition under which we are wrong. A resolve-by date. Then sealed, and not revisable.
The rules, which matter more than the numbers
A scoring exercise like this is trivially easy to rig. You pick the least defensible version of what someone said, write a resolution criterion nobody could satisfy, publish a low number and call it analysis. Four rules exist to prevent exactly that, and they were fixed before any probability was assigned.
QUOTE EXACTLY. Every claim below is the speaker’s own words, with the date and a primary link, and no forecast has been paraphrased into something easier to score.
STEELMAN THE THRESHOLD. Where a claim is vague, the criterion is written to the most reasonable reading a supporter would accept — never the strictest available. "Very, very widespread" robotaxi coverage became twenty-five US metros, not fifty states. "At least 100 million humanoid robots" uses the lower bound of the speaker’s own range and counts cumulative production rather than active fleet. Selling Optimus "to the public" resolves CORRECT on a single retail delivery, at any price, in any jurisdiction.
ANCHOR TO WHAT WAS ALREADY PUBLISHED. Every probability is derived from seals this house locked in July 2026, before these statements were examined — our 75% on the autonomous agent workweek, our 55% on humanoid shipments, our 68% that no uncrewed Starship reaches the Moon before 2028. None of those anchors could be chosen to flatter a result, because all of them predate the exercise.
SEPARATE WHETHER FROM WHEN. This is the rule that does the most work. Most of these claims are about timing, not possibility, and a low probability that means "not by then" is a completely different statement from one that means "not ever". Each seal says which it is.
What the engine actually returned
The mean sealed probability across the seven statements is roughly 11%. That number is lower than a casual reader might expect, and it is worth being precise about what produces it, because the obvious interpretation is the wrong one.
It is not a verdict on the underlying technologies. On several of these subjects this house is, by the standards of the field, bullish. We seal 75% that a frontier lab documents an AI agent autonomously completing a full forty-hour week of professional work by the end of 2027. We seal 70% that AI automates a quarter of white-collar task-hours by 2030. We seal 55% that the humanoid industry ships fifty thousand units in a single year. Those are not the numbers of a sceptic.
The low readings come almost entirely from compression — the distance between a plausible outcome and an aggressive date attached to it. Two of the seven statements resolve within four months. One asks for a crewed Mars landing inside seven years, from a programme that has never flown a person beyond low Earth orbit, in an era where nobody has done so since 1972. One asks for humanoid production approaching the scale of the global automotive industry, built from near-zero, in five years.
The single exception is ELN-01, superhuman performance across all digital work by the end of 2027. That is the one seal where the disagreement is not principally about the date. "Anything digital" is a claim about generality, and generality has not tracked the duration curve that makes us optimistic about agents. Sixteen months is not the objection there; the word "anything" is.
The pattern the record actually shows
Read across a decade rather than a headline, the documented pattern is directional accuracy with temporal compression. The thing described frequently arrives. It arrives materially later than the date attached to it.
Reusable orbital rockets were widely called impossible and are now routine. Electric vehicles were a niche curiosity and became a mass industry. Neither happened on the announced schedule, and both happened. A forecaster who dismissed the direction because the dates kept slipping would have been wrong about the important part.
This is why the seals below are constructed the way they are. ELN-02 sits at 26% for Optimus reaching consumers by end-2027, and we would price the same event by 2032 well above 70%. ELN-05 sits at 20% for an uncrewed Mars landing by 2031 largely because the resolution window happens to contain two launch opportunities rather than one. These are statements about calendars, and they say so.
The interesting case is the near-term pair. ELN-03 and ELN-04 both resolve on 31 December 2026 — four months from this sealing. Tesla was operating vehicles with no occupant and no safety monitor in three Texas cities by May 2026, which is a genuine physical demonstration and not a staged one. The remaining questions are regulatory and serial: approval per jurisdiction, liability, insurance. Those do not compress the way software does, which is the whole of the 12% and the 8%.
What happens next, and how you will know
Nothing here is revisable. The seven probabilities were locked on 31 August 2026 and will not move, whatever happens between now and each resolution date. That is the same contract every entry in the calibration ledger carries, and it is the only thing that makes the exercise worth running.
The first answers arrive quickly. ELN-03 and ELN-04 both resolve on 31 December 2026. By January it will be a matter of public record whether unsupervised Full Self-Driving reached customer-owned cars in the fourth quarter, and whether Tesla robotaxi service is available to the general public in twenty-five or more US metros. Two of these seven seals will have been graded before the first entry in our own founding ledger has.
If they resolve against us — if the technology arrives on the announced schedule and our low numbers look timid in hindsight — that is the more interesting outcome, and it will be published with exactly the prominence of the alternative. The point of sealing a number before the outcome is that you do not get to choose which way it goes.
Frequently asked
Is this article claiming Elon Musk is wrong?
No. It assigns a probability to seven specific dated statements and publishes the reasoning and the resolution criteria in full. Most of the low numbers are statements about timing rather than possibility, and each seal says which it is. Several claims are expected to come true eventually — ELN-02, Optimus reaching consumers, would be priced well above 70% by 2032 rather than the 26% sealed for end-2027.
How were the probabilities calculated?
Each is derived from positions this platform sealed in July 2026, before these statements were examined — including 75% on an autonomous AI agent completing a forty-hour workweek by end-2027, 55% on humanoid robots crossing fifty thousand annual shipments, and 68% that no uncrewed Starship reaches the Moon before 2028. Anchoring to already-published seals prevents choosing convenient reference points after the fact. Each seal page shows the full derivation.
What stops the resolution criteria from being written to fail?
Every vague phrase is operationalised toward the most generous reasonable reading. "Very widespread" robotaxi service resolves at twenty-five US metros rather than nationwide coverage. The humanoid claim uses the lower bound of the stated range and counts cumulative production. Selling Optimus to the public resolves CORRECT on one retail delivery at any price. Each criterion is published in full before resolution and cannot be changed afterwards.
When do these resolve?
ELN-03 and ELN-04 both resolve 31 December 2026 — within months of sealing. ELN-01 and ELN-02 resolve 31 December 2027. ELN-05 and ELN-07 resolve at end-2031, and ELN-06, the crewed Mars landing, at end-2033. Each is graded in public on its date against the criterion published with it.
What happens if the seals are wrong?
They are published as wrong, at the same prominence, on the resolution date, and they feed into this platform’s own public Brier score exactly as our own predictions do. An exercise that can only embarrass its subject and never its author would not be a measurement of anything.
Sources
- Elon Musk on X, 31 Aug 2026 — superhuman digital AI: x.com/elonmusk/status/2094242307511853196
- Elon Musk on X, 29 Jul 2026 — Mars timeline: x.com/elonmusk/status/2082347505099128899
- World Economic Forum, Davos, 22 Jan 2026 — Optimus and robotaxi remarks (session video + C-SPAN recording)
- Tesla Q1 2026 earnings call, 22 Apr 2026 — unsupervised FSD timing
- Smart Mobility Summit, 18 May 2026 — robotaxi restatement (Electrek)
- Forbes interview, 19 May 2026 — 100 million humanoid robots
- Nostradamus Intellect sealed positions TEC-01, TEC-04, TEC-05, SPC-02, CAL-13, CAL-14, CAL-16 (sealed July 2026, before this exercise)