AI predictions? Not even wrong
An audit of seventy-six years of AI prophecy has just been published. Nineteen out of twenty-four cannot even be contradicted. What that says about prediction, and what it says about our own trade.
Wolfgang Pauli, a physicist with a famously short temper, is said to have been handed a paper by a younger colleague. He read it and delivered his verdict: that isn't even wrong.
The phrase stuck, and it now has a settled meaning, as Wikipedia puts it: an explanation that looks scientific but rests on reasoning that doesn't hold, or on assumptions that can be neither proved nor disproved.
A false claim can be refuted, and often corrected. A claim that isn't even wrong offers nothing to correct, because it was never put in a form that could be contradicted in the first place.
At best it takes up room and adds to the noise. At worst it becomes a self-fulfilling prophecy, pushing the people who hear it in directions that run against their own interests.
Luciano Floridi, Jessica Morley and Claudio Novelli have just published a paper that applies this diagnosis, methodically, to the entire public practice of predicting the arrival of artificial general intelligence. Sixty-six pages, twenty-four predictions retained, from Turing in 1950 to Musk in 2026. It is the first prediction-by-prediction audit of that record, and the result deserves an audience beyond philosophers of science.
The method the researchers used
To judge what an AI prediction is worth, the three researchers began by turning the question around. They don't grade predictions on whether they were right. They grade them on whether they could ever be proved wrong: a prediction that no outcome could contradict can't be credited with having got anything right either.
Their protocol comes in two examinations, in this order.
The first examination asks what the prediction rules out. A prediction can only be refuted if it makes certain futures incompatible with it: if a given state isn't reached by a given date, it's wrong. One that rules nothing out accommodates whatever happens, and that is the point where it stops being checkable. Four tests answer the question, and they stand together: a single failure brings down the examination, and with it any claim the prediction had to being refutable.
- Is the announced state defined? Precisely enough for two specialists to agree on what is being claimed, without having to call the author.
- Is the measurement named? A protocol, or a measuring apparatus that already exists.
- Is the date closed?
- Can a third party settle it? Without recourse to the interpretation the predictor will offer after the fact.
Each test is worth two points, one when it is only half satisfied. Every prediction in the record therefore carries a score out of eight and, where relevant, the name of the criterion that gave way. One that fails here isn't wrong: it isn't even wrong, and the examination stops there.
Here is an example, and it sets one man against himself. In 1958, Herbert Simon and Allen Newell announced that a computer would be world chess champion within ten years. In 1968 you only had to look at who held the title: the prediction ruled out a future, the one where the champion is still human, and that is the future that arrived. Seven years later, the same Simon wrote that within twenty years machines would be capable of doing any work a man can do. This time nothing is ruled out: no list of jobs, no threshold, no referee, so that in 1985 as today one can argue the sentence came true just as easily as the opposite. The first could lose. The second can only endure.
The second examination applies only to survivors whose deadline has passed. Two questions: did the announced state occur? And was the test that would have shown it really measuring what it claimed to measure?
This second point is the sharpest in the paper: you can be right for the wrong reasons, when the test comes out in your favour but was measuring something other than what it advertised. The Turing test pass claimed in 2014 shows the shape of it. The program played the part of a thirteen-year-old Ukrainian boy with poor English: the character absorbed the clumsiness that would have given the machine away, and the threshold was crossed without the imitation proving anything. A score reached by routes that have nothing to do with what the test claims to name isn't weak evidence of the ability in question. It isn't evidence at all.
The revealing part is that this criterion has never had to be used. The only two predictions that made it to the second examination failed at the first question, the one about whether the thing happened, and a state that never occurred raises no question about the test that would have established it. The first occasion may come from AI 2027, one of the three predictions still open. It is a scenario published in April 2025 by the AI Futures Project, a group founded by a former member of OpenAI's governance team, which narrates month by month the arrival of superhuman AI up to the runaway of late 2027. The text has circulated widely, and it is currently the most discussed dated prediction there is. Its fulfilment would be established on scores obtained in batteries of tests, which is to say exactly the kind of exam a machine is known to pass without holding the ability it claims to measure.
The results
Nineteen predictions out of twenty-four, or 79 %, aren't even wrong. Two, or 8 %, could be refuted and were. Three, or 13 %, remain open.
And even that scale is lenient. A criterion half satisfied earns half a point and doesn't eliminate: an approximate date, a measurement mentioned without a protocol, and the prediction stays in the running. If each "half" were counted as a failure, a single prediction in the record would pass the examination, the chess one, and the other twenty-three would fall in a block. The authors preferred leniency, so that the result would come from the predictions themselves rather than from the severity of their grid. The 79 %, in other words, is what you get on the reading most favourable to those doing the predicting.
The failure always takes the same form: a date laid over a state nobody has defined, nothing measures, and no third party could confirm had arrived. The date supplies the appearance of rigour; there is nothing underneath it. And it is that last point, the referee, that is missing most often: in a single case out of twenty-four, somebody other than the author could have settled the matter. In seventy-six years, almost nobody has agreed to hand the verdict to anyone else.
Two details say more than the headline percentage.
The only two predictions ever refuted date from the first decade: Turing in 1950, Simon and Newell in 1958. This isn't a tribute to the old guard, it's an awkward finding. The pioneers took testable risks, and they lost them. Their successors learned not to expose themselves. Deep Blue beat the world champion twenty-nine years after the announced date, and by brute calculation rather than by the human reasoning Simon and Newell imagined they were reproducing. We can say they were wrong, and wrong about what exactly, only because they committed.
Quality, for its part, isn't improving. The authors say so carefully, twenty-four cases being too few to establish a trend, but the average after 2000 is no better than the average before. The other forecasting trades did learn how to work in the meantime. Meteorology announces probabilities and checks afterwards whether they landed. Epidemiology gives ranges rather than dates. Intelligence services score their analysts on the accuracy of what they announced, so as to know who is wrong and by how much. None of that has reached the people predicting AI.
The life cycle of a prediction
The most useful part of the paper, for anyone who makes decisions, describes what becomes of an under-specified prediction as it circulates.
There is first what the authors call the chain of conflations. It isn't a reasoning error committed by someone, it's a distortion that happens while information travels. Each person generalises slightly what they read, and states it slightly more firmly than their source did.
The sequence is always the same. Take an example. A lab publishes that its model solves seven programming problems out of ten in a battery of tests. An article concludes that AI can program. A second concludes that machines now write code without human supervision. An analyst infers that they will be deployed across companies. A report puts a number on the developer jobs about to disappear. A panel show concludes that we are losing control. A technical remark has become a question of the survival of the species in six sentences.
Yet not one of those steps is a deduction. Passing a battery of tests doesn't prove a program can program anywhere but on those tests. Being able to program doesn't mean anyone will let it work with nobody reviewing it. And that it could doesn't say that it will, nor that it would be reliable, legal or profitable. Nobody in that chain has lied, though: each of them can point to the sentence they started from, and it says roughly what they made of it. The distance between the score at the start and the conclusion at the end, on the other hand, none of them crossed alone. None of them ever had to justify it, and none of them can be held responsible for it.
Then comes the modal slide. What was offered as possible is received as probable, and a conditional is received as a forecast. The authors note a perverse effect: the most exposed to this slide are those who took the trouble to state a probability or a set of conditions, since they supplied exactly what the slide strips away, and end up held to the inflated version rather than to their own.
Then comes renewal by displacement. A prediction nobody can settle doesn't die on its date. It simply stops being current, and a new date succeeds the old one with no retraction, no analysis of the gap, no argument justifying the revision. The paper's own method illustrates the point: Musk's statements between 2020 and 2026 count as one entry, not five, because each repeats the same claim with the date moved. Counting them separately would have let a single practice of reissuing occupy five lines of the record.
Why does this practice still go on?
The paper explicitly sets aside the hypothesis of individual bad faith, and that is what makes it credible.
It draws first on the literature on expert judgement. The conditions under which an expert is reliable have been known since Shanteau: a stable environment, agreement on the object being judged, a repeated task, feedback that allows calibration. Applied to forecasting AI timelines, eight of those ten conditions are unfavourable and two partial. None is favourable. The problem isn't the people, it's the kind of task.
It then names what it calls the believers' fallacy. The procedure has two stages: you survey a population that already subscribed to a proposition, then you present their agreement as evidence that the proposition is true.
This is what happens when you ask the heads of AI labs how many years away AGI is. Nobody takes charge of a company whose stated purpose is to build general intelligence without first believing that it is possible and that it is close: conviction is part of the job. Asking them is therefore polling the converted about the very object of their conversion, and the average of their answers then circulates as an estimate of how likely the event is.
Asking the faithful whether they believe in miracles tells you about the faithful.
It notes, lastly, that the quality of a prediction varies with how much its author stands to gain from being believed. Predictions issued by companies in the sector get the lowest scores in the record. Academics do better, independent research organisations better still. Put plainly: the more money there is to raise on an announcement, the less checkable that announcement is. The authors take care not to read deliberate calculation into it, and note only that the ranking is consistent with their hypothesis. The restraint is to their credit, because a correlation like that lends itself far too easily to accusations of motive.
And it adds a cause we rarely look at squarely: demand. Regulators, institutions and the press need objects of forecast in order to function. Governing presupposes something to govern, legislating something to regulate, reporting something to report on. "AGI in 2027" is something you can write about, even if nobody can say what its arrival on that date would mean. The people who speak adjust to that demand. Whoever announces a date gets invited on the programme. Whoever answers that they don't know, and offers a range with conditions attached, does not. At that game, false precision wins the airtime and rigour loses it.
The study's blind spot
The instrument has a blind spot, and the authors know it. They state in the introduction that these predictions don't merely describe possible futures but are institutional facts that help determine them, by conferring legitimacy, informing public policy and steering investment. Then they stop there. A footnote points to the work that studies the effect of technological promises on the real world, and specifies that their own work concerns the form of the statements, not what those statements produce. You can't hold it against them: a method that measures refutability measures only that, and going further would be committing the very fault they charge others with. The fact remains that the effect of these predictions on the world is what interests all of us, and nobody attends to it here.
Their method can't capture that effect, and this isn't an oversight on their part: it's structural. Their grid examines a property of the sentence, what it forbids from happening. But a prophecy that fulfils itself acts elsewhere: it moves the world against which it will be compared on the day of reckoning. That can't be read in the sentence, it happens between the sentence and the world. None of the four criteria can see it.
And that is precisely what occurs. When the head of a lab announces a date, they aren't describing a state of the world observed from outside. They are emitting a signal that moves capital, redirects careers, opens public funding lines and concentrates research effort on the announced trajectory. They hold a significant share of the means of making their own claim come true. The prediction is part of the causal mechanism it purports to anticipate.
What would be needed, then, is a fifth criterion the paper doesn't have: does the prediction act on the conditions of its own fulfilment, and does the person making it control those conditions? A prediction issued by an actor with no grip on the ground is simply uncheckable. The same prediction issued by someone holding a fraction of the sector's capital and talent is partly performative, and the fact that it eventually comes true would no longer prove anything at all.
The irony is that the one historical case the paper documents runs the other way. In 1973, James Lighthill audited for the British government the unkept promises of AI research. His report triggered the withdrawal of funding, and the British field failed to deliver what it had announced, having lost the means to do so. The first AI winter. The prophecy of failure produced the failure. The authors open their paper on that episode without ever turning it on themselves. An audit concluding that the sector's promises are worthless is, after all, a statement that acts too: taken seriously by those who fund, it could bring about the second winter it would merely be describing. What holds for the prophecy of abundance holds for the prophecy of collapse.
What cybersecurity should take from it
I could stop here and let the reader smile at other people's prophecies. That would be dishonest.
Offensive security produces exactly the same kind of statements, and we make a living from them. "It's not a question of if, but when" has the textbook structure of the claim that isn't even wrong: an undefined state, no protocol, no date, no referee. It can't be contradicted, and that is precisely why it sells. "Cyber risk is going to explode within eighteen months" doesn't say which indicator, measured how, and who will confirm it. Our annual threat reports, our composite scores, our attack surface forecasts often belong to the same regime. We put dates on things we haven't defined, and we're rarely the ones who bear the cost of the vagueness.
The second lesson is even more direct. A good score on a test doesn't prove the machine can do what the test claims to measure. We know this in our field better than anyone, since our whole trade consists of showing that a compliant system isn't a secure system. The paper generalises the intuition: an indicator that can be produced without the target state being reached tells you nothing about that state. It's true of a compliance certificate and it's true of a score on a suite of tasks. Reading the arrival of AGI off a benchmark curve is repeating the very error benchmarks were created to prevent.
What the authors ask for
The authors aren't asking anyone to be more careful about the future. They're asking that we do the easy part, the part that costs nothing: say what you're announcing, how it will be measured, and who will confirm it.
They draw seven principles from climate science, economic forecasting and medical research. Three carry the weight. Announce a range with a probability attached, rather than a single date that gives itself the air of certainty. Name the referee before the answer is known, because anyone judging their own case after the fact always finds that events proved them right. And never mix up how likely an event is with how bad it would be.
That last point is worth pausing on, because it aims at the alarmist as much as at the prophet. To decide whether to act against a risk, you weigh two things: how likely it is to happen, and what it would cost if it did. A very unlikely but catastrophic danger can therefore justify acting as much as a likely but moderate one.
Bostrom showed where that arithmetic breaks. Picture a stranger stopping you in the street and demanding ten euros, failing which he will use his powers to wipe out a billion lives. You don't believe a word of it. Yet you only have to grant his story the faintest chance for the arithmetic to tell you to pay up, and if he senses you hesitating, all he has to do is name a bigger number. This is the trap Bostrom calls Pascal's mugging: whoever makes the threat also sets the size of the stake, so he can always raise it until it overwhelms your scepticism, however firm.
The debate on AI takes this shape the moment the word extinction is uttered. Objecting that a scenario is unlikely no longer has any effect, since the reply is always available: even at one chance in a thousand, what's at stake forbids taking the risk. Because the objection can never win, nobody wastes time establishing the probability any more, and a conversation about what is going to happen becomes a conversation about what we can't afford to rule out. Hence the separation the authors call for, on the IPCC model: one group says what is going to happen, another what it would cost, a third what can be done about it. Disputing the first isn't denying the second.
The conclusion fits in one line, and it cuts both ways: the poor quality of dated predictions doesn't prove that AI is harmless, and the gravity of the risk doesn't make those predictions any better.
The authors draw the limits of their own scope, as it happens. Whether AGI arrives remains an open question, the risk may be serious, and their audit leaves both questions exactly where it found them. They even note that the strongest arguments about the danger rest on no date at all: they bear on the kind of system the discipline has chosen to build, and on how hard it is to supervise a machine whose reasoning you can't follow.
They go further, and this is what surprised me most: they exclude from their record the texts that warn without dating, Wiener, Bush, the OpenAI charter. Not out of leniency, but because refusing to manufacture false precision is, in their eyes, a virtue, one that leaves readers to judge the plausible timeline instead of handing them one that rests on nothing. That's the distinction I'm keeping. The problem isn't the alarm, it's the borrowed precision.
That leaves the question of how such rigour is obtained, and for this the authors have the best analogy in the paper. Trial registration, which obliges a lab to declare what it is looking for before it looks, didn't take hold by asking researchers to be more scrupulous. It took hold the day the journals refused to publish unregistered trials. Voluntary rigour never arrives on its own. It arrives when somebody makes it a condition of entry.
In the meantime, a prediction that cannot be refuted goes on being treated as evidence: in our investment committees, in our scoping notes, on our panels. And when it comes from someone with the means to make it come true, the day it does still proves nothing. The cost isn't paid by the people who make these claims. It's paid by the people who, in the same room, put forward something precise and checkable, and watch it read like all the rest.
Sources
- Luciano Floridi, Jessica Morley and Claudio Novelli, Not Even Wrong 2: An Audit of Public AGI Prediction, 1950-2026, 13 September 2026, freely available on SSRN.
- On Pauli's phrase and what it names: Not even wrong.
- Nick Bostrom, Pascal's mugging, Analysis 69(3), 2009, pages 443-445. DOI: 10.1093/analys/anp062.
- On the 1973 audit and the first AI winter: Lighthill report.
- The scenario mentioned above: AI 2027, AI Futures Project, April 2025, ai-2027.com.