Home/Blog/Evaluation & evidence/Superintelligence: the empirical record is zero
Superintelligence: the empirical record is zeroSixty years, zero instances of sustained open-ended self-improvementSixty years, zero instances of sustained open-ended self-improvement
Sixty years, zero instances of sustained open-ended self-improvement

Superintelligence: the empirical record is zero

Sixty years after the intelligence explosion was described, no system has demonstrated sustained open-ended self-improvement. The public forecasts come from five people with the same financial interest.

TL;DR. Sixty years after the intelligence explosion was first described, the empirical record contains zero instances of it. No architecture has demonstrated sustained, open-ended, autonomous self-improvement. That is not an argument that it cannot happen; it is a statement about what has been observed. Meanwhile the loudest forecasts come from five people who run companies whose valuations depend on those forecasts being believed, and surveyed researchers give substantially longer timelines. The deepest problem is not disagreement about dates. It is that "superintelligence" has no operational definition, so the question of when it arrives is not yet an empirical question at all and every argument about it rests on capability claims from a field that this site has spent a hundred articles showing cannot measure capability.

---

The idea is sixty years old. A machine capable enough to improve its own design would produce a successor more capable still, which would repeat the process, and the resulting escalation would leave human capability behind quickly enough that the transition could not be managed once begun.

It is a clean argument and its structure is worth respecting. It does not require any particular technology. It requires only that self-improvement be possible and that each round of it be at least as effective as the last.

As of the most recent survey of the literature, the empirical record contains zero instances of the phenomenon. No architecture has demonstrated sustained, open-ended, autonomous self-improvement.

That sentence is the whole of the evidence, and almost everything written about superintelligence is an argument about what to infer from it.

Two inferences are available and both are respectable. The first: an event with no precedent, no observed instance and no working mechanism deserves the treatment we give other unprecedented events, which is serious contingency planning rather than scheduling. The second: an event that would happen once, by construction, cannot be forecast from base rates, because the absence of prior instances is exactly what the theory predicts right up until the moment it fails.

This article does not pick between those. It maps what is actually known, separates the empirical claims from the definitional ones, and identifies which parts of the debate could be settled by evidence and which cannot.

The definitional problem, which comes first

Almost every public disagreement about superintelligence is a disagreement about thresholds conducted as though it were a disagreement about facts.

There is no operational definition. Not a vague one, not a contested one: none that would let two people watching the same system agree on whether it had arrived.

Consider what would have to be specified. Superintelligent at what? A system that exceeds every human at chess, protein folding and arithmetic already exists in pieces. One that exceeds every human at every task is a different claim and depends entirely on the task list. Exceeds which humans, the median or the best specialist? Measured how, on benchmarks the system may have seen, or on novel problems, and who writes them? And sustained over what period, since a system that is superhuman on a Tuesday and degraded by a model update on Thursday is a different thing from a permanent capability.

Without answers, "when will superintelligence arrive" is not a question about the world. It is a question about where a person chooses to place a line, and people place it differently, which is sufficient to explain most of the observed disagreement without anyone being wrong about any fact.

The serious literature noticed. The term "ultraintelligence", once central to these debates, has largely disappeared from technical writing. Contemporary work prefers "frontier AI", "general purpose AI" and "transformative AI", terms chosen because they attach to measurable capabilities rather than to a hypothetical threshold.

That vocabulary shift is the field voting with its language, and it is a stronger signal than any individual position. Researchers moved to terms they could operationalise, and the ones who kept the older vocabulary are largely writing for a public audience rather than for each other.

Who is forecasting, and why the forecasts are not independent

The public timelines cluster, and the clustering is usually presented as convergence.

The chief executive of one major laboratory has said systems broadly better than almost all humans at almost all things could arrive by 2026 or 2027. The head of another has given five to ten years, with AGI around 2030. A third has said AGI will probably be developed within the current US presidential term. A prominent technologist has predicted a system more intelligent than any single human by the end of 2025 and one exceeding all humans combined by 2030. A former chief scientist declines to give dates while having founded a company whose premise only makes sense on a short timeline.

Five forecasts, and they are not five pieces of evidence.

Every one comes from a person whose company's valuation, recruitment and access to capital improve when short timelines are believed. That does not make them wrong. Insiders frequently know things outsiders do not, and the people building the systems have information nobody else has.

But five correlated forecasts from parties with the same interest is one forecast repeated, and it should be weighted accordingly. The convergence that looks like independent confirmation is what you would expect from a shared incentive whether or not the underlying claim is true.

And the researchers who do not run laboratories give longer timelines. Surveys of the broader expert population consistently return more conservative estimates than executive statements, and one large survey placed a median 50% probability on superintelligence arriving within thirty years of human-level machine intelligence, which is a forecast conditioned on an event that has not happened either.

The most rigorous public forecasting exercise, a detailed month-by-month scenario built from tabletop exercises and feedback from more than a hundred experts, is worth reading precisely because it is explicit about its assumptions. It presents two endings rather than one. And several of its own contributors published their disagreements with it, which is more intellectual honesty than the genre usually contains.

The one strong datapoint, presented fairly

There is a real observation that supports short timelines and it deserves to be stated at full strength rather than dismissed.

At one frontier laboratory, the model reportedly writes over 80% of the code that gets merged.

That is a striking figure and it is the closest thing to evidence of the loop beginning to close. If AI research is bottlenecked on engineering throughput, and a system is doing most of the engineering, the argument that improvement compounds has an empirical foothold.

The same source identifies precisely what has not moved: research direction-setting. The systems execute a great deal and do not choose which problems matter. Deciding what to work on, judging which failed experiment is informative and which is noise, and knowing when a research direction is exhausted remain human.

Whether that gap closes is the actual open question, and it is a much narrower and more tractable question than "will superintelligence arrive". It is also the one to watch, because it is the first version of this debate that could be settled by observation rather than argument.

What "self-improvement" actually refers to

A survey of 1,250 papers on self-improvement published between 2024 and 2026 found the term covering at least four distinct things, and conflating them is responsible for a substantial share of the confusion.

Inference-time revision. A model reviews and rewrites its own output before returning it. Real, useful, and it improves a response rather than the model. Nothing persists.

Training on self-generated data. A model produces training examples, is trained on them, and improves measurably. This is real and it works, and it is bounded: quality degrades as the generated distribution drifts from anything grounded, and the bound is empirical rather than theoretical.

Agents that rewrite their own code. Systems that modify their own scaffolding, tooling or prompts. Genuine and narrow, since the modification is to the harness rather than to the weights.

Autonomous research. Systems that conduct AI research end to end, choosing problems, designing experiments and interpreting results. This is the one the intelligence explosion argument requires, and it is the one with the least demonstrated.

Reading a result about the first three as evidence about the fourth is the most common error in this discussion, and it happens constantly because the same phrase covers all four.

Why the goalposts keep moving, and why that is not dishonest

A recurring complaint is that the definition of AI keeps shifting: once a machine can do something, it stops counting as intelligence. Chess was the test until a machine won it, then it was a search problem. Translation, image recognition, natural conversation, competitive programming, each was a marker and each was reclassified afterwards.

The complaint is usually made as an accusation of bad faith. It is more interesting than that.

Each reclassification was correct. When a chess engine won, the field learned something real: that world-class chess requires less general intelligence than anyone had assumed. That is a finding about chess, not a retreat about intelligence. The same happened with translation, which turned out to need less understanding than expected, and with image classification.

So the moving goalposts are a genuine research result being mistaken for a rhetorical dodge. We keep discovering that tasks we used as proxies for general capability were poorer proxies than we thought. The honest summary is that we do not have a good test, and every candidate test has been shown inadequate by something passing it.

Which has an uncomfortable consequence for this whole debate. If every operationalisation of intelligence has failed on contact with a system that satisfied it, there is no strong reason to expect the next one to hold. A benchmark someone proposes today as the marker for general capability is, on the historical record, likely to be reclassified as a narrow skill within a few years of being beaten.

That is an argument for humility in both directions. It means claims that a system has crossed a general threshold should be doubted, because every previous such claim was withdrawn. It also means claims that a system is definitely narrow should be doubted, because the same reclassification that shrinks a capability after it is achieved makes it very difficult to notice generality accumulating.

The practical form: be sceptical of anyone who tells you a specific benchmark result settles this, in either direction. Sixty years of the goalposts moving is sixty years of evidence that no single result has settled it yet.

The bottleneck arguments

Three constraints get proposed. Their status is different in each case and worth separating.

Compute. AI research needs both cognitive labour and experimental compute. If compute is the binding constraint, then automating the cognitive half does not produce runaway improvement, because the experiments still have to run on physical hardware that must be built, powered and cooled. Whether a software-only explosion is possible is an active and unresolved technical debate, with serious people on both sides.

Experimental latency. This is the most concrete objection and it comes from people who contributed to the leading forecast. Real-world experiments take weeks, months or years to return results. A system that can generate a thousand hypotheses per hour still waits for the training run, the wet lab or the deployment to report back. Cognitive speed does not compress physical time, and a loop is only as fast as its slowest stage.

Context and judgement. Human researchers carry a great deal of unwritten knowledge about their organisation, their field's history, which approaches were already tried and why they failed. That knowledge is largely not in any document. Whether it can be acquired, and how quickly, is unknown.

None of these establishes that the escalation cannot happen. Each establishes that a specific mechanism people assume would produce it may not, and the honest position is that the bottlenecks are real and their size is unmeasured.

The measurement problem, which is this site's actual contribution

Here is the argument that follows from everything else on this site, and it is the reason this article exists.

Every claim about superintelligence is a claim about capability. And the field cannot measure capability reliably.

That is not a rhetorical flourish. It is the finding of a hundred articles that measured what happens when these systems meet a domain with an evidence base.

In medicine, 1,524 devices have regulatory clearance and 1.6% of a reviewed sample cited a randomised trial. Clearance certifies resemblance to an existing product, not benefit.

In education, AI tutoring moved satisfaction by 0.93 and confidence by 0.91 against 0.53 for knowledge. Students feel roughly twice the improvement they demonstrate, and the authors rated the certainty of all three as very low.

In customer service, a platform can report 90% deflection on a 40% resolution rate, and both figures are accurate measurements of different events.

In journalism, assistants misrepresented news content in 45% of responses when evaluated by journalists against professional criteria.

In science, a materials model predicted 380,000 stable compounds and 736 have been made.

Now consider what a capability forecast requires. It requires a measurement of current capability, a measurement of the rate of change, and confidence that both extend to the tasks that matter. The five findings above are all cases where a widely reported capability number turned out to measure something else.

A benchmark score is a measurement on a fixed test whose contents may be in the training data, administered under conditions chosen by the party reporting it, on tasks selected because they are measurable rather than because they matter. Extrapolating a curve through those points and concluding something about the year a system exceeds all human capability is applying enormous inferential weight to numbers that this corpus has repeatedly shown do not survive contact with a domain.

This does not argue that superintelligence will not happen. It argues that the quantitative case for any particular date is weaker than it appears, because the inputs are weaker than they appear.

And it cuts in both directions, which is the part usually left out. If capability measurement is unreliable, the reassuring numbers are as suspect as the alarming ones. A benchmark showing a system fails at long-horizon planning is the same kind of artifact as one showing it succeeds. Anyone concluding from current evaluations that the systems are safely limited is making the same error in the opposite direction.

What would count as evidence

The most useful thing an article like this can do is specify what would change the picture, because a position that no observation could alter is not a position about the world.

Evidence for. A system that autonomously identifies a research problem nobody assigned it, designs an experiment, interprets a negative result correctly, and changes direction based on it. Not a system that executes a specified research plan well. The signature is choosing what to work on, and it is observable.

A second: a documented instance of successive model generations where each was substantially designed by its predecessor and the improvement per generation did not shrink. The last clause is what matters, since diminishing returns per round is exactly what distinguishes an escalation from a plateau.

Evidence against. Sustained investment in autonomous research capability producing improvements that decline per unit of compute, over a period long enough to rule out a single architecture's limits. Or a well-specified capability that repeated scaling fails to reach.

And a caution about both. The International AI Safety Report has noted uneven advances including declining performance on longer tasks. That is a real observation and it is not decisive, since a plateau in one dimension has repeatedly preceded a jump from an architectural change rather than from more of the same.

The reason to write the criteria down now is that they are much harder to write honestly after the fact, and a field that has already revised what counts as AGI several times should be sceptical of its own capacity to notice a threshold it did not define in advance.

What is unresolved

Whether direction-setting is a capability or a position. Choosing which problems matter may be a skill that can be learned, or it may be a function of standing in a community, holding a budget and having something at stake. If the second, no amount of capability closes the gap, and the question stops being technical.

Whether the compute bottleneck binds. Serious technical work disagrees on whether a software-only intelligence explosion is possible, and the disagreement is about parameters nobody has measured rather than about principles.

Whether the concept survives. The terminology retreat from ultraintelligence to frontier AI may reflect intellectual maturation, or a field avoiding a question it cannot operationalise. Both stories fit.

And whether any of this is the right frame. A world where narrow systems become extremely capable across many domains without anything crossing a general threshold produces most of the practical consequences discussed under this heading, with none of the theoretical structure. That scenario is under-analysed precisely because it is less interesting to argue about.

The counter-argument

Absence of precedent is exactly what the theory predicts. Zero prior instances is not evidence against a one-time event. There were zero nuclear detonations until there was one, and the record of zero was not informative about the physics. Using the empty record as a reason for scepticism misunderstands what kind of claim is being made.

Insiders knowing more is a real consideration, not just an interest. The executives with short timelines have seen unreleased systems and internal results. Discounting their forecasts for conflict of interest is reasonable and it discards information nobody outside has, and the correct weight is not zero.

The definitional complaint can be a way of avoiding the question. Many important things lack operational definitions, including intelligence itself, and we reason about them anyway. Demanding a threshold before discussion is a standard that would have prevented most useful thinking about most novel risks.

And the measurement argument cuts against caution too. If capability measurement is unreliable, that unreliability is symmetric. It is not available as a reason for calm. An argument that current benchmarks cannot support a confident timeline is also an argument that they cannot support confident reassurance, and the honest reading of a poorly-measured domain is wider uncertainty in both directions rather than a shift toward the comfortable end.

The strongest version of the case for taking short timelines seriously is not any forecast. It is that the cost of being wrong is asymmetric, and a decision framework that requires proof before preparation is poorly suited to events that cannot be prepared for afterwards.

The short version

The intelligence explosion argument is sixty years old, structurally clean, and requires only that self-improvement be possible and that each round be at least as effective as the last. The empirical record contains zero instances. No architecture has demonstrated sustained, open-ended, autonomous self-improvement.

The definitional problem comes before the empirical one. There is no operational definition of superintelligence that would let two people watching the same system agree on whether it had arrived. Superintelligent at what, against which humans, measured how, sustained how long. Without answers, "when will it arrive" is a question about where someone places a line rather than about the world, which explains most public disagreement without anyone being wrong about a fact. The technical literature noticed and moved to frontier AI, general purpose AI and transformative AI, terms chosen because they attach to something measurable.

The public forecasts are not independent evidence. Five prominent short timelines come from five people whose companies benefit when short timelines are believed. That does not make them wrong, and insiders do know things, and five correlated forecasts from parties sharing an interest is one forecast repeated. Surveyed researchers outside the laboratories give longer estimates.

One datapoint truly supports the short case: at one frontier laboratory the model reportedly writes over 80% of merged code. The same source identifies what has not moved, which is research direction-setting. Systems execute; they do not choose which problems matter. Whether that gap closes is the real open question, and it is narrower and more answerable than the one usually asked.

And the argument this site is positioned to make: every claim about superintelligence is a claim about capability, and the field cannot measure capability reliably. Medicine clears devices on resemblance rather than outcome. Education measures satisfaction moving twice as far as knowledge. Support platforms report 90% deflection on 40% resolution. Assistants misrepresent news 45% of the time. A materials model predicted 380,000 compounds and 736 exist. Extrapolating a capability curve through benchmark points is applying enormous inferential weight to numbers that repeatedly fail on contact with a domain.

That cuts both ways, and the symmetry is the honest conclusion. If the measurements cannot support a confident date, they cannot support confident reassurance either. The correct response to a badly measured question is wider uncertainty in both directions, not a comfortable answer.

Common questions

What is superintelligence? A hypothetical system substantially exceeding human capability across essentially all domains, as distinct from narrow systems that already exceed humans at specific tasks. The central difficulty is that no operational definition exists. Nobody has specified which tasks, measured against which humans, under what conditions, sustained for how long, in a way that would let two observers of the same system agree on whether the threshold had been crossed.

Has any AI system improved itself? Not in the sense the intelligence explosion argument requires. A survey of 1,250 papers found "self-improvement" covering four distinct things: revising output at inference time, training on self-generated data, agents rewriting their own scaffolding, and autonomously conducting research. The first three are real and bounded. The fourth is the one the argument needs and has the least demonstrated. The empirical record contains zero instances of sustained, open-ended, autonomous self-improvement.

When will superintelligence arrive? Nobody knows, and the question is not currently answerable in the form it is asked, because the event has no agreed definition. The public forecasts range from 2026 to the 2030s and come predominantly from executives whose companies benefit from short timelines being believed. Surveyed researchers outside frontier laboratories give consistently longer estimates. One large survey placed a median 50% probability on superintelligence within thirty years of human-level machine intelligence, itself an event that has not occurred.

Why do AI company CEOs predict such short timelines? They have information nobody outside has, having seen unreleased systems and internal results, and they have a financial interest in the forecast being believed, since valuations, recruitment and capital access all improve with short timelines. Both are true simultaneously. The important point is that five such forecasts are not five independent pieces of evidence: correlated predictions from parties sharing an interest should be weighted as roughly one.

What is recursive self-improvement? The proposed mechanism by which a system capable enough to improve its own design produces a more capable successor, which repeats the process. It requires only that self-improvement be possible and that each round be at least as effective as the last. The second condition is doing more work than it appears, since diminishing returns per round is what separates an escalation from a plateau, and nothing currently establishes which would occur.

What would stop an intelligence explosion? Three bottlenecks are proposed and none is established. Compute, since AI research needs experimental hardware as well as cognitive labour, and whether a software-only explosion is possible is actively disputed. Experimental latency, since real-world experiments take weeks to years and cognitive speed does not compress physical time. And context, since human researchers carry substantial unwritten knowledge about what has already been tried and why it failed. Each identifies a mechanism that may not work as assumed rather than proving the escalation impossible.

What evidence would show superintelligence is close? A system that autonomously identifies a research problem nobody assigned it, designs an experiment, correctly interprets a negative result and changes direction because of it. The signature is choosing what to work on rather than executing a plan well. Alternatively, successive model generations each substantially designed by its predecessor, where the improvement per generation does not shrink. That last clause matters, since diminishing returns is the difference between an escalation and a plateau.

Why is it so hard to evaluate claims about superintelligence? Because every such claim is a capability claim, and capability measurement in this field is documented to be unreliable. Regulatory clearance in medicine certifies resemblance rather than benefit. Education trials show satisfaction moving twice as far as knowledge. Support platforms report deflection rates that count abandoned conversations as successes. Extrapolating a curve through benchmark scores requires trusting numbers that repeatedly fail when tested against a domain, and that unreliability is symmetric: it undermines confident reassurance exactly as much as confident alarm.

Learn the concepts

← All posts