What is the Turing test? And why it stopped mattering
Turing never proposed the imitation game as a definition of thinking. He proposed it to replace a question he considered meaningless. Machines have now passed versions of it, sometimes judged more human than actual humans, and the striking thing is how little that settled.
For seventy years the Turing test was the popular shorthand for machine intelligence: build something that can hold a conversation well enough to be mistaken for a person and you have built a mind. Machines have now done that, in experiments close to the form Turing described, and in at least one case the machine was judged to be human more often than the actual humans were. If the test meant what people took it to mean, that should have been the end of a long argument. It was not, and the reason is instructive. Turing never offered the imitation game as a definition of thinking. He offered it as a replacement for a question he considered meaningless, and now that machines can pass versions of it, we have learned that the substitution failed, because indistinguishable behaviour tells you about performance and about human credulity rather than about what is happening inside.
This guide covers what Turing actually wrote and the philosophical move most summaries omit, how the test works and how it has been run, what the results were, why being judged more human than humans is the finding that breaks the whole framing, the criticisms that turned out to matter, and what the field uses instead. The test's most valuable legacy is not the test.
What Turing actually proposed
The source is a 1950 paper, Computing Machinery and Intelligence, and its opening move is the part that matters most and gets quoted least. Turing began by saying he would consider the question "Can machines think?" and then almost immediately abandoned it, on the grounds that the terms were too ill-defined for the question to be settled by anything other than a survey of how people happened to use words. He judged it, in his phrasing, too meaningless to deserve discussion.
Rather than define thinking, he replaced the question with a different one that could actually be answered by running an experiment. That substitution is the whole intellectual content of the proposal. Turing was not saying that conversational indistinguishability constitutes thought. He was saying that the original question could not be productively investigated, and that here was a concrete alternative one could investigate instead, with the implication that if a machine performed well at it, our reluctance to use the word "thinking" would come to seem like stubbornness about vocabulary.
Read that way, the imitation game is a philosophical device for escaping an unproductive debate, not a criterion for consciousness or understanding. Almost every popular account inverts this, presenting the test as Turing's definition of intelligence, which is close to the opposite of what he wrote.
The game itself, and a detail usually dropped
Turing described the setup by analogy with a party game. Three participants: a man, a woman, and an interrogator who communicates with both by text and must work out which is which, while the man tries to be mistaken for the woman. Then he posed his question: what happens if a machine takes the part of the man? Will the interrogator misidentify as often as before?
Two features of that description deserve keeping. First, the game is three-party: the judge converses with a human and a machine at the same time and must distinguish them, which is a comparative judgment rather than an isolated one. Many later experiments used a simpler two-party form in which a judge speaks to one participant and guesses, and that is an easier test, because it removes the human baseline from view. Second, the medium is text specifically, which Turing chose to prevent the judge from being influenced by voice or appearance, a deliberate separation of the capacity being tested from its packaging.
Turing also made a prediction, which is where the familiar numbers come from. He estimated that within about fifty years machines would play the game well enough that an average interrogator would have no more than a seventy percent chance of correct identification after five minutes of questioning. That prediction has often been converted into a pass mark of thirty percent, which is a shaky reading, since a forecast is not a definition of success. A better threshold is chance: if judges cannot do better than a coin flip, they cannot distinguish at all.
Has it been passed?
The honest answer is that several versions have been passed, that the versions differ in difficulty, and that "the Turing test" is not one fixed thing.
Careful experiments in the mid-2020s produced clear results. In a two-party study, a leading model was judged human in roughly half of conversations, an older model performed at about chance, and a simple rule-based chatbot from the 1960s was judged human around a fifth of the time, showing the design could still discriminate. Real humans in the same study were identified correctly about two thirds of the time.
A subsequent study moved to the three-party structure Turing described, with judges comparing a human and a machine simultaneously across hundreds of participants. A leading model, prompted to adopt a persona, was judged to be the human in a clear majority of interactions, comfortably beating chance. That is a pass under any reasonable reading of the original proposal.
Two qualifications belong alongside that. Conversations were short, on the order of five minutes, and the models were prompted to play a character. Both matter, because a brief exchange limits how much probing is possible, and persona prompting means the experiment partly measures how well a system can be configured to perform humanness rather than what it does unprompted. Versions with expert judges and extended conversation remain unpassed, and the difference between a casual five-minute chat and an hour of adversarial questioning by someone who knows the failure modes is very large. So the accurate summary is that the test as Turing sketched it has been passed, while harder variants have not.
The result that breaks the framing
Among these findings sits one that deserves more attention than it usually gets, because it undermines the test more effectively than any philosophical argument.
The machine was judged to be human more often than actual humans were. In the three-party experiment, the model prompted with a persona was selected as the human at a rate exceeding what real people achieved.
Sit with that. If a test is meant to detect humanness, and a machine scores higher on it than humans do, the test is not measuring humanness. It is measuring something else: performed humanness, the successful production of the cues people associate with a person, which is a skill at which a system trained on an enormous quantity of human text can straightforwardly exceed any individual. Real people are idiosyncratic, distracted, uneven, and unwilling to perform; a system optimised to seem human has no such handicaps.
This is why the result settled nothing. It demonstrated that these systems are extremely good at a specific and narrow thing, producing text that reads as human, which we already knew, and it revealed that the measurement was always partly about the judges. A test that a machine can win by being more typically human than humans is a test of stereotype-matching and of what the audience expects, not a probe of what is happening inside.
Why indistinguishability was never enough
Behind the empirical result is a structural problem that critics identified long before any machine passed.
The core objection is that identical behaviour can arise from very different internal processes, so behaviour underdetermines the facts you actually want. The most famous version of this argument imagines someone locked in a room manipulating symbols in a language they do not understand, following rules well enough to produce responses indistinguishable from a fluent speaker's, and asks whether the room understands. Whatever you conclude, the argument establishes that passing a behavioural test does not by itself settle the internal question, which is precisely the question people thought the test answered.
There is also a narrower objection with more practical bite: the test rewards deception. Success requires not only competence but the concealment of competence, since a machine that answered arithmetic instantly and never made a typing error would give itself away. Turing anticipated this, noting that a machine might pause or make deliberate mistakes. That is a strange property for a measure of intelligence, since it means part of what is being scored is skill at pretending to be less capable. This is why the test has been described as one not for machines to pass but for humans to fail: it measures whether people can be fooled, and people can be fooled by considerably less than a mind.
The final objection is scope. Conversation is one behaviour among many. Assessing intelligence through a single channel ignores everything else, which is exactly the lesson of Moravec's paradox: a system can be superb at fluent language and hopeless at things a small child manages, and a test conducted entirely through text cannot see any of that.
The term is also routinely misused
Part of what makes public discussion confusing is that "passed the Turing test" gets applied to things that are not the Turing test at all. It has been used for an engineer's personal impression that a chatbot was sentient. It has been used for studies comparing a model's statistical behaviour on psychological questionnaires against human distributions. It has been used for any occasion on which someone was briefly fooled by generated text. None of these is an imitation game, and treating them as equivalent produces the impression of a threshold being crossed repeatedly and meaninglessly.
When you encounter the claim, three questions clarify it. Was the format two-party or three-party, since the latter is harder. How long were the conversations, since brief exchanges are much easier to pass. And were the judges naive or expert, since people who know what to probe for perform far better. The phrase alone conveys almost nothing without those answers.
What replaced it
The field largely moved on well before the test was passed, and knowing what it moved to is more useful than the test itself.
Capability benchmarks became the working standard: specific tasks with specific correct answers, measuring performance on mathematics, coding, reasoning, or knowledge rather than on the ability to seem human. This is a real improvement, because it measures capability directly rather than through an audience's perception. It also brought its own well-documented problems, including saturation, contamination, and optimising to the measure, and no single benchmark claims to capture intelligence.
The deeper replacement is interpretability: rather than inferring from behaviour, look inside and examine what the system actually represents and computes. This is the direct response to the underdetermination problem, since the reason behaviour cannot settle the internal question is that behaviour is the wrong evidence, and internal evidence is the right kind. It is difficult and incomplete, but it is aimed at the question the Turing test could not reach.
Underneath both is a shift in framing. Rather than asking whether a system is intelligent, researchers increasingly ask which specific capacities it has and how strongly, because "intelligence" bundles several things that these systems have shown can come apart. That is the same decomposition that makes the understanding debate tractable and that predicts what these systems can and cannot do. One test yielding one bit of information was never going to characterise a system with a jagged and uneven capability profile.
What Turing got right
None of this makes the paper a historical curiosity, and it would be ungenerous to leave it there.
Turing was right that "can machines think" is not a productive question as posed, and the subsequent seventy years have supported him: the argument has produced far more heat than resolution, and the useful progress came from replacing it with narrower questions, which is exactly his method even if his particular substitution did not hold up. He was right that behaviour is where evidence has to start, since we have no other access to other minds, including human ones. He was right to insist on separating the capacity from its packaging, and the choice of a text-only channel to avoid judging a machine on its voice or appearance was a careful piece of experimental design. He anticipated most of the objections to his own proposal and answered them in the paper, which is more than most of his critics did. And his timing prediction, roughly fifty years for conversational imitation at that standard, was wrong by only a couple of decades on a question where most forecasts have been wrong by much more.
What he could not anticipate was the specific way it would be passed: by a system trained on an enormous corpus of human text, which acquires the surface of human expression with extraordinary fidelity while remaining thin in ways his test could not detect. The imitation game assumed that conversation was hard enough to require the underlying capacities. That assumption turned out to be false, and discovering it was false is worth something.
The short version
The Turing test comes from Alan Turing's 1950 paper, in which he considered the question "Can machines think?", judged it too ill-defined to be settled, and replaced it with an experiment: a text-only imitation game in which an interrogator converses with a human and a machine at once and tries to identify which is which. That substitution, not a definition of thinking, was the point, and most popular accounts get this backwards. Turing predicted machines would play it well enough within about fifty years that interrogators would be right no more than seventy percent of the time after five minutes, a forecast that became an arbitrary pass mark. Experiments in the mid-2020s passed both the simpler two-party form and the three-party form Turing described, with the striking result that a leading model prompted with a persona was judged human more often than actual humans were. That finding undermines the test, because a measure of humanness that machines beat humans at is measuring performed humanness and audience expectation rather than anything internal. Underlying objections had long been raised: identical behaviour can come from different internal processes, so behaviour underdetermines the question; the test rewards concealment of capability, making deception part of what is scored; and conversation is one narrow channel. The field moved to capability benchmarks, which measure directly but bring their own problems, and to interpretability, which examines internal structure rather than inferring from behaviour.
The idea to hold onto is that the Turing test was a device for escaping an unanswerable question rather than a definition of intelligence, and passing it taught us that fluent conversation is easier to produce than anyone expected and tells us far less than anyone hoped, which is why the field replaced a single behavioural verdict with specific capability measurement and with looking inside.
Common questions
What is the Turing test? It is a test proposed by Alan Turing in his 1950 paper Computing Machinery and Intelligence, in which a human interrogator holds text-only conversations with a human and a machine simultaneously and tries to determine which is which. If the interrogator cannot reliably tell them apart, the machine is said to have passed. Turing specified text to prevent judgments based on voice or appearance. Importantly, he did not present it as a definition of thinking. He introduced it as a replacement for the question "Can machines think?", which he considered too ill-defined to answer, offering instead something that could actually be tested.
Has AI passed the Turing test? Yes, in versions close to Turing's description, though the answer depends on which version. Studies in the mid-2020s found that leading language models were judged to be human at rates well above chance in both the simpler two-party format and the three-party format Turing originally described, with hundreds of participants. Important qualifications apply: conversations were short, around five minutes, and models were prompted to adopt a persona, so part of what was measured is how well a system can be configured to perform humanness. Harder variants using expert judges and extended adversarial conversation have not been passed.
Does passing the Turing test mean AI is intelligent? No, and the researchers who ran the passing experiments generally said so themselves, describing the test as a measure of substitutability rather than of intelligence. The decisive evidence is that in the three-party study, the machine was judged to be human more often than actual humans were. A test of humanness that machines win against humans is measuring performed humanness and audience expectation, not any internal property. More fundamentally, identical behaviour can arise from very different internal processes, so passing a behavioural test cannot settle questions about what is happening inside a system.
What did Turing actually mean by the imitation game? He meant it as a substitute for a question he thought could not be productively investigated. His paper opens by considering "Can machines think?" and immediately setting it aside, on the grounds that the terms are too ill-defined for the question to be settled by anything other than a survey of ordinary usage. Rather than define thinking, he replaced the question with a concrete experiment that could be run. The implication was that a machine performing well at it would make our reluctance to say it thinks look like stubbornness about vocabulary. The philosophical move, not the criterion, was the substance.
Why do critics say the Turing test is flawed? Three main objections. First, behaviour underdetermines internal facts, since identical outputs can be produced by very different processes, which is the point of thought experiments about symbol manipulation without comprehension. Second, the test rewards deception, because a machine that answered instantly and never erred would expose itself, so part of what is scored is skill at concealing capability, which is a strange property for a measure of intelligence. Third, conversation is a single narrow channel, and a system can be superb at fluent text while failing at tasks a small child handles, none of which a text-only test can detect.
What replaced the Turing test? Mainly two things. Capability benchmarks became the working standard, measuring performance on specific tasks with known correct answers such as mathematics, coding, or reasoning, which assesses capability directly rather than through an audience's perception, though these bring their own problems including saturation and contamination. More fundamentally, interpretability research examines what a system internally represents and computes rather than inferring from behaviour, which addresses the underdetermination problem directly. Underneath both is a shift from asking whether a system is intelligent to asking which specific capacities it has and how strongly.
Why do people keep saying AI passed the Turing test when it did not? Because the phrase is applied loosely to things that are not the imitation game. It has been used for an individual's impression that a chatbot seemed sentient, for studies comparing a model's statistical responses on questionnaires against human distributions, and for any occasion when someone was briefly fooled by generated text. None of these is the test. When you see the claim, ask whether the format was two-party or three-party, how long the conversations ran, and whether judges were naive or expert, since those three factors change the difficulty enormously and the phrase alone conveys very little.
Is the Turing test still useful for anything? It remains useful as history, as a teaching device, and as a measure of something real but narrow: whether a system can substitute for a person in a text conversation without being noticed. That property matters practically, because it bears directly on impersonation, fraud, and the reliability of online interaction, which is a genuine social concern even though it says nothing about intelligence. Turing's deeper contribution also survives: his method of replacing an unanswerable question with a narrower testable one is exactly what the field now does, even though his particular substitution turned out not to measure what people hoped.
Sources & further reading
The primary literature behind the claims above, drawn from the concept entries this post links to, so a claim carries the same source here as it does there.
- Moravec, H. (1988), Mind Children — the original formulation of the observation. :: https://www.hup.harvard.edu/books/9780674576186 Moravec's Paradox
- Dell'Acqua et al. (2023), Navigating the Jagged Technological Frontier — the modern empirical form of uneven capability. :: https://www.hbs.edu/faculty/Pages/item.aspx?num=64700 Moravec's Paradox
- Brooks, R. (1990), Elephants Don't Play Chess — a contemporaneous argument for the primacy of perception and embodiment. :: https://people.csail.mit.edu/brooks/papers/elephants.pdf Moravec's Paradox
Related articles
- What can AI do, and what can't it? A predictive mapAny list of what AI can do is out of date before you finish reading it. What does not go out of date is the set of task properties that predict whether AI will be good at something, and they have almost nothing to do with how hard the task feels to a person.
- Does AI actually understand? Why the debate is stuckAsk whether an AI model understands anything and you get two confident answers from serious people. The reason neither side wins is that understanding was never one property. It is a bundle of capacities that always arrived together in humans and comes apart in these systems.
- AI alignment and safety, without the hype or dismissalAI safety is discussed either as impending doom or as overblown hype, and neither framing helps you understand it. The actual landscape: what alignment means, the concrete risks experts agree on, the speculative ones they don't, and why serious people land in very different places.
- Why AI benchmarks mislead: contamination, gaming, saturationEvery model launch leads with benchmark scores, and buyers read them like thermometer readings. They are closer to opinion polls: directionally useful, methodology-dependent, and easy to game. Here is why the numbers mislead, from training-data contamination to Goodhart's law to saturation, and how to read them without being fooled.