Does AI actually understand? Why the debate is stuck
Ask whether an AI model understands anything and you get two confident answers from serious people. The reason neither side wins is that understanding was never one property. It is a bundle of capacities that always arrived together in humans and comes apart in these systems.
You can watch a model produce a lucid explanation of a concept, apply it correctly to a new case, and then fail at something a child would find trivial. Ask a serious researcher whether it understands what it is talking about and you will get a confident answer. Ask a different serious researcher and you will get the opposite answer, equally confident. This has been going on for years without converging, which is a signal worth attending to: when a question resists resolution while the evidence accumulates, the problem is often the question. "Does AI understand" resists an answer because understanding was never one property. It is a bundle of capacities, fluency, generalization, internal models of a domain, grounding in reality, and reliable self-knowledge, that always arrived together in humans and come apart in these systems, so a model can be strong on some and nearly absent on others, which makes both "it truly understands" and "it is just autocomplete" confidently wrong.
This guide separates the question from the one it is usually confused with, states both sides at their strongest rather than as caricatures, sets out the decomposition that explains why they talk past each other, shows how the specific failures we observe follow from it, and describes what would actually settle the matter. It will not tell you what to conclude, because the honest position is that this remains open. It should leave you with a much better question than the one you started with.
First, separate two questions
Public discussion routinely fuses two things that need pulling apart, and no progress is possible until they are.
The first is understanding: whether a system's internal states have the right kind of structure and content to count as grasping meaning rather than manipulating symbols. This is a question about representation and function, and it is at least partly empirical, since you can examine what a system represents internally and how it behaves.
The second is consciousness: whether there is something it is like to be the system, whether it has subjective experience. This is a much harder question, arguably not currently answerable even in principle for any entity other than oneself, and it is not what most technical debate is about.
These come apart in both directions. A system could plausibly represent and manipulate meaning without any inner experience. A creature could have rich experience with limited conceptual grasp. Evidence for one is not evidence for the other, so treating a model's fluent self-description as bearing on its consciousness, or its lack of experience as settling whether it represents meaning, is a category error. This piece is about understanding. Consciousness is a separate and harder problem, and anyone claiming to have settled it is overreaching.
The case that it does not understand
The skeptical position has serious philosophical foundations and real empirical support, and it deserves stating properly.
The oldest form is the symbol grounding problem. A system trained purely on text learns relationships between symbols, but those symbols were never connected to anything outside the text. It knows that "rain" co-occurs with "wet" and "cloud" and "umbrella," but it has never been rained on. Meaning, on this view, requires some causal contact between a symbol and the thing it refers to, and statistical relations among symbols cannot bootstrap that contact no matter how many of them you have. A closely related argument, Searle's Chinese Room, holds that manipulating symbols according to rules is never sufficient for comprehension, however convincing the output.
The modern version is the stochastic parrot framing: a model stitches together sequences based on probabilistic patterns in its training data without any grasp of meaning, and its fluency is precisely what makes this hard to see. Fluent text is our strongest social cue for a mind behind it, and these systems produce fluent text by construction, so we are systematically primed to over-read them.
There is empirical support beyond argument. Controlled studies have found that a model can produce an accurate, textbook-quality description of a physical concept in words and then fail to interpret the same concept presented in a different format, such as a simple visual or grid representation. That pairing is telling, because if the model grasped the concept, the representation should not matter much. There is also suggestive evidence from neuroscience that in humans, language and reasoning are handled by dissociable systems, meaning fluency and thought are separable even in us. If linguistic competence can come apart from reasoning in the one case we understand well, fluent output is weak evidence of comprehension.
The case that it does understand
The opposing position is not naive enthusiasm, and it has evidence the skeptical case has to answer.
The strongest strand comes from interpretability. Researchers have trained models on nothing but sequences of moves in a board game, with no description of the board, and then found that the model has developed an internal representation of the board state, one that can be located, read out, and causally manipulated: change the representation and the model's subsequent moves change accordingly. This matters because it is exactly the thing skeptics said could not happen. Predicting the next symbol turned out to require building a model of the process generating those symbols, and the world model emerged from pure sequence prediction without anyone asking for it.
Once you see that, the argument from training objective weakens considerably. "It was only trained to predict the next token" describes the objective, not the solution. Predicting text well enough, at sufficient scale, appears to require internal structure that is not merely surface statistics, which is unsurprising in retrospect: the text was produced by people reasoning about a world, so modelling the text well pushes toward modelling what produced it.
There are further strands. Models generalize to combinations they never saw, which memorization alone does not explain, and the generalization mystery is itself an open problem rather than a solved case for either side. Embeddings show meaningful geometric structure, with relationships between concepts encoded as consistent directions. Multimodal training gives at least partial grounding, since a model that has processed images along with text has some contact between words and appearances. And the emergence argument holds that a system can exhibit properties its components lack, so pointing at the mechanism does not settle what the whole does.
Why both sides are right about their piece
Notice what happens when you put the two cases side by side. The skeptics point at grounding, brittleness across representations, and the training objective. The believers point at internal world models, generalization, and emergent structure. These are not contradictory findings. They are findings about different things, and each side is largely correct about the thing it examines.
That is the shape of a question that has been badly posed. Both camps are selecting one component of what we ordinarily mean by understanding, demonstrating something true about it, and then treating that component as the whole. The dispute persists not because the evidence is thin but because the word is doing too much work.
Understanding is a bundle, and it comes apart
Here is the decomposition that dissolves most of the disagreement. When we say a person understands something, we are simultaneously attributing several distinct capacities, which in humans reliably travel together:
Linguistic competence: they can talk about it correctly, fluently, and appropriately to context. Generalization: they can apply it to novel cases, not just repeat familiar ones. Internal modelling: they have some representation of how the thing works, not just what is said about it. Grounding: their concepts connect, through perception and action, to the things they refer to. Self-knowledge: they have reasonable access to the limits of their own understanding and can tell you when they are unsure.
Because these always came bundled, we compressed them into one word and one question. Language models are the first systems to pull them decisively apart. On linguistic competence they are extraordinary, at or beyond human level. On generalization they are strong but uneven, handling many novel combinations while failing on others in ways that are hard to predict. On internal modelling the evidence says there is real structure, patchy and demonstrably incomplete. On grounding they are thin, with text-only models having no causal contact with what their words denote and multimodal training providing partial contact at best. On self-knowledge they are weak, being poorly calibrated about their own limits and unable to reliably report what they do or do not know.
That profile is not "understanding" and it is not "no understanding." It is a configuration no previous entity had, which is why our vocabulary fails on it. Asking whether it understands forces a yes or no onto a five-dimensional answer, and whichever answer you give, you are throwing away most of the information.
Why this explains the failures we actually see
The decomposition earns its keep by predicting the specific ways these systems fail, which neither pure position does as well.
Strong fluency combined with weak grounding produces exactly hallucination: confident, well-formed, plausible statements that are false, generated by the same process that produces true ones, with nothing in the system marking the difference. If the model had grounding, false statements would collide with something. If it lacked fluency, they would not be persuasive. The failure mode is the signature of that particular combination.
Patchy internal modelling produces brittleness across representations, where a concept a model handles well in one format collapses in another, and it shows up in the way performance degrades on inputs that differ from the training distribution in ways that ought not to matter.
Weak self-knowledge produces sycophancy and poor calibration: a system without reliable access to its own confidence will take its cue from the user instead, agreeing when pushed because it has no internal check to hold onto. And it produces the failure to say "I do not know," which requires knowing that you do not.
Each of these is a well-documented behaviour of current systems, and each falls out of the profile above. That a decomposition predicts the observed failure signature is decent evidence that the decomposition is tracking something real.
What would actually settle it
If the question is bad, the productive move is to replace it with tests that isolate components, and this is roughly what serious research now does.
For internal modelling, the method is interpretability with causal intervention: locate a candidate representation, alter it, and see whether behaviour changes in the way it should if that representation were doing the work. For generalization, it is out-of-distribution testing designed to be uncontaminated by training data, which is harder than it sounds given how much text has been consumed. For cross-representation transfer, it is exactly the paired designs described earlier, presenting the same concept in different forms and comparing. For grounding, it is tasks requiring contact between language and perception or action. For self-knowledge, it is calibration measurement: does stated confidence track actual accuracy.
None of these answers "does it understand." All of them answer something specific, checkable, and useful. That trade is worth making, and it is the difference between a debate and a research programme.
The honest state of things
Nobody currently knows whether these systems understand in whatever sense would satisfy both camps, and it is not clear the question has a determinate answer. What can be said is more useful than it sounds. These are not simply lookup tables or surface pattern matchers, since internal models of domains demonstrably form. They are also not minds that grasp meaning the way we do, since grounding is thin and self-knowledge is poor. They occupy a novel position, competent in ways that used to imply understanding and deficient in ways that used to be impossible alongside that competence.
There is one more reason for humility, which cuts against both camps. We do not have a rigorous account of what understanding consists in for humans either. The concept was never operationalised, because it never needed to be: we inferred understanding from behaviour in creatures we already assumed had minds like ours. These systems break that inference, and part of the discomfort is that they force us to specify what we always took for granted. The AGI debate has the same structure, and it is why both arguments tend to end where they began.
The short version
The question of whether AI understands is stuck because understanding is not one property. It bundles at least five distinct capacities that in humans always came together: linguistic competence, generalization to novel cases, internal modelling of how something works, grounding that connects concepts to the world through perception and action, and self-knowledge about one's own limits. Language models pull these apart. They are extraordinary at linguistic competence, strong but uneven at generalization, demonstrably possessed of real but incomplete internal models, thin on grounding, and weak on self-knowledge. Skeptics point at grounding, brittleness, and the prediction objective, and are correct about those. Advocates point at emergent world models, shown by interpretability work where a model trained only on move sequences develops a manipulable representation of board state, and at generalization, and are correct about those. Both then treat their component as the whole, which is why neither wins. The decomposition also predicts the observed failures: fluency without grounding gives hallucination, patchy modelling gives brittleness across representations, and weak self-knowledge gives sycophancy and poor calibration. Understanding should also be kept separate from consciousness, which is a different and harder question.
The idea to hold onto is that these systems have some of the components of understanding and lack others, in a combination no previous entity displayed, so the useful question is not whether they understand but which capacities are present, how strongly, and how each one can be tested. The debate has lasted this long because both sides are right about the part they are looking at.
Common questions
Do AI models actually understand what they are saying? There is no settled answer, and the reason is that understanding is not a single property. It bundles several capacities that come apart in these systems: language models are extraordinary at producing fluent, correct, contextually appropriate language, they generalize to many novel cases, and interpretability research shows they build real internal models of domains. But they have little grounding, meaning almost no causal contact between their words and the things those words refer to, and weak self-knowledge, meaning poor access to their own limits and confidence. So they possess some components of understanding and lack others, which is why both "yes" and "no" are misleading.
What is the stochastic parrot argument? It is the position that language models produce text by stitching together probabilistic patterns from training data without any grasp of meaning, much as a parrot reproduces sounds it does not comprehend. The argument draws on the symbol grounding problem, which holds that meaning requires some causal connection between symbols and what they refer to, something a system trained purely on text never has. Its strongest empirical support includes findings that models can describe a concept accurately in words while failing to handle the same concept in a different representational format, which suggests competence tied to surface form rather than to the concept itself.
Is there evidence AI builds internal world models? Yes, and this is the strongest evidence against the pure pattern-matching view. Researchers trained models on nothing but sequences of board game moves, with no description of the board, and found the models developed internal representations of the board state that could be located, read out, and causally manipulated: altering the representation changed subsequent behaviour in the predicted way. This shows that predicting the next symbol well can require modelling the process that generates those symbols. The representations are real but patchy and incomplete, so the finding refutes the strongest skeptical claim without establishing full comprehension.
Is understanding the same as consciousness? No, and conflating them causes most of the confusion in public discussion. Understanding concerns whether a system's internal states have the right structure and content to count as grasping meaning, which is at least partly investigable by examining representations and behaviour. Consciousness concerns whether there is subjective experience, whether there is something it is like to be the system, which is far harder and arguably not answerable from outside. The two can come apart in principle in both directions, so evidence about one is not evidence about the other, and technical debates about model understanding are generally not claims about machine experience.
Why can't we just settle whether AI understands? Partly because the concept was never rigorously defined, even for humans. We inferred understanding from behaviour in creatures we already assumed had minds like our own, so the term bundled several capacities that never needed separating. AI systems break that inference by displaying some capacities at very high levels while lacking others entirely, a combination no previous entity presented. Any single test therefore measures one component and gets read as settling the whole. Progress comes from replacing the general question with specific ones about internal modelling, generalization, grounding, and calibration, each of which is testable.
Does the "it's just predicting the next word" argument work? Not as decisively as it sounds, because it describes the training objective rather than the solution the model found. Predicting text well at scale appears to require internal structure that goes beyond surface statistics, which is unsurprising given that the text was produced by people reasoning about a world, so modelling the text pushes toward modelling what produced it. Interpretability results showing emergent world models support this. The argument does correctly identify that nothing in the objective requires grounding or self-knowledge, which is why those components remain weak, but the objective alone does not establish that nothing else is happening.
How does this affect how I should use AI? Practically, expect the failure signature that follows from the profile. Because these systems have strong fluency and weak grounding, they will produce confident, well-formed statements that are false, with no internal signal distinguishing those from true ones, so verify anything that matters. Because their internal models are patchy, expect performance to vary in ways that seem arbitrary, including on rephrasings that should not matter. And because self-knowledge is weak, treat expressions of confidence as unreliable and expect the model to defer when pushed rather than hold a correct position. Competence in one respect does not imply competence in the others.
Sources & further reading
The primary literature behind the claims above, drawn from the concept entries this post links to, so a claim carries the same source here as it does there.
- Ha & Schmidhuber (2018), World Models — the canonical formulation; a controller trained entirely inside a learned simulation. World Model
- Hafner et al. (2023), Mastering Diverse Domains through World Models — DreamerV3; learning in imagination across 150+ tasks with fixed hyperparameters. World Model
- LeCun (2022), A Path Towards Autonomous Machine Intelligence — the architectural case for predicting in representation space, and the argument that current LLMs lack this. World Model
Related articles
- What is the Turing test? And why it stopped matteringTuring never proposed the imitation game as a definition of thinking. He proposed it to replace a question he considered meaningless. Machines have now passed versions of it, sometimes judged more human than actual humans, and the striking thing is how little that settled.
- Beyond vector search: how RAG actually works in 2026RAG stopped being "vector database plus a language model" a while ago. Here's how retrieval-augmented generation actually works now, chunking, embeddings, reranking, knowledge graphs, and agentic retrieval, and where each piece quietly breaks.
- Why AI hallucinates: the confident lie is a feature, not a bugAI models don't hallucinate because they're broken. They hallucinate because we trained and scored them in a way that rewards confident guessing over honest uncertainty, and that has a mathematical floor. The real mechanism, the 2026 research that pinned it down, and what actually reduces it.
- How to tell if your AI actually worksMost teams ship AI features on vibes, then argue about whether changes helped. An evaluation set is an afternoon of work and it settles every argument you're about to have.