Home/Blog/Foundations/Does it understand? The argument, properly stated

Does it understand? The argument, properly stated

Both sides of this debate are usually presented by their opponents. Here is the sceptical case at full strength, the case for at full strength, why the two keep missing each other, and what would actually settle it.

Ask whether a language model understands anything and you will get two confident answers, each delivered by someone who has heard the other side only in summary.

That is a shame, because both cases are stronger than their opponents represent them, and the disagreement between them is more interesting than either. Some of it is empirical and being actively resolved. Some of it is definitional and cannot be resolved by evidence at all. Most of the public argument fails to distinguish which is which, which is why forty-five years of it have produced very little movement.

The honest position is that the sceptical argument is valid and rests on a premise that has never been established, the affirmative argument is coherent and rests on a definition its opponents reject, and the empirical work now underway is producing findings that neither thought experiment anticipated. This is an attempt to state all three fairly.

The sceptical case, at full strength

The argument has three stages, developed across forty years, and each is stronger than the caricature.

Searle, 1980. Imagine a person in a room who receives Chinese characters, consults an enormous rulebook written in English, and passes back the characters the rules specify. To an outside observer the room converses fluently in Chinese. The person inside understands no Chinese whatsoever.

The point is not that the room is slow or that the rulebook is large. It is that syntax does not constitute semantics. Manipulating symbols according to their shape, however sophisticated the manipulation, is not the same as grasping what they mean, and no amount of additional rule-following converts one into the other.

Harnad, 1990. The symbol grounding problem sharpens this. If a symbol is defined only by its relations to other symbols, the definitions never terminate in anything outside the system. You can look up a word in a monolingual dictionary and find more words. Harnad's image is a merry-go-round: without some point where symbols attach to non-symbolic experience, meaning has nowhere to come from. Maps of maps do not become territory.

Bender and Koller, 2020. The octopus updates the argument for systems trained on text. Two people are stranded on separate islands, communicating through an underwater cable. A hyper-intelligent octopus taps the cable and, over years, learns the statistical structure of their exchanges well enough to respond convincingly when one of them goes quiet.

Then one islander is attacked by a bear and asks urgently how to build a weapon.

The octopus has never seen a bear, a stick or a coconut. It has learned the form of the conversation with great precision and has no access to what the forms are about. Bender and Koller's claim is that a system trained only on form cannot in principle learn meaning, because meaning requires a relation to something outside the language, and the training data contains no such relation.

This is the strongest version. It is not a claim that models are unimpressive, or that the outputs are bad. It is a claim about what kind of thing could possibly have been learned from that kind of input.

The affirmative case, at full strength

The response is not "but the outputs are so good." That reply concedes the argument and appeals to appearance, which is exactly what Searle designed his room to defeat.

The serious response attacks the premise: that meaning requires reference to things outside language.

Conceptual role semantics holds that a term's meaning is constituted by its inferential relations to other terms and to the system's dispositions, rather than by a causal chain to an object. On this view, understanding "bear" is a matter of being appropriately connected to danger, forest, large, fur, avoid, and thousands of other relations, and a system with rich enough relational structure has what meaning consists of. Reference is one way to acquire that structure, not the thing the structure is made of.

If this is right, the octopus argument does not go through. The octopus has been learning the relational structure the whole time, and the bear question is hard for it in the way that a question about an unfamiliar domain is hard for anyone, rather than impossible in principle.

Structural correspondence offers a different route. If a system's internal states stand in relations that mirror the relations among things in the world, then those states carry information about the world regardless of how they were acquired. A map made from other maps still corresponds to the terrain if the copying preserved the structure. Text is itself a product of the world, so structure in text is not arbitrary; it is a lossy projection of structure in what the text is about.

And the negative argument. Searle's room and Bender's octopus both invite you to notice that no understanding is present, and neither offers a criterion by which you could have detected understanding if it were, which is the same objection made to the Turing test. If the thought experiment would return the same verdict for a system that did understand, it is not measuring understanding. Critics have argued that both arguments assume their conclusion: they identify understanding with something the system in question does not have by construction, then observe that it does not have it.

This is the strongest version of the affirmative case. It is not a claim that models definitely understand. It is a claim that the arguments against have not established what they are taken to establish.

The systematicity objection, which is separate and sharper

One argument in this territory is frequently folded into the others and deserves its own treatment, because it makes a testable prediction rather than a definitional claim.

Fodor and Pylyshyn argued in 1988 that thought is systematic: anyone who can understand "John loves Mary" can understand "Mary loves John", because the capacity comes from grasping the parts and the way they combine rather than from having encountered the whole. They claimed neural networks lack this by construction, since a network associating inputs with outputs has no guarantee that mastering one combination confers mastery of another.

This is a better argument than the Chinese Room in one specific respect: it predicts something. If a system's competence comes from compositional structure, performance on a novel combination of familiar parts should resemble performance on familiar combinations. If it comes from having seen enough combinations, performance should fall on truly novel ones.

The evidence is mixed and the reason it is mixed is instructive. Large models handle many novel compositions well, which looks like systematicity. They also fail on some in ways that suggest memorised patterns, and the failures are hard to characterise because establishing that a combination is entirely absent from a training corpus of that size is close to impossible.

That last difficulty is worth dwelling on. The systematicity question is empirically well posed and practically unanswerable at current training scales, because the test requires knowing what the system has not seen, and nobody can enumerate what a multi-trillion-token corpus contains. Controlled experiments on small models trained on known data address this and raise the question of whether results transfer.

The methodological lesson generalises beyond this argument: several of the sharpest questions about these systems are blocked not by philosophy but by the impossibility of characterising the training distribution.

Why the two sides keep missing each other

Set out side by side, the disagreement is partly not about language models at all.

The sceptic holds that meaning is a relation between symbols and the world, so a system with no access to the world cannot have it, whatever its internal structure. The affirmative holds that meaning is constituted by relations among representations, so a system with sufficient internal structure has it, whatever its history.

These are positions in the philosophy of mind that predate the technology by a century, and the language model is a new place to have an old argument. No experiment on a model settles which theory of meaning is correct, because the theories disagree about what would count as evidence.

That said, the debate is not entirely definitional, and the part that is not is worth separating out. Three questions are empirical:

Does the system have internal representations that track features of the world rather than features of the text? Do those representations support inference the training data did not contain? And do they generalise across domains in ways that a lookup table could not?

Those have answers. They are being investigated. And they are not the same question as whether the answer amounts to understanding.

What the empirical work has actually found

Interpretability research has changed the terms of the debate in a way neither thought experiment anticipated, because both were designed for systems whose internals were assumed to be inscrutable. They are not entirely inscrutable now.

Three findings are relevant, and none is decisive.

Models contain identifiable internal features. Large-scale work extracting interpretable concepts from production models has found millions of features corresponding to recognisable things, including abstract ones, activated across contexts that share meaning rather than share wording. This is evidence of representational structure that is not merely surface statistics.

Capabilities transfer across domains in ways suggesting shared abstraction. Interpretability work on fine-tuning has found that training on one domain tends to strengthen existing circuits rather than build new ones, and that improvements appear in unrelated domains. The system appears to have located abstractions that span areas nobody told it were related.

Some internal structure appears to model the world rather than the text. Work probing for representations of spatial, temporal and relational properties has found internal states that track those properties, in systems trained only on text.

What none of this establishes is the thing under dispute. A sceptic can accept every result and maintain that structure correlated with world-features is still structure over symbols, and that correlation is exactly what training on text produces. The findings raise the cost of the strongest sceptical claim, that nothing world-involving could be learned from text, without touching the underlying disagreement about what meaning is.

The honest summary: the empirical work has moved the burden without settling the question. That is real progress and it is not resolution.

What would settle it

Worth asking directly, because a question no evidence could answer is a different kind of question and should be labelled as such.

For the empirical part, several things would count. A demonstration that a text-trained system can acquire a concept it could not have encountered in any form in training and use it correctly, which would be hard to explain as recombination. Evidence that internal representations support counterfactual reasoning about physical situations no text described. Or the reverse: a demonstration that apparently world-tracking representations are systematically confounded with textual co-occurrence, which would deflate the interpretability results.

For the definitional part, nothing would settle it, because the disagreement is about what the word picks out. Two people who agree on every fact about a system can disagree about whether it understands, in the way they might disagree about whether a virus is alive. That is not a failure of evidence and it will not be repaired by more of it.

The practical upshot is that "does it understand" is a poor question and it decomposes into better ones. Can it do this task reliably? Does it fail in ways that suggest it has represented the situation or only the phrasing? Will its competence extend to cases the training data did not contain? Those are answerable, they are what anyone deploying a system needs to know, and none of them requires resolving the philosophy.

The argument that we should hope the answer is no

One line in the recent philosophical literature is worth surfacing because it inverts the usual framing.

The sceptical position is normally treated as deflationary, as though establishing that models do not understand would be a disappointment. But if a system did have the kind of original intentionality that the strong version of understanding requires, it would arguably enter the space of things with interests, and therefore the space of moral concern.

On that reading, the widespread desire to establish that these systems understand is a desire to have created something we would then owe obligations to, at industrial scale, with no framework for discharging them. The sceptic's conclusion is convenient rather than sad.

This does not settle anything and it is not evidence. It is a reason to notice that the question is not neutral, and that people arguing for understanding are frequently arguing for a conclusion whose implications they have not examined.

What is unresolved

Whether structure acquired from text can be world-involving in the relevant sense. This is the crux and it remains open. Text is produced by people interacting with the world, so it carries the imprint of that world, and whether an imprint is enough is exactly what the two camps disagree about.

Whether interpretability findings mean what they appear to. A feature that activates on world-relevant contexts might represent the world or might represent the linguistic contexts in which that feature of the world is discussed, and separating these is difficult because the two are correlated in every available dataset.

Whether multimodal training changes the argument. Systems trained on images and audio have some causal connection to the world through those channels. Whether that satisfies the grounding requirement, or merely adds another layer of representation, is contested. Sceptics have argued that a pixel array is still a symbol.

Whether "understanding" is one thing. The debate assumes a single property that a system has or lacks. It may be a cluster of loosely related capacities that come apart, in which case both sides are right about different components and the yes-or-no framing has been the error throughout.

The counter-argument to this article

Presenting both sides equally may misrepresent the state of the field. If one position is substantially better supported, treating them as symmetric is a distortion in the guise of fairness. Some philosophers would say the sceptical argument has been answered and the appearance of live debate is manufactured; others would say the reverse. This piece has assumed symmetry and that assumption is itself a position.

The definitional framing can be a dodge. Saying the disagreement is partly about words is true and can be used to avoid committing. There are facts about what these systems do, and someone has to make decisions on the basis of them, so retreating to "it depends what you mean" has a cost.

And the practical reframing may concede too much. Replacing "does it understand" with "does it work reliably" is useful for deployment and abandons a question people care about for reasons that are not merely practical. Whether we have built something that understands is a question about what we have done, and its irrelevance to a product decision does not make it unimportant.

The short version

Both sides of this argument are stronger than their opponents present them. The sceptical case runs from Searle, that syntax does not constitute semantics and no amount of rule-following converts one into the other; through Harnad, that symbols defined only by other symbols never terminate in anything outside the system; to Bender and Koller, whose octopus learns the form of a conversation perfectly and has no access to what the forms are about. The claim is not that outputs are unimpressive. It is about what could in principle be learned from that kind of input.

The affirmative case does not appeal to output quality, which would concede the point. It attacks the premise that meaning requires reference outside language. Conceptual role semantics holds that meaning is constituted by inferential relations among representations, in which case rich enough relational structure is what meaning is made of. Structural correspondence holds that internal states mirroring worldly relations carry information about the world regardless of acquisition route, and that text is a lossy projection of the world rather than an arbitrary system. And the negative argument observes that neither thought experiment supplies a criterion by which understanding could have been detected had it been present.

The two positions are theories of meaning that predate the technology, which is why no experiment settles them. But three questions inside the debate are empirical: whether internal representations track world features rather than textual ones, whether they support inference absent from training, and whether they generalise across domains. Interpretability work has found millions of identifiable internal features, cross-domain transfer suggesting shared abstractions, and representations that appear to track spatial and relational properties. A sceptic can accept all of it and maintain that structure correlated with world-features is what training on text produces.

The useful move is that "does it understand" decomposes into better questions: can it do this task reliably, does it fail in ways suggesting it represented the situation or only the phrasing, and will its competence extend beyond the training distribution. Those are answerable, they are what anyone deploying a system actually needs, and none requires settling the philosophy first.

Common questions

Do large language models understand language? There is no settled answer, and part of the disagreement is not empirical. Sceptics hold that meaning requires a relation between symbols and the world, which text-only training cannot supply. The affirmative position holds that meaning is constituted by relations among representations, in which case sufficient internal structure is what meaning consists of. These are competing theories of meaning that predate the technology, and no experiment on a model adjudicates between them, because they disagree about what would count as evidence.

What is the Chinese Room argument? Searle's 1980 thought experiment. A person who knows no Chinese sits in a room receiving Chinese characters, consults an English rulebook, and returns the characters the rules specify. From outside, the room converses fluently. The person understands nothing. The point is that syntax does not constitute semantics: manipulating symbols by their shape, however sophisticated, is not grasping what they mean, and adding more rules does not convert one into the other.

What is the symbol grounding problem? Harnad's 1990 formulation. If symbols are defined only by their relations to other symbols, the definitions never terminate in anything outside the system, like looking up a word in a monolingual dictionary and finding only more words. Without some point where symbols connect to non-symbolic experience, there is nothing for meaning to consist in. It sharpens the Chinese Room by identifying what specifically is missing rather than only asserting that something is.

What is the octopus test? Bender and Koller's 2020 update. Two people on separate islands communicate by underwater cable; a hyper-intelligent octopus taps it and learns the statistical structure of their exchanges well enough to respond convincingly. When one is attacked by a bear and asks how to build a weapon, the octopus fails, having never encountered a bear or a stick. The argument is that a system trained only on form cannot in principle learn meaning, since the training data contains no relation to what the forms are about.

What is the strongest argument that language models do understand? Not that the outputs are good, which concedes the point Searle designed his room to make. The serious response denies that meaning requires reference outside language. On conceptual role semantics, a term's meaning is its inferential relations to other terms, so rich enough relational structure is what meaning is made of rather than a substitute for it. There is also a negative argument: neither the Chinese Room nor the octopus provides a criterion by which understanding could have been detected if present, so neither is measuring what it claims.

What has interpretability research shown? Three things, none decisive. Large-scale feature extraction has found millions of identifiable internal features, including abstract ones activating across contexts that share meaning rather than wording. Fine-tuning work suggests training on one domain strengthens existing circuits and improves unrelated domains, implying shared abstractions the system located itself. And probing has found internal states tracking spatial, temporal and relational properties in text-only systems. A sceptic can accept all of it and hold that structure correlated with world-features is exactly what text training produces.

Would multimodal training solve the grounding problem? Contested. Systems trained on images and audio have some causal connection to the world through those channels, which appears to address the objection that text alone provides no such connection. Sceptics reply that a pixel array is itself a representation, so the system is relating symbols to other symbols with an extra step, and the grounding regress has been moved rather than terminated. Whether perception of any kind can terminate it, for machines or for people, is an old question this does not resolve.

Does the question matter for building things? Less than it seems. The practically useful questions are whether a system performs a task reliably, whether its failures suggest it represented the situation or only the phrasing, and whether competence extends beyond the training distribution. Those are answerable without settling what understanding is, and they are what a deployment decision actually depends on. The philosophical question matters for other reasons, including what obligations we might incur if the answer were yes, which is a consideration the enthusiasm for an affirmative answer rarely examines.

Learn the concepts

← All posts