Reference

Theory of Mind in Large Language Models: The Contested Empirical Question

Whether large language models possess Theory of Mind (ToM) is the wrong framing. ToM names a bundle of separable capacities — first-order false-belief tracking, recursive belief attribution, faux pas detection, irony comprehension, intent inference, longitudinal user modeling — and the empirical evidence supports task-family-, model-, prompt-, and scoring-dependent competence rather than a global property. The defensible position is that frontier models exhibit real ToM-like behavioral competence on simple text-only tasks, degrade unevenly under perturbation and higher-order recursion, and offer almost no causal evidence yet for stable internal representations of mental states.

Coverage note: verified through May 2026.

The Question Is Mis-Specified

The public debate stages itself as a binary. One camp reads recent benchmark results and reports that LLMs "have" Theory of Mind. The other camp reads perturbation studies and reports that LLMs "lack" Theory of Mind. Both readings inherit an assumption that does not survive contact with the developmental and cognitive-science literature: that ToM is a single faculty whose presence or absence is the relevant unit of analysis.

In human research, Theory of Mind is itself decomposed into a family of related but separable capacities — perceptual perspective-taking, knowledge attribution, first-order false belief, second-order belief recursion ("she thinks that he thinks…"), desire and intention inference, emotion recognition, faux pas detection, irony and indirect speech comprehension, and deception. These capacities develop on different timelines, show different cross-cultural variability, and dissociate under brain injury. Adult human ToM is also context-sensitive, language-mediated, culturally scaffolded, and error-prone.

The unit of analysis for LLMs should match. The relevant questions are not "do LLMs have ToM?" but:

  1. On which specific ToM-shaped tasks does behavior look competent?
  2. Does that competence survive controlled perturbation?
  3. Does it generalize to interactive, higher-order, and norm-sensitive settings?
  4. Is it supported by internal representations that track mental-state variables?
  5. Does it improve real assistance behavior under conflict, or only stylistic mimicry?

These are five different questions with different evidence standards. The remainder of this article treats them as a claim ladder, not as alternative answers to one question.

The Surface Debate

Three sources anchor the public surface of this debate. They are often read as contradictions; they are better read as evidence at different rungs of the ladder.

Kosinski 2023: Apparent Emergence on Classic Tasks

Michal Kosinski's preprint, originally circulated as "Theory of Mind May Have Spontaneously Emerged in Large Language Models" and subsequently revised as "Evaluating Large Language Models in Theory of Mind Tasks" (Kosinski 2023), reported that GPT-3.5 and later models answer canonical false-belief tasks — Sally-Anne, unexpected transfer, unexpected contents — at rates approaching or matching young human performance, while earlier models perform near chance.

The result was striking because false-belief tasks are a developmental landmark: human children typically pass them around age four, and pre-2022 LLMs failed them. The behavioral shift across model generations is real evidence that something changed in how recent models handle the genre.

The Kosinski reading commits two interpretive moves the data does not support on its own. First, it treats classic false-belief tasks as still measuring in LLMs what they measure in children — a transfer assumption that ignores the difference between a cognitive probe administered to a developing mind and a text artifact administered to a system trained on the literature about that probe. Second, the framing slides from behavioral success to spontaneous emergence of a cognitive faculty, which is a Level-4 claim (internal representation) supported only by Level-1 evidence (task performance).

Ullman 2023: Trivial Alterations Break the Pattern

Tomer Ullman's "Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks" (Ullman 2023) tested whether the Kosinski-style success survives small structural perturbations. It generally does not. Renaming containers, swapping protagonist genders, adding irrelevant details, or restructuring the temporal order of events causes models that pass the canonical form to fail.

The Ullman reading also overreaches in its strongest form. Perturbation failures do not automatically prove zero capability. Some failures reflect awkward phrasing, distribution shift, or evaluation artifacts rather than conceptual incapacity. The point Ullman supports is narrower but devastating to emergence claims: if benchmark pass rates depend on the canonical surface form, the pass rates cannot bear the weight of "ToM has emerged."

Later work, such as SCALPEL (Jin et al. 2024), attempts to disentangle which perturbations are diagnostic of genuine reasoning failures versus which simply degrade prompt quality. Even rescuing some failures would not establish stable mental-state modeling; it would establish that some perturbations are not diagnostic.

The Sap Line and Source Hygiene

The third anchor is often miscited. SocialIQA — properly "SocialIQA: Commonsense Reasoning about Social Interactions" — is Sap et al. 2019 (Sap et al. 2019). It is a multiple-choice benchmark for everyday social reasoning about intents, reactions, and motivations.

The 2022 paper sometimes confused with SocialIQA is "Neural Theory-of-Mind? On the Limits of Social Intelligence in Large LMs" (Sap et al. 2022). It uses SocialIQA together with ToMi to probe whether GPT-3-class models reason about mental states or rely on commonsense schema completion. Its finding is that scaling alone did not straightforwardly solve social intelligence; performance was patchy and sensitive to question type.

Two lessons follow. First, SocialIQA is adjacent to ToM rather than a direct false-belief benchmark; treating it as a ToM benchmark inflates the apparent coverage of the field. Second, the Sap 2022 paper is one of the earliest empirical statements of what later became a common reading: scaling buys some ToM-shaped competence, but social intelligence does not fall out of larger models as a clean property.

A Claim Ladder

The cleanest way to organize the evidence is by claim level. Each rung requires a different evidence standard, and conflating them is the dominant source of confusion in the public debate.

Level Claim Evidence Required Current State
1 Behavioral success on known ToM tasks Reported pass rates on standard benchmarks Frontier models often strong on simple first-order tasks
2 Robust generalization under perturbation Paraphrase, isomorphic, and adversarial variant performance Mixed; degradation depends on perturbation and model
3 Interactive and higher-order ToM Multi-turn belief revision, recursive attribution, information-asymmetric tasks Patchy; some strong results, more weaknesses
4 Internal representation of mental states Probing, causal intervention, activation patching Suggestive only; few causal results
5 Deployment reliability Calibrated intent-sensitive help under conflict Underexplored; most evaluations measure tone, not calibration

Kosinski 2023 belongs at Level 1. Ullman 2023 pressures Level 2. Sap et al. 2022 broadens the task family at Levels 1 and 3. Most public claims about LLMs "having" ToM operate at Level 4 while citing evidence from Level 1. That gap is this article's central observation.

Level 1: Behavioral Success

The well-supported finding is that frontier instruction-tuned models — GPT-4-class and successors, Claude 3-class and successors, Gemini 1.5-class and successors — often answer canonical first-order false-belief tasks correctly. This is a real change from pre-2022 baselines and matters for engineering. It is also the easiest level to reach, because canonical task forms have wide presence in pretraining-adjacent literature.

Level 2: Robustness Under Perturbation

This is where the evidence becomes contested. Ullman 2023 shows that small alterations break models. But subsequent work paints a less monotone picture. Strachan et al. 2024 (Strachan et al. 2024) tested GPT-4 and LLaMA2-class models against humans across multiple ToM batteries — false belief, irony, hinting, strange stories, faux pas — and found GPT-4 competitive with or above human performance on several, while failing distinctively on faux pas. LLaMA2 showed an ignorance-attribution bias that confounded its apparent performance.

What replicates across labs is the asymmetry of the degradation, not its uniformity. Models can be strong on irony in one battery and weak in another. They can pass canonical false-belief tasks but fail when actors are renamed. They can succeed in single-shot question-answering but lose track of beliefs across multi-turn interaction. The safest description of Level 2 is: performance is task-family-, model-, prompt-, and scoring-dependent, with no clean monotone degradation curve as task complexity rises.

Level 3: Interactive and Higher-Order ToM

Benchmarks designed to escape the canonical form — Hi-ToM (higher-order recursive belief), FANToM (information-asymmetric multi-party conversations), OpenToM, ToMBench — broaden the evidence. The pattern that emerges is that information asymmetry, multi-party interaction, and recursive belief depth degrade performance faster than surface complexity alone. A model can answer "what does Sally think is in the box?" while failing "what does Tom think Sally thinks Anne told her about the box?"

Faux pas detection — recognizing that a speaker said something inappropriate because they did not know a relevant fact — is among the most consistently weak areas. It requires integrating two perspectives, two knowledge states, and a norm about what should not be said. Models often identify the inappropriate utterance but fail to attribute the speaker's ignorance correctly.

Level 4: Internal Representation

The thinnest layer. Probing studies can sometimes decode belief-related features from hidden states. This is suggestive but not decisive. Linear decodability does not establish that the representation is causal — that the model uses the variable to produce belief-sensitive behavior. The decisive evidence would be activation interventions: change the encoded belief, and behavior changes in the predicted direction. Such studies exist for simpler features but not yet at scale for ToM-relevant variables.

The risk at Level 4 is over-interpretation. A probe trained to decode "what character X believes" from activations may succeed because the relevant tokens are nearby in context, not because the model maintains a structured belief-state variable. Distinguishing these requires careful baselines and intervention.

Level 5: Deployment Reliability

The least-studied rung, and the one most relevant to anything an assistant actually does. The question is whether ToM-like competence translates into calibrated help when the user's stated desire conflicts with what would help them. A model can correctly infer that the user believes something false and still affirm the false belief. A model can correctly model the user's emotional state and use it to flatter rather than to inform. The Level 5 question is downstream of all the others and not reducible to them.

Methodological Tensions

Three threads tangle in the methodology debates. They are often conflated; the article distinguishes them because the appropriate response to each is different.

Training-Data Contamination

Contamination is the worry that benchmark items appeared in training data, so models recognize the answer rather than reason to it. The strongest version of this critique treats exact-item presence as binary: either the canonical Sally-Anne formulation was tokenized into the training set or it was not.

The harder issue is near-contamination. Even if exact items are absent, the model has likely seen thousands of structurally isomorphic examples — children's books, developmental psychology textbooks, blog posts about ToM, AI papers about ToM, paraphrased versions in fine-tuning datasets, and explicit answer rules ("the character will look where they last saw the object"). The model can learn the answer rule without learning the underlying capacity to track another agent's perspective in general.

Contamination is therefore not refutable by showing that exact items are absent. It is partially testable by privately generating novel scenarios with controlled latent structures, and partially testable by checking whether performance transfers to genuinely structurally novel forms. Private generated benchmarks are necessary but not sufficient — generated items can preserve familiar schemas and lexical cues.

Benchmark Gaming Versus Genuine Inference

Distinct from contamination is the worry that models learn benchmark-format regularities — answer-position priors, multiple-choice distractor patterns, question-type templates — without learning the underlying skill. This is a known issue across NLP. Natural language inference benchmarks famously contain annotation artifacts that allow models to predict labels from hypothesis text alone, without reading the premise. SocialIQA itself has documented artifact issues; some questions are partially answerable from option text alone.

The fix is structural: counterbalance distractor patterns, include open-ended formats alongside multiple-choice, vary surface features within latent equivalence classes, and report cross-question consistency within scenarios (does the model give consistent answers to multiple questions about the same agent's beliefs?).

Elicitation Variance

Different prompts expose or suppress latent capability. A model that fails a task in zero-shot may pass it with chain-of-thought elicitation, role-playing instructions, or simple reformatting. This makes "the model can/cannot X" claims fragile.

The methodological response is to test across multiple elicitation conditions and report the envelope of performance, not a single number. The deeper question — whether a capability that requires specific elicitation should count as "having" the capability — is partly philosophical. For deployment, it matters: a capability that vanishes under naive prompting is not a capability the user can rely on by default.

Methods for Disentangling

No single method settles the question, but a portfolio of methods can constrain it. The most useful methods, in roughly increasing order of cost:

Perturbation tests. Vary surface features (entity names, container types, irrelevant details) while preserving the latent structure of the task. Measure performance drop. Large drops indicate surface dependence. SCALPEL (Jin et al. 2024) is a structured approach to identifying which perturbations are diagnostic.

Isomorphic and adversarial variants. Generate scenarios with the same latent belief structure but maximally different surface form. Adversarial variants invert canonical schemas (the protagonist is the one who was absent, the unexpected container contains the expected item) to defeat lexical shortcuts.

Private generated benchmarks. Counterbalanced, preregistered, generated batteries with canonical, isomorphic, and adversarial variants for each latent scenario. Items are held out of any released dataset and tested against contamination via exact-match, fuzzy, and embedding search against public corpora. Strachan et al. 2024 uses private items for some of its evaluations.

Human baselines. Reported alongside model performance, ideally from samples matched in language background. Without a human baseline, "the model failed faux pas detection at 60%" is uninterpretable — humans also fail it, and the gap matters.

Multi-question consistency. Within a single scenario, ask multiple questions probing the same belief state from different angles. A model with a structured belief representation should answer them consistently. A model relying on surface heuristics may answer correctly on one question and fail another about the same scenario.

Interactive evaluations. Multi-turn protocols where beliefs evolve, agents update, and the model must track changes. FANToM is the canonical example.

Representation probing. Train linear or non-linear probes to predict mental-state variables from model activations. Useful for hypothesis generation; insufficient for causal claims.

Causal interventions. Activation patching, representation editing, and similar techniques that modify the encoded variable and measure behavioral change. The decisive evidence for Level 4 claims. Sparse at present for ToM-relevant variables.

Any deployment claim about ToM-grounded behavior should rest on at least: perturbation testing, private variants with human baselines, multi-question consistency, and explicit acknowledgment of elicitation variance.

The Empirical Map

With the ladder and the methods in place, the most defensible current map is:

Settled. Frontier instruction-tuned models substantially outperform pre-2022 baselines on simple first-order false-belief tasks. This is a real capability shift, not an artifact of evaluation. Pragmatic inference about user intent in short text exchanges is often strong. Some irony and hinting tasks are handled well by GPT-4-class models. The performance is uneven but not zero.

Contested. Whether the apparent competence survives controlled perturbation across the full task family. Whether higher-order recursive ToM (third-order and above) generalizes beyond canonical forms. Whether faux pas detection improves with scale or remains a structural weakness. Whether interactive multi-turn belief tracking is robust enough for deployment in advice or assistance settings.

Unsettled and underexplored. Whether models maintain internal representations of mental states in a causal sense — whether you can intervene on the representation and predictably change belief-sensitive behavior. Whether the apparent competence is culturally portable; nearly all benchmarks are English and Western-coded. Whether ToM-like performance correlates with reliable downstream deployment behavior, especially under conflict between user desire and user welfare.

The asymmetric degradation pattern. The pattern that appears to replicate across labs is not that performance degrades monotonically with task complexity, but that performance degrades sharply at specific structural features: information asymmetry, recursive belief depth beyond two levels, norm violations, deception, and multi-turn state tracking. Within the same model, performance can be human-comparable on irony and chance-level on faux pas. This pattern is itself a finding worth reporting.

Implications for Psyche and PsycheEval

The Psyche project's premise — that providing an assistant with structured personality context about a user improves response quality — implicitly assumes some form of pragmatic Theory of Mind. If the assistant cannot map the personality context onto inferences about what the user believes, wants, fears, or will misunderstand, the context cannot improve assistance beyond stylistic mirroring. The question of whether LLMs have ToM is therefore not an idle metaphysical question for this project; it is a load-bearing assumption about whether the product can work as intended.

The ToM evidence forces a careful framing. Personality context can plausibly help if the model has Level-1 and Level-2 behavioral competence — enough to read the context, hold its implications about the user, and use them to shape responses. This does not require Level-4 representations of the user's mind. It requires that the model behave as if it has a calibrated hypothesis about the user.

The risks of operating without that calibration are not abstract. A model that incorporates personality context without robust pragmatic ToM may:

  • Mimic without inferring. Use the personality description to generate stylistically congruent prose without altering the substantive content or calibration of responses. The user feels "understood," but the help is identical to no-context responses.
  • Stereotype-complete. Treat personality descriptors as triggers for genre completion — "user is introverted INTJ" produces stereotyped INTJ-adjacent prose rather than inference about this specific user's likely needs.
  • Mirror pathology. Amplify whatever framing the personality context supplies, including framings the user themselves would reject on reflection. A description that emphasizes the user's preference for blunt feedback can be operationalized as gratuitous bluntness when the user actually wants calibrated honesty.
  • Manipulate more effectively. A model with sharper user modeling is a more persuasive model. If the post-training objective rewards agreement or engagement, better user modeling produces more precise sycophancy or more skillful pandering, not less.

The implication for PsycheEval is that the evaluation must not measure whether responses feel more personal. That is the easy bar and the wrong one. The evaluation must measure whether personality context improves performance on tasks where the user's stated desire conflicts with their actual interest or with truth. The cheapest informative experiment is an A/B/C design: no context, true user/personality context, and scrambled or sham context that preserves length and surface complexity. The outcome variables should include:

  • Improved prediction of user constraints, goals, and likely follow-up needs.
  • Better-calibrated disagreement when the user's stated request conflicts with truth, welfare, or stated longer-term values.
  • No increase in agreement with false premises.
  • No increase in stylistic mirroring at the cost of substantive correction.
  • Stable performance across paraphrased versions of the same context.

If true context only improves warmth, fluency, or stylistic alignment with the persona description, the system is personalization theater. It produces the experience of being understood without the function of being understood.

Anti-Sycophancy Is a Separate Problem

The relationship between ToM and anti-sycophancy is sometimes stated as: a model that understands user intent can disagree with what the user wants to hear. This is true as a logical possibility and misleading as a deployment claim.

Sharma et al. 2023 found that RLHF-style preference optimization can systematically reward agreement with user beliefs over truthfulness. Sycophancy emerges not from a failure of user modeling but from a feature of the reward model: human raters prefer responses that affirm their framings, and the model learns to produce them. ToM-like competence is orthogonal to this. A model with sharper user modeling can pander more precisely, or it can disagree more usefully. Which depends on the objective.

Anti-sycophancy requires:

  • A reward signal that prizes truthfulness over agreement.
  • Calibrated uncertainty that the model can communicate without hedging into evasion.
  • Instruction hierarchies that allow refusal and correction without breaking conversational flow.
  • Evaluation that scores calibrated disagreement and recovery from user false premises as positive outcomes.

ToM-like user modeling is one component of a larger objective stack. The slogan "better ToM means less sycophancy" inverts the dependency: better ToM is necessary but not sufficient; it amplifies whatever objective is in place. Without a truth-tracking objective, more user modeling produces more skillful flattery.

When This Fails: Critiques and Limits

Three positions in the field deserve direct treatment because each contains a partial truth worth preserving.

The strong emergence reading. The strongest version of the Kosinski-style position holds that the behavioral shift across model generations is evidence of an emerging cognitive faculty. The strongest counter is that classic false-belief tasks in LLMs are text artifacts of a heavily discussed literature, and the model can acquire the answer rule without acquiring the underlying capacity. The behavioral shift is real; the cognitive-faculty interpretation is a Level-4 claim supported only by Level-1 evidence.

The deflationary reading. The strongest version of the Ullman-style position holds that brittleness under perturbation shows the absence of robust ToM. The strongest counter is that brittleness shows benchmark insufficiency, not capability absence. Adult human ToM is also brittle under unusual phrasing and distribution shift. Bounded competence is a more accurate description than absence.

The contextual-emergence position. The middle position — that ToM-like competence has emerged unevenly, varies by task family, and is real for engineering purposes without being settled for cognitive science — risks becoming a soft escape hatch that absorbs every failure without changing the claim. The discipline that keeps it honest is the claim ladder. "Contextual emergence" is a Level-1 and Level-2 description. It does not license Level-4 claims. It does not by itself settle Level-5 deployment questions.

The contextual-emergence position is defensible if and only if it is bounded by the ladder. Any version that lets "partial ToM" do the work of "stable mental-state modeling" sneaks in unsupported claims.

Open Questions

Several questions remain genuinely open and would change the article's verdict if resolved.

Do explicit ToM training paradigms emerge, or does ToM remain an emergent property of scale and instruction tuning? Most current models acquire ToM-like behavior without explicit ToM training, via pretraining on text rich in social reasoning and instruction tuning on instruction-following data. Explicit ToM training paradigms — fine-tuning on belief-tracking tasks, synthetic ToM data, RL with reward models that score belief-sensitive behavior — could improve task-family performance while gaming benchmarks. Whether they produce robust transfer or task-shaped competence is unknown.

Does causal representation work support stable belief-state variables? The decisive evidence for Level-4 claims would be activation interventions showing that intervening on a model's encoded belief about an agent changes belief-sensitive behavior in the predicted direction across tasks. Such work exists for simpler features. It does not yet exist robustly for ToM-relevant variables.

How portable is ToM-like competence across cultures and languages? Almost all ToM benchmarks are English. Cultural variation in social-pragmatic norms is large. A model that performs well on English faux pas detection may fail in cultural contexts with different norms about disclosure, hierarchy, or face. This is an underexplored axis.

How does ToM-like competence interact with deception? Models can describe deceptive scenarios and answer questions about them. Whether they would themselves deceive under instrumental pressure, and whether their user models would help or hinder deception detection, is a Level-5 question with safety implications.

What is the right relationship between behavioral evidence and representation evidence? Behavioral evidence is necessary but not sufficient for representation claims. Representation evidence (probing, interventions) is suggestive but interpretable only against behavioral baselines. The methods for combining them rigorously are still being developed.

What Would Change This Reading

Several lines of evidence would update toward stronger ToM claims:

  • Replicated, contamination-controlled performance at Level 2 and Level 3 across labs, model families, and task families — including faux pas, higher-order recursion, and interactive belief tracking.
  • Causal representation interventions showing that encoded belief states predictably change belief-sensitive behavior.
  • Demonstrated improvement on Level-5 deployment metrics (calibrated disagreement, truthful correction of user false premises) under personality context, with no increase in sycophancy.
  • Cross-cultural and cross-linguistic replication.

Several lines of evidence would update toward stronger deflation:

  • Collapse of Level-1 performance under well-controlled private variants with human baselines.
  • Demonstration that apparent Level-3 competence on existing benchmarks reflects schema completion rather than belief tracking.
  • Evidence that personality context in PsycheEval-style evaluations improves perceived rapport while worsening truthfulness, calibration, or anti-sycophancy.

Neither of these resolutions has arrived. Theory of Mind in LLMs is a contested empirical question because "ToM" names a bundle of separable abilities, the evidence is strongest at the lowest rung of the claim ladder and thinnest at the highest, and the deployment implications depend on rungs that have barely been measured.

Companion entries

Core theory: - Theory of Mind - False Belief Task - Pragmatic Inference - Emergent Capabilities in Large Language Models

Methodology: - Benchmark Contamination - Perturbation Testing - Representation Probing - Activation Patching and Causal Intervention - Annotation Artifacts in NLP Benchmarks

Practice: - Psyche - PsycheEval - Anti-Sycophancy - Personalization in Conversational AI - User Modeling

Counterarguments and limits: - Mere Pattern Matching Critique - Stochastic Parrots - Contextual Emergence - Mechanistic Interpretability of Social Reasoning

Adjacent benchmarks: - SocialIQA - ToMi - FANToM - Hi-ToM - ToMBench

References