The cognitive capacity to attribute mental states — beliefs, desires, intentions, knowledge, and emotions — to oneself and to others, and to recognise that others’ mental states may differ from one’s own and from reality. Central to human social cognition and classically probed with false-belief tasks, theory of mind has become a benchmark capability for artificial agents, where machine analogues are pursued to support cooperation, communication, and safe interaction with people.

Semantic Classification

Content

Definition

Theory of mind (ToM) is the ability to model other agents as having minds: to infer what someone believes, wants, intends, or knows, and to predict their behaviour from those inferred states rather than from the world as it actually is. The term originates with Premack and Woodruff’s 1978 question of whether chimpanzees possess it; the canonical diagnostic is the false-belief task (Sally-Anne), which children typically pass around age four by predicting that an agent will act on an outdated belief rather than on the true state of the world.

ToM is layered. First-order attribution (“she believes X”) extends to higher orders (“he thinks that she believes X”), and full competence spans belief, desire, intention, knowledge and ignorance, emotion, and pretence. In cognitive science it is variously explained by theory-theory (an intuitive folk psychology), simulation theory (running one’s own cognitive machinery offline in the other’s place), and Bayesian inverse-planning accounts that recover goals and beliefs as the latent causes best explaining observed behaviour.

For artificial agents, ToM is a load-bearing part of Common Sense Reasoning about people and a prerequisite for fluent Social Interaction: a system that cannot track what its interlocutor knows will over-explain, under-explain, or mispredict their actions. In Social Robotics and human-robot teaming, machine ToM supports intention recognition, legible motion, and knowing when a human collaborator holds a stale or false belief that the robot should correct.

Current Landscape

  • Computational models: Bayesian inverse planning (Baker, Saxe, Tenenbaum) treats observed action as approximately rational given latent beliefs and goals; ToMnet (Rabinowitz et al., 2018) meta-learns agent models from behavioural traces; recursive reasoning appears in multi-agent RL as I-POMDPs and cognitive hierarchy models.

  • Language models: large models pass many classic false-belief vignettes, but performance degrades under adversarial rewording, prompting active debate over whether this constitutes robust ToM or pattern matching; benchmarks include ToMi, BigToM, FANToM, and OpenToM.

  • ToMBench (Chen et al., ACL 2024): a systematic, build-from-scratch bilingual benchmark of 2,860 items across 8 tasks and 31 ATOMS social-cognition abilities finds that even GPT-4 trails human performance by more than 10 points on average, though it surpasses humans on 9 of the 31 specific abilities.

  • Kosinski (2023) reported ChatGPT-4 solving ~75% of false-belief tasks (comparable to a six-year-old), but true-belief controls and adversarial rewording collapse accuracy, underlining fragility.

  • 2025 benchmarks broaden scope: ToMATO (AAAI 2025) evaluates first- and second-order ToM across beliefs, intentions, desires, emotions and knowledge; XToM/Multi-ToM add multilingual coverage and MuMA-ToM adds multimodal reasoning — all showing GPT-4o-class models still lag humans, especially on false beliefs.

  • Applications: assistive and service robots, dialogue systems tracking interlocutor knowledge state, pedagogical agents modelling learner misconceptions, and negotiation or teaming agents anticipating human partners.

  • Safety relevance: ToM cuts both ways — agents that model human beliefs can cooperate and communicate better, but the same capability underlies deception; alignment research therefore treats machine ToM as both a tool and a monitored risk.

    Sources:

  • https://aclanthology.org/2024.acl-long.847.pdf

  • https://arxiv.org/html/2402.15052v1

  • https://openreview.net/pdf?id=pzS8MJkf06

  • https://aclanthology.org/2025.acl-long.1522.pdf

Provenance