Do AI-based Technologies Have the Capacity to Empathize with Humans? Evaluating Theory of Mind in Large Language Models

19.02.2026

One of the defining features of human social cognition is Theory of Mind (ToM): the ability to infer other people’s beliefs, intentions, emotions, and knowledge states. ToM appears as an essential capacity driving everything from interpersonal communication and empathy to complex social decision-making. Consequently, this domain has remained a focal point for developmental, social, and clinical psychologists for decades. Additionally, recent advances in Large Language Models (LLMs), particularly GPT-based systems, have sparked debate about whether these models exhibit genuine ToM abilities or merely reproduce learned linguistic patterns.

To address this question, Strachan et al. (2024) conducted one of the most comprehensive psychological evaluations of ToM ever performed on LLMs, comparing their performance directly with human participants. In this study, published in Nature, these researchers compared GPT-4, GPT-3.5, and LLaMA2-70B with 1,907 human participants using established psychological tests, including false-belief tasks, irony comprehension, indirect intention recognition, and faux pas detection.

Can LLMs keep up the pace with human reasoning?

Strachan et al. (2024) obtained these interesting results from comparing the aforementioned LLMs:

  • GPT-4 performed at or above human levels on most ToM tasks, including understanding beliefs, intentions, and irony.
  • GPT-3.5 showed weaker but generally competent performance.
  • Both GPT models performed worse than humans on faux pas detection (recognizing when someone unintentionally says something inappropriate). However, further analyses suggested that GPT’s poorer performance reflected a cautious reasoning style rather than a lack of social understanding.
  • LLaMA2 achieved very high faux pas scores, but this appeared to result from a response bias rather than superior social reasoning.
  • GPT models often understood the situation correctly but were reluctant to commit to a conclusion when evidence was incomplete. GPT-4 used a hyper-conservative strategy that avoided making any conclusions not explicitly stated.

An opportunity to bolster LLMs-based tools to enhance computational ToM

The authors of this study suggest that modern LLMs can produce behavior highly consistent with many forms of human ToM reasoning. However, these authors argue that matching human answers does not necessarily imply human-like cognition. LLMs may arrive at correct responses through mechanisms fundamentally different from those used by people.

The findings highlight an important distinction between competence – possessing the computational ability to generate appropriate inferences -, and performance – how those inferences are expressed in specific situations -. In this regard, GPT models may possess substantial ToM competence while displaying non-human patterns of responding under uncertainty.

In addition, these findings may have important implications for AI and human interaction, computational psychology, and clinical applications. Indeed, LLMs may already possess social-reasoning capabilities sufficient for many conversational applications, making them increasingly useful in education, customer support, and mental health contexts. Furthermore, this kind of studies may contribute to the emerging field of machine psychology but, most importantly, these advanced LLMs may eventually contribute to clinical decision-support systems.

Read the full article

What are your thoughts on machine’s ToM reasoning? Is this type of application a worthy research area in clinical psychology? Leave a comment below and do not hesitate to read the full article to know more about the study!

Leave a Comment

Your email address will not be published. Required fields are marked *

Scroll to Top