What happened
Researchers evaluated whether large language models can use information about other players’ beliefs and intentions to guide decisions. They tested 2,099 individual agents from four model families—DeepSeek, GPT-4.1, GPT-5 and Gemini 2.0 Flash—in two economic games against opponents with different levels of sophistication, then compared the results with 251 human participants.
The preprint, submitted to arXiv on Aug. 26, examines mentalization, which the authors define as inferring other people’s beliefs and intentions in order to guide one’s own choices. The research question is narrower than whether models possess human consciousness or subjective understanding: it asks whether their decisions show behavioral and computational signatures consistent with using representations of other agents’ mental states.
The researchers used two economic games and cognitive computational modeling to study the latent strategies behind model decisions. They tested individual agents from four model families—DeepSeek, GPT-4.1, GPT-5 and Gemini 2.0 Flash—against opponents described in the abstract as having varying sophistication. The total model sample was 2,099 agents. A separate comparison involved 251 human participants.
The study also evaluated a prompting strategy intended to elicit more deliberate strategic reasoning. According to the abstract, strategic prompting generally improved performance, although the size of the benefit differed between the two games. The authors therefore report an effect on both outcomes and inferred reasoning strategy, rather than a uniform improvement across all settings.
The clearest model-specific result concerned GPT-5. The authors say GPT-5 agents flexibly adjusted the recursive depth of their reasoning as opponents became more sophisticated and achieved better performance than the human participants in these games. The abstract does not state the numerical scores, the margin of difference, the identities or versions of the tested systems beyond the labels listed, or how the human participants were recruited and instructed.
Why it matters
The study addresses a central question about AI systems that interact with people: whether behavior that resembles theory-of-mind performance can support adaptive strategy. Its results suggest that models can exhibit different, measurable forms of mentalizing, while also showing why apparently similar language models should not be treated as cognitively interchangeable.
Mentalization is relevant to AI systems that must respond to people, competitors or other software agents. A system that can estimate what another party knows, believes or intends may make different decisions from one that only matches surface patterns in text. The study’s contribution is to frame this ability as something that can be measured through choices and computational models, not only through conversational answers to theory-of-mind questions.
The reported differences among model providers and sizes are important because they challenge the idea that a strong result on one social-reasoning test describes all large language models. The abstract says the systems showed “clear behavioural and computational signatures” of mentalizing, but that those signatures differed markedly across providers and model sizes. In practical evaluation, this points toward testing specific models under specific interaction conditions rather than relying on a single general intelligence label.
The prompting result has a more limited but useful implication. Asking a model to engage in strategic reasoning may improve its behavior in some settings, but the uneven effects across the two games indicate that prompting is not a universal capability upgrade. Performance gains may depend on the task, the opponent and the way the model’s reasoning is elicited. The source does not establish that the models’ internal processes became more human-like; it reports changes in behavior and computationally inferred strategy.
The reported GPT-5 result could matter for research on negotiation, coordination and multi-agent systems, but it should not be read as evidence that GPT-5 is generally better than humans at social reasoning. The comparison was limited to two economic games, and the source gives no evidence about everyday relationships, high-stakes decisions, deception, cultural variation or real-world interaction. Nor does it establish that success in these games transfers to safe or reliable behavior outside the study.
What to watch next
The paper is an arXiv preprint, and the source page provides an abstract rather than the task-level scores, statistical uncertainty, prompts or full methodological details needed to assess the findings closely. Follow-up work should test whether the reported strategies generalize beyond controlled games and whether they reflect robust reasoning rather than learned patterns that fit particular benchmarks.
The most immediate question is whether the full paper supplies enough methodological detail to evaluate the comparison. Important missing information from the source includes the games’ rules, the sophistication levels assigned to opponents, the exact prompts, the number of trials per agent, model sampling settings, scoring procedures and uncertainty estimates. Those details determine whether the result is robust or sensitive to benchmark design.
Replication across independent model evaluations would help test whether the findings persist across model updates and different access conditions. The source identifies GPT-4.1, GPT-5 and Gemini 2.0 Flash, but does not specify deployment settings, system instructions or dates of access. Because model behavior can change with prompting, sampling and product configuration, later studies should document those variables and use preregistered evaluation protocols where possible.
Researchers should also examine whether the inferred recursive reasoning predicts behavior in unfamiliar tasks. A model may perform well in a game because its training data or prompt structure supports a familiar strategy, without possessing a broadly transferable representation of other minds. Useful extensions would include new games, mixed human-model groups, opponents who behave unpredictably, and settings where social cues are incomplete or misleading.
The practical boundary remains uncertain. This preprint does not test a deployed product, recommend using an AI system in consequential decisions or measure harms. Before applying such findings to negotiation, education, mental-health support or other human-facing uses, evaluators would need evidence about reliability, calibration, fairness across populations, resistance to manipulation and failure modes when the model’s assumptions about another person are wrong.

