Torna alle notizie
InnovazioneAI Understanding briefing

Preprint confronta il modo in cui i modelli linguistici di grandi dimensioni deducono le convinzioni degli avversari nei giochi economici

Una nuova prestampa riporta che i modelli linguistici di grandi dimensioni hanno mostrato segni misurabili di mentalizzazione in due giochi economici, ma le loro prestazioni e strategie differivano in base alla famiglia e alle dimensioni del modello. Gli autori affermano che gli agenti GPT-5 si sono adattati ad avversari sempre più sofisticati e hanno sovraperformato il gruppo di confronto umano in...

5 min readRead the primary source
Primary-source image accompanying Preprint compares how large language models infer opponents’ beliefs in economic games
Documento di origine primariaFonte registrata
Editore
arxiv.org
Collegamento alla fonte
arxiv.orghttps://arxiv.org/abs/2608.26291
Tipo di fonte
Documento principale: un annuncio ufficiale, un documento, un documento o una pagina proprietaria che leggiamo direttamente.
ContestoComprendilo in 60 secondi

Inizia qui

Termini chiave

Calibrazione
Quanto bene i punteggi di confidenza di un modello corrispondono alle probabilità di correttezza effettive.
Punto di riferimento
Un test o un set di dati standardizzato utilizzato per misurare e confrontare le prestazioni del modello.
Richiedi
Le istruzioni di input e il contesto forniti a un modello generativo.
Mettiti alla provaQuiz sulla spiegazione dei modelli di intelligenza artificiale

Cosa è successo

Researchers evaluated whether large language models can use information about other players’ beliefs and intentions to guide decisions. They tested 2,099 individual agents from four model families—DeepSeek, GPT-4.1, GPT-5 and Gemini 2.0 Flash—in two economic games against opponents with different levels of sophistication, then compared the results with 251 human participants.

The preprint, submitted to arXiv on Aug. 26, examines mentalization, which the authors define as inferring other people’s beliefs and intentions in order to guide one’s own choices. The research question is narrower than whether models possess human consciousness or subjective understanding: it asks whether their decisions show behavioral and computational signatures consistent with using representations of other agents’ mental states.

The researchers used two economic games and cognitive computational modeling to study the latent strategies behind model decisions. They tested individual agents from four model families—DeepSeek, GPT-4.1, GPT-5 and Gemini 2.0 Flash—against opponents described in the abstract as having varying sophistication. The total model sample was 2,099 agents. A separate comparison involved 251 human participants.

The study also evaluated a prompting strategy intended to elicit more deliberate strategic reasoning. According to the abstract, strategic prompting generally improved performance, although the size of the benefit differed between the two games. The authors therefore report an effect on both outcomes and inferred reasoning strategy, rather than a uniform improvement across all settings.

The clearest model-specific result concerned GPT-5. The authors say GPT-5 agents flexibly adjusted the recursive depth of their reasoning as opponents became more sophisticated and achieved better performance than the human participants in these games. The abstract does not state the numerical scores, the margin of difference, the identities or versions of the tested systems beyond the labels listed, or how the human participants were recruited and instructed.

Dettagli della fonte: arxiv.org ↗

Perché è importante

The study addresses a central question about AI systems that interact with people: whether behavior that resembles theory-of-mind performance can support adaptive strategy. Its results suggest that models can exhibit different, measurable forms of mentalizing, while also showing why apparently similar language models should not be treated as cognitively interchangeable.

Mentalization is relevant to AI systems that must respond to people, competitors or other software agents. A system that can estimate what another party knows, believes or intends may make different decisions from one that only matches surface patterns in text. The study’s contribution is to frame this ability as something that can be measured through choices and computational models, not only through conversational answers to theory-of-mind questions.

The reported differences among model providers and sizes are important because they challenge the idea that a strong result on one social-reasoning test describes all large language models. The abstract says the systems showed “clear behavioural and computational signatures” of mentalizing, but that those signatures differed markedly across providers and model sizes. In practical evaluation, this points toward testing specific models under specific interaction conditions rather than relying on a single general intelligence label.

The prompting result has a more limited but useful implication. Asking a model to engage in strategic reasoning may improve its behavior in some settings, but the uneven effects across the two games indicate that prompting is not a universal capability upgrade. Performance gains may depend on the task, the opponent and the way the model’s reasoning is elicited. The source does not establish that the models’ internal processes became more human-like; it reports changes in behavior and computationally inferred strategy.

The reported GPT-5 result could matter for research on negotiation, coordination and multi-agent systems, but it should not be read as evidence that GPT-5 is generally better than humans at social reasoning. The comparison was limited to two economic games, and the source gives no evidence about everyday relationships, high-stakes decisions, deception, cultural variation or real-world interaction. Nor does it establish that success in these games transfers to safe or reliable behavior outside the study.

Interactive Mechanism

Meccanismo interattivo: come funziona realmente

Esplora la tecnologia alla base di questo sviluppo in modo interattivo.

Thinking Budget (Test-Time Tokens):1,024 tokens
Complex Accuracy79%Math & Code Logic
Latency3.2sTime to first full output
Inference Cost$0.0092Per query estimated
Reasoning StyleStep VerificationInternal chain depth
Active Thinking Trace:
1Deconstruct user problem into formal constraints
2Propose candidate hypotheses & step-by-step calculation
3Self-correction: Backtrack and refute subtle edge cases
4Exhaustive consistency check & final output synthesis
Core takeaway: Test-time compute fundamentally changes AI economics. Instead of only scaling during pre-training, giving reasoning models more tokens at inference time allows them to systematically solve PhD-level STEM problems.
Verifica concettuale interattiva+10 Points
AI Models Explained Quiz

Which component of an AI application is the machine-learning model itself?

Cosa guardare dopo

The paper is an arXiv preprint, and the source page provides an abstract rather than the task-level scores, statistical uncertainty, prompts or full methodological details needed to assess the findings closely. Follow-up work should test whether the reported strategies generalize beyond controlled games and whether they reflect robust reasoning rather than learned patterns that fit particular benchmarks.

The most immediate question is whether the full paper supplies enough methodological detail to evaluate the comparison. Important missing information from the source includes the games’ rules, the sophistication levels assigned to opponents, the exact prompts, the number of trials per agent, model sampling settings, scoring procedures and uncertainty estimates. Those details determine whether the result is robust or sensitive to design.

Replication across independent model evaluations would help test whether the findings persist across model updates and different access conditions. The source identifies GPT-4.1, GPT-5 and Gemini 2.0 Flash, but does not specify deployment settings, system instructions or dates of access. Because model behavior can change with prompting, sampling and product configuration, later studies should document those variables and use preregistered evaluation protocols where possible.

Researchers should also examine whether the inferred recursive reasoning predicts behavior in unfamiliar tasks. A model may perform well in a game because its training data or structure supports a familiar strategy, without possessing a broadly transferable representation of other minds. Useful extensions would include new games, mixed human-model groups, opponents who behave unpredictably, and settings where social cues are incomplete or misleading.

The practical boundary remains uncertain. This preprint does not test a deployed product, recommend using an AI system in consequential decisions or measure harms. Before applying such findings to negotiation, education, mental-health support or other human-facing uses, evaluators would need evidence about reliability, , fairness across populations, resistance to manipulation and failure modes when the model’s assumptions about another person are wrong.

Guide e quiz correlati

Spiegazione dei modelli di intelligenza artificialeChatGPT e LLMEtica dell'IAFuturo dell'IAMetti alla prova ciò che sai: prova un quiz gratuito sull'intelligenza artificialeCerca un termine AI nel nostro glossarioSegui il tracker del rilascio del modello AI
Lo hai trovato utile?