返回新聞
創新AI Understanding 簡報

預印本比較了大型語言模型如何推斷對手對經濟博弈的信念

一份新的預印本報告稱,大型語言模型在兩種經濟遊戲中顯示出可測量的心智化跡象,但它們的表現和策略因模型系列和規模而異。作者表示,GPT-5 智能體適應了日益複雜的對手,並在…方面優於人類對照組。

5 min readRead the primary source
Primary-source image accompanying Preprint compares how large language models infer opponents’ beliefs in economic games
主要來源文件來源記錄
出版商
arxiv.org
來源連結
arxiv.orghttps://arxiv.org/abs/2608.26291
來源類型
主要文件-我們直接閱讀的官方公告、文件、文件或第一方頁面。
背景60 秒內了解這一點

從這裡開始

關鍵術語

校準
模型的置信度分數與實際正確性機率的匹配程度。
基準測試
用於測量和比較模型性能的標準化測試或資料集。
提示
提供給生成模型的輸入指令和上下文。
測試一下自己AI 模型解釋測驗

發生了什麼事

Researchers evaluated whether large language models can use information about other players’ beliefs and intentions to guide decisions. They tested 2,099 individual agents from four model families—DeepSeek, GPT-4.1, GPT-5 and Gemini 2.0 Flash—in two economic games against opponents with different levels of sophistication, then compared the results with 251 human participants.

The preprint, submitted to arXiv on Aug. 26, examines mentalization, which the authors define as inferring other people’s beliefs and intentions in order to guide one’s own choices. The research question is narrower than whether models possess human consciousness or subjective understanding: it asks whether their decisions show behavioral and computational signatures consistent with using representations of other agents’ mental states.

The researchers used two economic games and cognitive computational modeling to study the latent strategies behind model decisions. They tested individual agents from four model families—DeepSeek, GPT-4.1, GPT-5 and Gemini 2.0 Flash—against opponents described in the abstract as having varying sophistication. The total model sample was 2,099 agents. A separate comparison involved 251 human participants.

The study also evaluated a prompting strategy intended to elicit more deliberate strategic reasoning. According to the abstract, strategic prompting generally improved performance, although the size of the benefit differed between the two games. The authors therefore report an effect on both outcomes and inferred reasoning strategy, rather than a uniform improvement across all settings.

The clearest model-specific result concerned GPT-5. The authors say GPT-5 agents flexibly adjusted the recursive depth of their reasoning as opponents became more sophisticated and achieved better performance than the human participants in these games. The abstract does not state the numerical scores, the margin of difference, the identities or versions of the tested systems beyond the labels listed, or how the human participants were recruited and instructed.

來源詳情: arxiv.org ↗

為什麼這很重要

The study addresses a central question about AI systems that interact with people: whether behavior that resembles theory-of-mind performance can support adaptive strategy. Its results suggest that models can exhibit different, measurable forms of mentalizing, while also showing why apparently similar language models should not be treated as cognitively interchangeable.

Mentalization is relevant to AI systems that must respond to people, competitors or other software agents. A system that can estimate what another party knows, believes or intends may make different decisions from one that only matches surface patterns in text. The study’s contribution is to frame this ability as something that can be measured through choices and computational models, not only through conversational answers to theory-of-mind questions.

The reported differences among model providers and sizes are important because they challenge the idea that a strong result on one social-reasoning test describes all large language models. The abstract says the systems showed “clear behavioural and computational signatures” of mentalizing, but that those signatures differed markedly across providers and model sizes. In practical evaluation, this points toward testing specific models under specific interaction conditions rather than relying on a single general intelligence label.

The prompting result has a more limited but useful implication. Asking a model to engage in strategic reasoning may improve its behavior in some settings, but the uneven effects across the two games indicate that prompting is not a universal capability upgrade. Performance gains may depend on the task, the opponent and the way the model’s reasoning is elicited. The source does not establish that the models’ internal processes became more human-like; it reports changes in behavior and computationally inferred strategy.

The reported GPT-5 result could matter for research on negotiation, coordination and multi-agent systems, but it should not be read as evidence that GPT-5 is generally better than humans at social reasoning. The comparison was limited to two economic games, and the source gives no evidence about everyday relationships, high-stakes decisions, deception, cultural variation or real-world interaction. Nor does it establish that success in these games transfers to safe or reliable behavior outside the study.

Interactive Mechanism

互動機制:它實際上是如何運作的

以互動方式探索這項發展背後的基礎技術。

Thinking Budget (Test-Time Tokens):1,024 tokens
Complex Accuracy79%Math & Code Logic
Latency3.2sTime to first full output
Inference Cost$0.0092Per query estimated
Reasoning StyleStep VerificationInternal chain depth
Active Thinking Trace:
1Deconstruct user problem into formal constraints
2Propose candidate hypotheses & step-by-step calculation
3Self-correction: Backtrack and refute subtle edge cases
4Exhaustive consistency check & final output synthesis
Core takeaway: Test-time compute fundamentally changes AI economics. Instead of only scaling during pre-training, giving reasoning models more tokens at inference time allows them to systematically solve PhD-level STEM problems.
互動式概念檢查+10 Points
AI Models Explained Quiz

Which component of an AI application is the machine-learning model itself?

接下來看什麼

The paper is an arXiv preprint, and the source page provides an abstract rather than the task-level scores, statistical uncertainty, prompts or full methodological details needed to assess the findings closely. Follow-up work should test whether the reported strategies generalize beyond controlled games and whether they reflect robust reasoning rather than learned patterns that fit particular benchmarks.

The most immediate question is whether the full paper supplies enough methodological detail to evaluate the comparison. Important missing information from the source includes the games’ rules, the sophistication levels assigned to opponents, the exact prompts, the number of trials per agent, model sampling settings, scoring procedures and uncertainty estimates. Those details determine whether the result is robust or sensitive to design.

Replication across independent model evaluations would help test whether the findings persist across model updates and different access conditions. The source identifies GPT-4.1, GPT-5 and Gemini 2.0 Flash, but does not specify deployment settings, system instructions or dates of access. Because model behavior can change with prompting, sampling and product configuration, later studies should document those variables and use preregistered evaluation protocols where possible.

Researchers should also examine whether the inferred recursive reasoning predicts behavior in unfamiliar tasks. A model may perform well in a game because its training data or structure supports a familiar strategy, without possessing a broadly transferable representation of other minds. Useful extensions would include new games, mixed human-model groups, opponents who behave unpredictably, and settings where social cues are incomplete or misleading.

The practical boundary remains uncertain. This preprint does not test a deployed product, recommend using an AI system in consequential decisions or measure harms. Before applying such findings to negotiation, education, mental-health support or other human-facing uses, evaluators would need evidence about reliability, , fairness across populations, resistance to manipulation and failure modes when the model’s assumptions about another person are wrong.

相關指引和測驗

人工智慧模型解釋ChatGPT 與大型語言模型AI 倫理AI 的未來測試你所知道的—嘗試免費的人工智慧測驗在我們的詞彙表中尋找人工智慧術語關注 AI 模型發布追蹤器
覺得有用嗎?