返回新聞
創新AI Understanding 簡報

研究發現法學碩士可以辨識不可能的問題,但常常無法迴避

EMNLP 2026 接受的一篇論文報告稱,語言模型代表某些數學和代碼問題在結構上是否可以回答,但無法將該訊號轉化為拒絕行為。

5 min readRead the primary source
Source-provided image accompanying Study finds LLMs can recognize impossible questions but often fail to abstain
主要來源文件來源記錄
出版商
arxiv.org
來源連結
arxiv.orghttps://arxiv.org/abs/2608.29109
來源類型
主要文件-我們直接閱讀的官方公告、文件、文件或第一方頁面。
背景60 秒內了解這一點

從這裡開始

關鍵術語

API(應用程式介面)
一種軟體系統向另一個系統發送請求並接收回應的結構化方式。
培訓後
預訓練後應用的訓練步驟,例如指令調整、偏好最佳化和安全調整。
校準
模型的置信度分數與實際正確性機率的匹配程度。
測試一下自己AI 模型解釋測驗

發生了什麼事

Researchers Yucheng Du and Xiyang Hu examined why instruction-tuned language models sometimes answer questions that have no valid answer, such as computing cot(-540°) or evaluating (1).startswith("1"). Across models ranging from 1.7 billion to 70 billion parameters, they report finding a hidden-state direction that separates answerable prompts from structurally impossible ones. The paper argues that this indicates recognition is present, but is not reliably connected to abstention.

The paper focuses on a specific reliability failure: large language models answer prompts that are structurally unanswerable instead of declining to answer. Its examples include the trigonometric expression cot(-540°) and the code expression (1).startswith("1"). The authors distinguish this failure from ordinary factual error: the issue is whether a valid answer exists under the structure of the question at all.

Du and Hu report experiments across instruction-tuned models spanning 1.7 billion to 70 billion parameters. They say a single linear direction in the models’ hidden states separates answerable prompts from structurally impossible math and code prompts. In the authors’ interpretation, this separation shows that the models represent impossibility before they generate a response.

The reported recognition direction was nearly orthogonal to what the paper calls the canonical safety-refusal direction, which the authors associate with trained refusal of harmful content. An invalidity-aware direction defined from in-domain behavior was closer to the recognition direction, but only partially aligned with it and remained nearly orthogonal to the safety-refusal direction.

The authors also report that generation-time steering along the recognition direction changed invalidity-aware behavior in both directions and in a dose-responsive manner on structural math and code cells. Random directions did not produce the same reported effect. Comparisons between base and instruction-tuned models further suggested that the low-cosine geometry was already present at the pretraining endpoint. The paper concludes that the failure is better explained as a routing failure than an encoding failure: the model has a usable signal that there is no admissible answer, but its refusal pathway is not aligned to use it.

來源詳情: arxiv.org ↗

為什麼這很重要

The result challenges a simple explanation that models answer impossible questions because they cannot recognize the problem. If the paper’s interpretation holds, improving reliability may require routing or decision-making mechanisms that convert an internal impossibility signal into an explicit refusal or clarification. The study is a preprint-level research report, however, and its abstract does not establish how broadly the findings generalize beyond the tested math and code prompts or models.

The practical importance of the paper’s claim is its proposed diagnosis. If a model cannot represent that a prompt is impossible, developers might need better training data or representations. But if the model already carries that signal and fails to act on it, the intervention point may instead be the connection between internal recognition, response selection, and abstention.

That distinction matters because a fluent answer to an impossible question can appear more useful than a refusal while being fundamentally invalid. A system that recognizes structural impossibility could instead state that the prompt has no admissible answer, explain the relevant limitation, or ask the user to correct the input. The source does not report a deployed product or a demonstrated improvement in user outcomes, so those applications remain implications rather than established results.

The comparison with safety refusal is also consequential. The paper reports that the internal direction associated with recognizing impossible prompts is nearly orthogonal to a direction associated with harmful-content refusal. That suggests that treating all refusals as one generic capability may miss important distinctions between refusing because content is unsafe and abstaining because a question is malformed or has no valid solution.

The findings should be read with clear limits. The source provides no model names, benchmark sizes, accuracy tables, error rates, steering costs, or independent replication. It does not show that every model has the same geometry, that the reported directions are stable across languages or modalities, or that steering avoids unwanted changes in helpfulness. The paper is identified as accepted to EMNLP 2026, but the supplied material remains an arXiv abstract and does not by itself establish the full strength of the claims.

Interactive Mechanism

互動機制:它實際上是如何運作的

以互動方式探索這項發展背後的基礎技術。

Agent Lifecycle Stage:
1
User Intent & Planning: "Audit customer refund request #4092 and settle payment."
2
Tool Calling: Emits structured JSON call crm_get_transaction(id='4092').
3
Guardrail & Verification:🛡️ Paused: High-value action requires human operator sign-off.
4
Final Settlement: Refund recorded, email receipt dispatched, and audit log stored.
Core takeaway: An AI agent is not just a language model—it is a closed loop of planning, tool invocation, and environment feedback. Production systems require self-healing retries and strict human approval guardrails.
互動式概念檢查+10 Points
AI Models Explained Quiz

Which component of an AI application is the machine-learning model itself?

接下來看什麼

The key follow-up is whether the reported recognition signal survives broader testing and can be used in deployed systems without suppressing legitimate answers. The paper reports that steering along the direction changed behavior in both directions and in a dose-responsive way, while random directions did not. Further work would need to test whether this intervention improves real-world , how often it produces unnecessary refusals, and whether the geometry changes after later training.

A central question is scope. The reported experiments cover structural math and code prompts, but the abstract does not say whether the same recognition-and-routing pattern appears in everyday factual questions, ambiguous instructions, tool use, or multimodal systems. Testing those settings would show whether the proposed diagnosis is a narrow property of selected prompts or a broader model behavior.

The steering result deserves careful replication. The authors report bidirectional, dose-responsive behavior when steering along the recognition direction and no comparable effect from random directions. Follow-up studies should measure whether such steering improves abstention at a useful operating point, or whether it creates false refusals, changes the style of answers, or becomes fragile across model versions and decoding settings.

Developers and evaluators should also watch for better metrics that separate recognition from action. A model might internally detect an impossible prompt but still answer it, or refuse without correctly identifying why. Evaluations that record both the model’s recognition signal and its final behavior could test the paper’s routing-failure hypothesis more directly.

Important unknowns remain about training and deployment. The source does not establish whether can align the recognition and refusal pathways reliably, whether the intervention requires access to hidden states unavailable through an API, or how the approach would interact with safety safeguards. Those questions will determine whether the result becomes a practical reliability technique or remains an explanatory finding about model internals.

相關指引和測驗

人工智慧模型解釋ChatGPT 與大型語言模型AI 倫理人工智慧培訓測試你所知道的—嘗試免費的人工智慧測驗在我們的詞彙表中尋找人工智慧術語關注 AI 模型發布追蹤器
覺得有用嗎?