返回新闻
创新AI Understanding 简报

论文提出版本化合约以防止人工智能代理内存过时

一篇新的 arXiv 论文报告称,版本标记和行级驱逐可以帮助人工智能代理在服务器端更改后重用恢复建议,而无需默默地应用过时的修复程序。

6 min readRead the primary source
Source-page capture accompanying Paper proposes versioned contracts to keep AI agent memory from going stale
主要来源文件来源记录
出版商
arxiv.org
来源链接
arxiv.orghttps://arxiv.org/abs/2609.00243
来源类型
主要文件——我们直接阅读的官方公告、文件、文件或第一方页面。
背景60 秒内了解这一点

从这里开始

关键术语

内存(代理内存)
AI 代理跨步骤或会话使用存储的上下文来提高连续性。
人工智能代理
一种可以观察、推理并采取行动来实现目标的软件系统,通常使用工具和内存。
API(应用程序编程接口)
一种软件系统向另一个系统发送请求并接收响应的结构化方式。
测试一下自己AI 代理测验

发生了什么

Researchers Michael Wu and Arquimedes Canedo propose “invalidation contracts,” a protocol layer for AI agents that cache recovery suggestions from API errors across episodes. The contracts attach version stamps and cacheability hints to each suggestion so clients can remove stale entries after server-side data drift while preserving entries that remain valid.

The paper addresses AI agents that retain recovery suggestions after encountering API errors. Its central concern is that a suggestion that worked in one episode may become wrong after server-side data changes. Re-deriving a fix on every episode can avoid that failure mode, but the authors say it gives back the savings from caching. Their proposed solution, called invalidation contracts, adds protocol metadata to each recovery suggestion rather than treating the cache as an undifferentiated block.

The contract attaches version stamps and cacheability hints to suggestions. The client can then evict entries associated with changed data while keeping entries that are still valid. The authors separate the resulting savings into validity—the fraction of cached suggestions that remain correct after a drift event—and compliance—the fraction that a planner applies correctly on its first attempt. The paper says validity is determined by the protocol and is vendor-independent, while compliance depends on the planner model.

The evaluation covered seven models, three serving paths, two domains and approximately 9,400 episodes. The abstract reports that row-level invalidation increased compliance by between 0 and 66.7 percentage points across the models; three models saw gains between 55.6 and 66.7 points. Four of the seven models recovered 29% to 33% of baseline token cost. The paper also reports perfect eviction precision, 1.00, at row granularity under the row-level oracle described in Section 4.1.

The results varied sharply by model. The authors report 100% first-try compliance for Claude Haiku 4.5 with identical wire bytes, compared with 11% or less for Claude Sonnet 5. They attribute Sonnet’s behavior to input-schema conservatism: refusing fixes that add fields absent from the original request. By contrast, table-level invalidation destroyed co-located entries and reduced post-drift first-try rates to 0% on five of seven models. The contract added 15% to the response payload, and the authors report zero contract failures across the evaluation.

来源详情: arxiv.org ↗

为什么这很重要

The paper frames persistent agent memory as a reliability and efficiency problem: recomputing fixes every time avoids stale advice but sacrifices the token and model-call savings that caching can provide. Its reported results suggest that invalidating memory at row level can preserve useful cache entries, although the findings come from one preprint’s evaluation and do not establish production performance.

The practical issue is not simply whether an agent can remember. It is whether the memory remains safe to use when an external service changes. A cached recovery suggestion can reduce repeated reasoning, token use and model calls, but the same shortcut can become a silent failure if the surrounding API or data has drifted. The proposed protocol makes freshness information part of the interaction between the service and the client.

The paper’s distinction between validity and compliance is useful because it separates two different failure sources. A protocol can identify which cached entries should be discarded, yet an agent may still fail to apply a valid suggestion. The reported contrast between Haiku 4.5 and Sonnet 5 indicates that protocol design alone may not determine whether a planner benefits from preserved memory. Model behavior remains a separate operational constraint.

The reported row-level results suggest a possible design principle for systems that store multiple recovery suggestions together: invalidate only the affected entries when the protocol can identify them. The table-level comparison is consequential within the paper’s tests because broad invalidation eliminated unrelated entries and produced zero post-drift first-try rates on most tested models. That finding, if replicated, would matter for the cost and reliability of long-running agents.

The efficiency claim is also bounded. The authors report recovery of 29% to 33% of baseline token cost on four models, not across all seven, and the protocol adds 15% to response payloads. The source does not say how those costs translate into latency, bandwidth, infrastructure expense or user-visible reliability in production. Nor does the abstract establish that the tested domains or serving paths represent the range of APIs used by deployed agents.

This is a preprint rather than an independently validated production result. The source identifies the evaluation size and headline outcomes but does not provide the identities of five of the seven models, the names of the two domains, the exact drift scenarios, or uncertainty estimates. The perfect eviction precision is explicitly tied to a row-level oracle, so it should not be read as evidence that deployed clients will always know which rows are stale.

Interactive Mechanism

互动机制:它实际上是如何运作的

以交互方式探索这一发展背后的基础技术。

Thinking Budget (Test-Time Tokens):1,024 tokens
Complex Accuracy79%Math & Code Logic
Latency3.2sTime to first full output
Inference Cost$0.0092Per query estimated
Reasoning StyleStep VerificationInternal chain depth
Active Thinking Trace:
1Deconstruct user problem into formal constraints
2Propose candidate hypotheses & step-by-step calculation
3Self-correction: Backtrack and refute subtle edge cases
4Exhaustive consistency check & final output synthesis
Core takeaway: Test-time compute fundamentally changes AI economics. Instead of only scaling during pre-training, giving reasoning models more tokens at inference time allows them to systematically solve PhD-level STEM problems.
交互式概念检查+10 Points
AI Agents Quiz

An agent must create a draft calendar event for Tuesday at 2 p.m. Which evidence would establish the requested result?

接下来看什么

The key open questions are whether the protocol works with more planners, services, domains and real-world drift patterns, and whether its 15% response-payload overhead is acceptable in deployed systems. Further scrutiny should also examine the paper’s oracle-based eviction evaluation, the unidentified models among the seven tested, and the causes of the large compliance gap between the named Claude models.

Replication should test whether version-stamp validity remains deterministic when services change schemas, semantics or data at different rates. The paper reports identical validity results across every model and serving path and zero contract failures, but the source does not describe the range of drift events in enough detail to determine how broadly that result applies.

Model-side compliance deserves separate investigation. The reported gap between 100% first-try compliance on Claude Haiku 4.5 and 11% or below on Claude Sonnet 5 suggests that agents may interpret the same protocol payload differently. Future evaluations should clarify whether the problem is specific to the named models, to prompt or schema design, or to a broader tendency among planners to reject structurally unfamiliar fixes.

The cost tradeoff should be measured beyond tokens. A 15% response-payload increase could be minor or material depending on episode length, network conditions and how often cached suggestions are reused. The source reports token-cost recovery for four models but does not provide latency, bandwidth, monetary or failure-rate results, so those remain meaningful unknowns for implementers.

The evaluation’s row-level oracle is another point to scrutinize. Perfect precision under an oracle demonstrates the behavior of the proposed granularity when the affected row is known, but it does not establish that an actual client can identify the correct row from ordinary service signals. Evidence about automatic detection, false negatives and false positives would determine how much of the reported benefit survives outside the experimental setup.

Finally, the work should be compared with other approaches to persistent agent state and memory invalidation under the same tasks and drift conditions. The source establishes a protocol proposal and a reported benchmark, but it does not establish deployment availability, adoption, independent replication or superiority over unspecified alternatives.

相关指南和测验

人工智能代理人工智能模型解释人工智能培训AI 的未来测试你所知道的——尝试免费的人工智能测验在我们的词汇表中查找人工智能术语关注 AI 模型发布跟踪器
觉得这有用吗?