返回新闻
创新AI Understanding 简报

预印本发现代理人工智能性能取决于任务成功和资源使用

一份新的 arXiv 预印本比较了 OpenClaw 和 NanoBot,在完整任务完成方面没有发现统计上确定的获胜者,但报告了时间和峰值内存方面的巨大差异。作者认为,代理评估应该将结果与产生这些结果的资源和执行记录联系起来。

5 min readRead the primary source
Source-provided image accompanying Preprint finds agentic AI performance depends on both task success and resource use
主要来源文件来源记录
出版商
arxiv.org
来源链接
arxiv.orghttps://arxiv.org/abs/2608.27886
来源类型
主要文件——我们直接阅读的官方公告、文件、文件或第一方页面。
背景60 秒内了解这一点

从这里开始

关键术语

内存(代理内存)
AI 代理跨步骤或会话使用存储的上下文来提高连续性。
基准测试
用于测量和比较模型性能的标准化测试或数据集。
工具使用
模型调用外部工具(例如搜索、计算器或 API)的能力。
测试一下自己AI 代理测验

发生了什么

An arXiv preprint submitted August 28 compares OpenClaw and NanoBot as complete agentic AI systems, including their language model, tools, memory, state management and multi-step execution. On a primary , OpenClaw completed 31% of tasks and NanoBot 25%, but the reported 95% task-bootstrap interval ran from -3 to 15 percentage points, so the study did not establish a full-completion advantage for either system.

The preprint evaluates OpenClaw and NanoBot as complete agentic systems rather than comparing language models in isolation. The authors describe agentic systems as combinations of a language model with tools, memory, state management and multi-step execution. That framing matters because each layer can affect both what an agent accomplishes and the operational resources required to attempt a task.

In the primary , OpenClaw achieved full task completion on 31% of tasks, compared with 25% for NanoBot. The six-percentage-point difference was accompanied by a 95% task-bootstrap interval ranging from minus 3 to 15 percentage points. Based on that interval, the authors say there was no statistically established full-completion advantage for either system.

The paper also reports a more detailed instrumented subset of paired prompts. In that layer, both systems reached full completion on 26% of prompts. NanoBot nevertheless reached at least partial completion on 43% of prompts, compared with 26% for OpenClaw, indicating that a full-completion-only score concealed a difference in intermediate outcomes in this subset.

Resource measurements favored NanoBot in the reported comparisons. OpenClaw took longer on 83% of prompts and recorded a higher peak-memory value on every prompt. The paper gives geometric mean ratios of 2.98 for wall time and 19.44 for peak memory. Among the ten detailed-layer prompts where at least one system achieved partial or full completion, NanoBot weakly dominated on eight. Across all 23 prompts, however, ten of its 18 dominance cases were cheaper joint failures, meaning lower resource use did not always accompany useful task progress.

来源详情: arxiv.org ↗

为什么这很重要

The paper shows why an agent that completes slightly more tasks may not be the more practical system if it consumes substantially more time or memory. Its results also suggest that conclusions can change when researchers inspect execution details and distinguish full success, partial progress and joint failure.

The central contribution is an evaluation principle: capability and cost should be measured together and tied to the specific execution that produced each result. For people choosing or deploying agentic systems, a completion percentage alone can obscure whether a system is fast enough, memory-efficient enough or consistently able to make useful partial progress.

The reported results make that trade-off concrete. OpenClaw's 31% primary- completion rate was only modestly higher than NanoBot's 25%, and the uncertainty interval did not establish a reliable winner. Yet the instrumented results show much larger operational differences, with OpenClaw taking longer on most prompts and using more peak memory on every prompt in that comparison.

The distinction between partial completion and joint failure is especially important. NanoBot's lower resource use helped it dominate several comparisons, but the authors say ten of its 18 dominance cases across all 23 prompts were cheaper joint failures. A system should not be judged as better simply because it fails at lower cost; resource efficiency has to be interpreted alongside the quality and usefulness of the outcome.

The paper also highlights a reproducibility and accountability issue. If scores are not linked to attempt-level execution records and scoring provenance, researchers and users may be unable to tell whether a result reflects full success, partial progress, a shared failure or a measurement artifact. The source presents this as a reason to make evaluation records more detailed, not as evidence that either system is generally superior.

Interactive Mechanism

互动机制:它实际上是如何运作的

以交互方式探索这一发展背后的基础技术。

Agent Lifecycle Stage:
1
User Intent & Planning: "Audit customer refund request #4092 and settle payment."
2
Tool Calling: Emits structured JSON call crm_get_transaction(id='4092').
3
Guardrail & Verification:🛡️ Paused: High-value action requires human operator sign-off.
4
Final Settlement: Refund recorded, email receipt dispatched, and audit log stored.
Core takeaway: An AI agent is not just a language model—it is a closed loop of planning, tool invocation, and environment feedback. Production systems require self-healing retries and strict human approval guardrails.
交互式概念检查+10 Points
AI Agents Quiz

An agent must create a draft calendar event for Tuesday at 2 p.m. Which evidence would establish the requested result?

接下来看什么

The main follow-up questions are whether the findings hold across larger and more varied task sets, different hardware and software configurations, and other agent frameworks. The source does not provide those details in its abstract, so the reported resource ratios should be treated as study-specific rather than universal rankings.

The immediate question is whether the reported gap in wall time and peak memory persists beyond the study's prompts and test configuration. The abstract does not state the exact task composition, hardware, software versions, number of repeated trials or measurement procedure. Those omissions limit how broadly the numerical ratios can be applied.

Readers should also watch for evaluations that report more than a single aggregate score. Useful follow-up studies would separate full completion, partial completion and joint failure; show how often each system wins on task quality; and publish the execution records needed to connect each outcome to resource consumption. The preprint's own disagreement between its primary and instrumented evidence layers makes this a practical research priority.

The authors' conclusion is methodological rather than a product recommendation. The source does not establish that NanoBot is the better general-purpose agent, nor that OpenClaw's higher resource use is unjustified for every workload. Further comparisons with other agent systems, larger samples and independent replications would be needed before treating the findings as a broad market or engineering ranking.

A meaningful unknown is how the systems' resource profiles interact with task difficulty and . The abstract reports aggregate ratios and prompt-level comparisons but does not identify which kinds of tasks caused the largest differences. That information would help determine whether the results reflect a general property of the systems or a pattern specific to the evaluated workload. Until then, the strongest supported takeaway is that agent evaluation should report verified outcomes together with observed resource use and scoring provenance.

相关指南和测验

人工智能代理人工智能模型解释人工智能培训测试你所知道的——尝试免费的人工智能测验在我们的词汇表中查找人工智能术语关注 AI 模型发布跟踪器
觉得这有用吗?