Quay lại Tin tức
Đổi mớiAI Understanding tóm tắt

Bản in trước nhận thấy hiệu suất của AI tác nhân phụ thuộc vào cả thành công của nhiệm vụ và việc sử dụng tài nguyên

Bản in trước arXiv mới so sánh OpenClaw và NanoBot không tìm thấy người chiến thắng được xác lập về mặt thống kê khi hoàn thành đầy đủ nhiệm vụ, nhưng báo cáo sự khác biệt lớn về thời gian và bộ nhớ cao nhất. Các tác giả lập luận rằng việc đánh giá tác nhân nên liên kết kết quả với các nguồn lực và hồ sơ thực thi đã tạo ra chúng.

5 min readRead the primary source
Source-provided image accompanying Preprint finds agentic AI performance depends on both task success and resource use
Tài liệu nguồn chínhNguồn đã ghi
Nhà xuất bản
arxiv.org
Liên kết nguồn
arxiv.orghttps://arxiv.org/abs/2608.27886
Loại nguồn
Tài liệu chính - một thông báo chính thức, giấy tờ, hồ sơ hoặc trang của bên thứ nhất mà chúng tôi đọc trực tiếp.
Bối cảnhHiểu điều này trong 60 giây

Bắt đầu ở đây

Thuật ngữ chính

Bộ nhớ (Bộ nhớ tác nhân)
Bối cảnh được lưu trữ mà tác nhân AI sử dụng qua các bước hoặc phiên để cải thiện tính liên tục.
Điểm chuẩn
Một bài kiểm tra hoặc tập dữ liệu được tiêu chuẩn hóa dùng để đo lường và so sánh hiệu suất của mô hình.
Sử dụng công cụ
Khả năng của mô hình để gọi các công cụ bên ngoài như tìm kiếm, máy tính hoặc API.
Tự kiểm traCâu đố về đại lý AI

Chuyện gì đã xảy ra

An arXiv preprint submitted August 28 compares OpenClaw and NanoBot as complete agentic AI systems, including their language model, tools, memory, state management and multi-step execution. On a primary , OpenClaw completed 31% of tasks and NanoBot 25%, but the reported 95% task-bootstrap interval ran from -3 to 15 percentage points, so the study did not establish a full-completion advantage for either system.

The preprint evaluates OpenClaw and NanoBot as complete agentic systems rather than comparing language models in isolation. The authors describe agentic systems as combinations of a language model with tools, memory, state management and multi-step execution. That framing matters because each layer can affect both what an agent accomplishes and the operational resources required to attempt a task.

In the primary , OpenClaw achieved full task completion on 31% of tasks, compared with 25% for NanoBot. The six-percentage-point difference was accompanied by a 95% task-bootstrap interval ranging from minus 3 to 15 percentage points. Based on that interval, the authors say there was no statistically established full-completion advantage for either system.

The paper also reports a more detailed instrumented subset of paired prompts. In that layer, both systems reached full completion on 26% of prompts. NanoBot nevertheless reached at least partial completion on 43% of prompts, compared with 26% for OpenClaw, indicating that a full-completion-only score concealed a difference in intermediate outcomes in this subset.

Resource measurements favored NanoBot in the reported comparisons. OpenClaw took longer on 83% of prompts and recorded a higher peak-memory value on every prompt. The paper gives geometric mean ratios of 2.98 for wall time and 19.44 for peak memory. Among the ten detailed-layer prompts where at least one system achieved partial or full completion, NanoBot weakly dominated on eight. Across all 23 prompts, however, ten of its 18 dominance cases were cheaper joint failures, meaning lower resource use did not always accompany useful task progress.

Chi tiết nguồn: arxiv.org ↗

Tại sao nó quan trọng

The paper shows why an agent that completes slightly more tasks may not be the more practical system if it consumes substantially more time or memory. Its results also suggest that conclusions can change when researchers inspect execution details and distinguish full success, partial progress and joint failure.

The central contribution is an evaluation principle: capability and cost should be measured together and tied to the specific execution that produced each result. For people choosing or deploying agentic systems, a completion percentage alone can obscure whether a system is fast enough, memory-efficient enough or consistently able to make useful partial progress.

The reported results make that trade-off concrete. OpenClaw's 31% primary- completion rate was only modestly higher than NanoBot's 25%, and the uncertainty interval did not establish a reliable winner. Yet the instrumented results show much larger operational differences, with OpenClaw taking longer on most prompts and using more peak memory on every prompt in that comparison.

The distinction between partial completion and joint failure is especially important. NanoBot's lower resource use helped it dominate several comparisons, but the authors say ten of its 18 dominance cases across all 23 prompts were cheaper joint failures. A system should not be judged as better simply because it fails at lower cost; resource efficiency has to be interpreted alongside the quality and usefulness of the outcome.

The paper also highlights a reproducibility and accountability issue. If scores are not linked to attempt-level execution records and scoring provenance, researchers and users may be unable to tell whether a result reflects full success, partial progress, a shared failure or a measurement artifact. The source presents this as a reason to make evaluation records more detailed, not as evidence that either system is generally superior.

Interactive Mechanism

Cơ chế tương tác: Nó thực sự hoạt động như thế nào

Khám phá công nghệ cơ bản đằng sau sự phát triển này một cách tương tác.

Agent Lifecycle Stage:
1
User Intent & Planning: "Audit customer refund request #4092 and settle payment."
2
Tool Calling: Emits structured JSON call crm_get_transaction(id='4092').
3
Guardrail & Verification:🛡️ Paused: High-value action requires human operator sign-off.
4
Final Settlement: Refund recorded, email receipt dispatched, and audit log stored.
Core takeaway: An AI agent is not just a language model—it is a closed loop of planning, tool invocation, and environment feedback. Production systems require self-healing retries and strict human approval guardrails.
Kiểm tra khái niệm tương tác+10 Points
AI Agents Quiz

An agent must create a draft calendar event for Tuesday at 2 p.m. Which evidence would establish the requested result?

Xem gì tiếp theo

The main follow-up questions are whether the findings hold across larger and more varied task sets, different hardware and software configurations, and other agent frameworks. The source does not provide those details in its abstract, so the reported resource ratios should be treated as study-specific rather than universal rankings.

The immediate question is whether the reported gap in wall time and peak memory persists beyond the study's prompts and test configuration. The abstract does not state the exact task composition, hardware, software versions, number of repeated trials or measurement procedure. Those omissions limit how broadly the numerical ratios can be applied.

Readers should also watch for evaluations that report more than a single aggregate score. Useful follow-up studies would separate full completion, partial completion and joint failure; show how often each system wins on task quality; and publish the execution records needed to connect each outcome to resource consumption. The preprint's own disagreement between its primary and instrumented evidence layers makes this a practical research priority.

The authors' conclusion is methodological rather than a product recommendation. The source does not establish that NanoBot is the better general-purpose agent, nor that OpenClaw's higher resource use is unjustified for every workload. Further comparisons with other agent systems, larger samples and independent replications would be needed before treating the findings as a broad market or engineering ranking.

A meaningful unknown is how the systems' resource profiles interact with task difficulty and . The abstract reports aggregate ratios and prompt-level comparisons but does not identify which kinds of tasks caused the largest differences. That information would help determine whether the results reflect a general property of the systems or a pattern specific to the evaluated workload. Until then, the strongest supported takeaway is that agent evaluation should report verified outcomes together with observed resource use and scoring provenance.

Hướng dẫn và câu hỏi liên quan

Đại lý AIGiải thích về mô hình AIĐào tạo AIKiểm tra những gì bạn biết — thử một bài kiểm tra AI miễn phíTra cứu một thuật ngữ AI trong bảng thuật ngữ của chúng tôiTheo dõi trình theo dõi phát hành mô hình AI
Tìm thấy điều này hữu ích?