Zpět na Novinky
InovaceInstruktáž AI Understanding

Preprint zjistí, že výkon agentní umělé inteligence závisí jak na úspěchu úkolu, tak na využití zdrojů

Nový předtisk arXiv porovnávající OpenClaw a NanoBot nenachází žádného statisticky ověřeného vítěze při úplném dokončení úkolu, ale uvádí velké rozdíly v čase a špičkové paměti. Autoři tvrdí, že hodnocení agentů by mělo spojovat výsledky se zdroji a záznamy o provádění, které je vytvořily.

5 min readRead the primary source
Source-provided image accompanying Preprint finds agentic AI performance depends on both task success and resource use
Primární zdrojový dokumentZdroj zaznamenán
Vydavatel
arxiv.org
Odkaz na zdroj
arxiv.orghttps://arxiv.org/abs/2608.27886
Typ zdroje
Primární dokument — oficiální oznámení, papír, podání nebo stránka první strany, kterou čteme přímo.
KontextPochopte to za 60 sekund

Začněte zde

Klíčové pojmy

Paměť (paměť agenta)
Uložený kontext, který agent AI používá v krocích nebo relacích ke zlepšení kontinuity.
Benchmark
Standardizovaný test nebo soubor dat používaný k měření a porovnávání výkonu modelu.
Použití nástroje
Schopnost modelu volat externí nástroje, jako je vyhledávání, kalkulačky nebo rozhraní API.
Otestujte seKvíz AI agentů

Co se stalo

An arXiv preprint submitted August 28 compares OpenClaw and NanoBot as complete agentic AI systems, including their language model, tools, memory, state management and multi-step execution. On a primary , OpenClaw completed 31% of tasks and NanoBot 25%, but the reported 95% task-bootstrap interval ran from -3 to 15 percentage points, so the study did not establish a full-completion advantage for either system.

The preprint evaluates OpenClaw and NanoBot as complete agentic systems rather than comparing language models in isolation. The authors describe agentic systems as combinations of a language model with tools, memory, state management and multi-step execution. That framing matters because each layer can affect both what an agent accomplishes and the operational resources required to attempt a task.

In the primary , OpenClaw achieved full task completion on 31% of tasks, compared with 25% for NanoBot. The six-percentage-point difference was accompanied by a 95% task-bootstrap interval ranging from minus 3 to 15 percentage points. Based on that interval, the authors say there was no statistically established full-completion advantage for either system.

The paper also reports a more detailed instrumented subset of paired prompts. In that layer, both systems reached full completion on 26% of prompts. NanoBot nevertheless reached at least partial completion on 43% of prompts, compared with 26% for OpenClaw, indicating that a full-completion-only score concealed a difference in intermediate outcomes in this subset.

Resource measurements favored NanoBot in the reported comparisons. OpenClaw took longer on 83% of prompts and recorded a higher peak-memory value on every prompt. The paper gives geometric mean ratios of 2.98 for wall time and 19.44 for peak memory. Among the ten detailed-layer prompts where at least one system achieved partial or full completion, NanoBot weakly dominated on eight. Across all 23 prompts, however, ten of its 18 dominance cases were cheaper joint failures, meaning lower resource use did not always accompany useful task progress.

Podrobnosti o zdroji: arxiv.org ↗

Proč na tom záleží

The paper shows why an agent that completes slightly more tasks may not be the more practical system if it consumes substantially more time or memory. Its results also suggest that conclusions can change when researchers inspect execution details and distinguish full success, partial progress and joint failure.

The central contribution is an evaluation principle: capability and cost should be measured together and tied to the specific execution that produced each result. For people choosing or deploying agentic systems, a completion percentage alone can obscure whether a system is fast enough, memory-efficient enough or consistently able to make useful partial progress.

The reported results make that trade-off concrete. OpenClaw's 31% primary- completion rate was only modestly higher than NanoBot's 25%, and the uncertainty interval did not establish a reliable winner. Yet the instrumented results show much larger operational differences, with OpenClaw taking longer on most prompts and using more peak memory on every prompt in that comparison.

The distinction between partial completion and joint failure is especially important. NanoBot's lower resource use helped it dominate several comparisons, but the authors say ten of its 18 dominance cases across all 23 prompts were cheaper joint failures. A system should not be judged as better simply because it fails at lower cost; resource efficiency has to be interpreted alongside the quality and usefulness of the outcome.

The paper also highlights a reproducibility and accountability issue. If scores are not linked to attempt-level execution records and scoring provenance, researchers and users may be unable to tell whether a result reflects full success, partial progress, a shared failure or a measurement artifact. The source presents this as a reason to make evaluation records more detailed, not as evidence that either system is generally superior.

Interactive Mechanism

Interaktivní mechanismus: Jak to vlastně funguje

Interaktivně prozkoumejte základní technologii tohoto vývoje.

Agent Lifecycle Stage:
1
User Intent & Planning: "Audit customer refund request #4092 and settle payment."
2
Tool Calling: Emits structured JSON call crm_get_transaction(id='4092').
3
Guardrail & Verification:🛡️ Paused: High-value action requires human operator sign-off.
4
Final Settlement: Refund recorded, email receipt dispatched, and audit log stored.
Core takeaway: An AI agent is not just a language model—it is a closed loop of planning, tool invocation, and environment feedback. Production systems require self-healing retries and strict human approval guardrails.
Interaktivní kontrola konceptu+10 Points
AI Agents Quiz

An agent must create a draft calendar event for Tuesday at 2 p.m. Which evidence would establish the requested result?

Na co se dále dívat

The main follow-up questions are whether the findings hold across larger and more varied task sets, different hardware and software configurations, and other agent frameworks. The source does not provide those details in its abstract, so the reported resource ratios should be treated as study-specific rather than universal rankings.

The immediate question is whether the reported gap in wall time and peak memory persists beyond the study's prompts and test configuration. The abstract does not state the exact task composition, hardware, software versions, number of repeated trials or measurement procedure. Those omissions limit how broadly the numerical ratios can be applied.

Readers should also watch for evaluations that report more than a single aggregate score. Useful follow-up studies would separate full completion, partial completion and joint failure; show how often each system wins on task quality; and publish the execution records needed to connect each outcome to resource consumption. The preprint's own disagreement between its primary and instrumented evidence layers makes this a practical research priority.

The authors' conclusion is methodological rather than a product recommendation. The source does not establish that NanoBot is the better general-purpose agent, nor that OpenClaw's higher resource use is unjustified for every workload. Further comparisons with other agent systems, larger samples and independent replications would be needed before treating the findings as a broad market or engineering ranking.

A meaningful unknown is how the systems' resource profiles interact with task difficulty and . The abstract reports aggregate ratios and prompt-level comparisons but does not identify which kinds of tasks caused the largest differences. That information would help determine whether the results reflect a general property of the systems or a pattern specific to the evaluated workload. Until then, the strongest supported takeaway is that agent evaluation should report verified outcomes together with observed resource use and scoring provenance.

Související průvodci a kvízy

Agenti AIVysvětlení modelů AIŠkolení AIOtestujte si, co víte – vyzkoušejte bezplatný kvíz AIVyhledejte si termín AI v našem slovníkuPostupujte podle sledování vydání modelu AI
Považujete to za užitečné?