Műszaki ÚTMUTATÓ

Agent Scaffolding and Harnesses

An agent harness, or scaffold, is the code around a language model that turns it into an agent: it runs the loop of calling the model, executing the tools it requests, feeding results back, and deciding when to stop.

  • 3 perc olvasás
  • Utoljára frissítve
Ezen az oldalon3 perc olvasás
  1. Áttekintés
  2. Mély merülés
  3. Stratégiai hatás
  4. The Future of Agent Scaffolding and Harnesses
  5. Valós megvalósítás
  6. Kockázatok és védőkorlátok
  7. Végrehajtási ütemterv
  8. Folytassa a felfedezést
  9. Gyakran ismételt kérdések

Áttekintés

It matters because the harness strongly shapes results, so the same model can score very differently on the same benchmark depending on how its tools, context and error handling are designed.

Mély merülés

A language model on its own produces text. An agent needs a program that repeatedly asks the model what to do, performs the action, and shows the model what happened. That program is the harness. Its core parts are the loop, tool definitions, tool dispatch, context management, stopping rules and error handling. The loop sends the conversation so far to the model. If the model requests a tool, the harness validates the arguments, runs the tool, and appends the result. If the model gives a final answer, the loop ends. Stopping rules add limits: a maximum number of steps, a token or cost budget, a timeout, or a check such as tests passing. Tool design has large effects. The SWE-agent project from Princeton researchers, published in 2024, argued that agents need an agent-computer interface designed for models, with concise file viewers, search commands and immediate feedback on bad edits, and showed this improved results over giving the model a plain shell. Output formatting matters too: truncating huge outputs, summarising errors clearly and returning structured results help the model act correctly. Context management decides what the model sees as the history grows: keeping recent steps, summarising older ones, or storing notes in files. Error handling decides whether a failed tool call crashes the run or is returned to the model as information it can react to. This is why benchmark numbers need context. On coding benchmarks such as SWE-bench, reported scores for one model vary across scaffolds, retry policies and step limits. A common misconception is that a leaderboard number measures the model alone; it measures the model plus harness plus settings. Products such as command-line coding agents are, in large part, carefully engineered harnesses around a model.

Stratégiai hatás

Költség és költségvetés

Az építészeti döntések évekig növelik a teljesítményt és a működési költségeket.

Tisztább döntések

A technikai oktatás segít a csapatoknak a megfelelő verem kiválasztásában, nem csak a legújabb készletben.

Minőségellenőrzés

A jobb mérnöki döntések csökkentik a termelés megbízhatósági incidenseit.

The Future of Agent Scaffolding and Harnesses

Model providers increasingly ship their own agent products and software development kits, which standardises some harness patterns such as tool schemas, permissions and context compaction. Protocols for connecting tools to models, such as the Model Context Protocol, make tool integration more reusable. Harness design is likely to remain a meaningful source of performance differences, and evaluations are moving toward reporting the scaffold alongside the model. For builders, careful tool design, clear stopping rules and good logging will keep mattering regardless of which model is underneath.

Valós megvalósítás

A coding agent harness gives the model tools to view files, edit specific line ranges and run tests, then loops until tests pass or a step limit is reached.

A research assistant harness stops the loop after 20 tool calls and asks the model for a summary with sources, preventing runaway costs.

A team replaces a raw shell tool with a file editor that shows a short window of lines and reports syntax errors immediately, and their agent's success rate on internal tasks improves without changing the model.

A customer operations agent harness intercepts any refund tool call above a set amount and pauses for human approval before executing it.

Kockázatok és védőkorlátok

  • Egy benchmark optimalizálása elrejtheti a rendszer általános hiányosságait.

  • Az infrastrukturális és karbantartási költségeket gyakran alábecsülik.

  • A biztonsági és megfigyelhetőségi hiányosságok a rendszerek bonyolultabbá válásával nőhetnek.

Végrehajtási ütemterv

  1. Határozza meg a késleltetési, minőségi és költségcélokat a megvalósítás előtt.

  2. Benchmark reális terhelési és adatviszonyok mellett.

  3. Műszerfigyelés a hibák, az eltolódás és a felhasználói hatások szempontjából.

  4. A méretezés előtt készítse elő a visszagörgetési és az incidensre adott válaszútvonalakat.

Folytassa a felfedezést

Free newsletter

Get the daily AI briefing

Three verified AI stories every weekday morning, written in plain English. Free forever, no ads.

One email each weekday. Unsubscribe in one click. We never sell or share your address.

Test yourself

Take the Agent Scaffolding and Harnesses quiz

Instant feedback on every answer, and a shareable certificate with a verifiable ID once you pass a course.

Kezdő kvíz

Support free AI education. AI Understanding is a 501(c)(3) nonprofit — no ads, no paywall, ever. Make a donation

Gyakran ismételt kérdések

What is Agent Scaffolding and Harnesses?

An agent harness, or scaffold, is the code around a language model that turns it into an agent: it runs the loop of calling the model, executing the tools it requests, feeding results back, and deciding when to stop. It matters because the harness strongly shapes results, so the same model can score very differently on the same benchmark depending on how its tools, context and error handling are designed.

What is an agent harness?

The harness is the surrounding program that turns a text-generating model into an agent.

Which is an example of a stopping rule?

Stopping rules end the loop based on limits like step counts, budgets, timeouts or success checks.

What did the SWE-agent project emphasise?

SWE-agent argued that tools designed for models improve results over a plain shell.

Why can the same model score differently on SWE-bench in different reports?

A benchmark score measures the whole system, not the model alone.

How should a harness usually handle a failed tool call?

Returning actionable errors lets the model correct its approach instead of ending the run.