技術指南

Designing Tools for LLM Agents

Designing tools for LLM agents means writing the names, descriptions, parameter schemas, outputs and error messages that a language model reads to decide which function to call and how to call it.

  • 4 分鐘閱讀
  • 最後更新
本頁4 分鐘閱讀
  1. 概述
  2. 深入探討
  3. 戰略影響
  4. The Future of Designing Tools for LLM Agents
  5. 現實世界的實施
  6. 風險與防護欄
  7. 實施路線圖
  8. 不斷探索
  9. 常見問題

概述

The model never sees your code, only this text, so a vague description or a sloppy schema produces wrong calls just as reliably as a bug does.

深入探討

A tool definition has four parts the model depends on: a name, a natural-language description, an input schema (usually JSON Schema), and the text the tool sends back. Each works like a prompt. The model decides whether to call a tool by matching the user's goal to descriptions, and it fills in arguments using parameter names, types, enums and examples. Good names are specific and grouped. Prefixes such as github_create_issue and jira_create_issue keep tools apart when an agent connects to many services. Good descriptions say what the tool does, when to use it, when not to, and what it returns, the way you would brief a new colleague. Parameters should say what they mean: user_id is better than id, and an enum of allowed values beats a free-text string the model has to guess at. Outputs matter as much as inputs. Returning thousands of rows or raw internal IDs fills the context window and hides the information that matters. Better tools return readable fields, paginate or truncate by default, and sometimes offer a concise or detailed mode. Error messages should tell the model what went wrong and what to try next, because the model will read them and act on them. One common misconception is that more tools means a more capable agent. Every definition takes up context on every turn, and overlapping tools make selection harder. Accuracy tends to drop as the tool list grows long, especially when tools look alike. No universal cutoff applies. Many practitioners start to see trouble somewhere in the tens of tools, and the usual fixes are consolidating tools, loading them on demand, or splitting work across specialized agents. Another misconception is that tools should mirror an existing API one-to-one. Agents usually do better with fewer, higher-level tools shaped around the tasks people actually do.

戰略影響

成本與預算

多年來,架構決策決定著效能和營運成本。

更明確的決策

技術教育幫助團隊選擇正確的堆疊,而不僅僅是最新的堆疊。

品質管控

更好的工程選擇可以減少生產中的可靠性事故。

The Future of Designing Tools for LLM Agents

Tool design is becoming its own discipline, with published guidance from model providers and shared conventions spreading through protocols such as the Model Context Protocol. Some platforms can now search for and load tool definitions on demand, which eases the pressure to keep lists short, although descriptions still have to be clear enough to find and choose. Models are also being used to critique and rewrite tool descriptions based on evaluation transcripts. The basic rule is unlikely to change: the model can only use a tool as well as its interface explains it, so clear naming, strict schemas and helpful errors will stay central.

現實世界的實施

A support agent had two tools, search_orders and get_order. Renaming them orders_search_by_customer and orders_get_by_id and saying in each description when to use the other one cut the number of calls sent to the wrong tool.

A calendar tool that took a free-text date field kept getting 'next Tuesday'. Changing the schema to require an ISO 8601 date string with an example in the description removed a whole class of parsing failures.

A database tool that used to return a raw 'Error 1064' now returns 'Column created_date does not exist. Available date columns: created_at, updated_at.' The agent fixes its query on the next try instead of giving up.

A team exposed 60 near-duplicate endpoint wrappers to one agent. Merging them into a dozen task-level tools, such as schedule_meeting in place of separate list_users, find_free_slots and create_event calls, made the agent faster and cheaper to run.

風險與防護欄

  • 優化一項基準測試可以隱藏更廣泛的系統弱點。

  • 基礎設施和維護成本常常被低估。

  • 隨著系統變得更加複雜,安全性和可觀察性差距可能會擴大。

實施路線圖

  1. 在實施之前定義延遲、品質和成本目標。

  2. 在實際負載和資料條件下進行基準測試。

  3. 儀器監控錯誤、漂移和使用者影響。

  4. 在擴展之前準備回滾和事件回應路徑。

不斷探索

Free newsletter

Get the daily AI briefing

Three verified AI stories every weekday morning, written in plain English. Free forever, no ads.

One email each weekday. Unsubscribe in one click. We never sell or share your address.

Test yourself

Take the Designing Tools for LLM Agents quiz

Instant feedback on every answer, and a shareable certificate with a verifiable ID once you pass a course.

開始測驗

Support free AI education. AI Understanding is a 501(c)(3) nonprofit — no ads, no paywall, ever. Make a donation

常見問題

What is Designing Tools for LLM Agents?

Designing tools for LLM agents means writing the names, descriptions, parameter schemas, outputs and error messages that a language model reads to decide which function to call and how to call it. The model never sees your code, only this text, so a vague description or a sloppy schema produces wrong calls just as reliably as a bug does.

What does the model actually use to decide how to call a tool?

The model never sees the implementation. It works only from the text interface: the name, the description, the input schema and whatever the tool sends back.

Why might an agent use prefixes such as github_create_issue and jira_create_issue?

Namespacing tells the model which service a tool belongs to. That matters when several services offer similar actions.

Which parameter design is most likely to produce correct arguments?

Enums and descriptive names remove guesswork. The model picks from valid options instead of inventing a format.

What makes a good tool error message?

The model reads the error and decides what to do next. A message that names the problem and a fix lets it recover on the following call.

Why can adding many tools reduce agent performance?

Every definition is sent with each request, which uses context. Similar-looking tools also make it harder for the model to pick the right one.