Als nächstesNächster Leitfaden
Bewertung von KI-Agenten
Technisch
Technischer Leitfaden
Designing tools for LLM agents means writing the names, descriptions, parameter schemas, outputs and error messages that a language model reads to decide which function to call and how to call it.
The model never sees your code, only this text, so a vague description or a sloppy schema produces wrong calls just as reliably as a bug does.
A tool definition has four parts the model depends on: a name, a natural-language description, an input schema (usually JSON Schema), and the text the tool sends back. Each works like a prompt. The model decides whether to call a tool by matching the user's goal to descriptions, and it fills in arguments using parameter names, types, enums and examples. Good names are specific and grouped. Prefixes such as github_create_issue and jira_create_issue keep tools apart when an agent connects to many services. Good descriptions say what the tool does, when to use it, when not to, and what it returns, the way you would brief a new colleague. Parameters should say what they mean: user_id is better than id, and an enum of allowed values beats a free-text string the model has to guess at. Outputs matter as much as inputs. Returning thousands of rows or raw internal IDs fills the context window and hides the information that matters. Better tools return readable fields, paginate or truncate by default, and sometimes offer a concise or detailed mode. Error messages should tell the model what went wrong and what to try next, because the model will read them and act on them. One common misconception is that more tools means a more capable agent. Every definition takes up context on every turn, and overlapping tools make selection harder. Accuracy tends to drop as the tool list grows long, especially when tools look alike. No universal cutoff applies. Many practitioners start to see trouble somewhere in the tens of tools, and the usual fixes are consolidating tools, loading them on demand, or splitting work across specialized agents. Another misconception is that tools should mirror an existing API one-to-one. Agents usually do better with fewer, higher-level tools shaped around the tasks people actually do.
Architekturentscheidungen beeinflussen über Jahre hinweg die Leistung und die Betriebskosten.
Technische Schulungen helfen Teams dabei, den richtigen Stack auszuwählen, nicht nur den neuesten.
Bessere technische Entscheidungen reduzieren Zuverlässigkeitsvorfälle in der Produktion.
Tool design is becoming its own discipline, with published guidance from model providers and shared conventions spreading through protocols such as the Model Context Protocol. Some platforms can now search for and load tool definitions on demand, which eases the pressure to keep lists short, although descriptions still have to be clear enough to find and choose. Models are also being used to critique and rewrite tool descriptions based on evaluation transcripts. The basic rule is unlikely to change: the model can only use a tool as well as its interface explains it, so clear naming, strict schemas and helpful errors will stay central.
A support agent had two tools, search_orders and get_order. Renaming them orders_search_by_customer and orders_get_by_id and saying in each description when to use the other one cut the number of calls sent to the wrong tool.
A calendar tool that took a free-text date field kept getting 'next Tuesday'. Changing the schema to require an ISO 8601 date string with an example in the description removed a whole class of parsing failures.
A database tool that used to return a raw 'Error 1064' now returns 'Column created_date does not exist. Available date columns: created_at, updated_at.' The agent fixes its query on the next try instead of giving up.
A team exposed 60 near-duplicate endpoint wrappers to one agent. Merging them into a dozen task-level tools, such as schedule_meeting in place of separate list_users, find_free_slots and create_event calls, made the agent faster and cheaper to run.
Die Optimierung eines Benchmarks kann umfassendere Systemschwächen verbergen.
Infrastruktur- und Wartungskosten werden oft unterschätzt.
Sicherheits- und Beobachtbarkeitslücken können größer werden, wenn die Systeme komplexer werden.
Definieren Sie vor der Implementierung Latenz-, Qualitäts- und Kostenziele.
Benchmark unter realistischen Last- und Datenbedingungen.
Instrumentenüberwachung auf Fehler, Drift und Benutzereinflüsse.
Bereiten Sie vor der Skalierung Rollback- und Incident-Response-Pfade vor.
Free newsletter
Three verified AI stories every weekday morning, written in plain English. Free forever, no ads.
One email each weekday. Unsubscribe in one click. We never sell or share your address.
Test yourself
Instant feedback on every answer, and a shareable certificate with a verifiable ID once you pass a course.
Support free AI education. AI Understanding is a 501(c)(3) nonprofit — no ads, no paywall, ever. Make a donation
Designing tools for LLM agents means writing the names, descriptions, parameter schemas, outputs and error messages that a language model reads to decide which function to call and how to call it. The model never sees your code, only this text, so a vague description or a sloppy schema produces wrong calls just as reliably as a bug does.
The model never sees the implementation. It works only from the text interface: the name, the description, the input schema and whatever the tool sends back.
Namespacing tells the model which service a tool belongs to. That matters when several services offer similar actions.
Enums and descriptive names remove guesswork. The model picks from valid options instead of inventing a format.
The model reads the error and decides what to do next. A message that names the problem and a fix lets it recover on the following call.
Every definition is sent with each request, which uses context. Similar-looking tools also make it harder for the model to pick the right one.
Lerne weiter
Weitere Leitfäden zu diesem Thema ausgewählt
Als nächstesNächster Leitfaden
Bewertung von KI-Agenten
Technisch