A seguirPróximo guia
Evaluating AI Agents
Técnico
GUIA Técnico
Designing tools for LLM agents means writing the names, descriptions, parameter schemas, outputs and error messages that a language model reads to decide which function to call and how to call it.
The model never sees your code, only this text, so a vague description or a sloppy schema produces wrong calls just as reliably as a bug does.
A tool definition has four parts the model depends on: a name, a natural-language description, an input schema (usually JSON Schema), and the text the tool sends back. Each works like a prompt. The model decides whether to call a tool by matching the user's goal to descriptions, and it fills in arguments using parameter names, types, enums and examples. Good names are specific and grouped. Prefixes such as github_create_issue and jira_create_issue keep tools apart when an agent connects to many services. Good descriptions say what the tool does, when to use it, when not to, and what it returns, the way you would brief a new colleague. Parameters should say what they mean: user_id is better than id, and an enum of allowed values beats a free-text string the model has to guess at. Outputs matter as much as inputs. Returning thousands of rows or raw internal IDs fills the context window and hides the information that matters. Better tools return readable fields, paginate or truncate by default, and sometimes offer a concise or detailed mode. Error messages should tell the model what went wrong and what to try next, because the model will read them and act on them. One common misconception is that more tools means a more capable agent. Every definition takes up context on every turn, and overlapping tools make selection harder. Accuracy tends to drop as the tool list grows long, especially when tools look alike. No universal cutoff applies. Many practitioners start to see trouble somewhere in the tens of tools, and the usual fixes are consolidating tools, loading them on demand, or splitting work across specialized agents. Another misconception is that tools should mirror an existing API one-to-one. Agents usually do better with fewer, higher-level tools shaped around the tasks people actually do.
As decisões de arquitetura impulsionam o desempenho e os custos operacionais durante anos.
A educação técnica ajuda as equipes a escolher a pilha certa, não apenas a mais nova.
Melhores escolhas de engenharia reduzem incidentes de confiabilidade na produção.
Tool design is becoming its own discipline, with published guidance from model providers and shared conventions spreading through protocols such as the Model Context Protocol. Some platforms can now search for and load tool definitions on demand, which eases the pressure to keep lists short, although descriptions still have to be clear enough to find and choose. Models are also being used to critique and rewrite tool descriptions based on evaluation transcripts. The basic rule is unlikely to change: the model can only use a tool as well as its interface explains it, so clear naming, strict schemas and helpful errors will stay central.
A support agent had two tools, search_orders and get_order. Renaming them orders_search_by_customer and orders_get_by_id and saying in each description when to use the other one cut the number of calls sent to the wrong tool.
A calendar tool that took a free-text date field kept getting 'next Tuesday'. Changing the schema to require an ISO 8601 date string with an example in the description removed a whole class of parsing failures.
A database tool that used to return a raw 'Error 1064' now returns 'Column created_date does not exist. Available date columns: created_at, updated_at.' The agent fixes its query on the next try instead of giving up.
A team exposed 60 near-duplicate endpoint wrappers to one agent. Merging them into a dozen task-level tools, such as schedule_meeting in place of separate list_users, find_free_slots and create_event calls, made the agent faster and cheaper to run.
A otimização de um benchmark pode ocultar fraquezas mais amplas do sistema.
Os custos de infraestrutura e manutenção são frequentemente subestimados.
As lacunas de segurança e observabilidade podem aumentar à medida que os sistemas se tornam mais complexos.
Defina metas de latência, qualidade e custo antes da implementação.
Benchmark sob condições realistas de carga e dados.
Monitoramento de instrumentos para erros, desvios e impacto no usuário.
Prepare caminhos de reversão e resposta a incidentes antes de escalar.
Free newsletter
Three verified AI stories every weekday morning, written in plain English. Free forever, no ads.
One email each weekday. Unsubscribe in one click. We never sell or share your address.
Test yourself
Instant feedback on every answer, and a shareable certificate with a verifiable ID once you pass a course.
Support free AI education. AI Understanding is a 501(c)(3) nonprofit — no ads, no paywall, ever. Make a donation
Designing tools for LLM agents means writing the names, descriptions, parameter schemas, outputs and error messages that a language model reads to decide which function to call and how to call it. The model never sees your code, only this text, so a vague description or a sloppy schema produces wrong calls just as reliably as a bug does.
The model never sees the implementation. It works only from the text interface: the name, the description, the input schema and whatever the tool sends back.
Namespacing tells the model which service a tool belongs to. That matters when several services offer similar actions.
Enums and descriptive names remove guesswork. The model picks from valid options instead of inventing a format.
The model reads the error and decides what to do next. A message that names the problem and a fix lets it recover on the following call.
Every definition is sent with each request, which uses context. Similar-looking tools also make it harder for the model to pick the right one.
Continue aprendendo
Mais guias escolhidos para este tópico
A seguirPróximo guia
Evaluating AI Agents
Técnico