When an agent does something dumb, the first instinct is to blame the model or rewrite the prompt. After building Andor (a coding agent for VS Code), an on-site assistant in Laravel and a set of MCP servers, I have learned to look somewhere else first: the tools.
A tool is an API whose only consumer is a language model. That consumer reads every word of your description, takes your parameter names literally, and cannot open your source code to see what you meant. Most of the failures I debug come down to one of the problems below.
Name tools like verbs a stranger would guess#
The model picks tools by name and description. Names should say what happens, in the vocabulary of the task:
search_knowledgebeatskb_querytransaction_statusbeatsget_txshow_pagebeatsnav
Avoid near-duplicates. If you have read_file, get_file and open_file, the model will use all three at random. One tool per intent, and if two tools really are different, the names should make the difference obvious.
Group related tools with a shared prefix only when the prefix carries meaning. In an MCP server that exposes payments and reconciliation, b2c_payout and match_ledger are clear on their own. Prefixing everything with mpesa_ adds tokens and no information.
Descriptions are prompts, write them like prompts#
The description is the single most important field. It is where you tell the model when to use the tool, when not to, and what it will get back. Compare:
{ "name": "search_knowledge", "description": "Searches the knowledge base." }
with:
{ "name": "search_knowledge", "description": "Search answers the site owner has reviewed and approved. Use this before saying you do not know something about services, pricing approach or availability. Returns up to 5 entries with a question, an answer and an optional link. Returns an empty list when nothing matches; do not invent an answer in that case." }
The second one does four jobs: says what the data is, says when to call it, describes the output shape, and tells the model what to do on an empty result. That last part matters more than it looks. Empty results are where models start making things up.
Keep schemas small, typed and hard to misuse#
Every optional parameter is a decision the model has to make. Every free-text field is a place for it to improvise. My rules:
- Use enums wherever the set of values is known.
- Prefer one required parameter over three optional ones.
- Use units in names:
amount_kes,timeout_seconds,since_iso. - Validate on the server and return a useful error, never trust the model's input.
Here is a tool from an MCP server, registered with the TypeScript SDK and Zod:
server.registerTool( "transaction_status", { title: "Check M-Pesa transaction status", description: "Look up the final status of one M-Pesa transaction by its receipt number (for example QGR7XK2L9P). Read-only. Use this before telling a user a payment failed.", inputSchema: { receipt: z.string().regex(/^[A-Z0-9]{10}$/, "10 uppercase letters and digits"), }, annotations: { readOnlyHint: true, idempotentHint: true }, }, async ({ receipt }) => { const result = await daraja.transactionStatus(receipt); return { content: [{ type: "text", text: JSON.stringify(result) }] }; }, );
One parameter, a format the model can check against, and annotations that tell the client this tool is safe to call freely.
Errors should teach the next step#
The model will read your error and try again. Give it something to work with. Bad:
{ "is_error": true, "content": "Invalid input" }
Better:
{ "is_error": true, "content": "Unknown page 'pricing'. Valid pages: home, services, products, open-source, about, contact. Pick the closest one or skip showing a page." }
The second error names the bad value, lists the valid ones and suggests a recovery. In my experience this is the single cheapest improvement you can make to an agent. It turns a retry loop into a single corrected call.
In PHP, the on-site assistant I built returns tool results in one consistent shape, which keeps the loop simple:
/** * @return array{content: string, is_error: bool} */ public function execute(string $name, array $input): array { return match ($name) { 'search_knowledge' => $this->searchKnowledge((string) ($input['query'] ?? '')), 'show_page' => $this->showPage((string) ($input['page'] ?? '')), default => ['content' => "Unknown tool '{$name}'.", 'is_error' => true], }; }
Unknown tools, bad input and upstream failures all come back as data the model can read, never as an exception that kills the conversation.
Separate reading from doing#
The most important design decision for any agent that touches real systems: which tools change the world, and which only look at it.
- Read-only tools can be called freely, retried safely and cached.
- Mutating tools need idempotency keys, confirmation, and usually a human in the loop.
In Andor, destructive terminal commands like rm or git push go through an approval step, while reads and safe commands can be allowlisted. In the payments MCP server, money-moving tools default to sandbox and need explicit opt-in for production. The assistant on my studio site can only suggest pages and flag questions for review; it cannot change anything on its own.
If you only take one thing from this post: make the default path read-only, and make the dangerous path loud.
Return less, and return it structured#
Tools that dump a whole API response into the context waste tokens and bury the useful field. Return what the model needs for its next decision:
{ "receipt": "QGR7XK2L9P", "status": "completed", "amount_kes": 1500, "completed_at": "2026-04-18T09:12:44+03:00" }
Four fields, clear names, units in the key. If the model needs more, give it a second tool to drill in.
Make tools boring to retry#
Models retry. Networks fail. Agent frameworks resend tool calls after timeouts. Any tool that creates something should accept an idempotency key or derive one from its input, so that a second call with the same input returns the first result instead of doing the work twice. For payments this is non-negotiable. For everything else it is still cheaper than cleaning up duplicates.
Test tools with the model, not only with unit tests#
Unit tests prove the handler works. They do not prove the model will call it correctly. I keep a small set of prompts per tool (a normal case, an ambiguous case, a case where the tool should not be called) and run them after every description change. When a model picks the wrong tool, the fix is almost always in the description or the name, not the prompt.
This is also why I started building agent-test-harness: to inject failing tools on purpose and see whether the agent recovers or bluffs. Good tool design makes recovery possible. Testing makes sure it happens.
A checklist#
Before I ship a tool now, I check:
- Would a stranger guess what this does from the name alone?
- Does the description say when to use it, when not to, and what empty results mean?
- Is every parameter required, typed and constrained where it can be?
- Does every error name the problem and suggest the next step?
- Is it read-only? If not, is it idempotent and gated?
- Does it return only what the next decision needs?
None of this is clever. It is API design for a reader who is very literal and very fast. Treat the model like that reader and most of your agent bugs go away.