Hi, it's John. Two months in. Here is what happened.
What I shipped#
The assistant on mutexai.net is live. It answers questions about the studio, suggests the right page with a button, and when it does not know something it says so and flags the conversation for me. I wrote it in plain Laravel with the HTTP client, no SDK.
The three constraints I set before writing any code:
- It must always answer. If the API key is missing, the provider is down or the budget is spent, a local responder answers from the site's own config and a reviewed knowledge base.
- It must have a ceiling. A daily cap on model-backed replies, a limit of four model calls per message, and a thirty-second deadline.
- It must not learn on its own. Unanswered questions go into a review queue. I write the answer, approve it, and only then does the assistant use it.
That third one gets the most questions when I describe it. Why not let it learn from conversations automatically? Because it speaks for a business. Every fact it can state should be one a person has approved. The full write-up is coming later this month.
Andor got structured execution. The VS Code agent now works through explicit stages (analysis, plan, execution, verification) and ends every task with a completion state: COMPLETE, INCOMPLETE or BLOCKED, with the blockers listed.
What I learned#
The completion state change made a bigger difference than any model upgrade I tried.
Before it, the agent would finish a turn with a cheerful summary whether or not the work was actually done. A test still failing? "I've updated the tests." A command that errored? "The build should now work." Nothing in the output distinguished "done" from "tried".
Forcing an explicit state, and asking for verification evidence before COMPLETE, changed the behaviour. When the model has to choose between COMPLETE and BLOCKED and list what blocked it, it is much more likely to be honest. It also made the UI better: a blocked task shows a Continue button and the blocker, instead of a summary I have to read carefully to find the problem.
The general lesson: give agents a vocabulary for failure. If the only way to end a task is to sound successful, they will sound successful.
A smaller lesson, from the assistant: put the parts of the prompt that change per request (the visitor's current page, the conversation) in the messages, never in the system prompt. The system prompt is cached, and one changing value at the top throws the cache away on every call.
What I'm reading and trying#
- Prompt caching in detail. The cost difference between a cached and uncached system prompt is large enough that it changes how I structure every agent.
- Chaos testing ideas for agents. I keep seeing agents fail in the same ways when tools break, and I want a repeatable way to test that instead of discovering it by accident.
- M-Pesa Daraja, again. I want agents to be able to check payment status and reconcile statements. Doing that safely is mostly a design problem, and I am sketching the tools now.
A question for you#
When an AI tool tells you it finished something, how do you check? Do you read the diff, run the tests, or just trust it? I am collecting habits for a post on verification, and I would like to know what actually works for you. Hit reply.
John