Why the assistant's tools default to read-only
A short note on a naming habit from the mutexai.net assistant. Read-only tools get plain verbs. Anything that writes gets a verb that says so, on purpose, so a reviewer never has to guess.
Long-form articles and short notes on AI agents, MCP, testing and full-stack engineering. Mostly things I learned the hard way.
Featured 6 min read
I run several coding agents that share one durable memory made of markdown, JSON and git commits. Here is why plain files beat a vector store, and how sessions, leases and decisions work.
Read the articleA short note on a naming habit from the mutexai.net assistant. Read-only tools get plain verbs. Anything that writes gets a verb that says so, on purpose, so a reviewer never has to guess.
Three agents on one checkout means trampled files and lost work. Giving each task its own git worktree, its own model and a deterministic merge step fixed that for me.
Happy-path demos tell you nothing about what an agent does when a tool returns a 500. I test agents like distributed systems: inject faults on purpose and score recovery against bluffing.
Clone, install, run the tests, then separate real regressions from environmental noise. Why I built repo-doctor as a deterministic pipeline first and an agent second.
A short note from building cron-for-agents. The first version kept its watermark in memory, a restart erased it, and an agent reprocessed a week of old data. State belongs on disk.
Plain cron runs your agent on time, then wastes tokens reprocessing old data and spams you with identical summaries. Watermarks and content-hash dedup fix both.
What MCP is, why payments are the hardest tools to give an agent, and how I am designing Daraja tools around sandbox defaults, idempotency keys, reconciliation and human approval.
Worktrees do not disappear when an agent's task ends. A short note on the cleanup step I almost skipped while building worktree-orchestrator, and the disk space and confusion it saves.
How I built the assistant on mutexai.net with Laravel's HTTP client: a bounded tool-use loop, a daily spend cap, a local fallback that always answers, and a human-reviewed learning queue.
Most agent failures I debug are tool design failures. Names, schemas, descriptions and error messages decide whether a model calls the right tool with the right input.
A quick note on M-Pesa Daraja's sandbox. It will silently accept a duplicate STK push request that production would reject, so idempotency keys need testing against real behaviour, not just the docs.