Skip to content

Build notes #3: cron jobs, sandbox payments, and a test harness

Two months of unglamorous infrastructure. A cron tool that does not spam duplicate summaries, sandbox M-Pesa tools for agents, and the start of a proper way to test how agents fail.

Hi, it's John. Since the last issue I have mostly been doing the unglamorous kind of work: the plumbing that makes agents boring in a good way. No launches this time, just three things maturing.

What I shipped#

cron-for-agents. I kept running the same agent job twice because plain cron does not know it already did the work, and I kept getting three identical summaries because the source data had not changed. So I built a small scheduler that tracks two things per job: a watermark (how far it got last time, passed to the job as $CFA_WATERMARK) and a content hash (so an unchanged result never gets delivered twice). State lives in a single .cfa/state.json file, no database. Delivery to Telegram, WhatsApp and email is the next piece; right now it prints, which is enough to prove the dedup logic is right before I wire up the channels.

Early mcp-african-markets tools. I started exposing M-Pesa Daraja to agents through MCP: a daraja server for payouts and transaction queries, and a reconcile server for matching statement lines. Everything defaults to Daraja's sandbox, and switching to production mode is a separate, explicit opt-in, not a config flag an agent could flip by accident. I am not trusting an agent with real money movement until the sandbox version has been wrong in every way I can think of first.

Starting agent-test-harness. This is the one I am most excited about. The idea: sit as a proxy between an agent and its tools, inject failures on purpose (a tool that times out, one that returns a plausible but wrong answer), and grade whether the agent notices, retries sensibly, or claims success anyway. Scenarios are plain YAML: a task, a fault to inject, and what counts as a pass. The scorer is still simple (did it finish, did it retry within a bound, did it avoid claiming success when it should not have) but even that simple version has already caught agents that "fixed" a test by not touching it.

What I learned#

Sandboxes lie in specific, useful ways. Daraja's sandbox will accept requests that production would reject, and it stays quiet about timing behaviour that only shows up under real load. That is not a reason to skip the sandbox, it is a reason to write down exactly what it does not tell you, so the list is ready before the first production test instead of discovered by an outage.

The other lesson, from cron-for-agents: state that only lives in memory is not state. The first version kept its dedup hashes in a running process, and the moment I restarted it during a deploy, it forgot everything and re-sent a batch of reports. A single JSON file that a restart cannot erase turned out to matter more than any cleverness in the dedup algorithm itself.

What I'm reading and trying#

  • Idempotency keys as a design pattern, not just a payments thing. I am starting to add them to every tool that can be called twice by accident.
  • YAML scenario formats from other testing tools, mostly to steal naming conventions rather than mechanics.
  • Reading Daraja's callback documentation slowly, twice, because the first read always misses a field name that matters later.

A question for you#

If you have ever given an agent access to a tool that moves money or sends a message a human will see, what stopped you from shipping it straight to production? I am collecting real answers, not hypotheticals, for a longer post on this. Reply and tell me.

John