In Kenya, a huge share of everyday business runs through M-Pesa. Rent, salaries, supplier payments, school fees, the kiosk down the road. If agents are going to do real operational work here, they need to be able to read and, carefully, act on mobile money.
That is the idea behind mcp-african-markets: focused MCP servers for African rails, starting with Safaricom's Daraja API and statement reconciliation. It is early (v0.x), and this post is mostly about the design decisions, because with payments the design is the product.
What MCP is, briefly#
The Model Context Protocol is a standard way for an AI client (Claude Desktop, Claude Code, an IDE, your own agent) to discover and call tools exposed by a separate server. The server declares tools with a name, a description and a JSON schema for the input. The client lists them, the model decides when to call one, and the server runs it and returns a result.
The value is in the separation. I write a Daraja server once. Any MCP-aware agent can use it, and the credentials and business rules stay on the server side, not in a prompt.
Wiring it into a client looks like this:
{ "mcpServers": { "african-markets": { "command": "npx", "args": ["-y", "mcp-african-markets", "--server", "daraja"], "env": { "DARAJA_CONSUMER_KEY": "your-key", "DARAJA_CONSUMER_SECRET": "your-secret", "DARAJA_ENV": "sandbox" } } } }
Why payments are different#
Most tools are forgiving. If an agent searches twice, you wasted a few tokens. If an agent sends money twice, someone has to call a customer and ask for it back.
Agents make this worse in specific ways:
- They retry. A timeout on the client side does not mean the payment failed on Safaricom's side.
- They are confident. A model will happily report "payment sent" based on the initial acknowledgement, which only means the request was accepted.
- They can be steered. Text in an email, a ticket or a web page can try to convince the agent to pay someone.
So the design principles for the server are strict:
- Sandbox by default. Production requires
DARAJA_ENV=production, set by a human in config. The model cannot flip it. - Read-only where possible. Status checks, phone validation and reconciliation never mutate anything.
- Idempotency keys everywhere. Every money-moving call requires one.
- Final state comes from callbacks or status queries, never from the first response.
- A human approves anything that moves money in production.
The tool list#
The Daraja server exposes a small set of tools:
| Tool | Mutates | Purpose |
|---|---|---|
validate_phone |
no | Normalise and validate a Kenyan number to 2547XXXXXXXX |
transaction_status |
no | Final status of a transaction |
stk_push |
yes | Ask a customer to approve a payment on their phone |
b2c_payout |
yes | Send money from a business to a customer |
The read-only tools are annotated so clients know they are safe:
server.registerTool( "validate_phone", { title: "Validate a Kenyan phone number", description: "Normalise a Kenyan mobile number to the 2547XXXXXXXX or 2541XXXXXXXX format Daraja expects. Accepts 07..., 01..., +254... and 254... forms. Read-only. Call this before stk_push or b2c_payout.", inputSchema: { phone: z.string().min(9).max(15) }, annotations: { readOnlyHint: true, idempotentHint: true }, }, async ({ phone }) => { const normalised = normaliseKenyanPhone(phone); if (!normalised) { return errorResult(`'${phone}' is not a valid Kenyan mobile number. Ask the user to confirm it.`); } return textResult({ phone: normalised }); }, );
STK push, and what "success" means#
An STK push is the prompt that appears on a customer's phone asking them to enter their M-Pesa PIN. Under the hood it is a POST to /mpesa/stkpush/v1/processrequest with an OAuth token and a password built from the shortcode, passkey and timestamp:
const timestamp = darajaTimestamp(new Date()); // YYYYMMDDHHmmss const password = Buffer.from(`${shortcode}${passkey}${timestamp}`).toString("base64"); const response = await daraja.post("/mpesa/stkpush/v1/processrequest", { BusinessShortCode: shortcode, Password: password, Timestamp: timestamp, TransactionType: "CustomerPayBillOnline", Amount: amountKes, PartyA: phone, PartyB: shortcode, PhoneNumber: phone, CallBackURL: callbackUrl, AccountReference: reference, TransactionDesc: description, });
A ResponseCode of "0" here means Safaricom accepted the request. It does not mean the customer paid. The customer might cancel, enter the wrong PIN, or ignore the prompt. The real outcome arrives later, at your callback URL:
{ "Body": { "stkCallback": { "MerchantRequestID": "29115-34620561-1", "CheckoutRequestID": "ws_CO_191220191020363925", "ResultCode": 0, "ResultDesc": "The service request is processed successfully.", "CallbackMetadata": { "Item": [ { "Name": "Amount", "Value": 1500 }, { "Name": "MpesaReceiptNumber", "Value": "NLJ7RT61SV" }, { "Name": "PhoneNumber", "Value": 254708374149 } ] } } } }
So the stk_push tool returns a status of pending and the CheckoutRequestID, and its description says so plainly: "The payment is not complete until transaction_status reports completed. Never tell the user the payment succeeded based on this result." That one sentence prevents the most common failure I have seen: an agent announcing success on an acknowledgement.
Idempotency: the non-negotiable part#
Every mutating tool requires an idempotency_key. The server stores the key with the request and the result. A second call with the same key returns the stored result and does not hit Daraja again:
async function once<T>(key: string, run: () => Promise<T>): Promise<T> { const existing = await store.get(key); if (existing) return existing.result as T; await store.reserve(key); // fails if another call holds the key const result = await run(); await store.complete(key, result); return result; }
The key should come from the business event, not from the agent's imagination: an invoice ID, a payroll run plus employee ID, an order number. If the agent generates a fresh random key on every retry, the protection disappears, so the tool schema asks for the business reference and derives the key from it.
Human approval for real money#
In sandbox, the tools execute directly so you can build and test flows. In production, the design is that b2c_payout does not send money. It creates a pending payout with the amount, recipient, reason and the conversation that requested it, and returns awaiting_approval. A person approves it in a separate interface, outside the agent's reach, and only then does the server call Daraja.
This is slower, and that is the point. The agent does the tedious work (collecting details, validating numbers, preparing the batch) and the human does the one thing that should never be automated away: saying yes to moving money.
Reconciliation is where agents shine#
The most useful work here is not sending money. It is matching it. Every business that takes M-Pesa has the same monthly chore: export the statement, compare it against the ledger or invoices, and find the rows that do not match.
The reconcile server is read-only by design. Its tools parse an M-Pesa statement export into rows, fuzzy-match them against a ledger CSV (by receipt number first, then by amount, date window and phone or account reference), and return the discrepancies:
{ "matched": 412, "unmatched_statement_rows": [ { "receipt": "QHK4TY7B2M", "amount_kes": 3200, "date": "2026-06-03", "reason": "no ledger entry within 3 days" } ], "amount_mismatches": [ { "receipt": "QHL1PX9D0A", "statement_kes": 5000, "ledger_kes": 500, "invoice": "INV-2291" } ] }
An agent is good at the next step: reading the discrepancies, checking context, and writing a short, human-readable report of what needs attention. Nothing is changed, so the worst outcome of a mistake is a wrong line in a report that a person reviews anyway.
Where it stands#
The project is honest about its status: the scaffold and tool registry exist, and the sandbox auth flow, STK push and status query are the first things landing. Reconciliation follows, then county public data and KRA PIN validation helpers.
If you work on African payments and have opinions about any of the above, I would like to hear them. Issues and sandbox fixtures are the most useful contributions right now.