Skip to content
  • Proxies agent tool calls and injects failures per scenario: tool errors, hallucinated results.
  • Scores recovery 0 to 100: completion, bounded retries, no false success claims, efficient tool use.
  • Works with any CLI or SDK agent; scenarios are plain YAML.
README

A proxy that sits between an agent and its tools, injects failures on purpose (tool errors, hallucinated results), and scores whether the agent recovers, retries sensibly, or claims success anyway.

Why#

Happy-path demos do not tell you what an agent does when a tool returns a 500. agent-test-harness runs agents against scenarios where something is deliberately broken, and grades the recovery instead of the demo.

How it works#

Scenarios are YAML: a task, a fault to inject (tool_error, hallucination), when to inject it, and what counts as a pass. Recovery is scored out of 100: task completion (50 points), bounded retries (20 points), no false claims of success (15 points), efficient tool use (15 points). It runs against any agent with a CLI or SDK: ath run scenarios/file.yaml --agent "command {task}".

Status#

v0.x. The scenario parser and tool-error injection work today. HTTP fault injection, hallucination generation and a CI mode for catching regressions are next.

More open source

All projects

Self-hosted, git-backed memory any AI agent can plug into

TypeScript 0 2026 - now

Parallel coding agents in isolated git worktrees, merged deterministically

TypeScript 0 2026 - now

MCP servers for African fintech: M-Pesa payments, reconciliation, sandbox by default

TypeScript 0 2026 - now

Scheduled agent jobs with dedup, watermarks and delivery

TypeScript 0 2026 - now