A proxy that sits between an agent and its tools, injects failures on purpose (tool errors, hallucinated results), and scores whether the agent recovers, retries sensibly, or claims success anyway.
Why#
Happy-path demos do not tell you what an agent does when a tool returns a 500. agent-test-harness runs agents against scenarios where something is deliberately broken, and grades the recovery instead of the demo.
How it works#
Scenarios are YAML: a task, a fault to inject (tool_error, hallucination), when to inject it, and what counts as a pass. Recovery is scored out of 100: task completion (50 points), bounded retries (20 points), no false claims of success (15 points), efficient tool use (15 points). It runs against any agent with a CLI or SDK: ath run scenarios/file.yaml --agent "command {task}".
Status#
v0.x. The scenario parser and tool-error injection work today. HTTP fault injection, hallucination generation and a CI mode for catching regressions are next.