Every time I pick up an unfamiliar repository, the first hour goes the same way. Clone it. Install dependencies. Run the tests. Watch forty of them fail. Scroll through stack traces trying to work out which failures are real bugs and which are a missing environment variable, a test server that is not running, or a lockfile that drifted.
That hour is tedious and mostly mechanical, which makes it a good job for software. repo-doctor is my attempt: point it at a repo, get back a triaged report you can paste into a PR or an issue.
npx repo-doctor check https://github.com/acme/widget --out report.md
The pipeline#
repo-doctor is a straight line:
clone -> install -> run tests -> diagnose -> report
- Clone shallowly into a temp directory. Your working copy is never touched.
- Install with the package manager the project already uses. npm today, pnpm and yarn next.
- Run tests with stdout, stderr and the exit code fully captured.
- Diagnose each failure: which test, which file, what the assertion said, and whether it looks environmental.
- Report as markdown.
You can pin a branch or tag with --ref, which is useful when checking whether a failure is new since a release.
Deterministic first, model second#
The obvious way to build this in 2026 is to hand the whole log to a model and ask what went wrong. I did not start there, on purpose.
Test output is highly structured. Jest, Vitest and Mocha all print failing test names, file paths, expected versus received values and stack frames in predictable formats. Parsing that with code is fast, free and exactly right every time. A model reading the same log is slower, costs money, and occasionally invents a test that does not exist.
So the current diagnosis is pattern-based:
const ENVIRONMENTAL: Array<{ pattern: RegExp; hint: string }> = [ { pattern: /ECONNREFUSED 127\.0\.0\.1:(\d+)/, hint: "A local service on port $1 is expected to be running." }, { pattern: /EADDRINUSE.*:(\d+)/, hint: "Port $1 is already in use; another test process may be running." }, { pattern: /Missing required env(ironment)? var(iable)?:? (\w+)/i, hint: "Set $3 before running tests." }, { pattern: /ETIMEDOUT|getaddrinfo ENOTFOUND/, hint: "The test needs network access." }, ]; function classify(failure: Failure): Diagnosis { for (const { pattern, hint } of ENVIRONMENTAL) { const match = failure.error.match(pattern); if (match) return { kind: "environmental", hint: expand(hint, match) }; } return { kind: "regression", hint: null }; }
This separates the two categories that matter most. An environmental failure means "your setup is incomplete", and the fix is a line in the README or a CI service. A regression means "the code is wrong", and that is where a human, or a model, should spend attention.
The LLM step on the roadmap sits after this, not instead of it. Once the pipeline has isolated one real failure with its test name, assertion diff and the relevant source file, the model gets a small, focused context and a narrow question: what is the likely root cause, and what would a fix look like? That is a much better prompt than 3,000 lines of CI log.
What the report looks like#
The output is meant to be pasted, not read in a terminal:
# repo-doctor report **Repo:** https://github.com/acme/widget **Verdict:** 3 failing test(s) ## Failures ### 1. parse > rejects malformed input - **Suite:** src/parser.test.ts - **Error:** Expected: 0, Received: 1 (src/parser.test.ts:42:5) ### 2. api > retries on 503 - **Suite:** src/api.test.ts - **Error:** fetch failed: connect ECONNREFUSED 127.0.0.1:3000 ## Suggested next steps - Review the failing tests above for a regression. - 1 failure looks environmental (connection refused); check whether the test server is expected to be running.
Short, grouped, and honest about what it knows. When a failure does not match a known pattern, it says so instead of guessing.
Sandboxing is not optional#
repo-doctor runs other people's code. npm install executes lifecycle scripts, and test suites can do anything. The minimum is a throwaway directory and no access to your working copy. The next step, which I consider required before the tool runs anything automatically on untrusted repos, is a container with no credentials, limited network and a hard timeout.
This is also why the tool is a CLI you run on purpose, not a bot that acts on every repo it sees. If you are going to give an agent the ability to run arbitrary test suites, the blast radius has to be small by construction.
Why an agent at all#
A fair question: if the deterministic parts do most of the work, where does the agent come in?
Three places, in order of how much I trust them:
- Summarising a set of failures into the "suggested next steps" section. Low risk, high value.
- Root-cause analysis on one isolated failure with a small context. Medium risk. Always shown as a hypothesis, never as fact.
- Proposing a fix as a draft PR. Highest risk. Planned, and it will always be a draft for a human to review.
The pattern is the same one I use everywhere: code does the parts that must be correct, the model does the parts that need judgement, and a person approves anything that changes the repository.
Flaky tests#
One feature I want soon is flaky-test detection: rerun the failing tests a few times and diff the results. A test that fails once and passes twice is a different problem from one that fails every time, and the report should say so. It is cheap to implement and it removes a whole class of wasted debugging.
How I use it#
Mostly in three situations:
- Before contributing to an open-source project, to know which failures already exist on
mainso I do not chase them in my own branch. - After an agent changes code, as a neutral check. When several coding agents work on a repository in parallel, I want one tool that reports test status the same way regardless of which agent did the work.
- When onboarding a repo that has not been touched in a while, to get a list of what broke while nobody was looking.
Status#
v0.x. The clone, install, test and report pipeline works end to end on npm-based repositories, and failure scanning handles Jest, Vitest and Mocha output. LLM-assisted root cause, pnpm and yarn support, flaky detection and a GitHub Actions mode that comments on PRs are next.
If it breaks on your repo, please open an issue with the repo URL. Real failures are the most useful test data this project can get.