LLM eval harnesses · AI agent evaluation · CI gates
Harness Engineering: LLM eval harnesses and AI agent evaluation with pass / fail / error verdicts
AI agents write code fast. A fast change that nobody checks is a fast regression. I needed proof for each change to code, to models and to data pipelines.
How it works
Speed is easy. Proof is hard.
Many AI agents work at the same time. Each agent gets one task and its own git worktree.
No change merges on trust. Each change must pass a harness that grades it.
The same idea checks code, models and data pipelines. In ten months, 5,300+ commits went through it.
Figure: dots leave one engineer, split into eight parallel agent lanes, and pass a gate (build, tests, smoke) before they merge. Some dots stop at the gate.
Interactive
Six LLM eval harnesses, one contract
Each harness grades one kind of AI work: code, models, matching or agents. Select a harness to see what it checks and an example verdict.
Every change, before push
Code gate
Five layers run in order: build and type check, unit tests, server parity tests, a browser smoke test with stubbed backends, and an API test suite in Docker.
Why: Agents push many small changes. Each change must prove that nothing else broke.
{
"status": "pass",
"blocking": [],
"metrics": {
"build": "ok",
"unit": "358/358",
"parity": "ok",
"smoke": "16/16",
"api": "77/77"
}
} Example verdict. The shape is real. Some values are examples or redacted.
The detector, inside the product
Model eval
Calls the live detection endpoint on full drawing pages. Scores precision, recall, F1 and count error for each symbol class.
Why: The training metric says that training converged. This eval says what an estimator gets when they click AI Assist.
{
"status": "fail",
"blocking": ["recall.return_grille < baseline"],
"metrics": {
"precision": 0.91,
"recall": 0.78,
"count_error": "+6%"
}
} Example verdict. The shape is real. Some values are examples or redacted.
Competing recognizers
A/B harness
Each candidate recognizer plugs in through an adapter. All candidates run on the same corpus and are compared to frozen baselines.
Why: A new model must beat the old one on the same data. Opinions do not count.
{
"status": "pass",
"blocking": [],
"metrics": {
"candidate_f1": 0.84,
"baseline_f1": 0.81,
"delta": "+0.03"
}
} Example verdict. The shape is real. Some values are examples or redacted.
Matching strategies
Identity eval
One corpus, one label scheme, one frozen catalog, one scoring function. Each strategy is a function: decide(row, pool) returns auto-accept, a pick and a shortlist. A holdout split stops overfitting. Model replies are cached, so runs repeat exactly.
Why: A wrong match is worse than no match. The eval measures correctness, not only how many rows were matched.
{
"status": "pass",
"blocking": [],
"metrics": {
"auto_precision": "[REDACTED]",
"coverage": "[REDACTED]",
"split": "holdout"
}
} Example verdict. The shape is real. Some values are examples or redacted.
Features, end to end
Journey probes
Drives the real client modules in a real browser against real backends. Each feature has one graded check with a known answer.
Why: Unit tests can pass while the product is broken. A probe uses the product like a person does.
{
"status": "error",
"blocking": ["backend unreachable"],
"metrics": {
"ran": "0/12"
}
} Example verdict. The shape is real. Some values are examples or redacted.
Model experiments
Training loop
An agent checks a training run on a rented GPU, reads the metrics, picks the next experiment and launches it.
Why: Training takes hours. The loop keeps the GPU busy and writes down each decision.
{
"status": "pass",
"blocking": [],
"metrics": {
"mAP50": 0.755,
"next": "more clean labels, same architecture"
}
} Example verdict. The shape is real. Some values are examples or redacted.
The contract
Three answers
Every harness returns the same JSON verdict. A loop, a hook or an agent reads it and acts.
passThe gate is clear. Go to the next step.
failA tracked metric got worse. Revert and try again.
errorThe check could not run. Stop and alert a person. This is not a fail.
This video did not load.
Video · 38 seconds
Many agents, one gate
- One task per worktree.
- Every branch meets the gate.
- Merge, ship, learn.
Transcript
- One developer starts many AI agents.
- Each agent works on one branch in its own worktree.
- A task must pass every gate.
- A failed task goes back to its agent.
- Passed tasks merge into main.
- The loop repeats every day.
- Review and tests set the speed, not typing.
What I built
The parts
- A five-layer code gate: build, unit tests, parity tests, browser smoke tests, API tests. Agents must pass it before they push.
- One verify command. It runs the same checks on a laptop and in CI, and prints one JSON summary for agents. Main does not merge without it.
- Model evals that grade the live product, per class, on full pages. Not only the training metric.
Show 7 more
- An identity eval: one frozen corpus, one scoring function. Every matching strategy is scored on the same rows.
- A/B harnesses that score competing recognizers against frozen baselines.
- End-to-end journey probes: real client code, a real browser and known-answer fixtures.
- An autonomous training loop: check a run, read the result, pick the next experiment, launch it.
- I treat a hand-off as a claim to check. My pickup tool checks each claim against the live repo before an agent resumes.
- Walk-and-work: voice control of coding agents from a phone, with spoken status back.
- Agent rules files, session hand-off and pickup skills, and 80+ isolated worktrees, so many agents can use the harnesses at once.
Results
By the numbers
- 6
- Harness types
- ~1,100
- Automated tests (one product)
- 5,300+
- Commits / 10 mo
Lesson
An audit found two silent failures. In both, the checks existed but nothing forced them to run. Now the gate runs itself, and a check that cannot run says so.