An evaluation harness is not a test suite with fuzzy matching. It is a labelled dataset, a scoring function per dimension, a threshold agreed before anyone is invested in the implementation, and a stored record of every run. Build it in week one, before the feature - a harness written afterwards gets written to pass. Score retrieval and generation separately, because they fail for different reasons. Run each case several times and report the worst result alongside the mean, because a feature that works on average and fails one time in five fails for a fifth of your users.
Why is a test suite not enough for an LLM feature?
Ordinary tests assert that a function returns a value. An LLM returns a different value every time, and most of those values are fine. That breaks the assertion model in two directions at once: exact matching fails on correct output, and anything loose enough to pass correct output also passes output that is subtly wrong.
So an evaluation harness is not a test suite with fuzzy matching. It is a different instrument: a labelled dataset, a scoring function per dimension, a threshold agreed before anyone is invested in the implementation, and a record of every run so you can see the trend rather than one number.
The practical consequence is that you build it in week one, before the feature. A harness written after the system exists gets written to pass, and it will.
What is the expensive part of an evaluation harness?
The labelled dataset. Everything else here is a weekend; collecting and labelling the cases is the work, and it is the part that cannot be delegated to whoever is least busy - it needs someone who knows what a right answer looks like.
Where cases come from
- Support tickets and search logs. Real phrasing, real ambiguity, real typos - none of which you will invent.
- The questions people ask in Slack because the documentation did not answer them.
- Every production failure, added the day it is found. This is how the regression set grows, and it is the highest-value half of the dataset.
- Adversarial cases someone writes deliberately: the prompt injection, the question about a competitor, the request for legal advice.
How many
Fifty is enough to catch a broken change. Two hundred is enough to trust a threshold. Beyond about five hundred you are usually paying for labelling that a stratified sample would have given you for less - it is better to have 200 cases covering ten categories evenly than 800 that are 90% easy lookups.
What a label looks like
Not a model answer to match against. A label records the facts that must appear, the facts that must not, the documents that should have been retrieved, and - where it matters - the action the system should take.
{
"id": "carrier-rate-2019-amendment",
"query": "did we amend the northbound rate with Meridian in 2019?",
"must_include": ["amended", "2019-11", "northbound"],
"must_not_include": ["I could not find"],
"expected_docs": ["contracts/meridian-2019-amendment-3.pdf"],
"category": "contract-lookup",
"difficulty": "medium"
}Why score retrieval and generation separately?
A single quality number tells you something regressed and nothing about what. Retrieval and generation fail for different reasons and get fixed by different people, so they are measured apart.
| Dimension | What it measures | What a drop means |
|---|---|---|
| Retrieval recall@k | Was the right document in the top k results | Chunking, embeddings or filters - never the prompt |
| Groundedness | Is every claim traceable to a retrieved passage | The model is filling gaps. Usually a retrieval failure wearing a generation costume. |
| Completeness | Are the required facts present | Context window, ranking, or an answer that hedges instead of answering |
| Refusal correctness | Does it decline when it should, and only then | Thresholds drifting. Both directions are bugs. |
| Tool selection | For agents: did it call the right thing with the right arguments | Tool descriptions, not model capability, nine times out of ten |
| Cost and latency per case | Tokens in, tokens out, wall time | Something got more expensive. Catch it here, not on the invoice. |
How do you evaluate a non-deterministic model?
The single most common mistake: running each case once and reporting the result. You are sampling from a distribution, so one draw tells you very little, and a change that looks like an improvement is frequently noise.
Run each case several times - three is usually enough, five for anything near a threshold - and report the worst result alongside the mean. A feature that works on average and fails one time in five is a feature that fails for a fifth of your users.
async function scoreCase(testCase, { runs = 3 } = {}) {
const results = [];
for (let i = 0; i < runs; i += 1) {
const output = await system.answer(testCase.query);
results.push({
grounded: isGrounded(output, testCase),
complete: hasRequiredFacts(output, testCase),
retrieved: recallAtK(output.sources, testCase.expected_docs, 5),
costUsd: output.usage.costUsd,
});
}
// The spread is the signal. A mean of 0.9 hiding a 0.4 is not a pass.
return {
id: testCase.id,
mean: mean(results.map(scoreOf)),
worst: Math.min(...results.map(scoreOf)),
spread: Math.max(...results.map(scoreOf)) - Math.min(...results.map(scoreOf)),
};
}Set temperature to zero for evaluation runs if your provider supports it. It reduces variance without eliminating it - identical inputs still drift across provider-side model updates - so it is a way to make the signal cleaner, not a reason to run each case once.
Why run deterministic checks before model-graded ones?
A model grading another model is expensive, slow and itself non-deterministic. Most of what you want to know can be checked without one, and those checks should run first because they are free.
- Required and forbidden strings. Crude, and catches an enormous share of real regressions.
- Retrieval recall against expected document IDs - pure set arithmetic, no model involved.
- Schema validity when the output is structured. Parse it; if it does not parse, the case failed.
- Citation validity: every cited document ID must exist in the retrieved set. This catches invented sources deterministically.
- Cost and latency ceilings per case.
Keep the model-graded checks for what genuinely needs judgement - groundedness, tone, whether an answer addresses the question rather than an adjacent one. When you do use a grader, give it the retrieved context and ask for a per-claim verdict rather than a score out of ten. Scores out of ten from a language model are close to noise; "is this specific sentence supported by this specific passage" is a question it answers well.
When should evaluation thresholds be agreed?
Write down what ships before the implementation exists. It is a different conversation once someone has spent two weeks on the thing being measured.
# thresholds.yaml - changing these is a reviewed decision, not a fix retrieval_recall_at_5: { min: 0.90 } groundedness: { min: 0.95 } completeness: { min: 0.85 } refusal_correctness: { min: 0.95 } p95_latency_ms: { max: 3000 } mean_cost_usd_per_query: { max: 0.02 } # Any single case scoring below this fails the run regardless of the mean. per_case_floor: { min: 0.50 }
That last line matters more than the averages above it. A mean is easy to satisfy while a specific category quietly collapses, and the category that collapses is usually the one a particular customer depends on.
Where and how often should an evaluation harness run?
In CI, on every change to a prompt, a model version, a chunking parameter, a retrieval filter or a tool description. Not nightly, and not manually before a release - by then the change that caused the regression is three commits back and someone has to bisect prompts.
The practical problem is cost and time: a 200-case suite at three runs each is 600 model calls. Two things make that workable.
- A smoke subset on every commit - 30 cases stratified across categories, under two minutes - and the full suite on merge to main and before release.
- Cache by a hash of the prompt, the model version and the retrieved context. Most pull requests change one thing, so most cases are unaffected and need no new call.
Store every run. The absolute number matters far less than the trend, and the trend is what tells you a provider changed a model under you - which happens, and which you will otherwise discover from a customer.
What should you do when evals disagree with users?
It will happen: the suite is green and someone reports that the system has got worse. The suite is not automatically wrong, and neither is the user. There are three real possibilities and they are distinguishable.
- 01
The dataset does not cover what they are doing
Most common by far. Your cases came from last quarter's tickets; their usage moved. The fix is not to argue - it is to take their examples, label them, and add them. A harness that only contains cases you thought of is a harness that measures your imagination.
- 02
You are measuring the wrong dimension
Groundedness at 0.97 and users still unhappy usually means the answers are correct and unhelpful - technically supported, badly structured, burying the answer in the third paragraph. Add a dimension for it rather than assuming the complaint is vague.
- 03
The failure is outside the system
A document that was never ingested, a permission that hides the right answer from that particular user, a stale index. The generation is fine; the input was wrong. This is why retrieval is scored separately - it localises the problem in one look.
In all three cases the outcome is the same: the reported case joins the dataset. Over a year that is what turns a harness from a gate into an asset, and it is the piece of a delivery with the longest useful life.
What we hand over
Every build ships with the harness, not just the code that passes it: the labelled dataset, the scoring functions, the thresholds file, the CI configuration, and a written note on how to add a case. It is documented for the team who will own it, because a harness nobody knows how to extend stops being run within two releases.
How that fits into a delivery is on AI integration and RAG systems; the agent variant, which scores tool selection rather than prose, is on AI agent development. It is roughly 20% of a build, and it is the 20% that decides whether the other 80% still works in month six.
