Most AI projects reach their first decision point in a meeting room. Someone shares a screen, types a few questions into the system, and the answers are good — sometimes remarkably good. The steering group nods, and the project moves from pilot to rollout on the strength of what amounts to a handful of examples chosen by the people who built it.
Nobody in that room is being dishonest. The demo is genuinely impressive. The problem is that it answers the wrong question. It shows that the system can produce a good answer. What the business needs to know is how often it does, on the cases that actually arrive, and what happens on the ones it gets wrong.
Those are measurement questions, and they need a measurement instrument. For AI systems, that instrument is a test set — a fixed collection of real cases with agreed correct outcomes, run against the system every time it changes. It is the least glamorous artifact in the project, and in our experience it's the one that most reliably separates programs that scale from programs that stall.
Why demos mislead even when nobody means them to
Demos are selected, not sampled. The examples shown are the ones the team has tried before and knows work. That's natural — nobody opens a steering meeting with the failure cases — but it means the audience is seeing the top of the distribution and drawing conclusions about the whole of it.
Language models also fail in ways that look like success. A wrong answer from traditional software is usually obviously wrong: an error, a blank field, a number that doesn't add up. A wrong answer from a model is fluent, confident, and correctly formatted. Without a reference answer to compare against, reviewers routinely accept output that a subject-matter expert would reject in seconds.
And the system keeps changing. Prompts get tuned, retrieval sources get added, the underlying model is upgraded by the vendor. Each change fixes something somebody noticed. Without a fixed baseline, nobody notices what each change broke — and in our experience, roughly one change in three that improves the case it was aimed at degrades something else.
A demo is evidence that the system has worked. A test set is evidence about whether it works. Only one of those should be deciding a rollout.
What a useful test set contains
A test set doesn't need to be large to be valuable. For most business use cases, 150 to 300 well-chosen cases are enough to start making decisions with. What matters far more than size is what goes in it.
- Real inputs, not invented ones. Pull cases from the actual queue — the emails, tickets, documents, or questions the system will face in production, with sensitive data handled appropriately. Synthetic examples written by the project team share the team's assumptions, which is exactly the blind spot you're trying to cover.
- The distribution you'll actually see. If 40% of incoming requests are one routine type, roughly 40% of the set should be too. A set dominated by interesting edge cases will make a good system look bad; one dominated by easy cases will make a weak system look ready.
- A deliberate slice of the hard cases. On top of the representative sample, add a tagged group of known difficult categories — ambiguous requests, missing information, cases that should be refused or escalated. Score these separately so they don't get averaged away.
- Answers agreed by the people who own the outcome. The reference answer for each case should come from the domain experts who would be accountable if the system got it wrong, not from the engineers building it. Expect this to surface disagreements between experts; resolving those is valuable work in its own right.
- The cases that should produce no answer. A system that answers everything is not a system that knows its limits. Include questions outside scope and inputs that should trigger a handoff to a person, and treat a confident answer to one of those as a failure.
Scoring that means something to the business
A single accuracy percentage is a starting point, not a verdict. The number that matters is accuracy weighted by consequence. A system that is 92% correct but whose errors are all minor formatting issues may be ready to ship. A system that is 97% correct but whose errors include confidently wrong policy guidance may not be.
So before scoring anything, agree an error taxonomy with the business owner: which failures are cosmetic, which create rework, and which create real exposure — financial, regulatory, or reputational. Then report results by category. "Three severe errors in 220 cases, all in the refund-eligibility slice" is a sentence a steering group can act on. "94.1% accuracy" is not.
Automated scoring helps with scale. For structured outputs — a classification, an extracted field, a routing decision — it's straightforward. For free text, a second model can grade answers against the reference, but treat that grader as a component that also needs checking. Have people score a sample of the same cases periodically and confirm that the automated grades agree with theirs. When they diverge, trust the people and fix the grader.
Running it every time something changes
A test set earns most of its value after launch. Wire it into the release process so that every prompt edit, retrieval change, configuration tweak, or model version runs against the full set before it reaches users, and the results are compared against the last approved baseline.
Set explicit release gates in advance. For example: no increase in severe errors, no slice dropping more than a few points, and overall accuracy at or above the baseline. A change that fails a gate doesn't ship until someone with business authority decides the trade-off is acceptable — and that decision gets written down.
Then keep the set alive. Every production failure that gets reported should be reviewed and, where it represents a pattern, added to the set with its corrected answer. Over a year this turns the test set into an institutional memory of everything the system has ever gotten wrong, which is precisely what you want guarding the next change.
A short worked example
An insurer had built an assistant to help claims handlers answer policy-coverage questions. The demo was excellent, and the plan was to roll it out to 300 handlers the following month. Before approving, the claims director asked for a test set: 240 questions pulled from three months of real handler queries, with answers agreed by two senior claims technicians.
Overall, the assistant answered 89% correctly. But the scoring by category told the real story. On straightforward coverage lookups it was 96% correct. On questions involving policy endorsements — amendments layered on top of the base wording — it fell to 61%, and most of those errors were confident and plausible. The cause was retrieval: endorsements were stored as separate documents and the system often found the base policy without the amendment that changed it.
The fix took three weeks. Endorsement accuracy rose to 91%, the rollout went ahead with endorsement questions flagged for a second check in the first month, and the test set now runs on every change. Two months later it caught a vendor model update that reduced accuracy on exclusions by eight points — before any handler saw it.
The honest takeaway
Building a test set is slow, unexciting work that requires scarce expert time, and it rarely features in anyone's AI strategy deck. It is also the only thing that lets you answer, with evidence, the question every executive eventually asks: how do we know this is working?
Keep the demo for building enthusiasm. Use the test set for deciding what ships.