You can have a perfect intended use statement and sensible acceptance criteria, and still end up with a validation that means nothing — because the test set was too easy. It happens more than anyone likes to admit. The vendor's demo documents are clean. Yours aren't.
The test set is where an AI validation earns its credibility. Here's how we put one together.
Use real material, including the ugly bits
Pull examples from your actual work. Not the best ones — a fair spread. That means including the scanned document that's slightly crooked, the requirements list written by someone who left three years ago, the template somebody modified without telling anyone.
A good rule: if a case gives your own team trouble, it belongs in the test set. The AI will meet those documents in real use, so it had better meet them in testing first.
- Normal cases, in roughly the mix you see day to day.
- Hard cases — messy, long, unusual.
- Edge cases right at the boundary of the intended use.
- A few out-of-scope cases, to check it says "not sure" instead of guessing.
Agree the right answer first
For every test case you need to know what a correct output looks like — before you see what the AI produces. This is ground truth, and it's the most time-consuming part. It's also the part you can't skip.
Have two qualified people prepare it independently where you can, then compare. Where they disagree, that's valuable: it tells you the task has some genuine judgement in it, and your acceptance criteria should allow for that. If experts can't agree, you can't fairly fail the AI for disagreeing with one of them.
Keep test data away from training data
If a document was used to train or tune the model, it can't be in your test set. The model has effectively seen the answers, so the result will look better than it really is. This sounds obvious. It's surprisingly easy to get wrong when the same folder of examples gets reused for everything.
Record where every test case came from, and lock the test set once it's approved. If you use a vendor's model, ask them to confirm your data isn't used for training — covered in data privacy in GxP AI — otherwise your test set may not stay independent.
Shortcuts that make results look better than they are
- Testing only what the vendor supplied. Curated examples predict curated performance.
- Too few cases. Ten documents can't tell you much about a rare but serious mistake. Size the set around the error you most need to catch.
- Marking by impression. "Looks good" isn't a result. Compare each output against the agreed answer, item by item.
- Quietly dropping failures. A case that fails is data. Removing it after the fact is exactly the kind of thing an inspector looks for.
- Running it once. Generative models can answer differently each time. Run key cases more than once and record whether the output stays consistent. See validating generative AI.
Keep the test set after go-live. You'll need it again the moment the model changes — which is the next step, change control and revalidation. The wider framework is in GxP AI.
Where to go next
Explore GxP Copilot for AI-native validation, TraceDraft for source-traceable clinical documentation, or book a demo to see either on your own data.
