AI Validation

Building the Test Set: How to Validate an AI With Your Own Data

Your test set decides whether an AI validation means anything. How to choose real, messy examples, agree the right answers, keep test data apart from training data, and avoid the shortcuts that make results look better than they are.

2026-09-22Cybroscape Technologies10 min read
Key takeaway

Your test set decides whether an AI validation means anything. How to choose real, messy examples, agree the right answers, keep test data apart from training data, and avoid the shortcuts that make results look better than they are.

You can have a perfect intended use statement and sensible acceptance criteria, and still end up with a validation that means nothing — because the test set was too easy. It happens more than anyone likes to admit. The vendor's demo documents are clean. Yours aren't.

The test set is where an AI validation earns its credibility. Here's how we put one together.

Use real material, including the ugly bits

Pull examples from your actual work. Not the best ones — a fair spread. That means including the scanned document that's slightly crooked, the requirements list written by someone who left three years ago, the template somebody modified without telling anyone.

A good rule: if a case gives your own team trouble, it belongs in the test set. The AI will meet those documents in real use, so it had better meet them in testing first.

  • Normal cases, in roughly the mix you see day to day.
  • Hard cases — messy, long, unusual.
  • Edge cases right at the boundary of the intended use.
  • A few out-of-scope cases, to check it says "not sure" instead of guessing.

Agree the right answer first

For every test case you need to know what a correct output looks like — before you see what the AI produces. This is ground truth, and it's the most time-consuming part. It's also the part you can't skip.

Have two qualified people prepare it independently where you can, then compare. Where they disagree, that's valuable: it tells you the task has some genuine judgement in it, and your acceptance criteria should allow for that. If experts can't agree, you can't fairly fail the AI for disagreeing with one of them.

Keep test data away from training data

If a document was used to train or tune the model, it can't be in your test set. The model has effectively seen the answers, so the result will look better than it really is. This sounds obvious. It's surprisingly easy to get wrong when the same folder of examples gets reused for everything.

Record where every test case came from, and lock the test set once it's approved. If you use a vendor's model, ask them to confirm your data isn't used for training — covered in data privacy in GxP AI — otherwise your test set may not stay independent.

Shortcuts that make results look better than they are

  • Testing only what the vendor supplied. Curated examples predict curated performance.
  • Too few cases. Ten documents can't tell you much about a rare but serious mistake. Size the set around the error you most need to catch.
  • Marking by impression. "Looks good" isn't a result. Compare each output against the agreed answer, item by item.
  • Quietly dropping failures. A case that fails is data. Removing it after the fact is exactly the kind of thing an inspector looks for.
  • Running it once. Generative models can answer differently each time. Run key cases more than once and record whether the output stays consistent. See validating generative AI.

Keep the test set after go-live. You'll need it again the moment the model changes — which is the next step, change control and revalidation. The wider framework is in GxP AI.

Where to go next

Explore GxP Copilot for AI-native validation, TraceDraft for source-traceable clinical documentation, or book a demo to see either on your own data.

ai validation test setai test data gxpground truth dataset validationindependent test data aihow to test an ai model pharma

Frequently Asked Questions

What should an AI validation test set contain?+

Real material from your own work in a fair mix: normal cases in roughly day-to-day proportions, hard and messy cases, edge cases at the boundary of the intended use, and a few out-of-scope cases to confirm the AI flags uncertainty rather than guessing.

What is ground truth in AI validation?+

The agreed correct output for each test case, prepared before seeing what the AI produces. Ideally two qualified people prepare it independently and compare; where they disagree, the task involves genuine judgement and the acceptance criteria should allow for it.

Why must test data be separate from training data?+

If the model was trained or tuned on a document, it has effectively seen the answer, so results look better than real performance. Record the source of every test case, lock the approved test set, and confirm with any vendor that your data is not used for training.

How many test cases do you need to validate an AI?+

Enough to detect the error you most need to catch. Ten documents cannot tell you much about a rare but serious failure. Size the set around your most important error type rather than a round number.

Should AI tests be run more than once?+

For generative models, yes. They can produce different output for the same input, so run key cases several times and record whether the output stays consistent. Keep the test set after go-live too — you will rerun it whenever the model changes.

Next step

Bring a system. We'll show you the package.