"What accuracy do we need?" It's the first question every team asks, and the answer is almost always a round number someone liked. Ninety-five percent. Ninety-nine. It sounds rigorous. It usually isn't, because nobody can say why that number and not another one.
An acceptance criterion has to survive one question from an inspector: why is this good enough? Here's how we build criteria that can answer it.
Start with the cost of being wrong
Before any number, ask what a mistake costs. Not in general — for this exact use. If the AI drafts a test script and a reviewer rewrites half of it, a mistake costs some review time. If it classifies a deviation as minor and nobody checks, a mistake could hide a real quality problem.
Those two uses shouldn't share a threshold. The first can live with a fair number of rough drafts. The second needs to be very reliable — or, better, needs a person in the path so the AI's accuracy isn't the only control. Your intended use statement already tells you which situation you're in.
Compare against today, not against perfect
Here's the part people forget. Your current manual process isn't perfect either. People miss requirements, copy the wrong template, misread a number at the end of a long day. So a fair question is: is the AI-plus-reviewer at least as good as the process it replaces?
That means measuring today's process first. It's more work, and it's worth it. Once you know your current error rate, the acceptance criterion stops being a guess — it becomes "no worse than what we do now, and here's the data." The draft Annex 22 leans the same way, expecting AI in critical use to perform at least as well as the process it replaces.
Count the right kind of mistake
One accuracy number hides the thing that matters most: which way the mistakes go. For most GxP uses, errors are not equal.
- Missed items — the AI leaves out a requirement, misses a risk, fails to flag a problem. These are usually the dangerous ones, because nobody sees what isn't there.
- False alarms — it flags something that's fine. Annoying, costs time, but a reviewer catches it.
- Confident and wrong — plausible, well written, and incorrect. The worst kind, because it gets past tired reviewers.
So set separate criteria. Something like: no more than a set rate of missed critical items, a looser limit on false alarms, and every output below a confidence level sent to a person automatically. That's far more defensible than one headline figure.
Write the rationale next to the number
Every criterion gets one or two sentences explaining it: what error it guards against, what today's process achieves, and what the human review step will catch. The sentence is really the criterion. The number is just shorthand for it.
And set the criteria before you run the tests. We've all seen criteria that happen to sit just below the result. Inspectors have seen it too. Write them down, get QA to approve them, then test. If the AI fails, that's a useful answer, not a paperwork problem.
The criteria only mean something if the test data is honest — that's the next step, building the test set. For how criteria fit the whole lifecycle, see AI across the validation lifecycle and GxP AI.
Where to go next
Explore GxP Copilot for AI-native validation, TraceDraft for source-traceable clinical documentation, or book a demo to see either on your own data.
