AI Validation

How Good Is Good Enough? Acceptance Criteria for AI in GxP

Setting acceptance criteria for an AI system without picking a number out of the air — linking accuracy to the cost of an error, comparing against today's manual process, and counting the right kind of mistake.

2026-09-22Cybroscape Technologies9 min read
Key takeaway

Setting acceptance criteria for an AI system without picking a number out of the air — linking accuracy to the cost of an error, comparing against today's manual process, and counting the right kind of mistake.

"What accuracy do we need?" It's the first question every team asks, and the answer is almost always a round number someone liked. Ninety-five percent. Ninety-nine. It sounds rigorous. It usually isn't, because nobody can say why that number and not another one.

An acceptance criterion has to survive one question from an inspector: why is this good enough? Here's how we build criteria that can answer it.

Start with the cost of being wrong

Before any number, ask what a mistake costs. Not in general — for this exact use. If the AI drafts a test script and a reviewer rewrites half of it, a mistake costs some review time. If it classifies a deviation as minor and nobody checks, a mistake could hide a real quality problem.

Those two uses shouldn't share a threshold. The first can live with a fair number of rough drafts. The second needs to be very reliable — or, better, needs a person in the path so the AI's accuracy isn't the only control. Your intended use statement already tells you which situation you're in.

Compare against today, not against perfect

Here's the part people forget. Your current manual process isn't perfect either. People miss requirements, copy the wrong template, misread a number at the end of a long day. So a fair question is: is the AI-plus-reviewer at least as good as the process it replaces?

That means measuring today's process first. It's more work, and it's worth it. Once you know your current error rate, the acceptance criterion stops being a guess — it becomes "no worse than what we do now, and here's the data." The draft Annex 22 leans the same way, expecting AI in critical use to perform at least as well as the process it replaces.

Count the right kind of mistake

One accuracy number hides the thing that matters most: which way the mistakes go. For most GxP uses, errors are not equal.

  • Missed items — the AI leaves out a requirement, misses a risk, fails to flag a problem. These are usually the dangerous ones, because nobody sees what isn't there.
  • False alarms — it flags something that's fine. Annoying, costs time, but a reviewer catches it.
  • Confident and wrong — plausible, well written, and incorrect. The worst kind, because it gets past tired reviewers.

So set separate criteria. Something like: no more than a set rate of missed critical items, a looser limit on false alarms, and every output below a confidence level sent to a person automatically. That's far more defensible than one headline figure.

Write the rationale next to the number

Every criterion gets one or two sentences explaining it: what error it guards against, what today's process achieves, and what the human review step will catch. The sentence is really the criterion. The number is just shorthand for it.

And set the criteria before you run the tests. We've all seen criteria that happen to sit just below the result. Inspectors have seen it too. Write them down, get QA to approve them, then test. If the AI fails, that's a useful answer, not a paperwork problem.

The criteria only mean something if the test data is honest — that's the next step, building the test set. For how criteria fit the whole lifecycle, see AI across the validation lifecycle and GxP AI.

Where to go next

Explore GxP Copilot for AI-native validation, TraceDraft for source-traceable clinical documentation, or book a demo to see either on your own data.

ai acceptance criteriaai accuracy threshold gxphow accurate does ai need to beai validation acceptance criteriaai performance criteria pharma

Frequently Asked Questions

How accurate does AI need to be for GxP use?+

There is no single number. The threshold depends on what a mistake costs in that specific use and whether a person reviews the output before it matters. A drafting tool with line-by-line human review can tolerate rougher output than a classification that goes straight into a record.

Should AI be compared against perfection?+

No — compare it against the process it replaces. Manual work has an error rate too. Measure today's process first, then set the criterion as 'at least as good as what we do now', which is defensible with data. The draft Annex 22 takes a similar view for critical uses.

Why use more than one accuracy number?+

Because errors are not equal. Missed items are usually most dangerous since nobody sees what is absent; false alarms cost time but get caught; confident wrong answers are worst because they pass tired reviewers. Separate limits for each, plus automatic human review below a confidence level, are far more defensible than one headline figure.

When should AI acceptance criteria be set?+

Before testing, approved by QA. Criteria written after the results are seen tend to sit just below the result, and inspectors recognise the pattern. If the AI fails pre-set criteria, that is a useful answer.

What should accompany each acceptance criterion?+

One or two sentences of rationale: the error it guards against, what today's process achieves, and what the human review step will catch. The rationale is the real criterion; the number is shorthand for it.

Next step

Bring a system. We'll show you the package.