Annex 22 · ICH Q9 · GAMP 5

GxP AI Validation

An AI system can stop being fit for use without anyone changing it. That is what makes validating AI different — and what the lifecycle has to account for.

  • FDA
  • EMA
  • MHRA
  • PMDA
  • CDSCO
  • GAMP 5
  • ICH Q9
  • 21 CFR Part 11
In short

GxP AI validation means proving an AI system is fit for a bounded intended use and keeping it that way as the model and its inputs change. It requires a held-out evaluation set, thresholds set before testing, human decision on anything consequential, drift monitoring in production, model change control, and a documented fallback so the process runs with the AI switched off.

Why the usual lifecycle is not enough

Computer system validation assumes that an unchanged system behaves consistently. Freeze the version, control the configuration, and what you tested is what you keep. That assumption is the foundation of every IQ, OQ and PQ ever executed — and it does not hold for AI.

A model degrades as the inputs it meets in production drift away from the data it was evaluated against. Nothing in the configuration changes. No alert fires. The validation package remains accurate about what was tested and silently wrong about what the system now does.

Two further problems follow. AI features arrive as updates to software you already validated, so nobody classifies them as an AI project before they are in production. And a vendor can change the underlying model without telling you — your validated state depended on behaviour you no longer control.

The assurance stack, six parts

These map onto what the draft EU Annex 22 expects. A vendor selling AI for regulated use should be able to show you their own answers to all six, for their own product — it is a fair question and a revealing one.

Bounded intended use

What the model does, on what inputs, for which decision — and explicitly what it must not be used for, so scope creep after go-live is visible rather than gradual.

A frozen evaluation set

Representative, held out from training and tuning, version-controlled. Thresholds tied to the consequence of error, agreed before the evaluation runs rather than after seeing results.

A release gate that holds

Every candidate release evaluated against the same set, with results recorded. A gate that can be waived informally is not a gate.

Human decision on anything consequential

Classification, risk and release computed from confirmed facts rather than generated; approval a re-authenticated electronic signature by a named person who saw the evidence.

Drift monitoring in production

Input drift, output acceptance and override rates, reviewed by someone with the authority to suspend the system — plus periodic re-evaluation whether or not anything looked wrong.

A fallback that actually runs

The process continues with the AI switched off. Written down, tested, and known to the people who would have to use it — not a paragraph in a plan.

Where the model must not operate

The most important decision in validating an AI system is deciding what it is not allowed to do. Risk scores, GAMP categories, traceability coverage and release verdicts should be computed in code from stated rules, not generated. An inspector asking how a requirement came to be classified high risk needs an answer of the form: these inputs, this rule, this result.

That is also what makes the system reproducible. Run the same inputs twice and the risk assessment must be identical. Generated prose can vary between runs because a human reviews and approves it; a risk score that varies between runs is a defect.

Human review that is genuinely review

A reviewer handed fluent output and no affordable way to check it will approve it. Not through negligence — verifying each statement means opening the source separately and repeating that dozens of times. This is the real risk of AI in regulated work, and it is a workflow problem rather than a model problem: a better model produces more convincing text, which makes shallow review more likely, not less.

The mitigation is structural. Show the source beside the draft, flag what the source does not support, enforce that the author cannot approve their own work, and require re-authentication at signing with the meaning of the signature recorded.

Where to start

With an inventory. Most organisations cannot say how many AI systems are in use or which touch a GxP record, partly because AI arrives embedded in software bought for another purpose. Every other readiness activity depends on knowing what is actually running and who owns it.

Then take the highest-consequence use case and work the six-part stack through it properly. Our AI validation framework sets it out step by step, and Annex 22 readiness is a checklist you can work through against your own estate.

Common questions

What is GxP AI validation?+

Proving that an AI system is fit for a defined, bounded intended use in a regulated process — and keeping it that way as the model and its inputs change. Beyond ordinary computer system validation it requires a held-out evaluation set, acceptance thresholds set before testing, human accountability for consequential decisions, drift monitoring in production, model change control, and a documented fallback.

How is validating AI different from validating software?+

Conventional validation rests on an assumption that does not hold for AI: that an unchanged system behaves consistently. A model degrades as live inputs drift away from the data it was evaluated against, while every configuration setting stays exactly as approved. That is why AI validation adds production monitoring and periodic re-evaluation to the usual lifecycle.

What does EU Annex 22 require?+

Annex 22 is still a draft — consultation closed in October 2025 and it has not been adopted — but it is the clearest available signal. It sets out a defined and bounded intended use, performance demonstrated against data independent of training, human accountability for GMP-critical decisions, explainability proportionate to risk, drift monitoring, model change control, and a documented fallback so the process continues with the AI disabled.

Can AI make a GxP-critical decision?+

No. A qualified person must remain accountable for decisions affecting product quality or patient safety. AI may draft, retrieve, summarise or propose, but the decision and the signature stay with a human who has seen the underlying evidence rather than only the model's conclusion. Risk scoring and classification should be computed deterministically so each result can be explained and reproduced.

What is model drift and why does it matter?+

Drift is the degradation of performance over time as live inputs diverge from the data the model was evaluated against. It matters because it breaks the usual validation assumption — the system can stop being fit for its intended use without anyone changing it. Input monitoring is usually the earliest warning available, ahead of any visible failure.

Does a vendor updating their model break our validated state?+

Potentially, yes. An AI component whose model changes underneath you is an uncontrolled change to a validated system. Pin the model version where the vendor allows it, require notification of changes contractually, and route any model update through change control with re-evaluation against your frozen test set.

What evidence will an inspector ask for?+

Expect questions on the intended use and what is explicitly out of scope, what the model was evaluated against and when, who reviews its output and what they see before signing, how you would know if performance degraded, how a model change is controlled, and what the process does with the AI unavailable. Each should have a document behind it.

Where to go next

For the wider topic see the GxP AI hub. For delivery, see our AI validation service or how GxP Copilot validates its own AI.

See it on your own data. In 30 minutes.

Bring a system, a URS, or an AE listing. We'll show you how GxP Copilot and TraceDraft compress the validation and clinical documentation cycle without compromising Part 11 or Annex 22 posture.

See the products