framework

AI Validation Framework — Annex 22 aligned

How to validate AI systems used in GxP-regulated processes: model card, eval set, deterministic guardrails, HITL policy, drift, fallback.

In short

Validating AI in a GxP process means proving the system is fit for a defined intended use and keeping it that way as the model changes. EU Annex 22 sets the expectations: a stated intended use, a held-out test set that gates every release, human oversight of consequential decisions, monitoring for drift, and a documented fallback so the process still runs with the AI switched off.

Define the intended use narrowly

Most AI validation difficulty traces back to an intended use written broadly enough to sound impressive. A narrow statement is easier to test and easier to defend.

  • 1

    Intended use statement

    What the model does, on what inputs, producing what output, used for what decision, by whom.

  • 2

    Explicit out-of-scope uses

    What the model must not be used for. Without this, scope creep happens silently after go-live.

  • 3

    Criticality determination

    Whether the output influences a GMP-critical decision. Annex 22, still in draft, is clear that AI should not make such a decision autonomously.

  • 4

    Deterministic alternative considered

    Whether a rules-based approach would do the job. If it would, use it — it is cheaper to validate and easier to explain.

Data and model documentation

The model card is the equivalent of a design specification. An inspector will ask what the model was trained on and what it was tested against.

  • 5

    Model card

    Model type, version, training data provenance, known limitations, and the conditions under which performance degrades.

  • 6

    Training data governance

    Where the data came from, what rights you have to it, how it was cleaned, and whether any regulated data is involved.

  • 7

    Test set held out and frozen

    Never used in training or tuning, representative of real inputs including awkward ones, and version-controlled.

  • 8

    Bias and edge-case analysis

    Where the model performs worse, assessed deliberately rather than discovered in production.

  • 9

    Version pinning

    The exact model version in production is recorded. A silently updated vendor model is an uncontrolled change.

Acceptance criteria and the release gate

This is the part most AI projects skip, and the part that makes the difference between a validated system and a pilot that escaped.

  • 10

    Performance thresholds set before testing

    Defined in advance and tied to the consequence of error, not chosen after seeing the results.

  • 11

    Error-type weighting

    A false negative and a false positive rarely cost the same. State which matters more and why.

  • 12

    Evaluation run against the frozen set

    Every candidate release evaluated against the same held-out set, with results recorded.

  • 13

    Release blocked on failure

    Enforced by the process, not by goodwill. A gate that can be waived informally is not a gate.

Human oversight

Annex 22 expects a person to remain accountable. The design question is where that person sits and whether they can realistically intervene.

  • 14

    Human-in-the-loop policy

    Which outputs require human review before use, which are advisory, and which are logged only.

  • 15

    Reviewer sees the evidence

    The reviewer gets the source material and the model's citation, not just its conclusion. Review without evidence is rubber-stamping.

  • 16

    Segregation of duties preserved

    The person approving AI-assisted output cannot be the person who generated it, exactly as with human-authored work.

  • 17

    Override is possible and recorded

    Reviewers can reject or amend, and the rate of override is tracked as a quality signal.

  • 18

    Deterministic guardrails on outputs

    Risk scores, classifications and calculations computed in code, not generated. An inspector will ask how a number was produced.

Monitoring, drift and fallback

The obligation that separates AI validation from software validation: the thing can get worse without anyone changing it.

  • 19

    Input drift monitoring

    Detects when live inputs stop resembling the test set — usually the earliest warning available.

  • 20

    Output and performance monitoring

    Tracks acceptance rates, override rates and, where ground truth arrives later, accuracy over time.

  • 21

    Re-evaluation cadence

    A defined period after which the model is re-run against the test set regardless of whether anything appeared wrong.

  • 22

    Documented fallback

    The process runs without the AI. Written down, tested, and known to the people who would have to use it.

  • 23

    Rollback capability

    A previous model version can be restored, and the decision to roll back has a named owner.

  • 24

    Change control covers model updates

    Retraining, fine-tuning and vendor model upgrades go through change control like any other change to a validated system.

How to use this

EU Annex 22 is the GMP annex covering artificial intelligence. Read it alongside Annex 11 rather than instead of it — the electronic records and signature obligations still apply to AI-assisted records.

The single most useful design decision is to keep generative output away from anything that must be reproducible. Let the model draft and summarise; compute classifications, risk scores and calculations deterministically so they can be explained and repeated.

This is a working aid, not a regulatory document, and it does not replace your own quality system procedures.

Get the working copy

The same checklist as a spreadsheet, with columns for applicability, owner, evidence location and status — the version you take into an audit. Downloads immediately.

We will not add you to a mailing list. The checklist above stays free either way.

Common questions

How do you validate an AI system for GxP use?

Define a narrow intended use, document the model and its training data, hold out a frozen test set, set acceptance thresholds before you evaluate, gate each release against that set, require human review of consequential outputs, monitor for drift in production, and maintain a documented fallback so the process runs with the AI disabled.

What is EU Annex 22?

Annex 22 is the EU GMP annex covering artificial intelligence in regulated manufacturing. It is still a draft — consultation closed in October 2025 and it has not been adopted — but it is the clearest signal of what regulators will expect. It sets out a defined intended use, demonstrable model performance against representative data, human oversight of GMP-critical decisions, monitoring for drift, change control over model updates, and a documented fallback. It complements Annex 11 rather than replacing it.

Can AI make a GMP-critical decision on its own?

No. Annex 22, still in draft, expects that a qualified person remains accountable for decisions affecting product quality or patient safety. AI may draft, summarise, retrieve or propose, but the decision and the signature stay with a human who has seen the underlying evidence rather than only the model's conclusion.

What is model drift and why does it matter for validation?

Drift is the degradation of model performance over time as live inputs diverge from the data the model was evaluated against. It matters because it breaks the usual assumption of software validation: the system can stop being fit for its intended use without anyone changing it. Monitoring and periodic re-evaluation are the controls.

Does a vendor updating their model break our validated state?

Potentially, yes. An AI component whose model changes underneath you is an uncontrolled change to a validated system. Pin the model version where the vendor allows it, require notification of changes contractually, and route any model update through change control with re-evaluation against your frozen test set.

What's inside

  • Model card + eval set
  • Deterministic guardrails
  • HITL policy design
  • Drift monitoring + rollback

Explore related resources, products and services: GxP Copilot, TraceDraft, validation services, AI transformation.

See it on your own data. In 30 minutes.

Bring a system, a URS, or an AE listing. We'll show you how GxP Copilot and TraceDraft compress the validation and clinical documentation cycle without compromising Part 11 or Annex 22 posture.

See the products