AI in GxP
The question is not whether AI is allowed in regulated environments. It is which decisions it may touch, and what you have to be able to show afterwards.
- FDA
- EMA
- MHRA
- PMDA
- CDSCO
- GAMP 5
- ICH Q9
- 21 CFR Part 11
AI is permitted in GxP environments. It suits high-volume, pattern-following work that a person reviews before it becomes a record — drafting documents, producing first-pass test cases, retrieving and citing prior work. It must not decide anything GMP-critical. Classification and risk should be computed deterministically, and every consequential output needs an accountable human who could realistically have checked it.
The permission question is settled. The design question is not.
Most discussion of AI in GxP still argues about whether regulators allow it. They do. The FDA, EMA and MHRA have all signalled acceptance of AI in regulated processes, and the draft EU Annex 22 sets out what that acceptance depends on.
The harder question is where AI belongs inside a process that has to be defensible two years later. That is a design decision, and it is the one most organisations skip — which is why so many pilots work and so few reach production.
Where it fits, and where it must not
The dividing line is not how capable the model is. It is whether a person reads the output before it counts, and whether the result has to be reproducible.
Drafting documents from structured source
Validation and qualification deliverables, protocols, safety narratives, study report sections. High volume, predictable structure, reviewed before it becomes a record.
First-pass test cases from approved requirements
Anchored to the requirement they verify, with acceptance criteria a person then adjusts and approves before execution.
Retrieval across your own material
Which prior study used this endpoint, where a safety statement originated, which documents cite a superseded source — answered with citations a person can check.
Summarising history for a human decision
Deviation trends, change history, incident patterns. The summary informs the judgement; it does not make it.
Batch release and product disposition
A qualified person decides, full stop. No draft annex, no guidance and no vendor claim changes this.
Risk scoring and classification
Must be computed from stated rules so the same inputs always produce the same answer. An inspector asking how a requirement became high risk needs a derivation, not a model output.
Final approval of any regulated record
A re-authenticated electronic signature by a named person who saw the evidence. The AI may have drafted every word; it may not sign.
Causality and clinical judgement
Causality, expectedness and seriousness belong to the investigator and medical reviewer. A system that quietly infers them is worse than no system.
Why pilots stall
The common pattern is a working pilot that nobody can promote. The model does the job. The demonstration goes well. Then quality asks what the intended use is, what it was evaluated against, who reviews the output, and what happens when it degrades — and nobody has answers, because those questions were never part of the pilot.
Answering them after the fact is far harder than deciding them beforehand. A pilot designed around a bounded intended use, a held-out evaluation set and a named reviewer is barely more work to run and can actually be promoted. GxP AI validation sets out the six parts that need to be in place.
The oversight that isn't oversight
"Human in the loop" appears in every AI governance policy and means very little on its own. A reviewer handed fluent output with no affordable way to check it against source will approve it — not through negligence, but because verifying each statement means opening the source separately and repeating that dozens of times.
A better model makes this worse, not better: more convincing text invites less scrutiny. So the test of a human-in-the-loop design is not whether a reviewer exists. It is whether checking a statement costs seconds or minutes, and whether the system flags what it could not support.
The question nobody can answer yet
How many AI systems are in use across your organisation, and which of them touch a regulated record? Most quality teams cannot say, because AI arrives embedded in software bought for another purpose — a feature update to a system that was validated two years ago.
An inventory is the least glamorous and most useful first step. Every other activity — classification, validation, monitoring, governance — depends on knowing what is actually running and who owns it. Our AI governance framework covers what the registry should hold, and Annex 22 readiness is a checklist to work through against your own estate.
Common questions
Is AI allowed in GxP environments?+
Yes. Regulators including the FDA, EMA and MHRA accept AI in regulated processes provided the system is validated for a defined intended use, monitored in production, and operates with human oversight proportionate to the risk. The draft EU Annex 22 is the most specific guidance available, though it has not yet been adopted.
Which GxP tasks is AI genuinely suited to?+
Tasks that are high-volume, pattern-following and reviewed before they become a record: drafting validation and qualification documents, producing first-pass test cases from approved requirements, writing safety narratives from structured datasets, summarising deviation history, retrieving and citing prior work. The common factor is that a human reads the output before it counts.
What must AI never do in a GxP process?+
Decide anything GMP-critical on its own. Batch release, product disposition, deviation criticality, causality assessment and final approval stay with a qualified person. Risk scores and classifications should be computed deterministically rather than generated, so the same inputs always give the same result and the derivation can be shown.
Why do most GxP AI pilots never reach production?+
Not because the model fails. Because nobody defined what the use case had to demonstrate before it would be allowed near a regulated process — the intended use, the evaluation set, who reviews the output, what happens when it degrades. Those questions are answerable, but answering them after a successful pilot is far harder than deciding them before one.
Does using AI mean we need to revalidate our systems?+
Adding an AI capability to a validated system is a change, and it goes through change control like any other. Whether it triggers revalidation depends on what the AI touches. The more common oversight is the reverse: AI arriving as a vendor feature update to software you already validated, reaching production without anyone assessing it.
How is AI in GxP different from AI anywhere else?+
Two things. Every consequential output needs a human who is accountable and could realistically have checked it. And the system must be shown to still work — an AI system degrades as live inputs drift from what it was evaluated against, while every setting stays exactly as approved. Ordinary software validation assumes that cannot happen.
Where to go next
For the full topic see the GxP AI hub. For how to prove a system is fit for regulated use, see GxP AI validation. For the validation approach underneath it, see CSV to CSA.
See it on your own data. In 30 minutes.
Bring a system, a URS, or an AE listing. We'll show you how GxP Copilot and TraceDraft compress the validation and clinical documentation cycle without compromising Part 11 or Annex 22 posture.
