AI Software

What Annex 22 Actually Requires From GxP AI Software — and How to Check

The six assurance elements EU Annex 22 mandates, mapped to what your GxP AI software vendor must be able to demonstrate today — not in a future roadmap.

2026-08-04Cybroscape Technologies12 min read
Key takeaway

The six assurance elements EU Annex 22 mandates, mapped to what your GxP AI software vendor must be able to demonstrate today — not in a future roadmap.

EU Annex 22 on Artificial Intelligence in GMP environments formally establishes AI as a validated subsystem within GxP operations. It is not a list of suggestions. For any AI capability used in a regulated process — validation, quality management, clinical documentation, pharmacovigilance — the annex creates six concrete assurance requirements that must be met today, not on a future compliance roadmap.

Requirement 1: Published AI governance

Annex 22 requires that the organisation using AI in a GxP context publishes — internally and to regulators on request — documentation of what the AI is used for, what it is not used for, the model or models involved, the training data basis (to the extent disclosable), and the risk assessment that determined the AI is appropriate for the regulated use case. This is the model card requirement in practice. For GxP software vendors, it means the model card must exist and be customer-accessible — not a marketing brochure but a technical governance document. GxP Copilot publishes this in Annex 22 assurance.

Requirement 2: A performance benchmark that gates every release

Every AI model release — including minor updates, fine-tunes, and provider-side changes to underlying models — must be evaluated against a frozen, versioned benchmark before deployment to regulated environments. The benchmark is a set of representative inputs with documented expected outputs and pass criteria. It is not a general LLM benchmark (MMLU, HumanEval); it is a domain-specific evaluation of the AI's performance on the actual regulated tasks it is deployed for. The benchmark result must be documented and retained as part of the change control record for the model update.

Requirement 3: Deterministic guardrails on AI outputs

The AI must be constrained by deterministic checks that prevent out-of-scope, harmful, or non-compliant outputs from reaching users. These guardrails are not the AI's own judgment — they are programmatic constraints that run regardless of what the AI generates. Examples: a guardrail that blocks any AI output containing a specific regulatory claim the product is not authorised to make; a format constraint that ensures AI-generated test cases conform to the required template; a length constraint that prevents an AI-generated executive summary from exceeding the permitted document section length. Guardrails must themselves be change-controlled.

Requirement 4: Human-in-the-loop at every critical decision

Annex 22 does not permit fully autonomous AI decision-making at critical GxP decision points — release decisions, approval of validation deliverables, sign-off on clinical narratives. A qualified human must review the AI's output and approve it with an electronic signature that meets 21 CFR Part 11 requirements. The HITL workflow must be enforced by the system — not just described in an SOP. The system must prevent author self-approval, must record the reviewer's identity and timestamp, and must preserve the original AI output alongside the final approved version. See Human-in-the-Loop reviews.

Requirement 5: Drift monitoring and rollback capability

AI model behaviour drifts over time — the model's outputs on the same inputs change as the underlying model is updated, as context windows change, or as the distribution of real-world inputs shifts away from the training distribution. Annex 22 requires that this drift be monitored, that tripwires trigger a review when drift exceeds a threshold, and that a rollback to a prior model version is technically feasible within a defined timeframe. Drift monitoring is not the same as user satisfaction monitoring — it requires structured evaluation of AI outputs against the established benchmark, run on a schedule and triggered by output-quality alerts.

Requirement 6: Documented fallback

The regulated process must continue to function when the AI is unavailable, degraded, or producing outputs below the quality threshold. The fallback is a documented, tested alternative path — typically a manual or template-driven process — that users can invoke without retraining or system configuration changes. The fallback must be tested at least annually as part of the AI system's periodic review. An AI capability that has no fallback is a single point of failure in a GxP process, which is not an acceptable design.

How to assess your current GxP AI software against Annex 22

  • Request the model card from your vendor today. If one does not exist, that is a gap under Annex 22.
  • Ask for the benchmark documentation — what it tests, what the pass threshold is, and what the current model scores.
  • Map each AI capability in the platform to a GxP process step and document the HITL gate for each.
  • Test the fallback: disable the AI capability in a staging environment and verify the workflow continues without errors.
  • Ask your vendor for their drift monitoring methodology and frequency — and request the last drift report.
  • Check that the vendor's change control process for model updates includes customer notification and a re-validation assessment.

Where to go next

Explore GxP Copilot for AI-native validation, TraceDraft for source-traceable clinical documentation, or book a demo to see either on your own data.

annex 22 ai softwaregxp ai annex 22 requirementseu annex 22 software complianceai gmp annex 22 validation
Next step

Bring a system. We'll show you the package.