There's a tempting idea going around: let the AI run the tests. It reads the protocol, drives the system, decides pass or fail, writes it up. Fully automatic. It's a great demo. In a regulated environment it's where teams get into real trouble.
The problem isn't that AI is bad at testing. It's that test evidence has to be repeatable and explainable, and an AI deciding things at run time is neither, reliably. There's a simpler pattern that gets you most of the speed and keeps the evidence solid. It's how we built our own approach.
What goes wrong when AI runs the test
- You can't repeat it. Run the same test tomorrow and a generative model may take a slightly different path. A test you can't repeat isn't much of a test.
- You can't fully explain it. Why did it click there? Why did it decide that was a pass? "The model judged it" doesn't satisfy an inspector.
- It can quietly improvise. If a step fails, an AI may try something else to get there — which is exactly what a tester must not do. A test should fail and be recorded, not be worked around.
- It marks its own homework. The same thing doing the work and deciding it passed is a separation-of-duties problem, not just a technical one.
It also sits badly with the draft Annex 22, which doesn't expect generative models to be making critical GMP decisions — and a pass/fail on a validation test is a decision.
The pattern: AI drafts, a person approves, a fixed engine runs
Split the job into three parts and keep them separate.
1. The AI drafts. It reads your requirements and writes test steps — but only using a fixed, approved list of actions the testing engine understands. Open a page. Enter a value. Check a field. Confirm a file exists. It can't invent new kinds of action, and it can't write free-form commands. That limit is deliberate: it keeps every step something a reviewer can read and understand.
2. A person approves. A qualified tester or validation lead reviews the steps, changes what's wrong, and approves the protocol. From here on it's their protocol, not the AI's. This is the step that makes it GxP — the human decision is recorded, with their name and signature.
3. A fixed engine runs it. A plain, non-AI engine executes exactly the approved steps, in order, and captures evidence — screenshots, values, timestamps. No AI is involved while it runs. Same input, same steps, same result, every time. If a step fails, it stops and records the failure. It doesn't try to be clever.
A failed step then goes where failures always go: a deviation, investigated by a person. Nothing about the AI changes that.
Why this holds up in an inspection
- Repeatable. Run it again and you get the same steps. The engine doesn't improvise.
- Explainable. Every step is a plain action from an approved list. Anyone can read it.
- A clear human decision. The approval is on record, before execution. See GxP roles and responsibilities.
- Honest failures. A failing test fails. That's what you want from evidence.
- A narrower thing to validate. You validate the engine once as ordinary software. The AI sits upstream as a drafting aid with a human check, which is a much lower-risk role.
You still save most of the time. The slow part of testing was always writing the scripts and capturing the evidence — the AI speeds up the first, the engine handles the second. What you give up is the idea that nobody needs to look. In GxP, somebody always needs to look.
Where to start
Pick one system with stable screens and plenty of repetitive tests — an OQ for a configured application is a good first candidate. Run the new approach alongside your normal process for a few cycles and compare. The steps are laid out in running a 90-day GxP AI pilot.
If you're following the whole series, the order is: intended use, acceptance criteria, test set, change control, and this. For the regulatory picture, see GxP AI.
Where to go next
Explore GxP Copilot for AI-native validation, TraceDraft for source-traceable clinical documentation, or book a demo to see either on your own data.
