Here's the uncomfortable thing about validating AI. The validation is only true for the exact model you tested. Change the model, and strictly speaking you're looking at a different system. Traditional software changes when someone ships a release. AI can change when a vendor swaps the model underneath, when a prompt or template is edited, or when the kind of documents going in slowly shifts.
So change control for AI needs a few extra triggers. Here's how we think about it.
What counts as a change
Add these to your change control procedure by name. If they aren't written down, they get treated as routine maintenance.
- A new model version — from your vendor or your own team.
- A change to prompts, instructions or templates that shape the output. Small wording changes can shift behaviour more than you'd expect.
- New types of input — a new document template, a new site, a new language.
- A wider intended use — using it for something it wasn't validated for.
- Monitoring going out of bounds — performance drifting below your limits. See continuous monitoring.
How much to retest
Not every change means starting over. This is where keeping your test set pays off — you already have approved cases and agreed answers, so retesting is a rerun, not a rebuild.
- New model version: rerun the full test set against the same acceptance criteria. Compare results side by side with the last validated run. Look at which cases changed, not just the overall score.
- Prompt or template change: rerun the cases the change could affect, plus a sample of the rest.
- New input type: add new test cases for it. The old set can't tell you about inputs it never contained.
- Wider intended use: this is really a new validation for the new use. Update the intended use statement first.
Watch the side-by-side closely. A new model can score the same overall while getting different cases wrong — maybe the ones that matter more. The average hides that.
When the vendor changes the model without asking
This is the one that worries quality teams most, and fairly so. Some AI services update the model behind the scenes. If that happens, your validated state has quietly expired and you didn't know.
Three things help. First, the contract: ask for version pinning, advance notice of changes, and a window to stay on the old version while you retest — covered in qualifying an AI vendor. Second, record the model version with every output, so you can always tell which version produced what. Third, monitoring: a small regular check against a few known cases will catch a silent change even if nobody tells you.
And if a vendor can't tell you which version you're running, that's your answer about whether you can use it for GxP work.
Keep the paper trail simple
For each change: what changed, the impact assessment, what you retested, the results next to the previous ones, and who approved it. That's it. Link it to the model version so an inspector tracing an old output can see exactly which validated version made it.
The last piece of the series is about keeping AI where it belongs in testing: why AI should draft your tests, not run them. For the full framework, see GxP AI.
Where to go next
Explore GxP Copilot for AI-native validation, TraceDraft for source-traceable clinical documentation, or book a demo to see either on your own data.
