Supplier qualification was designed for software that behaves the same way every time you run it. AI does not necessarily do that, and the standard questionnaire does not ask about the things that matter. A vendor can answer every question on your existing form truthfully and still leave you unable to defend the system in an inspection.
This guide covers what to ask an AI vendor, what evidence to demand rather than accept on assertion, how to assess a model you cannot inspect directly, and the contract terms that determine whether you can keep operating when something changes. It assumes the general context set out in GxP AI.
What the standard questionnaire misses
A conventional supplier assessment asks about the vendor's quality system, development lifecycle, testing practice, change control and security. All still necessary. None of it addresses the four questions specific to AI:
- Does the system's behaviour change without a release? If the underlying model can be updated by the vendor or a third party without a versioned, notified change, your validated state has an expiry date you do not control.
- Is the output reproducible? Can the vendor demonstrate that the same input produces the same output, and if not, what is the bounded variation? "Approximately the same" is not a validation acceptance criterion.
- Where does the reasoning come from? Can a reviewer see why the system produced this output — the source passages, the rule applied, the confidence — or only the output itself?
- What happens to your data? Processing boundary, retention, and whether your content is used to improve the vendor's models. Covered in depth in data privacy in GxP AI.
The questions to ask, and what a good answer sounds like
On change control
"If the model behind this feature is updated, how am I notified and what are my options?"
A good answer describes versioning, advance notice with a defined period, release notes that state what changed in behavioural terms, and the ability to remain on a prior version for a stated window while you reassess. A weak answer is "we continuously improve the model" — which means your validated state can change silently, and no validation package survives that.
On reproducibility
"Show me the same input run ten times."
Ask for this in the evaluation, with your own documents. Some variation in phrasing may be acceptable for a drafting task where a human approves the result. Variation in a classification, a risk score or a pass/fail determination is a different matter. Establish which of your use cases tolerate variance and which do not, then test exactly those.
On evidence and traceability
"For a given output, what can a reviewer see about how it was produced?"
You want citation to source, a record of the inputs used, the version of the system that produced it, a timestamp, and the identity of the person who reviewed it — all captured in a tamper-evident audit trail. If the vendor can only show you the answer, your reviewer cannot perform a meaningful check and your human oversight is ceremonial.
On performance claims
"What is the measured accuracy, on what data, and how was it measured?"
A credible answer names an evaluation set, describes how ground truth was established, and reports failure modes as well as success rates. A vendor who cannot tell you how their system fails has not characterised it. Marketing accuracy figures with no stated methodology should be treated as unsupported — the pattern examined in AI washing.
On the regulatory position
"Which specific requirements does your evidence address, and which remain mine?"
The right answer is a clear division, not a claim of total coverage. A vendor asserting their product is "fully compliant" and needs no validation is describing something that cannot exist — compliance depends on your configuration, your data and your process.
Assessing a model you cannot inspect
You will usually not be given model weights or architecture, and demanding them is rarely productive. You do not need them. What you need is behavioural evidence, and that you can generate yourself.
- Test with your own material. Vendor demos use curated inputs. Bring your genuinely messy documents — the badly scanned one, the one with an inconsistent template, the one written by someone who left. Performance on those is the performance you will get.
- Test the failure path deliberately. Give it something outside its competence and observe whether it declines, flags low confidence, or answers confidently and wrongly. The third behaviour is the dangerous one and it is the one to probe for.
- Test the review experience, not just the output. Time how long it takes a qualified reviewer to verify a generated document against source. If verification takes as long as authoring, the efficiency case does not hold regardless of output quality.
- Test at your data boundary. Documents containing product names, patient identifiers or commercially sensitive material — confirm the handling matches the contract before, not after, you sign it.
This behavioural approach is what AI validation services formalise, and it is the same evidence base your eventual validation package draws on.
Contract terms that matter later
- Notice period for model changes, with a right to remain on the prior version for a defined window.
- An explicit statement that your data is not used for model training, if that is your requirement — and it usually is. Silence is not a commitment.
- Audit rights. The right to audit the vendor, or to receive a recent independent audit report, in a form your QA will accept.
- Data export in a usable, documented format. Your records must outlive the relationship, and retention obligations are long — see GxP archiving and data retention.
- Subprocessor disclosure and notification. If your content is processed by a party the vendor contracts with, you need to know who and where.
- Incident notification within a defined period, covering security events and material defects affecting output correctness.
The qualification decision is ultimately about whether you can defend the system without the vendor in the room. If every answer to an inspector's question begins "we would need to ask the supplier", the qualification has not done its job. What inspectors ask about AI is a useful test to run against any vendor before signing.
Where to go next
Explore GxP Copilot for AI-native validation, TraceDraft for source-traceable clinical documentation, or book a demo to see either on your own data.
