AI Software

AI-Washing in GxP Software: Eight Questions That Separate Real AI From a Compliance Chatbot

The eight questions that expose whether a GxP vendor's AI actually does the work — or is a thin wrapper around an LLM bolted onto a legacy workflow engine.

2026-08-04Cybroscape Technologies11 min read
Key takeaway

The eight questions that expose whether a GxP vendor's AI actually does the work — or is a thin wrapper around an LLM bolted onto a legacy workflow engine.

AI-washing is the practice of attaching the label "AI" to a software product without the product actually using AI to do the work the label implies. In consumer software, AI-washing is a marketing problem. In GxP software, it is a compliance problem: if a regulated team selects a platform on the strength of AI capabilities it does not actually have, the validation programme for that platform is built on a false assumption. Eight questions expose AI-washing in GxP software before you commit to a platform.

Question 1: What does the AI actually generate?

A genuine AI-native platform generates regulated-grade content — validation documents, risk assessments, test cases, clinical narratives — from structured inputs. An AI-washed platform uses AI for search, tagging, classification, or routing: tasks that improve user experience but do not reduce the primary work of the regulated function. Ask the vendor to demonstrate the AI generating a complete IQ protocol from a URS. If they cannot, the AI is not a drafter — it is a search engine with a modern interface.

Question 2: Can you show me the model card?

EU Annex 22 requires published governance for AI used in GxP environments. A model card documents: what AI model is used, what it was trained on, what it is approved for use on, what it is explicitly not approved for, and what safeguards constrain its outputs. An AI-native vendor publishes this today — not as a roadmap item. An AI-washed vendor either does not have one or presents a generic LLM policy that does not address the regulated domain. See Annex 22 assurance.

Question 3: What is the benchmark and who ran it?

A credible GxP AI platform has a frozen, versioned evaluation benchmark — a set of representative inputs and expected outputs that tests the AI against the regulated domain. The benchmark gates every model release: if the new model does not meet the benchmark, it does not ship. Ask the vendor: what is your benchmark, what score does the current model achieve, and who ran the evaluation? "Our AI produces high-quality outputs" is not a benchmark. A numbered test suite with documented pass criteria is.

Question 4: How is a model update handled as a change?

An AI model update changes the behaviour of a validated system. Under GAMP 5 and Annex 22, that is a change control event. The vendor must run the model against the benchmark, produce a test summary, assess the impact on existing validated outputs, and communicate the change to customers before it goes live. Ask the vendor: walk me through your last AI model update — what was the change control process, what was the benchmark result, and how were customers notified? An AI-washed vendor will not have a clear answer.

Question 5: Is the HITL policy enforced by the workflow or by the SOP?

Human-in-the-loop control is required for AI in regulated environments. But there is a critical difference between "our SOP says humans must review AI outputs" and "our workflow enforces that no AI-generated document can be signed without a qualified reviewer who is not the author." The former can be bypassed; the latter cannot. Ask the vendor to show you what happens if a user tries to approve their own AI-drafted document. If the system allows it, the HITL policy is on paper only. See Human-in-the-Loop reviews.

Question 6: What is the fallback when the AI is unavailable?

Annex 22 requires a documented fallback — the regulated process must keep working when the AI is offline or produces an output below the quality threshold. In practice, this means the platform must function without the AI: users can draft documents manually, the workflow continues, the audit trail keeps writing. An AI-native platform designs for this from day one. An AI-washed platform that treats AI as the primary path often has no graceful degradation mode.

Question 7: Is the audit trail of AI interaction tamper-evident?

The audit trail for AI-generated content must record what the AI produced, who reviewed it, what changes the reviewer made (and why), and who approved the final version. That record must be tamper-evident — no path in the application should be able to retroactively clean up an AI output that the reviewer corrected. Ask the vendor: show me the audit trail entry for an AI-generated document that a reviewer substantially edited. If the original AI output is not preserved in the record, the audit trail does not meet the standard regulators are moving toward. See Audit Readiness.

Question 8: What is the GAMP 5 classification and can I see the validation package?

GxP AI software is itself a GxP system and must be validated by the customer under GAMP 5. A credible vendor has done the classification work, can articulate whether their product is Category 4 or 5 (and why), and provides a customer-facing validation package — at minimum an IQ/OQ protocol template and pre-written test cases. An AI-washed vendor may not have a GAMP 5 position at all, or may claim "the AI is SaaS so it does not need validation" — which is incorrect. The AI capability is a configurable function of the software system and requires the same validation rigour as any other configurable element.

Where to go next

Explore GxP Copilot for AI-native validation, TraceDraft for source-traceable clinical documentation, or book a demo to see either on your own data.

gxp ai software evaluationai washing life sciencesgxp ai platform evaluationcompliance ai software real vs fake
Next step

Bring a system. We'll show you the package.