Most AI pilots in regulated companies fail in a specific and avoidable way. They succeed technically, everyone agrees the output is impressive, and then the programme stalls because nothing produced during the pilot can be used to support validation. The pilot proved capability and generated no evidence.
A well-run pilot does both. This is a week-by-week plan for ninety days: scoping so QA can approve it, defining success before you start, running in parallel with the existing process, and finishing with a package that converts into a validation effort rather than restarting one. Context for the wider programme is in GxP AI.
Before day one: three decisions
Pick one use case, on one process, with one team. The instinct is to pilot broadly to see where value lands. Resist it. A narrow pilot produces a clean comparison and a defensible evidence set; a broad one produces anecdotes. Good first candidates are high-volume, rule-governed, and have a verifiable output — the criteria set out in what to automate.
Decide the GxP posture explicitly. Either the pilot output is used for real regulated work — in which case validation obligations apply from day one and QA must be involved before it starts — or it runs strictly in parallel with the existing process as the record of truth. Both are legitimate. Ambiguity between them is not, and it is the most common reason a pilot cannot be built on afterwards.
Define success numerically, in writing, before starting. Not "see if it helps". Specify the metric, the current baseline, and the threshold that constitutes success. If you cannot state the baseline, measure it in week one — that measurement alone is often the most valuable output of the whole exercise.
Weeks 1–2: baseline and setup
- Measure the current process honestly. Elapsed time, active effort hours, rework rate, review cycles, defect rate at QA review. Sample enough instances to be credible — a handful is not a baseline.
- Assemble the evaluation set: real work items, including the awkward ones. A pilot run on clean inputs predicts nothing about production.
- Establish ground truth for that set. What is the correct output? Agreed by whom? Without this you cannot measure accuracy, only impression.
- Agree the data boundary with IT and QA and configure it. See data privacy in GxP AI.
- Write the pilot plan: scope, success criteria, roles, data handling, duration, decision gate. Two to four pages. Have QA approve it. This document is what makes the pilot a controlled activity rather than an experiment, and it is the first thing an auditor will ask for if pilot output ever touches regulated work.
Weeks 3–6: parallel running
Run both processes on the same work. The existing process remains the record of truth. The AI output is produced alongside and compared, not substituted.
- Capture every comparison. For each item: AI output, human output, differences, whether the difference was a defect or a preference, and the reviewer's time. This dataset is the pilot's real deliverable.
- Record errors precisely. Not just that it was wrong, but how — omission, incorrect content, correct but unsupported, misread source. Error taxonomy drives the acceptance criteria later and tells you where human review must concentrate.
- Measure verification time, not generation time. The question is not how fast the system produced a draft. It is how long a qualified person took to confirm it was right. If those are similar, the business case fails whatever the output quality.
- Rotate reviewers. One enthusiastic reviewer produces unrepresentative results in both directions. Include someone sceptical.
Expect weeks 3 and 4 to look worse than week 6. Early results reflect learning the tool as much as the tool itself. Do not judge at week 4.
Weeks 7–10: tune and stress
- Address the top error categories. Usually this is configuration, prompt or template work, or supplying better source material — not a limitation of the technology.
- Deliberately test the edges. Feed it the genuinely difficult cases and observe whether it flags uncertainty or answers confidently and wrongly. The second behaviour determines how much human review the workflow needs.
- Exercise the fallback. Turn the system off for a day and confirm the team can proceed. An untested fallback is not a control.
- Draft the SOP changes the production workflow would require. Doing this during the pilot surfaces accountability questions while you still have time to resolve them. See updating SOPs and training for AI.
- Start the validation thinking: what would the acceptance criteria be, what evidence do you already have, what is missing? Much of the pilot data maps directly onto validation evidence if you planned for it.
Weeks 11–13: decide and document
Write the pilot report against the success criteria you set in week zero. Report the result you got, not the result you hoped for. A pilot that produces a clear, evidenced no is a successful pilot — it cost you ninety days instead of a two-year programme.
The report should contain: what was piloted and on what scope, the baseline, the measured result against each criterion, the error taxonomy with frequencies, verification time analysis, limitations identified, data handling as operated, and a recommendation with its rationale.
If proceeding, the pilot hands the validation effort a substantial head start: a characterised evaluation set, established ground truth, known failure modes, measured performance, and a tested fallback. That is most of what a validation package needs — the stages are mapped in AI across the validation lifecycle, and AI validation services covers formalising it.
The failure mode to avoid at this gate: expanding scope before validating the first use case. A pilot that succeeds and then grows to five processes before any of them is validated produces an unvalidated system in routine regulated use — which is a worse position than not having piloted at all.
Where to go next
Explore GxP Copilot for AI-native validation, TraceDraft for source-traceable clinical documentation, or book a demo to see either on your own data.
