A green tick is not evidence. It is a claim that evidence exists somewhere. The gap between those two things is where automated validation gets into trouble.
The standard to aim for is not "the tool recorded a result". It is: could a competent person, three years from now, read this record and tell what was tested, against what, and on whose authority — without asking anyone who was there.
The record every step should leave
Per step, not per run. This is the distinction that matters, because a run-level summary cannot tell you where something went wrong.
- The step as approved — its text, from the approved protocol version, not a paraphrase generated at run time.
- Expected and actual, both recorded, both verbatim. A step that records only "pass" has thrown away the thing that made it a pass.
- A contemporaneous timestamp, at the step. ALCOA+ asks for contemporaneous; a single timestamp for a forty-minute run is not.
- Captured evidence where the step type implies it — a screenshot for a UI assertion, the log region for a service check.
That is the per-step floor. Everything below is run-level context that makes the steps interpretable.
The run-level context
- Which protocol, at which version, and the signature that approved it. A result against "the OQ protocol" is ambiguous the moment the protocol is revised.
- Which system, at which version, in which environment. Test evidence from an environment nobody can identify is not evidence.
- Who initiated the run, and under what authority.
- Start and end, and the overall outcome.
- The tool and its version — relevant because the tool is itself qualified, and a result is only as good as the qualified version that produced it.
Immutability, and what it really requires
Under FDA 21 CFR Part 11 compliance, the record has to be protected from alteration, and "we do not allow editing in the UI" is not the same thing as the record being protected.
Three properties worth checking in any tool, including one you built:
- Results are written once. A completed run cannot be re-opened and amended. A correction is a new record that references the old one.
- Evidence is bound to the step. Screenshots and logs are stored so that they cannot be swapped for others without it being visible.
- The audit trail covers the record itself, not only the system under test — who viewed, exported, or signed a result, and when.
A quick way to test your own system: ask how you would correct a result entered against the wrong protocol version. If the answer involves editing the record, the record is not protected.
The readability test
All of the above can be satisfied by a system that produces something nobody can read. Evidence is for a human being, usually one under time pressure who was not involved.
So the last check is editorial rather than technical: print one execution record and give it to a colleague who was not on the project. If they can follow what was tested and what happened without a guided tour, the format works. If they need you to explain the columns, it does not — and an inspector will not ask you to explain them, they will simply write down that the records were difficult to follow.
Where to go next
Explore GxP Copilot for AI-native validation, TraceDraft for source-traceable clinical documentation, or book a demo to see either on your own data.
