A GxP data integrity pipeline is a validated data flow that keeps ALCOA+ evidence — attributable, legible, contemporaneous, original, accurate, complete — intact from the moment data is generated by an instrument or system through every transformation, every system boundary, and every downstream consumer. Most teams have validated the systems at either end of the pipeline. Few have validated the pipeline itself. That gap is where inspectors find data integrity failures.
The ALCOA+ requirements mapped to pipeline stages
- Attributable — acquisition layer. Every record must carry the identity of the system or person that created it, and that identity must be fixed at the moment of creation — not added later. Service accounts used for system-to-system transfers must be non-shared and auditable.
- Contemporaneous — landing zone. The system-assigned ingest timestamp must be written before any transformation occurs and must come from a synchronised, monitored time source. Records without a controlled timestamp are not contemporaneous.
- Original — binary retention. The raw instrument file (Waters .raw, Bruker .d, Agilent .D directory) must be retained alongside the extracted result. Retaining only the extraction is a loss of originality that generates 483 findings.
- Accurate — quality gates. Automated checks validate that extracted values fall within instrument-calibrated ranges, that transcriptions are arithmetically consistent, and that cross-system identifiers match. Failures quarantine the record.
- Complete — delivery guarantees. Message queues with at-least-once delivery and idempotent receivers ensure no record is silently lost at a system boundary. Completeness checks compare expected record counts against received record counts for every batch.
The three pipeline failure modes FDA finds most often
Based on published 483s and Warning Letters, three failure modes account for the majority of data integrity pipeline findings. First: the instrument data file is overwritten or deleted after extraction, losing the original. Second: the integration between systems uses a shared service account whose actions cannot be attributed to a responsible individual when a discrepancy appears. Third: the pipeline has no record of failed or rejected records — they are simply dropped, making the dataset appear complete when it is not. Each of these is preventable with deliberate pipeline design rather than remediation after the fact.
Immutable storage: the technical foundation
The foundation of a compliant data integrity pipeline is an immutable storage layer where records can be appended but never modified or deleted by any application path. In cloud environments, this is implemented with object storage configured for Write-Once-Read-Many (WORM) or object lock. On-premise, it requires a combination of filesystem permissions, database row-level security, and application-layer enforcement — none of which is sufficient alone; all three must be present. The Audit Readiness approach in GxP Copilot applies the same principle to validation documents: signed versions are immutable, and corrections create new versions.
Automated quality gates: design and validation
Quality gates are the automated equivalent of a human reviewer who checks a result before it enters the system of record. They must be designed with the same rigour as any other validated process: a specification of what is being checked and why, a set of acceptance criteria, and test cases that prove the gate catches known failures and passes known good records. Quality gate logic must be versioned and change-controlled — a threshold change from 99% to 98% purity is a change to a validated quality control process, not a configuration tweak. And the result of every gate execution — pass or fail, timestamp, record ID — must be written to the audit trail.
Lineage as the backbone of regulatory defence
When a regulatory submission contains a data point that is later questioned, the only defensible response is a complete lineage trail: this result came from instrument run X, performed by analyst Y on date Z, extracted by pipeline version A.B.C, checked by quality gate version D.E.F, transformed by function G on date H, and consumed by submission assembly job I at timestamp J. Without that lineage, the answer is "trust us" — which is not a regulatory posture. Lineage is not expensive to build when it is designed into the pipeline from day one. It is very expensive to reconstruct after the fact.
Validating the pipeline under GAMP 5
The pipeline itself must be validated under GAMP 5. The validation approach depends on how much of the pipeline is custom: configuration-driven orchestration with pre-validated connectors is Category 4; custom transformation code is Category 5 and requires unit test evidence, code review records, and source control history. The validation package includes a URS for the pipeline, an interface risk assessment for each source and destination system, IQ for the infrastructure, OQ for each transformation and quality gate, and PQ for end-to-end flows with real representative data including known-good and known-bad test cases. GxP Copilot produces the GAMP 5 packages; our data integrity (ALCOA+) team handles pipeline-layer validation.
Where to go next
Explore GxP Copilot for AI-native validation, TraceDraft for source-traceable clinical documentation, or book a demo to see either on your own data.
