A GxP DataOps architecture is not a single tool — it is a layered stack where every layer has a defined responsibility, a validation status, and an audit posture. Getting the layer boundaries right before implementation is the most consequential decision a life sciences data team makes, because the wrong boundary placement creates ALCOA+ gaps that are expensive to close after the fact.
The five-layer reference architecture
- Layer 1 — Acquisition. Raw data lands from instruments, LIMS, ELN, MES, or ERP into an immutable landing zone. Every record gets a content hash and a system-assigned timestamp at ingest. Nothing is transformed here. This is the ALCOA+ anchor point.
- Layer 2 — Staging & quality gate. Raw records are validated against schema contracts, instrument-range checks, and cross-system consistency rules. Failures are quarantined with a full audit record. Passes are promoted to the curated layer. Quality rules are versioned and change-controlled.
- Layer 3 — Curated / transformed. Passed records are enriched, joined across systems, and prepared for analytical consumption. Every transformation is a versioned, tested function. Lineage links every output record back to its source records in Layer 1.
- Layer 4 — Serving / analytical. Curated data is exposed to consumers — dashboards, submission assembly tools, reporting services. Read-only access. No modifications. Consumers receive a lineage pointer, not a data copy they can manipulate.
- Layer 5 — Audit & lineage. A dedicated, tamper-evident metadata store records every pipeline run, every quality gate result, every schema version, and every consumer access. This is the layer that an inspector actually reads.
Instrument-to-LIMS: the highest-risk boundary
The instrument-to-LIMS interface is where the majority of data integrity findings originate in FDA 483s and Warning Letters. Raw instrument data is often produced in proprietary binary formats (Waters .raw, Bruker .d, Thermo .raw) that require instrument-specific parsing before they can be ingested. The acquisition layer must parse the file, extract the result, hash the original binary, and write both to the immutable landing zone before any transformation occurs. Many teams write only the extracted result, losing the "original" element of ALCOA+. The raw binary must be retained and accessible. See data integrity (ALCOA+).
ELN-to-LIMS: lineage across system boundaries
Study data moving from an ELN (Benchling, Dotmatics, LabArchives) to a LIMS (LabWare, LabVantage, STARLIMS) typically crosses a REST API boundary. The architecture must log the exact API payload, the target system's acknowledgement, and the record IDs assigned at both ends — in both systems' audit trails. A message queue (Kafka, RabbitMQ) between the systems provides delivery guarantees and a replay capability when the target is temporarily unavailable. Without delivery guarantees, partial transfers create phantom gaps that are difficult to reconstruct under inspection.
LIMS-to-submission: the final mile
Data assembled for a regulatory submission — an IND, NDA, CTD module 2.7, or MAA — must carry its full provenance from source system to submission package. The submission assembly layer must be able to answer: which LIMS batch record does this analytical result come from, what instrument ran it, who reviewed it, and what was the audit trail entry at sign-off? These links must be preserved in the submission metadata, not just asserted in a cover letter. TraceDraft handles the clinical writing side of this with source-traceable claims — see TraceDraft.
Validation of the architecture under GAMP 5
Each layer of the architecture requires its own GAMP 5 classification and validation. Layer 1 (acquisition) is typically Category 4 — configurable, not custom; it requires OQ against each source system type. Layer 2 (quality gate) depends on how quality rules are authored: configuration-driven rules are Category 4; custom Python or SQL logic is Category 5 and requires unit testing and code review evidence. Layer 3 (transformation) follows the same split. Layers 4 and 5 are infrastructure and require IQ. GxP Copilot produces the validation packages for GxP systems feeding the architecture; our data integrity (ALCOA+) team validates the pipeline layers.
Change control for a live pipeline
A data pipeline in production is a validated system. Adding a new source system, changing a quality rule threshold, or updating a transformation function are all changes that require formal change control before deployment to production. The pipeline must have a staging environment where changes are tested against representative data before promotion, and the change record must include a test summary and a risk assessment. This is the discipline most DataOps teams underestimate — and the one that generates the most 483 findings when it breaks down. See Change controls.
Disaster recovery and data retention
GxP data retention requirements vary by regulation: 21 CFR Part 211 requires two years post-expiry for pharmaceutical manufacturing records; clinical trial data under ICH E6(R3) must be retained until at least two years post-marketing authorisation; EU GMP Annex 11 requires the computerised system to retain data for the entire retention period. The DataOps architecture must have a tested restore procedure — not just a backup policy. Every layer of the archive must be independently readable without the original application, because the application may be decommissioned before the retention period expires.
Where to go next
Explore GxP Copilot for AI-native validation, TraceDraft for source-traceable clinical documentation, or book a demo to see either on your own data.
