The shift from Computer System Validation (CSV) to Computer Software Assurance (CSA) is the most significant change in GxP validation methodology in two decades. For traditional software, CSA reduces unnecessary testing burden while maintaining compliance. For AI systems, the choice between CSA and CSV is not just about efficiency — it fundamentally changes what defensible validation looks like. This article explains the difference, where each applies to AI systems, and how to document the decision in a way that survives inspection.
CSV: what it is and why it struggles with AI
Computer System Validation, rooted in the 2002 FDA guidance and GAMP frameworks, follows a lifecycle approach: define requirements, design the system, test against requirements with scripted protocols, document the results, and maintain the validated state through change control. It works well for deterministic software where the same input reliably produces the same output.
CSV struggles with AI for three reasons. First, scripted testing with expected outputs does not work for non-deterministic systems. You cannot write an OQ protocol that says "given this input, the system shall produce exactly this output" when the AI may produce a different — but equally valid — output each time. Second, CSV treats validation as a point-in-time event; AI systems require continuous validation as models drift. Third, CSV's emphasis on exhaustive scripted testing for every requirement creates an enormous test burden for AI systems with complex, high-dimensional input spaces that cannot be exhaustively tested.
CSA: what it is and why it fits AI better
Computer Software Assurance, formalised in FDA's 2023 draft guidance, replaces the CSV emphasis on scripted testing with a risk-based emphasis on assurance activities. The core principles:
- Risk-based testing. Focus testing effort on high-risk functions. Low-risk functions can be assured through vendor evidence, ad-hoc testing, or documentation review rather than scripted IQ/OQ/PQ protocols.
- Critical thinking over checklists. Testers are expected to apply domain expertise and critical thinking rather than mechanically executing test scripts. This is directly relevant to AI, where the tester must evaluate whether an AI output is "good enough" rather than whether it matches an exact expected value.
- Unscripted testing. CSA explicitly allows unscripted (ad-hoc, exploratory) testing as a legitimate assurance activity. For AI systems, this means testers can evaluate AI outputs against professional judgment rather than predefined expected results.
- Leveraged vendor evidence. CSA allows using vendor-provided testing and validation evidence to reduce the customer's testing burden. For commercial AI platforms, this means the platform vendor's own validation evidence can be leveraged rather than duplicated.
Where CSA works for AI and where CSV still applies
CSA is the right approach for:
- AI document generation (validation deliverables, clinical documents, SOPs) — outputs are reviewed by experts; unscripted testing evaluates output quality against professional standards.
- AI-assisted risk classification — risk-based testing focuses on high-risk classification decisions; low-risk classifications can be assured through sampling and ad-hoc review.
- AI-assisted search and summarisation — low-risk informational use cases where the primary assurance is human review of the output.
- AI monitoring and alerting — medium-risk systems where testing focuses on alert accuracy and false-negative rates.
CSV (or CSV-depth testing within a CSA framework) is appropriate for:
- AI process control systems — high-risk systems that directly control manufacturing parameters require exhaustive testing of safety boundaries, fail-safe modes, and edge cases.
- AI batch release decisions — any system that influences whether a batch is released to patients requires the deepest level of validation evidence, including scripted protocols for critical decision points.
- AI systems that create or modify regulated data — where data integrity requirements (ALCOA+, Part 11) demand reproducible and verifiable system behaviour.
Documenting the CSA vs CSV decision for inspection
The decision to use CSA or CSV for an AI system must be documented and defensible. The documentation should include:
- A risk assessment of the AI system that classifies each function by its impact on product quality, patient safety, and data integrity. AI Risk Assessment in GxP Copilot produces this classification.
- A rationale for the validation approach chosen for each risk level — why CSA unscripted testing is appropriate for low-risk functions, and why scripted testing is used for high-risk functions.
- The specific assurance activities used for each risk level — vendor evidence leveraged, unscripted testing performed, scripted protocols executed.
- The acceptance criteria used to evaluate AI outputs — for CSA, these are often qualitative ("the output meets professional standards and contains all required sections") rather than quantitative ("the output matches the expected string").
This documentation becomes part of the Validation Plan and is referenced in the Validation Summary Report. GxP Copilot generates both with the CSA rationale built into the template structure.
The hybrid approach most organisations should use
In practice, most organisations should use a hybrid approach: CSA as the overall framework, with CSV-depth testing for high-risk AI functions and CSA-appropriate unscripted testing for lower-risk functions. The risk classification determines the boundary. This hybrid approach produces less testing volume than pure CSV (because low-risk functions get lighter assurance) while producing stronger evidence than pure CSA for high-risk functions (because critical decisions get scripted, reproducible tests). GxP Copilot implements this hybrid approach by default: the AI Risk Assessment classifies each requirement by risk, and the test generation engine sizes the test effort accordingly — scripted protocols for high-risk, structured unscripted testing for medium-risk, and leveraged vendor evidence for low-risk. book a demo to see this in practice on your own requirements.
Where to go next
Explore GxP Copilot for AI-native validation, TraceDraft for source-traceable clinical documentation, or book a demo to see either on your own data.
