AI Governance

Your Data Probably Isn't Ready for AI. Here's How to Tell

The step nobody budgets for. How to judge whether your deviations, batch records and documents can actually support an AI project, what to fix first, and why a model trained on inconsistent history learns the inconsistency.

2026-10-02Cybroscape Technologies11 min read
Key takeaway

The step nobody budgets for. How to judge whether your deviations, batch records and documents can actually support an AI project, what to fix first, and why a model trained on inconsistent history learns the inconsistency.

Most AI projects in regulated companies stall at the same place, and it is never the model. It is the moment someone exports two years of deviations and discovers that the same root cause is written eleven different ways, half the records are free text, and the category field was clearly filled in by whoever was closest to the keyboard.

Data readiness is the step nobody budgets for and the one that decides whether the project works. Here is how to judge where you stand before you buy anything.

What 'ready' actually means

Not perfect. Four properties, and you can assess each in an afternoon.

  • Consistent. The same thing is recorded the same way. If "equipment failure", "equip. fault" and "machine broke" all appear, a model will treat them as three different things — and so will your trend analysis.
  • Structured enough. Not everything needs to be a dropdown, but the fields you want to reason over should not live inside a paragraph of narrative.
  • Complete where it matters. A field populated 40% of the time cannot support a prediction, however important it is.
  • Linked. Can you get from a deviation to the batch, the equipment, the operator role and the resulting CAPA without a human joining them by hand? Most value comes from the links, not the records.

A one-day assessment

Pull the last two years of whichever records the project depends on and answer six questions honestly:

  • How many distinct values appear in the category field, and how many should there be?
  • What share of records have the key fields populated?
  • Pick twenty records at random — could a colleague who was not there understand what happened?
  • Do the same event types get recorded differently by different sites or shifts?
  • Can you join the records to related data programmatically?
  • How far back does the current format go, and when did it last change?

That last question catches people out. If your deviation form changed eighteen months ago, you do not have two years of comparable data; you have eighteen months, and a different dataset before it.

The part that actually matters

A model learns the patterns in your history, including the ones you are not proud of. If a large share of your deviations closed as "operator error, retrained", a system trained on them will propose operator error, fluently and consistently, and give your worst habit the appearance of analytical support.

So before any AI project touching investigations, look at your own closure patterns. That exercise is worth doing whether or not you ever buy anything — see AI for deviation investigations and CAPA.

The same applies to recruitment data in trials, release decisions, supplier scoring — anywhere the historical record encodes a judgement you would make differently today.

What to fix, and what not to

Fix going forward, not backward. Cleaning up five years of historical free text is rarely worth it. Tightening the form so the next two years are consistent usually is, and it costs a fraction as much.

Fix the field the project needs, not the whole system. Readiness is per use case. Your batch data can be excellent while your supplier data is unusable.

Do not let readiness become a reason never to start. Some uses tolerate messy data well — retrieval, summarisation, finding similar past cases. Those are good first projects precisely because they work on imperfect records and surface the inconsistencies as a by-product.

If you are scoping a first project, running a 90-day pilot covers the structure, and building the test set covers the data you will need to prove it works. Wider context in GxP AI.

Where to go next

Explore GxP Copilot for AI-native validation, TraceDraft for source-traceable clinical documentation, or book a demo to see either on your own data.

ai data readinessdata quality for ai pharmais our data ready for aimaster data life sciencesgxp ai

Frequently Asked Questions

How do you know if your data is ready for AI?+

Four properties, each assessable in an afternoon: consistent (the same thing recorded the same way), structured enough (the fields you want to reason over are not buried in narrative), complete where it matters (a field populated 40% of the time cannot support a prediction), and linked (you can get from a deviation to the batch, equipment and CAPA without joining by hand).

What should a data readiness assessment look at?+

Pull two years of the relevant records and ask: how many distinct category values exist versus how many should; what share of key fields are populated; could a colleague understand twenty random records; do sites or shifts record the same events differently; can you join records programmatically; and when did the form last change — because a change eighteen months ago means you have eighteen months of comparable data, not two years.

What is the biggest risk in poor historical data?+

The model learns your habits. If many deviations closed as 'operator error, retrained', a system trained on them will keep proposing operator error, fluently and consistently, giving your worst habit the appearance of analytical support. Check your own closure patterns before any project touching investigations.

Should you clean historical data before starting?+

Usually not. Cleaning five years of free text is rarely worth it; tightening the form so the next two years are consistent usually is, at a fraction of the cost. Fix the field the project needs rather than the whole system — readiness is per use case.

Can you start if the data is imperfect?+

Yes, with the right first project. Retrieval, summarisation and finding similar past cases tolerate messy records well, and they surface the inconsistencies as a by-product — which makes them good first projects rather than reasons to wait.

Next step

Bring a system. We'll show you the package.