AI Governance

Data governance before AI governance: the foundation insurers keep skipping

Explainability is a data property before it is a model property. The targeted data foundation insurers should build before scaling AI decisions.

By Jonas Osman AbdelghafourPublished 29 July 2026

Walk into any insurer's AI steering committee and you will find a mature conversation about models: which underwriting engine to buy, which claims triage system to pilot, which vendor's fraud detection to trust. Ask instead where the data feeding those systems lives, who owns its quality, and whether the same customer exists once or five times across core systems — and the conversation gets quieter. The industry is building AI governance on top of information architectures that were never governed in the first place.

This ordering error is becoming the sector's most predictable compliance failure. Regulators from the UK's FCA — which expects AI to reshape financial services by 2030 — to US state insurance departments are converging on demands for transparency and explainability in AI-driven decisions. But no institution can explain a decision built on data it cannot trace. Explainability is a data property before it is a model property.

Why insurance data is uniquely hard Insurers are, in effect, hundred-year-old data companies with hundred-year-old data. Policy administration systems from different eras, books of business absorbed through acquisitions, claims histories in formats designed for microfiche, and third-party data appended at every stage: the result is information scattered across disconnected systems with inconsistent definitions of even basic entities like "policyholder" and "claim". Machine learning does not clean this up. It launders it — transforming inconsistency into confident-looking predictions whose defects surface only when a regulator or plaintiff asks how a specific decision was reached.

The bias-testing gap illustrates the stakes. Industry surveys suggest a substantial share of insurers using AI do not regularly test models for unfair discrimination. Part of that is process immaturity, but a deeper cause is data immaturity: you cannot test for disparate impact across groups you cannot reliably identify in your own records, and you cannot fix a biased model whose training data no one can reconstruct.

The foundation, in buildable pieces The remedy is not a five-year data warehouse program. It is targeted governance around the data that feeds consequential decisions. Start with lineage for every AI system in production: what data goes in, from which systems, transformed how. Assign named owners for the critical data domains — customer, policy, claims — with authority over definitions and quality thresholds. Establish quality metrics that are measured, published internally, and tied to remediation budgets. And gate new AI deployments on a data-readiness review, so the organization stops launching models on foundations it knows are cracked.

Insurers that sequence it this way discover something pleasant: data governance done for AI pays for itself elsewhere — in reserving accuracy, in reinsurance negotiations, in regulatory reporting. Trusted information is not a compliance cost. It is the asset the entire AI strategy was always going to depend on.

Related reading - [From principle to proof: bias testing in AI underwriting and claims](/insights/ai-bias-testing-underwriting-claims-jonas-osman-abdelghafour) - [EIOPA's AI opinion is the Rosetta Stone between Solvency II and the AI Act](/insights/eiopa-ai-opinion-solvency-ii-ai-act-jonas-osman-abdelghafour) - [The NAIC's AI evaluation pilot is the new exam playbook](/insights/naic-ai-evaluation-pilot-insurers-jonas-osman-abdelghafour)

See also Insurance Risk, Governance, Risk and Compliance and Regulatory Compliance.

About the author Jonas Osman Abdelghafour writes about AI governance, risk management, and regulation in insurance and banking. Follow Jonas Osman Abdelghafour for analysis of how supervisors, carriers, and banks are adapting to artificial intelligence.

Frequently asked questions

Why insurance data is uniquely hard?

Insurers are, in effect, hundred-year-old data companies with hundred-year-old data. Policy administration systems from different eras, books of business absorbed through acquisitions, claims histories in formats designed for microfiche, and third-party data appended at every stage: the result is information scattered across disconnected systems with inconsistent definitions of even basic entities like "policyholder" and "claim". Machine learning does not clean this up. It launders it — trans...

What should risk leaders know about the foundation, in buildable pieces?

The remedy is not a five-year data warehouse program. It is targeted governance around the data that feeds consequential decisions. Start with lineage for every AI system in production: what data goes in, from which systems, transformed how. Assign named owners for the critical data domains — customer, policy, claims — with authority over definitions and quality thresholds. Establish quality metrics that are measured, published internally, and tied to remediation budgets. And gate new AI depl...

What should risk leaders know about related reading?

See also [Insurance Risk](/expertise/insurance-risk), [Governance, Risk and Compliance](/expertise/grc) and [Regulatory Compliance](/expertise/regulatory-compliance).

What should risk leaders know about about the author?

Jonas Osman Abdelghafour writes about AI governance, risk management, and regulation in insurance and banking. Follow Jonas Osman Abdelghafour for analysis of how supervisors, carriers, and banks are adapting to artificial intelligence.