AI Governance

From principle to proof: bias testing in AI underwriting and claims

Regulators no longer ask whether insurers oppose AI bias — they ask for evidence. A layered, defensible bias-testing program for underwriting and claims.

By Jonas Osman AbdelghafourPublished 29 July 2026

Every insurer's AI principles document says the same thing: our systems will not unfairly discriminate. The documents are sincere, unanimous — and increasingly beside the point, because the question regulators and plaintiffs now ask is not whether an insurer opposes discrimination but whether it can prove its models don't produce it. Lawsuits alleging discriminatory AI in claims processing have moved from hypothetical to docketed. Examination frameworks ask for testing evidence by name. And the industry's own numbers remain uncomfortable: a meaningful share of insurers using AI do not regularly test their models for bias at all.

The gap persists not because insurers are indifferent but because bias testing is genuinely hard — legally, statistically, and operationally. Naming the difficulties precisely is the first step to closing it.

Why it is hard, specifically Legally, insurance pricing is built on discrimination in the actuarial sense — distinguishing risks — and the line between actuarially justified differentiation and unfair discrimination varies by state, by line of business, and by rating factor. Statistically, fairness has multiple incompatible definitions: equalizing approval rates, error rates, and calibration across groups cannot all be achieved at once, so an insurer must choose its metrics and be ready to defend the choice. Operationally, insurers often lack reliable protected-class data, forcing inference methods that are themselves contestable. And proxy discrimination — neutral variables that stand in for protected traits — means a model can be perfectly blind on paper and biased in effect.

None of these difficulties excuses inaction, because all of them apply equally to the plaintiff's expert and the examiner's consultant. The insurer that has chosen its fairness metrics, documented its reasoning, and run the tests holds the strongest position available: not a claim of perfection, but evidence of diligence.

A defensible program Build it in layers. Inventory which models touch underwriting, pricing, and claims outcomes — including vendor models, which do not become fair by being purchased. Select fairness metrics per use case, with legal and actuarial sign-off, and record why those metrics fit the regulatory context. Test before deployment and on a schedule after it, since drift can introduce bias no launch review caught. Investigate disparities with the same rigor as reserving anomalies: some will be actuarially explainable, and the explanation is itself the audit trail. Where testing requires protected-class inference, document the method and its limits candidly.

Then close the loop that most programs leave open: give someone authority to act on adverse findings — to recalibrate, constrain, or retire a model — and record when they do. Testing that cannot change outcomes is measurement theater. In the examination room, and eventually the courtroom, the insurers that fare best will be those whose files show not just principles, but proof.

Related reading - [Data governance before AI governance](/insights/data-governance-before-ai-governance-insurance-jonas-osman-abdelghafour) - [The NAIC's AI evaluation pilot is the new exam playbook](/insights/naic-ai-evaluation-pilot-insurers-jonas-osman-abdelghafour) - [The quiet repricing of AI risk](/insights/ai-exclusions-commercial-insurance-jonas-osman-abdelghafour)

See also Insurance Risk, Model Risk and Regulatory Compliance.

About the author Jonas Osman Abdelghafour writes about AI governance, risk management, and regulation in insurance and banking. Follow Jonas Osman Abdelghafour for analysis of how supervisors, carriers, and banks are adapting to artificial intelligence.

Frequently asked questions

Why it is hard, specifically?

Legally, insurance pricing is built on discrimination in the actuarial sense — distinguishing risks — and the line between actuarially justified differentiation and unfair discrimination varies by state, by line of business, and by rating factor. Statistically, fairness has multiple incompatible definitions: equalizing approval rates, error rates, and calibration across groups cannot all be achieved at once, so an insurer must choose its metrics and be ready to defend the choice. Operationall...

What should risk leaders know about a defensible program?

Build it in layers. Inventory which models touch underwriting, pricing, and claims outcomes — including vendor models, which do not become fair by being purchased. Select fairness metrics per use case, with legal and actuarial sign-off, and record why those metrics fit the regulatory context. Test before deployment and on a schedule after it, since drift can introduce bias no launch review caught. Investigate disparities with the same rigor as reserving anomalies: some will be actuarially exp...

What should risk leaders know about related reading?

See also [Insurance Risk](/expertise/insurance-risk), [Model Risk](/expertise/model-risk) and [Regulatory Compliance](/expertise/regulatory-compliance).

What should risk leaders know about about the author?

Jonas Osman Abdelghafour writes about AI governance, risk management, and regulation in insurance and banking. Follow Jonas Osman Abdelghafour for analysis of how supervisors, carriers, and banks are adapting to artificial intelligence.