All insights

    Part of our work on insurance

    Insurance

    What a regional carrier should ask an AI underwriting vendor about its own historical bias

    Big Sky Consulting Group · October 5, 2026 · 7 min read

    The vendor will answer the wrong question very well

    You are three meetings into an AI underwriting evaluation. The vendor has a fairness slide. It describes their training data, their testing approach, and a governance framework with a name. Your compliance lead nods. The demo goes well.

    Nobody in the room has asked the question that decides whether this deal is safe to sign: what will this model do with your book?

    A regional carrier rarely buys a model cold. It buys a model that gets tuned, calibrated, or retrained on its own policy and loss history, because that is the whole selling point. The model learns how you write business. That is the feature. It is also the exposure, because your book is a record of every declination, every surcharge, and every territory your underwriters quietly stopped writing in 2014. A model trained on those decisions does not neutralize them. It learns them, scales them, and makes them look like math.

    The vendor's fairness slide is about the vendor's model. Your regulator is going to ask about your outcomes.

    Responsibility does not transfer with the invoice

    This is the part buyers most often get wrong, and the regulators have been unusually clear about it.

    The NAIC Model Bulletin on insurers' use of AI systems, now adopted in 25 states and the District of Columbia, puts responsibility for third-party AI squarely on the insurer. It expects diligence on the vendor, contract terms that include audit rights and cooperation with regulatory inquiries, and ongoing monitoring after go-live (NAIC bulletin summary, state adoption status). "The vendor won't show us the internals" is not an answer an examiner accepts.

    Colorado goes further. SB 21-169 and Regulation 10-1-1 require insurers using external consumer data, algorithms, and predictive models to run quantitative testing for unfair discrimination, document the methodology and results, remediate what they find, and monitor for drift. The amended regulation, effective October 15, 2025, extends that framework from life insurance to private passenger auto and health benefit plans (Colorado DOI, Faegre Drinker). The obligation sits with the insurer even when the model belongs to someone else (SB 21-169 overview).

    If you write only commercial lines in states that have not adopted the bulletin, you may feel some distance from all this. We would not lean on that distance. The direction of travel is one way, and a model you deploy in 2026 will still be in production when your state catches up.

    So the liability is yours. The question is whether your diligence matches it.

    Why "is your model fair" is the wrong question

    Ask a vendor whether their model is fair and you will get a confident yes, because the vendor is answering about the thing they control: their base model, their training corpus, their test suite. All of that can be genuinely sound and still produce a problem on your book.

    Here is the mechanism. Your historical data encodes your historical judgment. If your underwriters in a given period declined more small contractors in certain ZIP codes, or priced a class of older properties with a heavy hand that was never revisited, those patterns are in the outcomes the model learns from. Bias rarely shows up as a protected attribute. It shows up as a proxy: territory, building age, credit-adjacent signals, the kind of prior carrier an applicant came from. A model that never sees race can still reproduce a pattern that tracks it closely, which is exactly the thing Colorado's quantitative testing exists to detect.

    A vendor's generic fairness testing does not see your proxies, because it was not run on your data. That is not a vendor failing. It is the limit of what they could have tested before they met you.

    This is the general shape of the problem. Which parts apply to your process depends on answers only your systems can give.

    Put us on it, from $5,000

    The question that actually decides the deal

    The useful diligence question is comparative and specific:

    Show us what your model does differently from our underwriters, and on which segments.

    Run the model against a historical slice of your own book, a year or two your team considers representative, and compare its decisions to the decisions your people actually made. Where does it decline more? Where does it price higher? Where does it agree with you so completely that it has clearly learned your habits rather than the risk?

    Every one of those divergences is information. Some will be the model catching risk your underwriters missed, which is the business case. Some will be the model amplifying a pattern you would never defend to a regulator if you saw it written down. Some will be the model faithfully copying something you did not know you were doing. You need to know which is which before production, not after a market conduct exam.

    The tell is simple. A vendor who cannot produce this comparison has not run it. Some will offer to run it after contract signature, as part of implementation. That sequencing is backwards. By then the purchase is a sunk cost and every finding becomes a negotiation about who pays to fix it.

    We are not suggesting the comparison is easy. Deciding which period is representative, which segments to cut, which differences are material, and how to separate real risk signal from inherited habit is careful work, and it is specific to your book. But the willingness to do it before signing is the cleanest filter we know for separating vendors who understand regulated deployment from vendors who understand demos. It sits alongside the general vendor screens in our guide to how to evaluate AI vendors, and it is the one that matters most in this category.

    The second question is in the contract

    Assume the comparison looks acceptable. The next exposure is time.

    Models drift. Your book changes, the economy shifts, the vendor ships a new version. A model that was clean at go-live can stop being clean eighteen months later, and the Colorado framework explicitly expects monitoring for that drift. You cannot monitor what you cannot see.

    So the second question is whether the contract gives you what the NAIC bulletin assumes you have: the right to audit, the right to receive testing results, the vendor's cooperation when a regulator asks, and notice before a model change reaches your production decisions. In our experience these terms are rarely in a vendor's default paper. They have to be asked for, and the time to ask is while you are still a prospect rather than a customer.

    A vendor who resists audit rights is telling you something about the next three years of the relationship. Listen to it.

    What a regional carrier is actually buying

    It helps to be honest about the shape of the trade.

    A large carrier has a model risk function, an internal data science team, and counsel who has seen these contracts before. A regional carrier usually has a capable underwriting leader, a stretched compliance officer, and an IT team supporting a policy admin system older than some of its staff. The vendor knows this. The sales motion is built for it.

    That is why the buy-side questions matter more here, not less. The regional carrier carries the same regulatory responsibility as the national one with a fraction of the internal capacity to check the work. The answer is not to avoid AI underwriting. Plenty of regional books would genuinely benefit from faster, more consistent risk selection. The answer is to make the vendor do the proving before the contract, while you still have leverage, rather than discovering in an exam what the model learned from you.

    There is a familiar pattern underneath this. A model faithfully reproducing decisions nobody wants to defend is the same failure we described in the renewal process that quietly determines your loss ratio: a selection effect running without anyone choosing it. The tool changes. The blind spot does not. And, like the triage problem in MGA submission triage, speed applied to an unexamined decision is how a manageable issue becomes an expensive one on a delay long enough that nobody connects the two.

    You could say an underwriting model with unexamined history has a pre-existing condition. We would just prefer you find out before it is your policy.

    Where this stops being an article

    An article can tell you which two questions decide the deal. It cannot tell you which segments of your book carry the inherited patterns, what a defensible comparison period looks like for your lines, or which contract terms your current vendor shortlist will actually concede.

    If you are evaluating an AI underwriting vendor and nobody on your side has yet asked what the model does differently from your own underwriters, talk to us before you sign. We sit on the buyer's side of that table, and we have no model to sell you.

    AI underwritingalgorithmic biasvendor due diligenceNAIC model bulletinregional carriers

    Want this applied to your business?

    A paid working session on one of your processes, ending in a written implementation plan. Not a sales call. Consults start at $5,000, and you see the price before you commit.

    Book a consult