AI Operations Roadmap for Insurance: From Mandate to Measurable Impact
Published · Wendy Kinney
Last updated
Published · Wendy Kinney
Last updated
An effective AI operations roadmap for insurance moves through four phases: establish a ground-truth baseline of what your claims, underwriting, and policy teams actually do; prioritize the high-volume, low-judgment, low-regulatory-risk work AI can absorb; deploy and measure impact against the baseline; then expand carefully into higher-judgment work with the right guardrails. The phase most carriers skip is the first one. They start with technology selection and automate against assumptions about the work, which is why so many insurance AI initiatives stall after the pilot.
If you run operations at a carrier, you have read the roadmaps already. Pick a use case, run a pilot, scale, govern. They are not wrong, exactly. They just start one step too late.
This roadmap starts where the others assume you already are: knowing, with evidence, what the work in your operation actually consists of. In insurance, where the line between automatable processing and regulated judgment runs through almost every workflow, that step is not optional. Most of what follows is about that first phase, because it decides whether the other three are worth running.
Key Takeaways
The standard insurance AI roadmap opens with a use-case workshop: brainstorm where AI could help, score the ideas, pick a pilot. It feels rigorous. It is built on a foundation of estimates.
The problem is that “where AI could help” is answered from intuition and vendor decks, not from data about how your claims examiners or underwriters actually spend their hours. So the pilot gets aimed at the workflow that sounds most impressive, or the one a vendor demoed well, rather than the one where the data says the most reducible work actually sits. When the pilot delivers less than promised, and most do, the whole program loses momentum.
Readiness scoring has the same defect one level up. Most carriers already hold a readiness score. It came from a vendor questionnaire or a consulting sprint, it rated the data estate, the core systems, and the governance framework on paper, and it came back green. Nothing in it was wrong. It measured the carrier, and the decision depends on the work.
Three failures follow, and they are specific to insurance. The unit of measurement is too coarse, because “claims processing” names dozens of distinct tasks, some pure data movement and some binding the carrier. The same task is often not the same task, because a carrier built through acquisition can run three versions of policy servicing under one name.
The third is the number itself. Across the US economy, the McKinsey Global Institute estimates that up to 30% of hours worked could be automated by 2030, and mandates get built out of figures like it. But a macro benchmark tells a carrier nothing about which 30% of its claims operation is the automatable part. The roadmap’s job is to turn that percentage into a named list of tasks, or to show with your own data that the list falls short. Skip that and you are sequencing AI in the dark, which is the same trap behind the finding that 55% of the companies that cut staff for automation now regret it: confident action on unverified assumptions about the work.
Insurance operations are unusually well-suited to this kind of analysis because the work divides so cleanly, once you can see it. At the coarse level, the split looks like this. The task-level version is what Phase 0 produces.
Claims processing. Intake, document classification, data entry, and routing at the front of the lifecycle are high-volume and largely rules-based. Coverage determination, complex and disputed claims, and anything involving fraud judgment or policyholder hardship are judgment-intensive and often regulated.
Underwriting support. Data gathering, completeness checks, and routine risk-scoring inputs can be assisted or automated. The underwriting decision itself, especially on non-standard risks, stays human and frequently carries regulatory weight.
Policy administration. Endorsements, renewals, and routine servicing transactions are heavy with repetitive, automatable steps. Exceptions and complex policy changes are not.
Customer service. Status inquiries and simple transactions are automatable; complex coverage questions and complaint handling are not, and getting this line wrong is exactly how carriers end up rehiring after over-automating, the insurance version of the Klarna story.
Compliance and reporting. Data aggregation and report generation can be automated; interpretation and regulatory judgment cannot.
In every one of these, the automatable and the protected work sit side by side in the same role. A claims examiner spends part of the morning re-keying an estimate from a PDF and part of the afternoon deciding whether a policy responds. Both are “claims processing” on an org chart, and no org chart can tell them apart. Only activity-level data can.
Insurance adds a constraint most AI roadmaps underweight: regulation does not just slow automation, it changes what is allowed. It is also the vertical where your readiness evidence has a second audience. The first is your budget meeting. The second is a state examiner.
The National Association of Insurance Commissioners adopted its Model Bulletin on the Use of Artificial Intelligence Systems by Insurers on December 4, 2023. It is guidance rather than law: it states a department’s expectations, and it reaches carriers only once their own state adopts it. Where it has, insurers “are expected to develop, implement, and maintain a written program (an ‘AIS Program’) for the responsible use of AI Systems that make, or support decisions related to regulated insurance practices,” spanning the life cycle “including areas such as product development and design, marketing, use, underwriting, rating and pricing, case management, claim administration and payment, and fraud detection.”
Adoption is uneven, and it is worth checking against your own footprint rather than assuming. The NAIC’s implementation map, status as of April 1, 2026, lists 24 states plus the District of Columbia as having adopted the bulletin. Four more, California, Colorado, New York, and Texas, address insurers’ use of AI through their own bulletins or regulations instead. A multi-state carrier is therefore working to more than one instrument, and the map is the place to check which.
The bulletin’s risk management section asks for the “identification of constraints and controls on automation and design,” and among the factors insurers are asked to weigh is “the extent to which humans are involved in the final decision-making process.” Data practices get their own line: “data currency, lineage, quality, integrity, bias analysis and minimization, and suitability.” The bulletin is careful about its own reach. Its stated goal “is not to prescribe specific practices or to prescribe specific documentation requirements,” but to make insurers aware of what a department may request during an investigation or examination.
Read that back as an operations leader rather than as a compliance officer. You cannot describe controls on automation without describing the work being automated, and you cannot state the extent of human involvement without knowing how the decision is currently made, step by step. The baseline is not only a planning artifact. It is the first exhibit in your governance file.
That is why the roadmap has to score each candidate workflow on two axes at once: how automatable it is, and how much regulatory risk automating it carries. A workflow can be technically automatable and still be off-limits, or permitted only with human-in-the-loop review and a documented audit trail. Generic “claims automation” is a slogan. “This specific data-entry step within first-notice-of-loss, which involves no coverage judgment,” is something you can actually defend to a regulator.
Phase 0: Establish the ground-truth baseline. Before selecting a single tool, capture what your claims, underwriting, and policy teams actually do at the activity level. The output is a precise map of automatable versus judgment-or-regulation-bound work across the operation. This is the phase everyone skips and the one everything else depends on, so the next section covers it in full. See how the baseline is built.
Phase 1: Prioritize low-risk, high-volume automation. Using the baseline, sequence the work that is both highly automatable and low in regulatory risk. These are your early wins, claims intake, document routing, policy-servicing transactions, that build credibility and capacity without touching protected work.
Phase 2: Deploy and measure against the baseline. Implement, then measure actual impact against the Phase 0 baseline rather than against projections. Because you have the original activity data, you can prove what changed in capacity, unit cost, and cycle time. See what the measurement looks like.
Phase 3: Expand into higher-judgment work, with guardrails. Only after the foundation is proven do you approach the harder workflows, and only with human-in-the-loop controls, audit trails, and explainability that satisfy your regulators. The baseline keeps updating, so each expansion is evidence-based.
This sequence works because it is grounded before it is ambitious. It also slots directly into a broader AI readiness assessment for operations if you are evaluating the whole operation rather than one line.
Phase 0 is the phase the plan above names and does not explain, so this is what it consists of. Most carrier assessments cover the data estate and the core systems, assume the workflow evidence, and rate a governance framework instead of producing the evidence that framework is meant to hold. The distance between those two is where a readiness score stops being a measurement and starts being an opinion with a number on it.
Every carrier assessment starts here, and most stop here.
The data estate question is not whether you have data. It is whether the data describes the work. Policy, claim, and billing records tell you what was produced: claims closed, submissions quoted, premium applied. They do not tell you what was done to produce it. For each decision you intend to hand to AI, which record carries the input, how stale is it by the time the work is done, and who reconciles it when systems disagree? Is the document estate structured, or is it images in a repository with metadata typed by a person?
The core system question is not whether it integrates. It is where the work actually happens. Much of a carrier’s day happens beside the administration platform rather than inside it: re-keying between the policy system and a document repository, a spreadsheet tracking what the system cannot, an email thread that is the real workflow. Integration readiness measured at the system boundary misses all of it, because none of it is in a log. Automating what the platform records while the work lives around it is how a project delivers a technically successful pilot and no capacity.
This is what separates a carrier assessment from a questionnaire, and it is what generic frameworks assume rather than measure. Seven measures apply across the operation.
In claims, separate intake and document classification, coverage verification, reserve setting and changes, assignment and routing, vendor and adjuster coordination, invoice and medical bill review, diary management, subrogation identification, salvage, settlement authority checks, statutory and index reporting, and litigation referral. The signal that matters is the ratio of file assembly to file decision, split by severity band and by line. Assembly is where AI has room. Decision is where it does not. A carrier that cannot produce that ratio has assessed its claims system, not its claims operation.
In underwriting support, break out submission clearance and duplicate checks, prefill and re-keying from broker email and application forms, loss run ordering and chasing, ordering motor vehicle records and inspections, exposure schedule entry, completeness checks, referral packaging, quote issuance, and declination correspondence. The signal here is how much of the assistant’s day is data movement between systems that do not talk to each other. That work is usually the largest and least visible automation candidate in the function, because it never appears as a queue. It appears as somebody being busy.
In servicing and billing, measure the share of transactions that are genuinely identical rather than identical-looking: two endorsements that read the same on a report can land in different administration systems with different exception rates. Suspense and unapplied cash resolution is often the clearest automation candidate and the least measured, because it is handled in the gaps of the day by whoever notices it first. Non-payment cancellation notices carry statutory timing and belong on the protected list until someone proves otherwise.
Then measure the exception layer. It is what the other functions generate when something does not fit, and it is usually the least measured work in the building. A flow that clears 85% of transactions without a human can still consume most of a team’s day if the remaining 15% takes 20 times longer, and that arithmetic is invisible in any report showing only the straight-through rate.
To see what a baseline of your claims, underwriting, and servicing teams would show, book a 30-minute strategy call. The wider version of this discipline across the vertical is in workforce intelligence for insurance.
The overlay above sets an evidence standard, and it is stricter than the one most assessments apply.
This is what we mean by ground truth workforce data, and the capture method is in our approach: click-region activity rather than full screens or keystrokes, classified at the activity level, with the carrier owning the data.
Run it before the number is set. Once a headcount target exists, the same evidence becomes a defense rather than a decision input, a weaker position even when the data is identical.
Then protect the measurement window, because insurance has more ways to poison one than most industries. A period dominated by catastrophe response, a renewal peak, open enrollment on health lines, or the weeks either side of a system conversion will not describe the normal operation. If the window has to include one, run long enough to isolate it and label it in the output rather than letting it average in silently.
A finished Phase 0 leaves you with six artifacts. If it leaves a number instead, the evidence was not there.
With those in hand, Phases 1 through 3 are sequencing. Without them, sequencing is guesswork with a Gantt chart attached.
The reason carriers skip Phase 0 is that they assume it requires a consulting firm shadowing examiners for a year. It does not anymore.
The Ground Truth AI² Platform™ captures individual-level activity across your operation automatically and combines it with 20-plus years of operational expertise to produce a consulting-grade analysis in a fixed 90-day engagement, with an initial findings summary at weeks three to four. See the platform. For insurance, that means a documented map of which claims, underwriting, and policy tasks AI can absorb, which are protected by judgment or regulation, and in what sequence to proceed, before you commit budget to a single tool.
If your AI mandate spans banking lines as well as insurance, the same approach applies there; see the roadmap for banks and credit unions.
The measure-first discipline behind the 90-day baseline has already been applied inside a carrier. At Nationwide, our consultants ran more than 10,000 side-by-side observations across two back-office processing sites and standardized the work they mapped. The result was a 22% reduction in unit costs, the elimination of overtime, and $1.6M in savings. Read the Nationwide back office case study.
Where does AI fit in insurance operations today? Strongest in high-volume, rules-based work: claims intake, document classification, policy-servicing transactions, routine reconciliations, and parts of regulatory reporting preparation. Weaker, and often off-limits, in coverage decisions, complex claims judgment, and underwriting calls on non-standard risks.
How do insurance regulators view AI in operations? The NAIC model bulletin sets out what a department expects of an insurer’s AI governance, but it is guidance and reaches you only once your state adopts it. As of the NAIC’s April 1, 2026 implementation map, 24 states plus the District of Columbia had adopted it, and California, Colorado, New York, and Texas address insurer AI through their own instruments. Fair-claims-handling rules and growing demand for explainability apply regardless. The roadmap has to score automation potential and regulatory risk together, not separately.
What is the difference between the Phase 0 assessment and the roadmap itself? The assessment establishes what is true now: what the work is, how much of it AI can absorb, and what the evidence is. The roadmap sequences what happens next. A roadmap built without an assessment sequences assumptions.
What does a carrier readiness assessment actually measure? Time and volume per task by role and line of business, straight-through and exception rates with causes, system switches per transaction, the judgment content of each step, whether it feeds a regulated decision, and how much the same task varies between the people performing it.
We already have process documentation and a workflow tool. Is that enough? Process documentation records how work is supposed to happen, and workflow tools record what the systems logged. Neither captures the work performed outside them, which is where the exceptions, the re-keying, and the spreadsheets live. That gap is usually the difference between the projected saving and the realized one.
Does this apply equally to P&C, health, and life lines? The four-phase model applies in all of them. The specific mix of automatable versus judgment-bound work differs by line (claims structure in P&C, underwriting in life, medical necessity in health), so the prioritization differs even when the framework is the same.
How long does a baseline-grade insurance AI roadmap take to build? The ground-truth baseline is fixed at 90 days, with a minimum of 50 employees in the target area. A questionnaire review takes days, but it tells you about the carrier rather than about the work. Deploying against the baseline, measuring, and expanding into more regulated work is multi-phase and continues from there.
If you are being asked what AI can absorb in claims, underwriting, or servicing, the honest answer today is probably an estimate wearing a percentage sign. Closing that gap takes a quarter rather than a year.
Book a 30-minute strategy call and we will walk through which functions would be measured first, what a ground-truth baseline of your claims and underwriting operations would reveal, and what you could put in front of your board and your examiner when it is done.
Ready to Help Your Team Reach the Peak? See us in Action.