Workforce Intelligence for Insurance: What to Measure Before You Automate
August 14, 2026 — Amelia
August 14, 2026 — Amelia
Workforce intelligence for insurance means capturing what your operations workforce actually does, activity by activity, before you commit budget to automation. Not what the policy administration system logs. Not what the org chart implies. What claims adjusters, underwriting assistants, policy-service reps, and compliance analysts actually spend their day doing. In insurance there is a second reason to measure first that most industries do not carry: a workflow can be perfectly automatable and still be constrained by regulation, so the baseline has to score what the work is and what the rules allow at the same time. Skip that step and the pilot stalls, or worse, it ships and a market-conduct exam finds it.
If you run operations at a carrier, you do not need another AI plan that opens with “pick a use case.” You need to know, at the activity level, which work is high-volume and rule-based, which is judgment-bound or regulated, and which quietly depends on one adjuster who has handled a particular claim type for twenty years. That is a measurement problem, and it comes before the roadmap.
Key Takeaways
- Roughly 30% of work hours across the economy are automatable by 2030 (McKinsey Global Institute), yet RAND research finds more than 80% of AI projects fail to deliver their intended value. In insurance the failure usually traces to acting on an unverified picture of the work, with a regulator watching.
- System-level data (“6 hours in the claims platform”) tells you nothing an automation decision can use. Activity-level data (“44 minutes handling coverage exceptions the system flagged but could not resolve”) tells you whether AI can help.
- The automatable work and the protected work sit inside the same roles. An org chart will never separate them. Activity-level data will.
- Five things to measure before you automate: time allocation by activity, exception and surge volume, process variation across lines of business, automation adjacency scored against regulatory risk, and institutional knowledge density.
- A 90-day ground-truth baseline turns “automate the claims shop” into a phased plan with defensible numbers, and gives model-governance and compliance the documentation they will ask for anyway.
Carriers are not short on AI ambition. The standard operations AI initiative opens with a use-case inventory and a technology evaluation, runs a pilot on the cleanest, most visible workflow, and shows an early win on straight-through claims or routine renewals. Then the savings flatten at a fraction of the target, and the program quietly stalls.
The reason is rarely the model or the vendor. RAND Corporation research found that more than 80% of AI projects fail to deliver their intended business value, roughly double the failure rate of comparable non-AI IT projects, and the recurring cause is that the organization did not know, in measurable terms, what it was automating before it started. Insurance adds a second failure mode on top of the first: the workflow that looked automatable turns out to be governed by unfair-claims-practices rules, a model-governance expectation, or an adjuster judgment call that nobody scoped.
Insurance operations run on two layers. The transaction layer is visible: FNOL intake logged, payments issued, policies bound, endorsements processed, renewals generated. Every core system can report volume and throughput.
The exception layer is nearly invisible. The adjuster working a coverage question the automated adjudication could not clear. The underwriting assistant chasing missing application data across three sources that do not reconcile. The policy-service rep handling an endorsement governed by state-specific rules. The compliance analyst triaging a fraud-flag that is probably noise but cannot be closed without documentation. These people make hundreds of judgment calls a day, and their work determines whether the transaction layer actually closes clean. In most carriers, nobody can say with evidence what they do all day.
That is precisely the population most exposed when an automation or headcount mandate arrives, and the one for which the carrier typically has no activity-level data at all. Seeing that invisible layer is what the Capture, Classify, Insight methodology was built to do.
At a carrier, workforce intelligence gets confused with two things it is not: what the claims and policy systems already report, and what monitoring software produces. The difference is what decides whether an automation plan survives contact with a market-conduct exam.
Core-system and BI reporting measures outcomes: claims closed, policies bound, calls handled, cycle time. Useful, but it measures the throughput of the transaction, not the work of the people around it. It cannot tell you why exception handling eats a third of the claims team’s week, or which part of that is automatable.
Monitoring measures presence: logged in, active, idle. An adjuster who is idle in the claims platform because they are reading a state bulletin scores badly and is doing precisely the right thing, which is why a productivity score cannot tell you what is safe to automate.
Workforce intelligence measures the work itself. Not that an adjuster spent six hours in the claims platform, but that 44 minutes went to coverage-exception research, 31 minutes to a status update the portal should have handled, 39 minutes to a subrogation referral that requires judgment and cannot be automated, and 51 minutes to a reserve-adequacy tally that two other adjusters also build separately because nobody ever standardized it.
Only that level of detail supports an automation decision in a regulated environment. System-level data captures the container. Activity-level data captures the work, and separates the work AI can absorb from the work regulation protects.
In an unregulated back office, shallow data is merely wasteful. At a carrier it is dangerous, because so much of the work exists to satisfy rules that no application log records. State Departments of Insurance (DOI) set market-conduct and fair-claims-handling expectations that vary jurisdiction by jurisdiction. NAIC model governance and model-audit expectations mean any AI that influences decisions needs validation, monitoring, and documentation. Unfair-claims-practices rules constrain how claim decisions are made and communicated. And the newer wave of AI bulletins and model-explainability scrutiny is raising the bar on how carriers justify automated decisions.
The practical consequence: every candidate workflow has to be scored on two axes at once, how automatable it is and how much regulatory risk automating it carries. A task can be perfectly automatable and still require human-in-the-loop review, or be off-limits entirely. You cannot make those calls from a slogan like “automate claims operations.” You can only make them from a precise view of what each task actually involves, which is exactly what the Ground Truth AI² Platform™ captures: click-region activity, no keystrokes, classified by vision AI into the actual activity being performed.
This is the data an operations leader needs before any insurance AI investment conversation. Not which vendor to pilot. What to measure first.
“Claims adjuster” is not a unit of measurement. “29% of adjuster time goes to clearing coverage and documentation exceptions the automated adjudication flags but cannot resolve” is a unit of measurement.
Insurance roles blend transaction processing, exception handling, customer contact, and regulatory work in proportions that vary by line and channel. FNOL intake, adjudication, and exception handling in claims. Data gathering, application completeness, and routine risk-scoring inputs in underwriting support. Endorsements, renewals, and servicing in policy administration. Until you know the actual breakdown, you cannot identify automation candidates, and you cannot size the benefit case with anything but vendor benchmarks drawn from carriers that do not look like yours.
Insurance has a measurement trap: the operation has more than one state. There is the steady state, when work is planned and exceptions are moderate, and there are the surges, catastrophe-event claim waves, renewal cycles, open enrollment for health lines, and exam or audit preparation, when the same operations staff pivot to work that never appears in a blue-sky sample.
A baseline captured only in the steady state sizes automation to steady-state volume. Then a CAT event floods FNOL, or a renewal cycle peaks, or an exam request lands, and the automated workflow meets exceptions it was never trained on, while the people who used to absorb the surge have been reassigned or reduced. Measure long enough to see the peaks, and treat surge workload as a first-class input to the automation decision, not an outlier to be cleaned from the data. High-exception workflows are high-knowledge workflows; automate them without a baseline and you learn what you lost when cycle times, complaint volumes, and exam findings start moving.
Two claims teams, two underwriting units, or two acquired books may run the “same” process very differently: different products, different state rules, different core-system configurations, different inherited workarounds. P&C, life, and health lines carry genuinely different requirements, and a rule that is legitimate in one state is a workaround in another. When a process is performed differently across the carrier, that variation is information. Either a best practice is waiting to be standardized, or the variation is legitimate and any automation must accommodate it.
You cannot know which without measuring all of them. Skip this step and you build the automation around one team’s version of the process, then discover in rollout that the other teams were handling real edge cases, like state-specific claim-notice timing or a legacy-product servicing rule, that the automated version ignores.
Within a single role, some activities are structured and rule-based, exactly what AI handles well. FNOL intake, document classification, data entry, reconciliation, and routine policy transactions automate cleanly. But the exception layer, coverage determinations, fraud-triage judgment, subrogation calls, and claim decisions governed by unfair-claims-practices rules, carries judgment, regulatory exposure, and policyholder risk.
Automation adjacency scoring maps this at the activity level: automatable now, automatable after a training period with controls, or off-limits by regulation. That is what turns a vague mandate into a phased roadmap with defensible numbers, and it is the input the AI operations roadmap for insurance sequences against.
Every carrier has workflows that function because one person knows something written down nowhere. The adjuster who knows which fraud flags are false alarms after a system change. The policy-service lead who knows how a specific state’s endorsement timing actually works. The compliance veteran who knows which patterns a market-conduct examiner cares about and which are noise.
Automate around that person without measuring what they know and the gap surfaces on the first unusual claim after they leave, with nothing on record for a model to have learned from. Measured first, the same knowledge becomes what you capture and transfer before the workflow changes, and what keeps the claims-handling audit trail defensible after it does.
Facing an automation or headcount mandate for your claims, underwriting, or policy operations? Book a 30-minute strategy call with Wendy Kinney. No pitch. A clear conversation about where your data gaps are and what a 90-day baseline would take.
Here is how this plays out without a baseline.
A VP of claims operations at a mid-sized carrier gets a mandate: reduce operating cost 12% through automation, with claims as the target because a consultant benchmark says peers run it leaner. Her team maps the process in workshops, builds the business case, and automates FNOL intake, document handling, and completeness checks. The pilot works. Routine intake automates cleanly.
Two quarters later the savings are a third of the target. The gap is the exception layer: coverage questions that fail the automated adjudication, claim decisions governed by unfair-claims-practices rules, and a CAT surge nobody had measured. The two senior adjusters who handled those exceptions through a mix of system access and years of policy knowledge were redeployed early, and one took a package. Escalations now route to a thinner team, cycle times climb, complaint volumes rise, and the next market-conduct review flags weaknesses in the automated claim-handling workflow that no one had documented.
The technology did not fail. The decision was made without activity-level data on what those two adjusters actually did, and without scoring the regulatory risk of the workflows around them.
The pressure does not always arrive as a technology project. Sometimes it arrives as a number: a board sets the carrier’s operations headcount against a peer cohort and asks why claims and policy service run heavier than the benchmark says they should.
With ground-truth data, the response is specific: exception volumes by category, CAT-surge and exam-prep workload the benchmark cohort does not carry, activity-level evidence of what the “extra” headcount actually does, and a phased plan for which roles can shrink once specific workflows are stabilized and their controls are proven. Without it, the response is instinct against a spreadsheet, and instinct loses that meeting. The playbook is in how to respond to an AI headcount mandate.
The Ground Truth AI² Report™ is a data document, not an opinion document. Over a 90-day engagement, the platform captures activity-level data across your operations workforce, claims, underwriting support, and policy administration alike, and the operational overlay interprets what it means for your specific mandate and your regulators. It is the version of an AI readiness assessment for operations that actually reaches the activity level: time allocation, exception and surge profiles, process variation, automation adjacency scored against regulatory risk, and knowledge-density mapping, for every role in scope, with initial findings at weeks three to four. The same “what to measure” discipline runs through the whole series, including workforce intelligence for financial services; the workflows and regulators change, the measurement discipline does not.
That changes what there is to argue about. The question stops being “which claims workflows should we automate” and becomes “this is what the work is, this is what can be automated now, what can be automated later with controls, and what cannot be touched at all, and this is what a regulator expects at each step, so where do you want to begin?” See what the output looks like.
What data does a carrier need before automating operations?
Activity-level data on the operations workforce: how time is actually allocated by task, where exceptions concentrate (including CAT events, renewal cycles, open enrollment, and exam prep), how much process variation exists across lines of business and states, which workflows depend on undocumented institutional knowledge, and how much regulatory risk each carries. Core-system reporting measures throughput, not this workforce activity layer.
Do our claims and policy admin systems already tell us this?
No. Core and BI data measure transaction outcomes: claims closed, policies bound, cycle time. They do not measure the exception and judgment layer where most automation mandates actually land, and they cannot separate rule-based work from regulation-bound work within a single role.
How is workforce intelligence different from employee monitoring at a carrier?
Monitoring tells you whether someone is present and active and produces productivity scores. Workforce intelligence classifies what work is being done at the activity level and produces automation adjacency scores. For an AI roadmap, only the second is useful. The capture is also less invasive than tools many carriers already run: click-region only, no keystrokes, with the customer owning the data.
How does regulation change what you can automate in insurance?
Every candidate workflow has to be scored on automation potential and regulatory risk together. State DOI market-conduct and fair-claims-handling rules, NAIC model-governance and model-audit expectations, unfair-claims-practices requirements, and the newer wave of AI bulletins and model-explainability scrutiny all mean a task can be automatable and still require human-in-the-loop review, documentation, and an audit trail, or be off-limits. Activity-level data is what lets you make those calls workflow by workflow.
How long does it take to build a workforce activity baseline for a carrier?
Ninety days for an operations area of 50 or more people, with the first findings landing at weeks three to four. The interview-and-sampling approach a consulting engagement runs takes 12 to 18 months to approximate the same picture, and it still stops above the activity level where claims exceptions actually sit.
The insurance AI conversation gravitates to the visible layer: the models, the vendors, the throughput dashboards. Those matter. But the automation and headcount mandates now reaching operations leaders mostly target the exception and judgment layer, the least-measured and most-regulated work in the building.
Before you automate, measure what that workforce actually does. Time allocation by activity across claims, underwriting, and policy. Exception and surge volume across CAT events, renewal cycles, and exam prep. Process variation across lines of business and states. Automation adjacency scored against regulatory risk. Institutional knowledge density, on a tenure clock. Carriers do not have an AI ambition problem. They have a measurement problem, and the ones who close it before committing budget are the ones whose roadmaps survive contact with an examiner.
Book a 30-minute strategy call with Wendy Kinney. Half an hour is enough to establish which parts of your claims, underwriting support and policy administration work you can currently evidence, which parts you are estimating, and what a 90-day baseline of your operation would actually involve. No pitch, and no obligation.
Ready to Help Your Team Reach the Peak? See us in Action.