A self-assessment for operations leaders deciding what AI should take over, and what it should not touch.
Most AI readiness assessments measure whether AI can run in your environment. This scorecard measures something different: whether you are ready to decide what AI should do. Infrastructure readiness is an IT question. Decision readiness is yours. This is the tool for the second one.
Score one operational area at a time (claims processing, loan operations, policy service, back-office support). Do not score “the company.” AI decisions are made at the level of specific work, so readiness has to be measured there too.
How to Use This Scorecard
- Pick one target operation. Minimum useful scope is a team or function with a shared workflow, roughly 20 or more people. Score each area separately if you have several candidates.
- Score all 24 criteria across the six dimensions below. Each criterion scores 0, 1, or 2. The rubric next to each criterion tells you exactly what each score means.
- Apply the evidence rule. A score of 2 requires evidence you could put in front of your board this week: a report, a dataset, a documented process. If your answer is “I’m confident, but I’d have to pull something together,” that is a 1. If your answer relies on what you believe rather than what you can show, that is a 0. Self-assessments fail when leaders grade their instincts instead of their evidence.
- Total your score (maximum 48), then read your band in the interpretation table.
- Check the red-flag rule. Any single dimension scoring 3 or less caps your overall verdict, regardless of total. The bands explain how.
Be honest. Nobody sees this but you, and an inflated score here becomes an expensive assumption later.
Dimension 1: Process Visibility
Can you see the work itself, not just its outputs? Dashboards tell you how many claims were processed. They do not tell you the sequence of tasks a person performs to process one. AI decisions live in that second layer.
| Criterion | 0 | 1 | 2 | Your score |
|---|---|---|---|---|
| 1.1 Task inventory | No documented list of the tasks this operation performs day to day | A high-level process doc exists but is more than a year old or covers only the happy path | A current, task-level inventory exists, including the unofficial workarounds people actually use | ___ |
| 1.2 Time allocation | You could not say how work time splits across tasks for a given role | You have estimates from interviews, surveys, or manager judgment | You have measured time allocation per role at the task level, from actual activity data | ___ |
| 1.3 Workflow mapping | No end-to-end map of how work moves through the operation | A map exists but predates your current systems or staffing | A current map exists, including handoffs, queues, and where work stalls | ___ |
| 1.4 Shadow work | You have no view of rework, status-chasing, tool-switching, and duplicate entry | You know these exist and have anecdotal examples | You can quantify how much capacity shadow work consumes | ___ |
Dimension 1 subtotal: ___ / 8
Dimension 2: Data Ground Truth
Is the data behind your AI decisions activity-level or app-level? “The team spends six hours a day in the claims system” is app-level. “Within the claims system, 40% of that time is rote data entry and 25% is coverage interpretation” is activity-level. Only the second supports an automation decision.
| Criterion | 0 | 1 | 2 | Your score |
|---|---|---|---|---|
| 2.1 Data granularity | Your workforce data is output metrics only (volumes, handle times, SLA attainment) | You have application-level data: which tools people use and for how long | You have activity-level data: what people actually do inside those tools | ___ |
| 2.2 Data currency | Your last workforce analysis is more than a year old, or has never been done | You have a study or time analysis from the last 6 to 12 months | You have data collected within the last 90 days that reflects current systems and staffing | ___ |
| 2.3 Coverage | Data covers a sample of people or a snapshot in time (one workshop, one week of shadowing) | Data covers most roles but was collected once, not continuously | Data covers every role in scope, collected across enough days to smooth out anomalies | ___ |
| 2.4 Decision lineage | Automation candidates so far have come from vendor pitches, benchmarks, or intuition | Candidates come from manager nominations backed by partial data | Every automation candidate can be traced to measured activity data | ___ |
Dimension 2 subtotal: ___ / 8
Dimension 3: Volume and Variability
AI absorbs high-volume, repeatable work. Low volume means low payback; high variability means high failure rates. This dimension scores whether the work itself is a good candidate.
| Criterion | 0 | 1 | 2 | Your score |
|---|---|---|---|---|
| 3.1 Transaction volume | You do not know transaction counts per work type | You know totals for the operation but not per work type | You know volumes per work type, per period, and their trend | ___ |
| 3.2 Standardization | The same work item is handled differently depending on who picks it up | Standard procedures exist but adherence is unmeasured | Work is performed consistently, and you have data showing it | ___ |
| 3.3 Input predictability | Inputs arrive in unstructured, inconsistent formats requiring interpretation | Inputs are mixed: some structured, some free-form | Inputs are largely structured and predictable, or you know exactly which share is not | ___ |
| 3.4 Demand pattern | Volume swings are large and unpredicted | Seasonality is understood roughly, from memory | Demand patterns are quantified, so you can size automation against real peaks and troughs | ___ |
Dimension 3 subtotal: ___ / 8
Dimension 4: Exception Rates and Judgment Intensity
The most expensive automation mistake is deploying AI on work that looks routine but is quietly full of exceptions. This is the dimension that caught the companies now rehiring the people they cut. 55% of companies regret AI-driven layoffs, and exception-heavy work misjudged as routine is a leading reason why.
| Criterion | 0 | 1 | 2 | Your score |
|---|---|---|---|---|
| 4.1 Exception rate | You do not know what share of work items deviate from the standard path | You have a rough sense of exception frequency from team feedback | Exception rates are measured per work type | ___ |
| 4.2 Judgment share | You cannot separate rules-based work from judgment calls within a role | You can name the judgment-heavy tasks but not their share of time | You can quantify judgment-intensive versus rules-based time per role | ___ |
| 4.3 Escalation paths | When AI or automation fails on an item, there is no defined human fallback | Escalation paths exist informally, relying on individual initiative | Escalation and human-review paths are defined and staffed | ___ |
| 4.4 Error cost | You have not assessed what a wrong automated decision costs (rework, customer harm, regulatory exposure) | Error costs are understood for the biggest risks only | Error cost is assessed per work type and factored into what you would automate first | ___ |
Dimension 4 subtotal: ___ / 8
Dimension 5: Workforce Knowledge Concentration
Restructure too aggressively and you destroy institutional knowledge you cannot rebuild. Before AI absorbs any work, you need to know where the irreplaceable knowledge sits, because that is the work you deliberately protect.
| Criterion | 0 | 1 | 2 | Your score |
|---|---|---|---|---|
| 5.1 Key-person risk | You could not list which individuals hold process knowledge that exists nowhere else | You can name the critical people, but what exactly they know is undocumented | Key knowledge holders are identified and what they uniquely handle is documented | ___ |
| 5.2 Knowledge capture | Critical procedures live in people’s heads | Documentation exists but lags reality | Procedures are documented, current, and used | ___ |
| 5.3 Cross-training | Most critical tasks have exactly one person who can do them | Backup coverage exists for some critical tasks | Every critical task has at least one trained backup | ___ |
| 5.4 Protected-work clarity | You have not identified which work must stay human regardless of automation potential | You have an informal sense of what to protect | You have a named, defensible list of human-critical work and the reasons for each entry | ___ |
Dimension 5 subtotal: ___ / 8
Dimension 6: Compliance and Risk Constraints
In regulated operations (insurance, banking, utilities), what AI is allowed to do matters as much as what it can do. Scoring low here does not mean AI is off the table. It means you do not yet know where the lines are, and finding out mid-deployment is the expensive way.
| Criterion | 0 | 1 | 2 | Your score |
|---|---|---|---|---|
| 6.1 Regulatory mapping | You have not mapped which tasks carry regulatory or contractual constraints on automation | Constraints are understood at the department level, not the task level | Constraints are mapped task by task, so you know exactly which work is restricted | ___ |
| 6.2 Auditability | Automated decisions in your environment could not currently be explained or reconstructed for an auditor | Some audit trail exists, but coverage is partial | Decision logging and explainability requirements are defined and achievable | ___ |
| 6.3 Data privacy | You have not assessed the privacy implications of collecting workforce activity data or feeding operational data to AI | Privacy review is in progress or handled ad hoc | Privacy requirements (GDPR, CCPA, internal policy) are defined for both workforce data collection and AI processing | ___ |
| 6.4 Accountability | If an automated decision goes wrong, it is unclear who owns the outcome | Ownership is assumed to sit with the operation but is not written down | Decision accountability is explicitly assigned for every process AI would touch | ___ |
Dimension 6 subtotal: ___ / 8
Score Sheet
| Dimension | Subtotal | Max |
|---|---|---|
| 1. Process Visibility | ___ | 8 |
| 2. Data Ground Truth | ___ | 8 |
| 3. Volume and Variability | ___ | 8 |
| 4. Exception Rates and Judgment Intensity | ___ | 8 |
| 5. Workforce Knowledge Concentration | ___ | 8 |
| 6. Compliance and Risk Constraints | ___ | 8 |
| Total | ___ | 48 |
Interpreting Your Score
Red-flag rule first: if any single dimension scored 3 or less, treat that dimension as your verdict regardless of total. A 40-point operation with a 2 in Compliance is not a 40-point operation. It is a compliance problem with good data. Close the weakest dimension before acting on the strongest ones.
| Total score | Verdict | What it means | Recommended next step |
|---|---|---|---|
| 0 to 15 | Not decision ready | Any AI or headcount decision made now would rest on instinct and benchmarks, not on your operation’s reality. This is the profile behind most regretted AI layoffs. | Do not commit numbers to your board yet. Start with visibility: build the task inventory and activity baseline before evaluating a single AI tool. |
| 16 to 27 | Directionally aware, not defensible | You know your operation well enough to have good instincts about what to automate, but you could not defend the specifics line by line in a budget meeting. | Pick your single best-scoring candidate area and close the data gap there first. Depth in one area beats shallow readiness everywhere. |
| 28 to 38 | Conditionally ready | You can scope credible AI pilots. Your risk is precision: which specific tasks, what volume they represent, and what happens to the people and SLAs attached to them. | Validate your highest-priority automation candidates with measured activity data before committing budgets or headcount numbers. |
| 39 to 48 | Decision ready, pending validation | You have the visibility, the data discipline, and the guardrails to make automation decisions responsibly. | Pressure-test the self-score. High scorers usually hold their rating on infrastructure and compliance and lose points on Dimension 2 when self-reported data meets measured data. Validate, then move. |
One pattern to watch for
The most common profile we see: strong scores on Dimensions 3 and 6, weak scores on 1, 2, and 4. That is an operation that knows its volumes and its rules but not its work. It feels ready because the infrastructure conversation has gone well. It is the exact profile that automates the wrong tasks first, because it has to guess which tasks those are.
What Your Score Is, and What It Is Not
This scorecard is a structured way to find your gaps. It is still a self-assessment, which means every score in it is a hypothesis. You graded your own operation from what you believe about the work. The entire lesson of the last three years of AI-driven restructuring is that what leaders believe about the work and what the work actually is are two different datasets.
There is one way to convert the hypothesis into evidence: measure the work itself. The Summit Trails 90-day assessment captures individual-level activity data across your target operation and turns it into the Ground Truth AI² Report™: a measured version of every dimension you just scored by hand, plus a prioritized view of what AI can absorb at acceptable risk. Most operations find their self-score was off in both directions: readier than they thought in some areas, and guessing in others they had marked as strengths.
Scored your operation and want to know how it holds up against real activity data? Book a 30-minute strategy call and bring your score sheet. We will walk through where the numbers usually move.
Use this scorecard with: the AI readiness assessment guide for operations, how to respond to an AI headcount mandate, and how to know what to automate in operations.