An AI readiness scorecard rates an operation on what it can actually hand to AI, not on whether its technology is modern. The Summit Trails scorecard rates six dimensions: process visibility, data ground truth, volume and variability, exception rates and judgment intensity, workforce knowledge concentration, and compliance and risk constraints. You score 24 criteria from 0 to 2, total them out of 48, and read the number against a band that tells you what to do next. It is free, it runs on this page, and nothing here is gated behind an email address.
A self-assessment for operations leaders deciding what AI should take over, and what it should not touch.
Most AI readiness assessments measure whether AI can run in your environment. This scorecard measures something different: whether you are ready to decide what AI should do. Infrastructure readiness is an IT question. Decision readiness is yours. This is the tool for the second one.
Score one operational area at a time (claims processing, loan operations, policy service, back-office support). Do not score “the company.” AI decisions are made at the level of specific work, so readiness has to be measured there too. If you want the reasoning behind that, and the method for building a scorecard of your own, it is set out in how to build an AI readiness scorecard for operations.
How to Use This Scorecard
- Pick one target operation. Minimum useful scope is a team or function with a shared workflow, roughly 20 or more people. Score each area separately if you have several candidates.
- Score all 24 criteria across the six dimensions below. Each criterion scores 0, 1, or 2. The rubric next to each criterion tells you exactly what each score means.
- Apply the evidence rule. A score of 2 requires evidence you could put in front of your board this week: a report, a dataset, a documented process. If your answer is “I’m confident, but I’d have to pull something together,” that is a 1. If your answer relies on what you believe rather than what you can show, that is a 0. Self-assessments fail when leaders grade their instincts instead of their evidence.
- Total your score (maximum 48), then read your band in the interpretation table.
- Check the red-flag rule. Any single dimension scoring 3 or less caps your overall verdict, regardless of total. The bands explain how.
Be honest. Nobody sees this but you, and an inflated score here becomes an expensive assumption later.
Dimension 1: Process Visibility
Can you see the work itself, not just its outputs? Dashboards tell you how many claims were processed. They do not tell you the sequence of tasks a person performs to process one. AI decisions live in that second layer.
| Criterion | 0 | 1 | 2 | Your score |
|---|---|---|---|---|
| 1.1 Task inventory | No documented list of the tasks this operation performs day to day | A high-level process doc exists but is more than a year old or covers only the happy path | A current, task-level inventory exists, including the unofficial workarounds people actually use | ___ |
| 1.2 Time allocation | You could not say how work time splits across tasks for a given role | You have estimates from interviews, surveys, or manager judgment | You have measured time allocation per role at the task level, from actual activity data | ___ |
| 1.3 Workflow mapping | No end-to-end map of how work moves through the operation | A map exists but predates your current systems or staffing | A current map exists, including handoffs, queues, and where work stalls | ___ |
| 1.4 Shadow work | You have no view of rework, status-chasing, tool-switching, and duplicate entry | You know these exist and have anecdotal examples | You can quantify how much capacity shadow work consumes | ___ |
Dimension 1 subtotal: ___ / 8
Dimension 2: Data Ground Truth
Is the data behind your AI decisions activity-level or app-level? “The team spends six hours a day in the claims system” is app-level. “Within the claims system, 40% of that time is rote data entry and 25% is coverage interpretation” is activity-level. Only the second supports an automation decision.
| Criterion | 0 | 1 | 2 | Your score |
|---|---|---|---|---|
| 2.1 Data granularity | Your workforce data is output metrics only (volumes, handle times, SLA attainment) | You have application-level data: which tools people use and for how long | You have activity-level data: what people actually do inside those tools | ___ |
| 2.2 Data currency | Your last workforce analysis is more than a year old, or has never been done | You have a study or time analysis from the last 6 to 12 months | You have data collected within the last 90 days that reflects current systems and staffing | ___ |
| 2.3 Coverage | Data covers a sample of people or a snapshot in time (one workshop, one week of shadowing) | Data covers most roles but was collected once, not continuously | Data covers every role in scope, collected across enough days to smooth out anomalies | ___ |
| 2.4 Decision lineage | Automation candidates so far have come from vendor pitches, benchmarks, or intuition | Candidates come from manager nominations backed by partial data | Every automation candidate can be traced to measured activity data | ___ |
Dimension 2 subtotal: ___ / 8
Dimension 3: Volume and Variability
AI absorbs high-volume, repeatable work. Low volume means low payback; high variability means high failure rates. This dimension scores whether the work itself is a good candidate.
| Criterion | 0 | 1 | 2 | Your score |
|---|---|---|---|---|
| 3.1 Transaction volume | You do not know transaction counts per work type | You know totals for the operation but not per work type | You know volumes per work type, per period, and their trend | ___ |
| 3.2 Standardization | The same work item is handled differently depending on who picks it up | Standard procedures exist but adherence is unmeasured | Work is performed consistently, and you have data showing it | ___ |
| 3.3 Input predictability | Inputs arrive in unstructured, inconsistent formats requiring interpretation | Inputs are mixed: some structured, some free-form | Inputs are largely structured and predictable, or you know exactly which share is not | ___ |
| 3.4 Demand pattern | Volume swings are large and unpredicted | Seasonality is understood roughly, from memory | Demand patterns are quantified, so you can size automation against real peaks and troughs | ___ |
Dimension 3 subtotal: ___ / 8
Dimension 4: Exception Rates and Judgment Intensity
The most expensive automation mistake is deploying AI on work that looks routine but is quietly full of exceptions. This is the dimension that caught the companies now rehiring the people they cut. In Orgvue’s 2025 workforce research, 39% of business leaders said they had made people redundant because of AI, and 55% of those said they made the wrong decision. Exception-heavy work misjudged as routine is one of the reasons why.
| Criterion | 0 | 1 | 2 | Your score |
|---|---|---|---|---|
| 4.1 Exception rate | You do not know what share of work items deviate from the standard path | You have a rough sense of exception frequency from team feedback | Exception rates are measured per work type | ___ |
| 4.2 Judgment share | You cannot separate rules-based work from judgment calls within a role | You can name the judgment-heavy tasks but not their share of time | You can quantify judgment-intensive versus rules-based time per role | ___ |
| 4.3 Escalation paths | When AI or automation fails on an item, there is no defined human fallback | Escalation paths exist informally, relying on individual initiative | Escalation and human-review paths are defined and staffed | ___ |
| 4.4 Error cost | You have not assessed what a wrong automated decision costs (rework, customer harm, regulatory exposure) | Error costs are understood for the biggest risks only | Error cost is assessed per work type and factored into what you would automate first | ___ |
Dimension 4 subtotal: ___ / 8
Dimension 5: Workforce Knowledge Concentration
Restructure too aggressively and you destroy institutional knowledge you cannot rebuild. Before AI absorbs any work, you need to know where the irreplaceable knowledge sits, because that is the work you deliberately protect.
| Criterion | 0 | 1 | 2 | Your score |
|---|---|---|---|---|
| 5.1 Key-person risk | You could not list which individuals hold process knowledge that exists nowhere else | You can name the critical people, but what exactly they know is undocumented | Key knowledge holders are identified and what they uniquely handle is documented | ___ |
| 5.2 Knowledge capture | Critical procedures live in people’s heads | Documentation exists but lags reality | Procedures are documented, current, and used | ___ |
| 5.3 Cross-training | Most critical tasks have exactly one person who can do them | Backup coverage exists for some critical tasks | Every critical task has at least one trained backup | ___ |
| 5.4 Protected-work clarity | You have not identified which work must stay human regardless of automation potential | You have an informal sense of what to protect | You have a named, defensible list of human-critical work and the reasons for each entry | ___ |
Dimension 5 subtotal: ___ / 8
Dimension 6: Compliance and Risk Constraints
In regulated operations (insurance, banking, utilities), what AI is allowed to do matters as much as what it can do. Scoring low here does not mean AI is off the table. It means you do not yet know where the lines are, and finding out mid-deployment is the expensive way.
| Criterion | 0 | 1 | 2 | Your score |
|---|---|---|---|---|
| 6.1 Regulatory mapping | You have not mapped which tasks carry regulatory or contractual constraints on automation | Constraints are understood at the department level, not the task level | Constraints are mapped task by task, so you know exactly which work is restricted | ___ |
| 6.2 Auditability | Automated decisions in your environment could not currently be explained or reconstructed for an auditor | Some audit trail exists, but coverage is partial | Decision logging and explainability requirements are defined and achievable | ___ |
| 6.3 Data privacy | You have not assessed the privacy implications of collecting workforce activity data or feeding operational data to AI | Privacy review is in progress or handled ad hoc | Privacy requirements (GDPR, CCPA, internal policy) are defined for both workforce data collection and AI processing | ___ |
| 6.4 Accountability | If an automated decision goes wrong, it is unclear who owns the outcome | Ownership is assumed to sit with the operation but is not written down | Decision accountability is explicitly assigned for every process AI would touch | ___ |
Dimension 6 subtotal: ___ / 8
Score Sheet
| Dimension | Subtotal | Max |
|---|---|---|
| 1. Process Visibility | ___ | 8 |
| 2. Data Ground Truth | ___ | 8 |
| 3. Volume and Variability | ___ | 8 |
| 4. Exception Rates and Judgment Intensity | ___ | 8 |
| 5. Workforce Knowledge Concentration | ___ | 8 |
| 6. Compliance and Risk Constraints | ___ | 8 |
| Total | ___ | 48 |
The AI Readiness Scorecard Template, to Copy or Rebuild
Most people scoring an operation want an AI readiness scorecard template they can take into a room and fill in offline. Copy the block below into a document, a notebook, or a spreadsheet, and fill it in offline. It carries all 24 criteria in the same order as the rubrics above, so you can score from memory once you have read them and check the borderline ones against the rubric as you go.
AI READINESS SCORECARD Operation scored: ______________________
Scale: 0 = belief 1 = partial or stale evidence 2 = evidence you could show your board this week
Date scored: ___________
D1 PROCESS VISIBILITY
1.1 Task inventory [ ]
1.2 Time allocation [ ]
1.3 Workflow mapping [ ]
1.4 Shadow work [ ]
D1 subtotal ___ / 8
D2 DATA GROUND TRUTH
2.1 Data granularity [ ]
2.2 Data currency [ ]
2.3 Coverage [ ]
2.4 Decision lineage [ ]
D2 subtotal ___ / 8
D3 VOLUME AND VARIABILITY
3.1 Transaction volume [ ]
3.2 Standardization [ ]
3.3 Input predictability [ ]
3.4 Demand pattern [ ]
D3 subtotal ___ / 8
D4 EXCEPTION RATES AND JUDGMENT INTENSITY
4.1 Exception rate [ ]
4.2 Judgment share [ ]
4.3 Escalation paths [ ]
4.4 Error cost [ ]
D4 subtotal ___ / 8
D5 WORKFORCE KNOWLEDGE CONCENTRATION
5.1 Key-person risk [ ]
5.2 Knowledge capture [ ]
5.3 Cross-training [ ]
5.4 Protected-work clarity [ ]
D5 subtotal ___ / 8
D6 COMPLIANCE AND RISK CONSTRAINTS
6.1 Regulatory mapping [ ]
6.2 Auditability [ ]
6.3 Data privacy [ ]
6.4 Accountability [ ]
D6 subtotal ___ / 8
TOTAL ___ / 48
Lowest dimension: D__ at ___ / 8 Any dimension at 3 or less caps the verdict.
Rebuilding the template in a spreadsheet
One row per criterion, one column for the score, and three columns holding the 0, 1, and 2 wording so nobody has to remember what a 2 means at the moment they are tempted to award one. Two formulas do the rest: a subtotal per dimension block, and a total. Add a third that returns your lowest dimension subtotal, because that number decides your verdict more often than the total does.
The reason to rebuild it rather than score in your head is the audit trail. Six months from now, someone will ask why claims processing was scored ahead of policy service, and the answer needs to be a sheet with dates on it rather than a recollection.
Changing the template without breaking it
Adapt it. Two rules keep the result readable. First, if you add or remove criteria, keep every dimension at the same number of criteria as the others, so the subtotals stay comparable and the floor rule still works. Second, keep the 0, 1, and 2 definitions anchored to evidence rather than to confidence. The moment a 2 means “we are sure” instead of “we can show it,” the scorecard stops measuring anything.
If a criterion genuinely does not apply, score it against the dimension’s intent rather than deleting it, and write down the reason next to it. A deleted criterion silently raises your percentage. A scored one with a note attached does not.
How an AI Readiness Score Is Calculated
The arithmetic is deliberately plain. There are 24 criteria, each scored 0, 1, or 2, grouped four to a dimension. Each dimension therefore totals 8, and the six dimensions total 48. Your AI readiness score is the sum of all 24 criterion scores, and your six subtotals are the part that tells you where the number came from.
Nothing is weighted, nothing is normalized, and there is no hidden model behind it. That is the point: a score you cannot reconstruct on paper is a score you cannot defend in a budget meeting, which is the only room where it matters.
Why the scale is 0, 1, and 2
Three points force a decision. On a five- or seven-point scale, most operations park almost everything in the middle, and the result is a score that moves smoothly and says nothing. With three points, every criterion is a call.
More importantly, each point is defined by evidence rather than by feeling. A 0 means you are working from belief. A 1 means you have something partial, stale, or estimated. A 2 means you have evidence you could put in front of your board this week. So the number measures how much of your operation you can currently prove, rather than how sure you feel about it.
That distinction is what makes the score useful. Confidence is high in exactly the operations that get automation wrong.
Turning your score into a percentage, and when that misleads
Divide your total by 48 and you have a percentage. A 31 out of 48 is 65%. It is a convenient way to talk about the number in a slide.
Use it for one thing only: comparing your own operational areas against each other, scored with the same 24 criteria, in the same window, by people applying the same evidence rule. Used that way it ranks your candidates, which is exactly what you want before you pick where AI goes first.
Do not use it as a benchmark. There is no industry average for this score, and any percentage you compare it against was produced by a different instrument measuring a different thing. A percentage invites that comparison, which is why the raw number out of 48 is the safer one to carry into a conversation.
Why every dimension carries the same weight here
All six dimensions are worth 8 points. That is deliberate. Weighting lets a strong dimension pay for a weak one, and the whole failure pattern this scorecard exists to catch is a healthy total sitting on top of one fatal gap.
Some operations do have a genuine case for weighting, usually a heavily regulated function that wants compliance to count double. If that is you, weight it, but weight it before you score rather than after you see the result, and keep the unweighted total alongside it. The method for setting weights on a scorecard of your own is covered in how to build an AI readiness scorecard for operations.
The floor rule: your weakest dimension outranks your total
A total is an average wearing a disguise, and averages hide variance. Two operations can both score 36 out of 48: one with every dimension at 6, another with five dimensions at 7 and Compliance at 1. They are not the same operation and they do not get the same advice.
So the rule is simple. If any single dimension scores 3 or less, that dimension is your verdict, whatever the total says. Automation fails at the weakest link, and the weakest link is usually the one nobody scored honestly because it belonged to a different department.
Why your score cannot be compared to another framework’s score
Most of the AI readiness scores you will be offered measure infrastructure: cloud posture, data platform maturity, governance, model tooling, skills. They answer whether AI can run in your environment. This one answers whether you can decide what AI should run on. Both are real questions and they have almost no overlap.
The practical consequence is that an organization can hold a strong score on a cloud-vendor readiness assessment and a weak one here, at the same time, without either being wrong. If you are trying to work out which instrument you are actually holding, the landscape is broken down in the honest guide to AI readiness assessment tools and in who does AI readiness assessments for operations.
Interpreting Your Score
Red-flag rule first: if any single dimension scored 3 or less, treat that dimension as your verdict regardless of total. A 40-point operation with a 2 in Compliance is not a 40-point operation. It is a compliance problem with good data. Close the weakest dimension before acting on the strongest ones.
| Total score | Verdict | What it means | Recommended next step |
|---|---|---|---|
| 0 to 15 | Not decision ready | Any AI or headcount decision made now would rest on instinct and benchmarks, not on your operation’s reality. This is the profile behind most regretted AI layoffs. | Do not commit numbers to your board yet. Start with visibility: build the task inventory and activity baseline before evaluating a single AI tool. |
| 16 to 27 | Directionally aware, not defensible | You know your operation well enough to have good instincts about what to automate, but you could not defend the specifics line by line in a budget meeting. | Pick your single best-scoring candidate area and close the data gap there first. Depth in one area beats shallow readiness everywhere. |
| 28 to 38 | Conditionally ready | You can scope credible AI pilots. Your risk is precision: which specific tasks, what volume they represent, and what happens to the people and SLAs attached to them. | Validate your highest-priority automation candidates with measured activity data before committing budgets or headcount numbers. |
| 39 to 48 | Decision ready, pending validation | You have the visibility, the data discipline, and the guardrails to make automation decisions responsibly. | Pressure-test the self-score. High scorers usually hold their rating on infrastructure and compliance and lose points on Dimension 2 when self-reported data meets measured data. Validate, then move. |
One pattern to watch for
The most common profile we see: strong scores on Dimensions 3 and 6, weak scores on 1, 2, and 4. That is an operation that knows its volumes and its rules but not its work. It feels ready because the infrastructure conversation has gone well. It is the exact profile that automates the wrong tasks first, because it has to guess which tasks those are.
What to Do at Each Score Band
The bands above give you a verdict. This is the version with dates on it: what the number is telling you, what to do in the next 30 days, and what not to do yet.
0 to 15: Not decision ready
Your score is telling you that the operation is currently described rather than measured. Every automation candidate you could name today came from somewhere other than your own data.
Next 30 days: build the task inventory for one area, at the level of “what does a person do between opening a case and closing it,” not at the level of process documentation. Interview four people doing the same role and write down where their accounts differ, because the differences are the shadow work in Dimension 1.
Not yet: vendor demos, tool shortlists, and any number given to a board. A shortlist built on this score is a shortlist of solutions to problems you have not confirmed you have, and a number given now becomes the number you are held to later.
16 to 27: Directionally aware, not defensible
You almost certainly know which area you would start with, and you are probably right. What you cannot do is show the working, and that is the gap that matters when finance or the board pushes back on a specific line.
Next 30 days: stop trying to raise the score everywhere. Pick the one area you would automate first and take it deep: measured time allocation for its roles, exception rates per work type, and a named list of the work that has to stay human. Depth in one area is what converts an instinct into a proposal.
Not yet: a company-wide readiness program. Scoring five more areas at this level produces five more directional answers, which is the same answer you already have, five times.
28 to 38: Conditionally ready
You can scope a credible pilot. The risk at this band is precision rather than direction, and precision is where the money is: which specific tasks, what share of hours they represent, and what happens to the people and the service levels attached to them when those hours move.
Next 30 days: take your top two automation candidates and try to write the business case at task level, with hours and volumes attached. Wherever you have to write “approximately” or “we believe,” you have found the exact place your score is soft, and it is almost always inside Dimension 2.
Not yet: committing headcount numbers. At this band the direction holds and the magnitude does not, and magnitude is the part that is difficult to walk back once it has been said out loud. If a mandate has already landed on your desk, the mandate response framework is the sequence for answering it without committing to a number you cannot support.
39 to 48: Decision ready, pending validation
The verdict is genuine and the caveat is real. Operations scoring in this band are usually right about their strengths and consistently over-scored on Dimension 2, because self-reported data quality is the hardest thing to assess about yourself.
Next 30 days: pressure-test the score against evidence rather than against agreement. Take three criteria you scored 2 and ask for the artifact: the dataset, the report, the current map. If it arrives within a day, the 2 stands. If it takes a week to assemble, it was a 1.
Not yet: treating the score as the baseline you will measure results against. A self-score cannot be a baseline, because you cannot show a board an improvement measured against your own earlier opinion. The baseline has to be measured.
What Your Score Is, and What It Is Not
This AI readiness scorecard is a structured way to find your gaps. It is still a self-assessment, which means every score in it is a hypothesis. You graded your own operation from what you believe about the work. The entire lesson of the last three years of AI-driven restructuring is that what leaders believe about the work and what the work actually is are two different datasets.
There is one way to convert the hypothesis into evidence: measure the work itself. The Summit Trails 90-day assessment captures individual-level activity data across your target operation and turns it into the Ground Truth AI² Report™: a measured version of every dimension you just scored by hand, plus a prioritized view of what AI can absorb at acceptable risk. Most operations find their self-score was off in both directions, readier than they thought in some areas, and guessing in others they had marked as strengths.
The dimension that moves most between the self-score and the measured version is Dimension 2, and it moves in one direction. Data granularity, currency, coverage, and decision lineage are the four things that cannot be improved by documenting harder. They are the four that have to be measured, which is why a low Dimension 2 sits underneath most low totals on this scorecard.
Questions About AI Readiness Scores
What is a good AI readiness score? On this AI readiness scorecard, anything at 39 or above out of 48 means you can make automation decisions responsibly, provided no single dimension is at 3 or less. But “good” is the wrong frame for a first score. The useful reading is which dimension is lowest, because that is what decides your next 90 days regardless of the total. An operation moving from 22 to 31 by closing its process visibility gap has changed more than one that started at 38.
Is this a free online AI readiness score tool? Yes. The scorecard runs on this page, there is nothing to install, and no email address is required to use it or to copy the template. It scores one operational area, meaning the work your people do inside it. Scoring a website or a cloud and data stack is a different assessment, and the options are compared in the honest guide to AI readiness assessment tools.
Our infrastructure readiness score is high and this one is low. Which one is right? Both. They measure different things and a gap between them is the normal result, not a contradiction. A high infrastructure score says AI can run in your environment. A low score here says you do not yet know what it should run on. The second gap is the more expensive one, because infrastructure problems surface during deployment while decision problems surface after it, in the form of work you automated that should not have been.
Does a low AI readiness score mean we should not use AI? No. It means you should not yet commit to specific automation targets or headcount numbers. The work that raises the score, task inventory, measured time allocation, exception rates, a named list of protected work, is the same work that makes a good AI deployment possible. A low score tells you about sequence. It says nothing about whether AI belongs in your operation.
Who should fill in the scorecard? The person accountable for the operation, with the people who run it in the room. Scoring alone produces a cleaner sheet and a less accurate one, because the criteria most often over-scored, shadow work, exception rates, and key-person risk, are precisely the ones a frontline supervisor can correct in a sentence. Score it together and record where you disagreed, because the disagreements are findings.
What if a criterion does not apply to our operation? Score it rather than delete it, and write down why. Deleting a criterion quietly raises your percentage while lowering what the number is based on. If a criterion genuinely has no meaning in your context, score it against the intent of its dimension and annotate it, so the next person reading the sheet knows what happened.
Turn the Self-Score into Measured Data
Scored your operation and want to know how it holds up against real activity data? Book a 30-minute strategy call and bring your score sheet. We will walk through where the numbers usually move, and where they usually do not.
It is a working session, not a pitch. You will leave with a clearer view of which dimension is actually blocking your next decision, whether or not you ever measure anything with us. If you want to see what measured data looks like first, our approach sets out how activity is captured and classified, and the results page shows what comes out the other end.
Use this scorecard with: the AI readiness assessment guide for operations, the platform evaluation checklist for when you are ready to compare vendors, how to respond to an AI headcount mandate, how to know what to automate in operations, and why companies regret AI-driven layoffs. Both free tools are listed on the tools page.
Prefer to work through it live? Wendy runs a 30-minute session on answering an AI mandate with evidence, and the next date is listed there.
Source for the redundancy figures in Dimension 4: Orgvue, Human-first, machine enhanced: From optimism to pragmatism in AI-driven workforce transformation, Spring 2025. Research conducted by Vitreous World among 1,163 business leaders across eight countries, fieldwork February to March 2025. Read the report.