AI Operations Pilot: How to Structure a 90-Day Proof of Concept That Actually Gets Approved

July 16, 2026 — Wendy Kinney

AI Operations Pilot: How to Structure a 90-Day Proof of Concept That Actually Gets Approved, Summit Trails

An AI operations pilot gets approved when it is structured as a decision, not an experiment. That means five things on paper before anyone writes a check: a bounded scope (one workflow, not “our operation”), a measured baseline of what that work looks like today, success criteria agreed with the budget holder in advance, kill criteria that state exactly when you will shut it down, and a decision date. Structure the 90 days in four phases, baseline, instrument, measure, decide, and the approval conversation gets dramatically easier, because you are no longer asking the board to fund curiosity. You are asking them to fund an answer.

Before you scope a pilot, it pays to assess AI readiness first so you know which processes are worth testing.

Most pilots are not structured this way, and it shows in the outcomes. MIT’s 2025 NANDA research found that roughly 95% of enterprise generative AI pilots produce no measurable P&L impact. S&P Global found 42% of firms scrapped most of their AI initiatives in 2025, up from 17% the year before. Those pilots did not fail in month three. They failed at the design stage, before kickoff. This article covers how to design an AI pilot for an operations environment, claims, loan processing, back-office work, so that it gets funded, survives scrutiny, and produces a result you can defend either way.

Key Takeaways

  • Boards approve bounded decisions with dates attached. They reject open-ended “explore AI” asks and success criteria that cannot fail.
  • Structure the 90 days in four phases: baseline the work, instrument the pilot, measure against the baseline, decide.
  • Success criteria only survive scrutiny if the “before” number is measured, not estimated. The baseline is the pilot’s foundation, and most teams skip it.
  • Pre-agreed kill criteria are not an admission of doubt. They are what make your funding ask credible.
  • The deliverable of a pilot is not a working tool. It is a defensible scale, fix, or kill decision.

What Boards Approve, and What They Reject

Sit through enough budget reviews and the pattern is obvious. Rejected AI pilot proposals share a shape, and so do the funded ones.

What gets rejected:

  • Open-ended exploration. “We want to pilot AI in operations” is not a proposal, it is a hobby with a budget line. No scope boundary means no cost boundary, and finance knows it.
  • Tool-first pitches. A pilot built around a vendor demo answers the vendor’s question (does the product work) instead of yours (does this change our numbers).
  • Success criteria that cannot fail. “Improve efficiency” and “build organizational learning” are unfalsifiable. A board that cannot picture the failure condition assumes you have not either.
  • No decision date. Pilots without an end date become permanent line items. Everyone has watched that movie.

What gets approved:

  • One workflow, named. Claims intake triage. Loan document verification. Meter-to-cash exception handling. A boundary the CFO can see the edges of.
  • A measured baseline. Not “we think the team spends a lot of time on this,” but the actual hours, volumes, and cycle times the workflow consumes today.
  • A named owner. One accountable operations leader, not a committee. Executive sponsorship matters, but accountability is what boards fund.
  • A pre-committed decision. “On day 90 we scale, fix, or kill, against thresholds we are agreeing to now.” That sentence, more than any ROI projection, is what separates funded pilots from rejected ones.

You cannot be accountable for decisions you do not have the data to make. The whole design below exists to put that data in your hands by day 90.

The Four-Phase Structure of a 90-Day AI Pilot

Most 90-day AI templates you will find run diagnose, build, optimize. Notice what is missing: the baseline. If you do not measure the work before the AI touches it, you will spend day 90 arguing about what the “before” was. Here is the structure that avoids that.

Phase Days Purpose Output
1. Baseline 1 to 30 Measure the workflow as it runs today Activity-level “before” picture
2. Instrument 31 to 45 Deploy the AI on the scoped slice, define measurement Live pilot with agreed metrics
3. Measure 46 to 80 Run at steady state, capture results against baseline Before/after evidence
4. Decide 81 to 90 Score against pre-agreed thresholds Scale, fix, or kill decision

Phase 1: Baseline (Days 1 to 30)

Before the AI does anything, capture what the target workflow actually is: how many hours it consumes, across how many people, what portion is the rules-based work the AI is supposed to absorb, and what the current cycle times and error rates look like. This is activity-level measurement, not a survey and not a manager’s estimate. In a claims operation, that means knowing what intake triage really costs in hours per week and where the exceptions go, not what the process map says.

Thirty days feels long to teams eager to deploy. It is the highest-leverage month of the pilot; every number you present on day 90 stands on it.

Phase 2: Instrument (Days 31 to 45)

Deploy the AI on the scoped slice of work, and set up the measurement alongside it. Decide now how you will count: same metrics as the baseline, same definitions, same data source. Keep humans in the loop for exceptions and regulated decision points, and log where the AI hands work back, because that handoff rate becomes one of your most honest metrics.

Phase 3: Measure (Days 46 to 80)

Run at steady state and resist the urge to tune the narrative. You are capturing the same activity-level picture as Phase 1, with the AI in place: hours absorbed, cycle time, error and rework rates, exception volume, and whether the freed capacity actually shows up or gets quietly reabsorbed. Five weeks is short, which is exactly why the measurement has to be continuous rather than anecdotal. A good week proves nothing. A measured month is evidence.

Phase 4: Decide (Days 81 to 90)

Score the results against the thresholds you agreed to at approval. Three outcomes, all legitimate: scale it (criteria met, production path already defined), fix it (specific criteria missed for identifiable, correctable reasons, one bounded iteration), or kill it (thresholds missed, no correctable cause). Then present the decision, not just the data. The pilot’s deliverable is the decision.

Success Criteria That Survive Scrutiny

Here is where most pilots quietly rig themselves to fail later. The success criteria sound rigorous in the proposal, “reduce processing time by 30%,” and then day 90 arrives and someone asks: 30% of what? Measured how? Against which period? If the answer is an estimate someone made in a planning meeting, your result is an argument, not a proof.

Criteria that survive a skeptical CFO share four properties:

  1. Baseline-anchored. The “before” number is measured activity data from Phase 1, not an estimate. “Document verification consumed a measured 340 hours per week across the loan ops team; the pilot absorbed 31% of it” survives scrutiny. “We believe we saved significant time” does not.
  2. Outcome-based. Hours absorbed, cycle time, unit cost, error rate. Not model accuracy, not usage counts. A tool can be used daily and change nothing.
  3. Time-bound. Measured over a defined window at steady state, not cherry-picked from the best week.
  4. Pre-agreed with the budget holder. If finance signs the criteria before the pilot, finance owns the definition of success. Nobody moves the goalposts on their own numbers.

Note what this implies about sequencing: the activity-based baseline is not a nice-to-have inside the pilot. It is the thing that makes every criterion falsifiable. This is the same measurement discipline that lets you respond to an AI headcount mandate with data instead of instinct , applied at pilot scale.

Why Most AI Pilots Fail

The failure statistics are grim, but the failure modes are boringly consistent. When a pilot shows no measurable impact, it is rarely because the model was bad. It is because the pilot was never designed to produce a measurement.

  • No baseline. The most common and most fatal. Without a measured “before,” even a genuinely successful pilot cannot prove it, and an unsuccessful one cannot be caught.
  • The demo masquerading as a pilot. Two enthusiastic users and a vendor success manager produce a great story and zero operational evidence.
  • Scope picked for flash, not measurability. The most impressive-sounding project is usually the hardest to measure and the riskiest to run first. Scope selection is its own discipline, covered in how to prioritize AI projects in operations ; the short version is that the first pilot should be high enough in value to matter and clean enough to measure.
  • Goalpost drift. Criteria defined loosely at kickoff get renegotiated at day 85, usually downward, and everyone in the room knows it.
  • No production path. The pilot succeeds and then dies anyway, because nobody scoped what scaling requires, so “success” leads to a second discovery project instead of a rollout.
  • Success theater. The pilot is declared a win because too much was spent for it to be a loss. This is how organizations accumulate a portfolio of “successful pilots” and no measurable P&L change, which is precisely the 95% pattern.

Every one of these is a design failure, addressable before approval. Which is the point: the approval package below is not bureaucracy, it is the failure-mode checklist in document form.

Kill Criteria Are a Feature, Not an Admission

Operations leaders often hesitate to put kill criteria in a proposal, on the theory that naming the failure condition invites it. The opposite is true, for three reasons.

First, kill criteria are what make your ask credible. A leader who says “if the AI absorbs less than 15% of the baselined work by day 80, or exception rates exceed the current error rate, we shut it down” is demonstrably not asking for a blank check. Boards fund people who can articulate the downside.

Second, kill criteria cap the real risk, which is not the pilot budget. It is the scaled deployment of something that does not work. 55% of companies say they regret AI-driven layoffs, and the recurring theme is acting on confident projections instead of measured results. A pilot killed cleanly at day 90 is cheap. A failed capability scaled across an operation, with headcount decisions made on top of it, is not. Measure twice, cut once.

Third, a clean kill builds institutional trust for the next pilot. The first time an organization watches a leader shut down their own project on pre-agreed thresholds, the second proposal gets approved faster. Killing a pilot on criteria is not a failure of the program. Funding four pilots and declaring all four ambiguous successes is.

The Approval Package: What to Put in Front of the Board

Everything above compresses into a one-page pilot charter. If you cannot fill in every line, you are not ready to ask.

  • Scope: the single workflow, and what is explicitly out of scope
  • Baseline plan: how the work will be measured in Phase 1, and by what method
  • Success criteria: the specific thresholds, anchored to the baseline, signed by the budget holder
  • Kill criteria: the specific thresholds that end it, and who pulls the trigger
  • Owner: the one accountable name
  • Timeline: the four phases, with the decision date on the calendar
  • Production path: what scaling would require if the pilot succeeds, in one paragraph
  • Cost: the pilot budget, all-in

A charter like this changes the meeting. You are presenting a bounded decision with a date, an owner, and a defined downside. That is a shape boards recognize and fund.

Where the Baseline Comes From

The honest obstacle to all of this is Phase 1. Most operations do not have activity-level data about their own workflows, and gathering it manually, through shadowing and interviews, is slow, biased, and exactly the method that produces the estimates a skeptical CFO will not accept.

This is the problem the Ground Truth AI² Platform exists to solve: it captures individual-level activity data automatically, across everyone in scope, and turns it into a measured picture of what the work actually is, which workflows are genuinely automatable, and what a defensible pilot scope looks like. Run as a 90-day assessment , it produces the baseline, the prioritized pilot candidates, and the measurement infrastructure your pilot will report against, so your first pilot starts with its hardest phase already done.

Expensive assumptions are still assumptions. A pilot built on a measured baseline is the difference between proving your AI decision and defending a guess.

FAQ: Structuring an AI Operations Pilot

How long should an AI pilot take? 90 days is the practical standard for operations: 30 to baseline, 15 to instrument, 35 to measure at steady state, 10 to decide. Shorter pilots skip the baseline or the steady-state window and produce evidence too thin to fund a scale decision.

What is the difference between an AI proof of concept and a pilot? A proof of concept tests whether the technology works, usually offline. A pilot tests whether it changes your operation’s numbers under real conditions, with real work and real exceptions. Boards fund production decisions, so structure a pilot, not a science project.

What are good success criteria for an AI pilot? Outcome metrics anchored to a measured baseline: share of baselined hours absorbed, cycle time change, error and rework rates, exception volume. Defined before launch, measured at steady state, and signed off by the budget holder. Model accuracy and usage statistics are not success criteria.

Should an AI pilot have kill criteria? Yes, and they should be in the approval document. Pre-agreed kill thresholds make the funding ask credible, cap the real risk of scaling something that does not work, and build the trust that gets your next pilot approved faster.

Why do most AI pilots fail? MIT research puts the share of enterprise GenAI pilots with no measurable P&L impact around 95%. The dominant causes are design failures: no measured baseline, unfalsifiable success criteria, scope chosen for impressiveness, no pre-committed decision point. All fixable before approval.

Planning an AI pilot you will have to defend? Book a 30-minute strategy call and we will show you what a measured baseline of your operation would change about the proposal.

Mail Signup Section

Ready to Help Your Team Reach the Peak? See us in Action.