Setting up an AI pilot that does not stall: scope, success and stop criteria

Drie collega's werken op laptops aan een vergadertafel
Date
September 23, 2026
Author
Isatis Group
Category
AI & Innovation
Read time
9 min read

An AI pilot succeeds when it ends in a real decision. Learn how to fix scope, a baseline, success and stop criteria, and a test set before you start.

An AI pilot works when it ends in a real decision. To get there, agree on five things before you start: one process, a measured baseline, success and stop criteria, real data with a fixed test set, and one owner who decides on scaling up. Then build the pilot on foundations that can stay in production, so a successful trial doesn't have to be rebuilt.

Most organizations don't get that far. IDC, working with Lenovo, looked at how many AI proofs of concept reach production: for every 33 launched, four made it, and 88% never reached wide deployment. In July 2024, Gartner predicted that at least 30% of generative AI projects would be abandoned after proof of concept, citing poor data quality, escalating costs, and unclear business value among the causes.

You can prevent most of those causes with choices you make before the pilot starts. Here they are, one by one.

Why an AI pilot stalls

AI is now part of everyday work. In the Netherlands, where our head office is, statistics agency CBS found that one in six businesses used AI in 2025, twice as many as two years earlier. Among companies with 250 or more employees, the figure was 66%.

Use says little about results. In McKinsey's global survey of 1,993 respondents, 88% say their organization regularly uses AI in at least one business function. About one third have begun to scale their AI programs, and 39% report any impact of AI on EBIT at the enterprise level.

A pilot starts with an impressive demo. A few months later there's a working prototype, and it turns out nobody agreed on when it would be good enough. Three causes keep coming back:

  • No yardstick. Without a baseline and criteria agreed in advance, the evaluation becomes a matter of taste.
  • Demo data. A model that scores well on cleaned sample data behaves differently on the data in your own systems.
  • No owner. When nobody has the authority to decide on scaling up, the decision slides into the next quarter.

"After last year's hype, executives are impatient to see returns on GenAI investments, yet organizations are struggling to prove and realize value." Rita Sallam, Distinguished VP Analyst at Gartner

Pick one process and keep the scope small

A good AI pilot tests one use case in one process with one team. Choose a process with plenty of repetition, a clear start and end, and an outcome you can check. Think of classifying incoming documents, drafting a reply to a recurring customer question, or checking an application for completeness. Processes you've already mapped for business process automation are often a good place to start.

Write the scope on a single page that everyone on the project can read:

  1. The process. Which step, owned by which team, with what input and what output.
  2. The role of AI. Whether the system makes a suggestion that an employee reviews, or carries out a step on its own. For a first pilot, choose the suggestion.
  3. The boundaries. What falls outside the pilot: other departments, rare exceptions, extra integrations.
  4. The decision date. The day the result goes on the table.

That second choice reflects how we see AI: it works best when people stay close to it. An employee who reviews every suggestion also gives you measurement data for free. Each correction shows where the system goes wrong. The idea is close to a minimum viable product: as small as possible, as real as necessary.

Start with a baseline

Without a baseline, you can't say whether the pilot improves anything. Measure how the process performs today, before the pilot starts, using the same definitions you'll apply to the pilot later. Three or four numbers are enough:

  • lead time per case or request
  • the percentage of errors or rework
  • minutes of manual work per case
  • weekly volume, so you know how many cases the pilot touches

Measure over a representative period, including busy weeks and exceptions. If the numbers come from an existing system, that source is also where you'll measure the pilot. If nothing is recorded in a system, a few weeks of manual tallying beats an estimate.

The baseline also tells you whether the process is worth the effort. A task that takes half an hour a week won't return much, even with the best model.

Agree on success and stop criteria up front

Success criteria describe when the pilot has worked. Stop criteria describe when you end it early or decide not to continue. Agree on both before the team writes the first line of code, each with a number, a measurement method, and a data source. In the table, x is a value you set from the baseline.

Type of criterion What you record Example wording
Quality How often the output is correct, measured on the test set At least x% of suggestions are correct according to the team's review
Time Gain compared with the baseline Lead time per case drops by at least x%
Adoption Whether employees actually use the suggestion Employees accept at least x out of 10 suggestions without major edits
Stop on quality When further tuning makes no sense After two improvement rounds, the pilot still scores below the baseline
Stop on data When the foundation is missing The required data isn't available in usable form within the pilot period

Stop criteria turn stopping into a result. A pilot that ends on a criterion agreed in advance has done its job: you know something you didn't know before, at limited cost. Also agree that criteria can only change during the pilot through the owner, with a short note explaining why.

Use real data and a fixed test set

Demo data is clean, complete, and predictable. Real data rarely is: fields are missing, text sits in the wrong field, scans are skewed, and abbreviations differ between departments. Gartner lists poor data quality first among the reasons projects are dropped after proof of concept, and IDC points to data that isn't ready for AI. So work from the first week with a representative extract from the systems the process uses today, within the data access arrangements your organization already has.

Next, build a fixed test set: historical cases where the correct outcome is known, reviewed by people who work in the process. Four rules keep that set useful:

  • Keep the test set separate from any data you use to configure or train the system.
  • Make sure difficult cases and exceptions appear in the same proportion as in daily work.
  • Have two employees review borderline cases independently to see where even people disagree.
  • Don't change the test set during the pilot, and add new cases as a separate set.

After the pilot, you run every change to the model, instructions, or integrations against the same set. That makes it one of the most important deliverables.

Decide up front who owns the scaling decision

Appoint one owner at the start, with the mandate and budget for the phase after the pilot. Usually that's the manager responsible for the process. IT and the innovation team advise, and the owner decides. Put the decision date in the calendar on day one.

On that date, there are three possible outcomes:

  1. Scale up. The success criteria are met. Team, planning, and budget for the production phase have already been discussed, so there's no gap of several months.
  2. Adjust. Results are close to the bar and there's an identifiable cause. You plan one extra round with a new date and the same criteria.
  3. Stop. A stop criterion has been hit. You record what you learned about the process, the data, and the use case, and you close the pilot.

Build the AI pilot so it can grow into production

Many pilots are built as an island: a script, a test environment, and an export from the source system. If the result is good, the team starts over for production. So build a few parts the way you'd want them in production from day one:

  • An integration with the source systems through an API, in place of manual exports.
  • Logging of every input, output, and correction, so you can trace why the system made a suggestion.
  • A screen where employees review suggestions, accept them, or change them.
  • Version control for the model, instructions, and settings, linked to test set results.
  • An automated test that runs the test set again after every change.
  • An off switch, so the process keeps running without AI when needed.

This is the part of an AI pilot where software engineering makes the difference. We approach data and AI from the business value it delivers, and we build the pilot as the first version of custom software that can grow from there. For more than 30 years, Isatis has built software that stays in production for years, from maintenance software for aviation to a subsidy administration platform and RFID in the supply chain. With 30+ engineers in Nijmegen and Sarajevo, 100+ projects, and ISO 9001 and ISO 27001 certification, we know what a system needs to keep standing after the trial. Book a call with no obligation and we'll look together at which process suits a first pilot.

Frequently asked questions

What is the difference between a proof of concept and an AI pilot?

A proof of concept shows that something is technically possible, usually with a limited or cleaned dataset. An AI pilot tests whether the use case delivers enough in a real process, with real data and real users, to justify going further. That's why a pilot ends with a decision: scale up, adjust, or stop.

Which success criteria suit an AI pilot?

Criteria you can measure against the baseline: output quality on a fixed test set, time saved per case, less rework, and the share of suggestions employees accept. For each criterion, record a number, a measurement method, and a data source before the pilot starts.

When should you stop an AI pilot?

As soon as a stop criterion agreed in advance is hit. For example, when quality on the test set stays below the baseline after an agreed number of improvement rounds, or when the required data doesn't become available. Stopping on such a criterion is a valid outcome, because you've learned at limited cost that this use case doesn't pay off right now.

How many AI pilots reach production?

A minority. IDC found that for every 33 AI proofs of concept, four went into production, and 88% didn't make it. According to McKinsey, about one third of organizations have begun to scale AI.

Share this article

Jack van Poll

Jack van Poll

Co-Founder, Isatis

Writes about nearshore engineering,
software partnerships and building teams that last.

Get in touch

Read next.

All articles