6 minute read. Published 2026-10-08.

Why AI pilots fail in production, and what to build before launch

A pilot proves a model can do something once. Production needs it to do it every day for strangers.

A demo answers friendly questions

Pilots are usually built and judged by the people who made them. They ask questions they know the system can answer, with clean inputs, one at a time. Real users are different. They misspell, paste whole documents, ask about things you did not plan for and sometimes try to break the system on purpose. The first production week is the first time the system meets that population.

Five gaps worth closing before launch

1. No evaluation set

If quality is judged by trying a few prompts, nobody can say whether a change made things better. Build a set of real and adversarial cases with expected outcomes, score every change against it and set a release threshold. A hundred well chosen cases beat a thousand arbitrary ones.

2. Unversioned behaviour

Prompts, model versions, retrieval settings and tool definitions all change behaviour. Keep them in version control, pin model versions where the provider allows it and run the evaluation on every change. Then a bad Wednesday has a diff.

3. Unmeasured cost

Cost scales with tokens, retries and context size. Meter by feature and tenant from the first release, set budgets and alerts and route simple tasks to cheaper models. A business case that ignores run cost is not finished.

4. Unbounded actions

If the system can act, it needs narrow tools, permissions tied to the user, limits and approval for consequential actions. Treat any text the model reads, from documents or emails, as untrusted input that may contain instructions.

5. No owner

Someone must be on call for the system, review sampled outputs and decide when to roll back. Decide who before launch, and write the runbook while the pilot team still remembers how it works.

A short readiness check

  • Do we have a labelled evaluation set and a threshold that blocks release?
  • Can we reproduce yesterday's behaviour from version control?
  • Do we know cost per request and per successful task?
  • Can the system do anything harmful, and who must approve it?
  • Who is paged when it fails?

If you cannot answer one of these, that is where to spend the next sprint. References for risk categories include the OWASP Top 10 for LLM Applications and the NIST AI Risk Management Framework.

Related services

Want help applying this?

Tell us where you are. We will suggest the first check to run.