5 minute read. Published 2026-10-08.
Release readiness for AI features: seven checks beyond the usual test pack
Conventional tests cannot tell you if an assistant got worse. Add these seven checks to the release gate.
Why the usual pack is not enough
Unit, API and end to end tests check that behaviour stays the same. AI features are probabilistic and their quality is a matter of degree, so a change can pass every conventional test and still make answers worse. Release criteria need measures of quality, safety and cost as well as pass or fail checks.
Seven checks
- Quality score on a labelled evaluation set meets the threshold and does not fall against the last release.
- Groundedness: claims are supported by retrieved evidence, and the system refuses when evidence is missing.
- Retrieval quality: expected sources appear in the top results for known questions.
- Access boundaries: users cannot retrieve content outside their permissions.
- Prompt injection: instructions hidden in documents or messages do not change behaviour or trigger tool calls.
- Latency and cost per request are within budget at expected load.
- Rollback: the previous prompt, model version and index can be restored quickly.
Making it routine
Run the evaluation in CI on every change to a prompt, model, retrieval setting or tool. Show the results beside the conventional test results on the readiness dashboard, and record any accepted risk with a name against it. The Intelligent Quality Engineering Platform demonstration lets you adjust thresholds and see how the gate decides, using simulated data.