A product team should compare the same measurement frame during planning and after a pilot: what each metric means, what data shows performance against it, and whether the evidence was produced under conditions similar to deployment. Planning establishes that frame; the post-pilot review checks whether the evidence answers the same questions.
How to check it
The GOV.UK Service Manual says to design metrics with a clear meaning and collect data that shows how a service is performing against them. The NIST AI RMF says that AI system performance or assurance criteria are measured qualitatively or quantitatively and demonstrated for conditions similar to deployment settings.
Those points can be applied as a direct review comparison:
| What to compare | Planning review | Post-pilot review |
|---|---|---|
| Metric meaning | State what each metric means and what it is intended to show. | Check whether the definition used in the pilot matches the planned definition. |
| Performance data | Specify the data that will show performance against each metric. | Compare the collected data with the planned performance question. |
| Measurement method | Decide whether relevant performance or assurance criteria will be measured qualitatively or quantitatively. | Record how each criterion was measured and what the resulting evidence shows. |
| Conditions | Define the conditions that should resemble the intended deployment setting. | Compare the pilot conditions with the intended deployment setting and record any differences. |
Using the same comparison fields at both stages helps the team distinguish a change in results from a change in the question being evaluated. If a metric definition, measurement method, or test condition changes after planning, the post-pilot review should record that change rather than treating the results as directly comparable.
What the team must still confirm
The cited guidance does not specify a universal set of AI metrics, numerical thresholds, minimum data requirements, review deadlines, or pass/fail rules. Those choices are not supplied by the sources and should not be attributed to them.
The team must confirm its own:
- intended deployment setting and the conditions that can reasonably represent it;
- metric definitions and data-collection plan;
- qualitative or quantitative measurement approach for each relevant criterion;
- way to document differences between the pilot and deployment conditions; and
- decision rule for interpreting the evidence.
Without those confirmations, a post-pilot result can be described as evidence from the tested setting, but not as a general conclusion about deployment.