An AI team should document each metric before results arrive by recording the service purpose it supports, its precise meaning, and the data that will show performance against it. GOV.UK’s Service Manual provides the sequence: define the service’s purpose first, then design metrics with a clear meaning and collect data that shows how the service is performing against them.
What should the team document?
A metric name alone is not a definition. Before any results exist, the documentation should connect the measure to a defined service purpose and make clear what evidence will count as performance.
The following pattern is a practical way to organize that record; it is not an AI-specific format prescribed by GOV.UK:
| Documentation item | What to state before results arrive |
|---|---|
| Service purpose | The service objective the metric is intended to inform |
| Metric meaning | The precise concept being measured, including its scope and unit |
| Measurement boundary | Which observations, events, or cases are included or excluded |
| Evidence plan | The data to be collected, where it will come from, and how it will be checked |
| Comparison basis | The baseline, reference condition, numerator, or denominator relevant to the metric |
| Interpretation boundary | What the metric can show—and what conclusions it cannot support |
| Ownership and version | Who owns the definition, its approval status, and which version was agreed |
Actual values can remain pending. Ambiguity about the metric’s meaning or the planned evidence should be resolved before results are reviewed.
How can the team check that a definition is ready?
A pre-results review can apply several checks:
- Purpose alignment: Does the metric support the defined purpose of the service, rather than merely describe an activity?
- Consistent application: Could another reviewer apply the definition without asking the author to reinterpret it?
- Traceable evidence: Is it clear what data would demonstrate performance against the metric?
- Unambiguous scope: Are relevant units, inclusions, exclusions, and calculation boundaries stated where needed?
- Controlled comparisons: Is the intended baseline or reference condition explicit, so different results are not presented as equivalent?
- Visible changes: Will an earlier definition remain identifiable if it is later revised?
The team should also preserve the original definition. Any later revision should be recorded with its rationale and approval rather than silently replacing the pre-results version.
What must the team still confirm?
The GOV.UK guidance cited here should not be treated as settling an AI-specific template, target, threshold, data source, owner, review schedule, or legal, privacy, or security conclusion. The team must confirm those details against its own service and applicable requirements.
In particular, it must establish the intended decision, operational definitions, available evidence, comparison method, acceptance rules, ownership, and review arrangements. It should also state the limits of each metric so that a measure of performance is not presented as proof of a broader business outcome.