A post-launch review should decide whether the next experiment is warranted, what service purpose it serves, which unresolved question it will test, and what evidence would change the team’s decision. The purpose should be defined before metrics, as GOV.UK advises. Testing should not end at launch: NIST states that AI systems should be tested before deployment and regularly while in operation.
How to check the decision
The review can examine the proposed experiment through four connected questions:
-
Does it serve a defined purpose?
The team should be able to state the service’s purpose before selecting a metric. GOV.UK specifically advises defining that purpose first, so the experiment can be evaluated against a clear reason rather than an arbitrary measure. -
Does it address an unresolved question?
The proposed experiment should respond to an evidence gap identified after launch, such as an untested assumption or a question left by operational testing. A preferred feature or another release is not, by itself, sufficient justification. -
Does it fit the testing lifecycle?
The review should check whether the AI system was tested before deployment and whether regular testing while in operation is planned or taking place. NIST supports both points, but the cited statement does not prescribe a testing method or interval. -
What will the result change?
Before interpreting the experiment’s findings, the team should record what evidence would support continuing, revising, or ending the current approach. This is a practical review control, not an additional requirement stated by either cited source.
An experiment with a clear purpose but no identified evidence gap remains incomplete. Likewise, an evidence gap without a defined decision consequence does not yet establish why the experiment should run.
What the team must still confirm
The cited statements do not establish a universal metric, testing interval, experiment duration, threshold, approval role, or sequence for future experiments. The team must confirm those elements against its own service context, testing records, operational evidence, and any applicable official materials.
It must also determine whether another experiment is actually needed, which decision the experiment will inform, and whether its measures can be traced back to the defined service purpose. Those conclusions should remain team-specific rather than being presented as GOV.UK or NIST requirements.