A team should treat the decision as an evidence review rather than treating a launch score as a verdict. Continue when the evidence still supports the current scope and further learning remains justified; narrow when usefulness appears limited to a defined context; redesign when the objective remains relevant but the current system or workflow needs material change; stop when the evidence or unresolved risk does not justify further work.
The common decision rule is to combine performance metrics with user research. The GOV.UK Service Manual advises combining metrics with user research and iterating based on insight from both. Testing provides a second check: the NIST AI RMF says AI systems should be tested before deployment and regularly while in operation.
How the four decisions differ
| Decision | Evidence pattern | Appropriate action |
|---|---|---|
| Continue | Metrics and user research still support the intended use, and the remaining questions can be addressed through another review cycle. | Keep the current scope, continue measurement and research, and define what evidence will trigger the next decision. |
| Narrow | The evidence supports the use for a particular user group, task, or operating context, but not for the broader scope. | Restrict the pilot to that context and test whether the narrower application remains supportable. |
| Redesign | The intended outcome remains relevant, but weaknesses in the system, workflow, interface, or measurement approach require a material change. | Revise the design, repeat relevant testing and user research, and then reassess the evidence. |
| Stop | The current evidence does not support further work, or unresolved risks and uncertainties exceed the team’s stated tolerance. | End the current pilot, document the reasons and evidence, and avoid carrying unsupported assumptions into the next initiative. |
These categories describe different responses rather than automatic pass-or-fail outcomes. For example, weak evidence across most use cases may justify narrowing, while a mismatch caused by the workflow may call for redesign. Stopping the current pilot does not, by itself, settle whether a materially different concept should be explored later.
What to check at each review
Frame the decision before interpreting the evidence
The team should state the intended use, current scope, affected users, workflow, and relevant failure consequences. Without that context, a metric can appear acceptable or unacceptable without showing whether the pilot is fit for its intended purpose.
Read metrics alongside user research
Metrics show measured performance, while user research can reveal whether the observed behavior matches actual needs and expectations. Neither should be interpreted in isolation.
The review should record:
- What each metric measures and what it does not capture.
- How the results were collected and where uncertainty or measurement limitations remain.
- What users experienced, including confusion, workarounds, unmet needs, or unintended consequences.
- Which contexts and user groups the research covers, as well as important gaps in coverage.
A metric that appears favorable can still require caution if user research identifies a material problem. Conversely, favorable feedback should not erase weak or incomplete performance evidence.
Check testing across the system’s lifecycle
Testing should occur before deployment and continue regularly during operation, following the NIST AI RMF statement. The team must translate that principle into context-specific checks by identifying relevant conditions, testing what happens under those conditions, examining failures, and checking whether system changes alter earlier results.
Neither cited statement supplies a universal test method, test set, or review interval. The team must therefore determine what “regularly” means for the pilot and what evidence would require additional testing.
Look for evidence that could change the decision
The review should not focus only on results that support continuation. It should also examine contradictory findings, user-reported problems, untested conditions, and assumptions that have not been checked. Any redesign or scope reduction should identify what must be retested rather than carrying forward conclusions drawn from the previous design.
Record the decision and its basis
The decision record should identify the chosen response, the metrics and research behind it, important unknowns, and the conditions that would cause the team to revisit the decision. This makes clear whether the team is continuing because the present evidence remains persuasive or because further learning still appears justified.
What the team must still confirm
The cited guidance establishes the need to combine metrics with user research and to test both before deployment and during operation. It does not provide project-specific thresholds or a universal decision rule. The team must confirm:
- The definitions, collection methods, and limitations of the selected metrics.
- Whether user research adequately represents the people and situations affected by the pilot.
- Which tests are relevant before deployment and which should be repeated as operation continues.
- What performance and risk levels are acceptable for the intended use.
- Who has authority to continue, narrow, redesign, or stop the pilot.
- When the next review is needed and what new evidence could change the decision.
Until those details are resolved, the team should state them as explicit assumptions rather than infer a score, threshold, or outcome that the available evidence does not support.