Evidence should connect a change in the user’s final action to the AI workflow—not merely show that the output was viewed, praised, or produced quickly. A credible case combines a clearly defined decision metric, a suitable comparison, and evidence that the user interpreted the output in the actual decision context.
GOV.UK says service metrics need a clear meaning and data about performance. NIST says AI system output is interpreted within its context. Those points provide measurement discipline, but neither source alone proves that a particular workflow changed a particular decision.
What evidence matters?
| Evidence | What it can establish | What it cannot establish alone |
|---|---|---|
| Recorded final action | What the user selected after using the workflow | Why the user selected it |
| Defined comparison | Whether the final action differed from a baseline | That the workflow caused the difference if other conditions also changed |
| Workflow trace | Whether the output reached the decision point and was considered | What the user would have decided without it |
| Decision rationale | How the output related to the user’s goal, constraints, and available options | A general causal effect across other users or tasks |
| Repeated comparable cases | Whether the same pattern appears under similar conditions | Whether confounding factors explain the pattern |
A matching decision does not automatically prove that the output influenced the user. Likewise, an overridden output does not automatically prove that it had no effect. The process record and the user’s explanation are both relevant.
How to check the evidence
1. Define the decision event
The team should identify when the decision is made, which actions count as possible decisions, and which recorded action represents the final choice. Viewing, accepting, or clicking an AI-generated recommendation is not necessarily the same as making the final decision.
2. Give the metric a clear meaning
A decision-change metric could compare the final action in an AI-assisted case with the action in a defined baseline case. The team should state which cases are included, how the baseline was created, and how missing observations are handled.
This follows the GOV.UK measurement principle: a metric is useful only when its meaning is clear and the collected data show how performance compares with it.
3. Make the comparison credible
The user’s goal, available options, instructions, timing, and evaluation conditions should remain as comparable as possible. If they change at the same time as the workflow, a before-and-after difference cannot be attributed confidently to the AI output.
When a controlled comparison is not possible, the result should be described as an observed association rather than causal proof.
4. Connect the output to the decision point
The record should distinguish among whether the output was delivered, viewed, considered, accepted, modified, or rejected. It should also preserve the subsequent action and the stated reason for it.
Applied to this question, NIST’s context principle means that interpretation should not be separated from the circumstances in which the decision was made.
5. Check alternative explanations
A changed action may also reflect different inputs, a changed process, different participants, or another change in the surrounding conditions. Those possibilities need to be examined before attributing the difference to the workflow.
6. Separate change from improvement
A different decision is not automatically a better decision. Evidence that the action changed answers the question asked here, but it does not by itself establish greater accuracy, effectiveness, fairness, or any other improvement. Such conclusions require their own metrics and objectives.
What the team must still confirm
Before making the claim, the team still needs to confirm that:
- the decision metric represents the intended decision rather than an earlier interaction;
- the comparison is valid for the stated task and population;
- the evidence connects the AI output to the decision point;
- material alternative explanations have been considered;
- the observed result is not being generalized beyond the tested context; and
- the wording distinguishes a recorded decision change from a proven causal or sustained effect.
Until those checks are complete, the defensible conclusion is that the workflow was associated with a recorded decision change, not that it necessarily caused the change.