Review questions should ask what users do differently, where the workflow path changes, and how those actions relate to the needs the service is designed to meet. Metrics should be paired with user research because observed patterns and reported reasons answer different questions. The GOV.UK Service Manual advises combining performance metrics with user research, iterating from insight from both, and choosing metrics that reflect user needs.
Which questions expose behavior changes?
| Review question | Behavior to examine |
|---|---|
| What has changed since the previous review? | Differences in task completion, step sequence, retries, corrections, verification, handoffs, or stopping. |
| Which expected actions no longer occur? | Steps that users now skip, repeat, reorder, or add. |
| How do users continue, revise, verify, or stop? | Actions that may indicate reliance, resistance, correction, or disengagement. |
| Where does the workflow change direction? | Workarounds, tool changes, pauses, handoffs, abandonment, or re-engagement. |
| Which metric changed alongside the behavior? | Patterns tied to the user needs the service is designed to meet. |
| What do users say is happening? | Reported context and reasoning that may explain—or complicate—the metric pattern. |
How to check the answers
A useful review keeps four elements separate:
- Observed behavior: what users actually did.
- Metric: how that pattern is represented.
- Interpretation: what reviewers think the change may mean for a user need.
- User research: what users report about their actions and context.
Comparing metrics with research can expose where an interpretation does not fit the evidence. A changed metric may have several explanations, while a reported frustration may describe an experience without showing how widely it occurs. Neither observation alone provides a complete basis for iteration.
Reviewers should therefore avoid labeling a behavior change as beneficial, harmful, or causal until the relevant context and user needs support that conclusion.
What teams must still confirm
The cited guidance does not define an AI-specific checklist, a single correct metric, or a universal review interval. Each team must confirm the relevant user needs, comparison baseline, behavior context, and whether user research supports the interpretation. Without that confirmation, a review can report a behavior signal, but it cannot establish a verified explanation of user motivation or impact.