AI Aimaiaim.org

What should count as meaningful improvement over the current workflow?

Meaningful improvement should count when a change in the current workflow is measurable, relevant to a clearly defined purpose, and important enough to support the intended decision. A higher benchmark, greater automation, or a successful demonstration is not enough on its own; the evidence must show how the actual workflow performs against the chosen metrics.

The GOV.UK Service Manual advises teams to define the service’s purpose before designing metrics. It also recommends giving each metric a clear meaning and collecting data that shows how the service performs against it.

What should meaningful improvement demonstrate?

A practical evaluation should answer four connected questions:

Question Evidence to examine
Is the measure relevant to the purpose? Each metric should connect to an explicit service or workflow objective rather than exist only because it is easy to measure.
Does the metric have a clear meaning? The team should be able to state exactly what is being measured and how the result would be interpreted.
Does the data show better workflow performance? Collected evidence should support a comparison between the current workflow and the proposed one.
Is the change important in practice? The size of the observed improvement should be evaluated against an explicit threshold or decision rule established for that workflow.

The first two questions follow directly from the cited measurement guidance. The comparison and practical-importance tests turn those principles into a workflow evaluation method.

Depending on the stated purpose, relevant measures might include completion time, error or defect rates, failure rates, human review, rework, or workflow cost. These are examples of possible measures, not a universal measurement checklist. Their relevance depends on the purpose and the way the current workflow operates.

How to check the improvement claim

Before accepting an improvement claim, a team should be able to answer the following:

  1. What is the current baseline?
    The existing workflow needs a documented performance reference. Without one, a later result has no clear comparison point.

  2. How is each metric defined?
    Terms such as “accuracy,” “speed,” or “reliability” are not sufficient by themselves. Each needs an operational definition tied to the workflow.

  3. Are the measurements comparable?
    The evidence should not obscure material differences in workload, operating conditions, or cases handled. Otherwise, an observed difference may not represent a valid workflow comparison.

  4. Was the importance threshold set before evaluation?
    A predefined threshold makes it harder to redefine success after seeing the results. If no formal threshold exists, the team should record and justify the basis for its decision.

  5. Does the improvement survive the broader workflow?
    A gain in one measure may be accompanied by additional review, rework, failure handling, cost, or another downstream effect. These effects belong in the assessment when they affect the stated purpose.

  6. Is the result stable enough to rely on?
    A limited test may show that improvement is possible without establishing that it persists across normal variations in the workflow.

What still needs confirmation

The cited guidance does not establish a universal threshold for meaningful improvement. Each team must still confirm:

  • the purpose that matters most;
  • the current performance baseline;
  • metric definitions and data quality;
  • the comparison conditions;
  • the threshold for an important change; and
  • whether any trade-offs change the overall conclusion.

Without those checks, the available evidence may show activity or potential, but it does not yet establish meaningful improvement over the current workflow.

Sources