AI Aimaiaim.org

How can a team turn a broad efficiency goal into a testable AI problem?

A team can turn a broad efficiency goal into a testable AI problem by rewriting it as a bounded hypothesis about a user, task, workflow, and measurable outcome. The hypothesis should specify the proposed AI intervention, the current process it will be compared with, the evidence required, and the conditions that would count as improvement.

Start with the user’s problem

The GOV.UK Service Manual advises focusing on the user’s problem rather than a possible solution. Applied to AI planning, this means beginning with work rather than with a model or vendor.

A broad goal such as “improve efficiency” does not identify whose work needs to change or what friction matters. A more useful starting question is:

Which users experience what friction while completing which task, and what evidence shows that this friction matters?

The answer should remain specific to a real workflow. It may concern time spent searching for information, repeated reworking, delayed handoffs, or difficulty finding an appropriate response. These are possible categories to investigate, not assumptions about what a particular team needs.

Define the hypothesis before testing it

A testable problem should answer several connected questions:

Element What the team should specify
User and task Who performs the work, and what are they trying to complete?
Context Under what conditions does the task occur?
Current condition What happens in the present workflow, including where time, effort, or rework accumulates?
AI intervention What specific AI-assisted step is being considered?
Comparator What current or feasible alternative workflow will it be evaluated against?
Primary measure Which observable dimension should improve?
Guardrails Which quality, risk, or responsibility requirements must remain intact?
Decision rule What evidence would lead the team to adopt, revise, or reject the approach?

A reusable statement is:

For [users] completing [task] in [context], test whether [specific AI intervention] improves [efficiency measure] compared with [current workflow], while meeting [defined quality and risk requirements].

For example:

For staff classifying incoming requests, test whether AI-assisted classification suggestions reduce review time and rework compared with the current workflow, while meeting separately defined quality and risk requirements.

This is a hypothesis to test, not a claim that the intervention will work. The team must still determine the relevant cases, measures, and acceptance conditions.

Choose evidence that answers the question

The efficiency measure should reflect the actual bottleneck. Depending on the workflow, it might concern elapsed time, active handling time, repeated work, wait states, or the volume completed under the same conditions. A measure should be defined clearly enough that different reviewers would interpret it consistently.

Efficiency should also be evaluated alongside relevant guardrails. Faster work is not sufficient if the output introduces unacceptable omissions, quality failures, or unmanaged risk. The appropriate guardrails depend on the task and the team’s operating context.

For AI risk measurement, the NIST AI RMF describes quantitative, qualitative, and mixed-method tools, techniques, and methodologies. In practical terms:

  • Quantitative evidence can record observable changes in time, volume, rework, or other defined measures.
  • Qualitative evidence can reveal why a workflow is difficult, how users interpret a result, or where an intervention creates new friction.
  • Mixed-method evidence can connect an observed change with the reasons and conditions behind it.

A mixed-method test is particularly useful when an operational measure shows that something changed but does not explain why. The cited risk-measurement guidance does not establish a universal metric or pass mark; the team must define those elements for its own context.

What the team must still confirm

Before running a test, the team should verify:

  • The problem is evidenced. User research or workflow observation supports the selected user, task, and friction.
  • The baseline is defined. The current process, measurement method, and test cases are documented.
  • The intervention is bounded. The test evaluates a specific capability rather than an unspecified promise to “use AI.”
  • The measures are operational. Everyone knows what will be recorded, how it will be interpreted, and what quality or risk checks apply.
  • The decision rule is set in advance. Thresholds and stop conditions are not chosen merely because the results look favorable.
  • The test remains in context. Findings apply to the tested users, tasks, and conditions and should not automatically be generalized elsewhere.
  • Relevant responsibility requirements are addressed. The team confirms which review, escalation, or oversight conditions apply to its particular use case.

The resulting testable problem is therefore not simply, “Can AI improve efficiency?” It is: For this user and task, under these conditions, does this bounded intervention improve this measure against this baseline without violating this guardrail?

Sources