AI Aimaiaim.org

How can a team review an AI use case without quietly changing the task?

A team can review an AI use case without quietly changing the task by freezing a written baseline, checking every new metric or test against that baseline, and treating changes to the task’s purpose, context, inputs, outputs, or downstream decision as version changes. The GOV.UK Service Manual advises defining the service’s purpose before designing metrics, while NIST AI RMF states that AI output is interpreted within its context. A warning sign appears when a proposed metric or test no longer traces to that purpose or changes the context in which results are judged.

Start with a fixed task baseline

Before evaluating new results, the team should record what the use case is meant to do. A practical baseline can contain the following:

Baseline field What to record Possible task-drift warning
Purpose The service outcome or decision the use case is meant to support A new metric rewards a different outcome
Context The setting in which the output will be interpreted and used The test changes the conditions that give the output meaning
Inputs The information the use case is expected to receive Data selection or preprocessing changes the underlying task
Outputs The result or form the use case is expected to produce An easier-to-score output replaces the intended result
Downstream decision What action the output supports Evaluation begins optimizing a different action
Exclusions What remains outside the use case’s scope New scope is treated as a clarification rather than a change

The baseline should have a version and retain the exact wording used for review. Each later review can then compare the current proposal with that record instead of assuming the task is still the same.

A recurring review should reopen the baseline before discussing performance. Otherwise, a familiar score can gradually become a substitute definition of the task.

Check every metric against the original purpose

The GOV.UK Service Manual’s sequence is straightforward: define the service’s purpose clearly before starting metric design. Applied to an AI use-case review, each proposed metric should therefore have an explicit connection to the baseline purpose.

The team can ask:

  • Which baseline purpose does this metric support?
  • Is it measuring the intended service result, or only the system’s output?
  • If it is a proxy, what evidence shows that changes in the proxy correspond to changes in the purpose?
  • Could the result improve because the task, input, output, or evaluation conditions became easier?
  • Does the metric answer the original operational question, or has it become the new objective?

A metric does not have to remain unchanged for the task to remain stable. The key is whether the new measure represents the same purpose without expanding, narrowing, or redirecting the use case.

Treat context as part of the task

NIST AI RMF’s point about context matters because identical output does not necessarily have the same meaning in every setting. A review can therefore alter the task even when the original wording remains intact if it changes who receives the output, how it is interpreted, or what happens next.

A context map should contain only factors relevant to the actual use case. Possible candidates include the operating setting, recipient, handoff, downstream action, and exceptions, but they are not universal requirements. The team should record which factors affect interpretation and which remain unknown.

For each proposed test, the review can ask:

  • Will the output be interpreted under the same conditions as in the baseline?
  • Will it lead to the same downstream action or decision?
  • Has an exception, recipient, or operating condition been added or removed?
  • Is the proposed result still comparable with earlier results?

If the answer to any of these questions is unclear, the difference should remain visible in the review record rather than being absorbed into the score.

Separate clarification from task change

Every review should distinguish wording changes, measurement changes, and changes to the work itself.

Observed change Classification Review consequence
Wording becomes clearer without changing the baseline fields Clarification Preserve both versions and document that the meaning is unchanged
A metric, test, or evidence source is added while purpose and context remain intact Review expansion Check whether it still measures the same task
Purpose, inputs, outputs, context, downstream decision, or exclusions change Task change Create a new baseline and make the change explicit before treating results as comparable

Results from different task versions should not be presented as a continuous performance history unless the team documents why they remain equivalent. Otherwise, an apparent improvement may reflect a changed task rather than better execution of the original one.

What the team must still confirm

The review can end with one of four evidence-based conclusions: the use case still matches its baseline; the proposal changes measurement but not the task; the proposal changes the task; or the available evidence is insufficient to decide.

The team must separately confirm the project-specific purpose, context map, metric relationship, approval requirements, unresolved assumptions, and review interval from its own approved records. The two cited statements do not establish a universal review form, pass threshold, risk tolerance, approval rule, or legal conclusion. Missing items should remain explicitly unresolved rather than filled in by assumption.

Purpose first, context explicit, baseline fixed, changes visible: that is the practical safeguard against evaluating a different task under the name of the original one.