AI Aimaiaim.org

Which AI outcomes are visible to users, and which remain internal?

User-visible AI outcomes are effects a person can observe or act on in the live service, such as content they receive, a recommendation they can follow, a classification shown in an interface, or a workflow action that changes the service state. Internal outcomes are model- or system-level results that the user does not encounter, such as a backend confidence score or an offline evaluation result. A visible AI output is not automatically an outcome: the outcome is its effect on the user’s task or need.

How to check the boundary

Observation Classification
A person receives or encounters an AI-produced result User-visible output
The result changes what the person can see, do, or decide User-visible outcome
The result exists only within the model pipeline or an evaluation process Internal outcome or indicator
An internal measure influences a user-facing result Internal indicator linked to a user-visible outcome

The location of a measure helps make the initial distinction, but the user experience determines the final classification. A confidence score may remain internal when it is used only by the system. If an uncertainty warning or another expression of that score appears in the interface, the user-facing warning becomes visible even though the underlying score remains internal.

Why an output is not necessarily an outcome

A generated draft is a visible output. The relevant outcome concerns whether and how that draft helps a person complete the intended task. Similarly, a recommendation is visible, while the outcome may be the decision or action a person takes after receiving it. An internal classification can also produce a visible effect when it determines what the person sees.

This prevents an AI capability from being treated as a user benefit without examining its actual effect. The measure should describe what the person experiences against the need the service is designed to meet, not merely that an AI component produced something.

How to select measures

Product teams can use three complementary views:

  • User-visible outcome: What changed in the person’s ability to complete the task, receive the intended result, or use the service?
  • Visible output: What did the AI system return or present?
  • Internal indicator: What model or system measure may help explain that output or effect?

Internal indicators can be useful, but they should not replace measures of the user-facing effect. They should be connected to a visible result and interpreted through the context in which the service is used.

The GOV.UK Service Manual states that performance metrics should reflect the user needs a service is designed to meet. It also says that metrics should be combined with user research, with the service iterated using insight from both.

What still needs confirmation

Before classifying a measure as an achieved outcome, the product team still needs to establish:

  • which user need the result is intended to serve;
  • where the result becomes visible or actionable;
  • whether the measure captures an output, an outcome, or an internal indicator;
  • what evidence shows that the result meets the need in context; and
  • whether user research supports the interpretation of the metric.

This classification provides a boundary, not a universal success threshold. A specific team must define what constitutes a meaningful user-visible result and validate that interpretation through both measurement and user research.

Sources