AI Aimaiaim.org

What outcome measure fits a recommendation feature?

The closest fit is a user-outcome completion rate: the share of eligible users who meet a predefined success criterion for the task the recommendation is intended to support. Clicks, views, and accepted recommendations can provide useful diagnostics, but they do not by themselves show that the user’s need was met.

The cited public-service guidance states that metrics need a clear meaning, performance data, and a connection to the user needs the service is designed to meet. Applied to a recommendation feature, that means starting with the user’s intended outcome rather than with an available system event.

How to define the measure

A product team can make the outcome measure testable by answering five questions:

  1. What user need is the feature meant to support?
    The measure must correspond to that need. A measure of completed purchases, for example, is inappropriate if the intended need is to help someone identify a suitable option.

  2. What observable event defines success?
    “The recommendation was useful” is not measurable by itself. The team must specify what the user does—or what can be verified about the result—that demonstrates successful completion.

  3. Who is included?
    The denominator must match the intended user population. Users who were not eligible for a recommendation should not be mixed into the completion rate unless the measure explicitly covers them.

  4. When is completion assessed?
    The observation window must reflect the task. A decision made during the interaction and a decision completed later require different measurement periods. No universal window is established here.

  5. Which supporting measures are needed?
    Interaction and operational measures should remain separate from the primary outcome so that activity is not mistaken for user value.

Measure What it indicates Appropriate role
Successful task completion Whether the defined user need was met Primary outcome
Recommendation acceptance or click-through Whether users selected the recommendation Diagnostic
Time to completion Efficiency or friction Primary only when efficiency is the intended need; otherwise secondary
Failure, error, or abandonment rate Unsuccessful or difficult experiences Guardrail
Recommendation coverage or availability Whether the system could present recommendations Operational measure

How to check whether the measure is meaningful

The definition should remain stable long enough for the team to interpret the collected data. The team can also review the measure by user segment, task type, and other relevant context. A single aggregate rate may hide materially different experiences.

Any baseline, comparison group, or target should be documented separately from the metric definition. A change in completion after launch does not, by itself, prove that the recommendation feature caused the change; that interpretation requires an appropriate evaluation design.

What the product team must still confirm

Before implementation, the team must confirm:

  • the exact user need and scope of the feature;
  • the observable success criterion;
  • the user or interaction unit used for measurement;
  • the denominator and observation window;
  • data availability, quality, and governance;
  • applicable privacy, consent, and security requirements;
  • relevant guardrails and unintended consequences;
  • whether a baseline, comparator, or target is required for a performance decision.

The central choice is not which interaction occurs most often. It is whether users reliably achieve the outcome the recommendation was designed to support.

Sources