A review can distinguish the two by holding one side constant and changing the other: compare the same model configuration across different product contexts, then compare different model configurations within the same product context. Evidence points toward product drift when results change mainly with the operating context, and toward model behavior when results change under equivalent conditions.
The distinction is not always conclusive. If both the product context and the model configuration change together, the observed difference is confounded until the review team separates those changes.
How to check the distinction
Map the product context
The review should first document the conditions surrounding the system: intended use, user inputs, workflows, policies, data sources, instructions, interface changes, downstream decisions, and other relevant constraints. It should then compare the current map with the earlier version.
Changes in these conditions can produce different outputs even when the underlying model remains unchanged. That pattern is consistent with product drift, although it does not prove that the product context is the sole cause.
Hold the model configuration constant
Run representative cases from both product contexts with the same model version, instructions, configuration, and supporting system conditions. A material difference in this comparison points toward product drift.
If the team cannot establish that the model and its configuration were unchanged, the attribution remains uncertain.
Hold the product context constant
Next, compare the previous and current model configurations using the same mapped context and comparable cases. Changes in output content, structure, error patterns, consistency, or other relevant behavior under those controlled conditions point toward model behavior.
This step still requires a valid context map. A changed result should not automatically be assigned to the model merely because its version changed.
Use a controlled comparison
A change matrix makes the evidence easier to review:
| Product context | Model configuration | What the comparison indicates |
|---|---|---|
| Changed | Held constant | Evidence consistent with product drift |
| Held constant | Changed | Evidence consistent with model behavior |
| Changed | Changed | Confounded comparison; separate the changes before attributing the result |
| Apparently unchanged | Apparently unchanged | Unexplained variation that requires further investigation |
“Points toward” and “consistent with” matter here. A controlled comparison narrows the explanation; it does not automatically establish a single cause.
Track risks as the review develops
NIST guidance says AI output should be interpreted within its context. Its risk-tracking guidance also calls for regular identification and tracking of existing, unanticipated, and emergent AI risks.
Applied to this review, that means the evidence log should not stop at whether the product or model changed. It should also record unexpected output patterns, newly observed failure modes, changes in risk, and limitations that were not apparent during the initial assessment.
What the review team must still confirm
Before assigning a cause, the team must confirm several things independently:
- The product-context map reflects the conditions in which each result was produced.
- The compared cases are sufficiently comparable for the intended assessment.
- Model versions, instructions, configuration, and connected system dependencies have been recorded.
- Other changes in the operating environment have not been mistaken for model behavior.
- An observed difference matters to the product’s intended use and risk profile.
- Mixed or inconclusive evidence remains labeled as unresolved rather than presented as a definitive cause.
The most defensible conclusion is therefore conditional: identify which factor changed, document the controlled comparison, and state what evidence remains missing.