The AI pilot passes offline tests but fails during customer shadow use. How would you find what changed?

Instruction: Use the fictional offline/live trace below. Diagnose from the actual model input, separate authorized access from source freshness, and keep the isolation test free of live side effects.

Context:

This is a fictional customer trace, not a report about a real deployment. The operations owner confirms that v3 is active. Reading archived v2 is authorized, but v2 is no longer authoritative for this decision. The trace is one example; it does not prove every live failure has the same cause.

EvidenceOffline evaluationCustomer shadow run
User requestCan order O17 dispatch tomorrow?Same request
Model / prompt versionsm2 / p8m2 / p8
Order inputSnapshot 42: stock shortageSame snapshot 42
Effective policy confirmed by ownerv3: inventory-manager approval requiredSame active policy v3
Identity and document accessNorth operations; v2 and v3 authorizedSame identity; v2 and v3 authorized
Retriever selectionPolicy v3, chunk 17Policy v3, chunk 17
Context in captured final model requestv3 clause: approval requiredv2 clause: dispatch without that approval
Context-cache traceMiss; assembled v3 bundleHit; bundle labeled v2
OutputHold dispatch; obtain approvalDispatch may proceed
External writesNoneNone; shadow output only

Your task: State a leading hypothesis, one minimal isolation test, and the result that would weaken your hypothesis. Propose the smallest justified correction and the evidence needed before changing the release recommendation. Keep identity permissions fixed and keep the test free of live writes.

Updated

Official answer available

Read the opening below, then unlock the full answer and practical guidance.

In this fictional trace, my leading hypothesis is stale context assembly: retrieval selects v3, but the actual model request contains v2...

Related Questions