The sponsor wants a production launch next week, but the pilot has no credible failure-rate baseline. What would you recommend?
Instruction: Next week is a fictional deadline. Use the customer workflow owner’s approved failure categories and current-process comparison; no universal acceptable error rate is implied.
Updated
Example Answer
I'd ask the customer's workflow owner which errors matter and how the current process handles them. We need a credible comparison, not just a demo accuracy number. I'd agree representative cases, the kinds of failure that block launch, and who can decide that the remaining uncertainty is acceptable. I'd keep missing baseline evidence explicit.
My initial recommendation would be a limited shadow or assisted trial where operators retain the established process and the new system cannot make unreviewed consequential changes. We'd compare both paths on the same permitted cases, including exceptions, and record actual correction effort. The scope must fit the customer's capacity to inspect results; a small deployment is not automatically safe.
I'd explain what that trial can establish before next week's deadline and what it cannot. If it supports an approved limited launch, I'd define the exposed workflow, stop conditions and next review. If we cannot evaluate the essential failure cases or staff the fallback, I'd recommend delaying those actions and offering a smaller demonstrable deliverable. The sponsor would receive a concrete decision and its cost, rather than an unsupported production promise.
Make it your own
Next week is a fictional deadline. Use the customer workflow owner’s approved failure categories and current-process comparison; no universal acceptable error rate is implied.
Why this works
Tests a launch recommendation grounded in customer acceptance and a credible current-workflow comparison rather than a generic canary plan.
Interviewer follow-up
What if the current manual process also has errors?
I'd include those errors in the comparison instead of assuming the old process is perfect. I'd separate their consequences and the safeguards already in place. An overall improvement may still introduce an unacceptable new failure. I'd ask the owner to approve the tradeoff using relevant cases, rather than treat a better average as automatic permission to launch.
Pressure test follow-up
The customer cannot staff shadow review next week, and the new workflow can change payment records. What would you recommend instead?
I'd withdraw the assisted-launch recommendation because its operating dependency is unavailable. I'd keep payment writes disabled and recommend an approved demonstration using permitted fixtures while the owner establishes review coverage and acceptance evidence. If a separate read-only workflow has usable support and approved access, it could progress on its own. I'd reset the delivery commitment explicitly rather than relabel an unreviewed payment workflow as a pilot.
What this tests
The candidate must revise the release scope when the customer cannot provide a safeguard that the original recommendation depended on.
Assessment criteria
These are practice criteria for this scenario, not an employer's scoring rubric.
- Strong: Defines consequential failure categories and an independent current-process comparison; chooses exposure that actual review capacity can support.
- Adequate: Keeps missing evidence visible and proposes a bounded trial or narrower deliverable.
- Weak: Uses an unsupported average score or the sponsor's deadline as permission for unreviewed consequential writes.
A tempting weak answer
"We can launch to a few users and collect the baseline afterwards."
Why it fails: Small exposure does not by itself control consequences or supply the review capacity on which a responsible trial depends.
Other defensible approaches
A staffed shadow comparison, a customer-approved assisted trial with retained fallback, or a smaller read-only deliverable can be reasonable under different capacity and consequence assumptions. Delaying the affected write scope is also defensible. Name which dependency your recommendation needs and what new evidence would let it expand.
References
Related Questions
-
easy
-
easy
-
easy
-
easy
-
medium
-
medium