The customer needs strict data residency and low response latency, but the preferred AI service cannot meet both. How would you choose a design?
Instruction: Use the fictional workload, evaluation and delivery constraints below. Keep customer residency fixed, state your priority assumption and recommend a feasible initial scope; do not combine independent p95 values.
All numbers below are fictional case inputs from a representative customer replay, not universal AI benchmarks. Both options are approved for the customer-defined residency requirement across prompts, retrieved content, outputs, telemetry, retention, backups and support access. That requirement stays fixed. Option B is the preferred higher-quality service; it misses the urgent latency requirement.
- During the modeled peak hour, 100 requests are urgent dispatch decisions and 300 are planning requests. Operators review recommendations; neither path may write automatically.
- For this pilot, the customer agreed at least 95 correct urgent cases out of 100 and 90 correct planning cases out of 100, with no designated critical failure in the assessed set. These are local acceptance checks, not proof of future accuracy.
- Urgent end-to-end p95 must be at most 2 seconds. Planning currently uses a synchronous interface with end-to-end p95 at most 6 seconds.
- The latency measurements are complete request timings from 1,200 replayed calls per option at the stated workload and observed burst pattern; they are not sums of component percentiles or guarantees for another traffic mix.
- One engineer is available for a five-working-day delivery window. Each four-day estimate includes integrating and accepting one option in the current synchronous workflow. Integrating both requires an estimated eight engineer-days; no extra staff is committed.
- The model-spend ceiling for the stated peak hour is USD 20. Infrastructure, operator time and future demand are outside these model-only estimates.
- Both workflows currently have a manual path. Dispatch timing and planning backlog have different costs, but the accountable customer owner has not ranked them. Continuing either manual path needs an explicit operating agreement.
| Observed input | Option A lightweight regional model | Option B larger regional service |
|---|---|---|
| Observed correct urgent cases | 96 of 100 | 98 of 100 |
| Observed correct planning cases | 84 of 100 | 95 of 100 |
| Designated critical failures observed | 0 in this case set | 0 in this case set |
| Measured urgent end-to-end p95 | 1.2 seconds | 4.6 seconds |
| Measured planning end-to-end p95 | 1.4 seconds | 4.8 seconds |
| Estimated model cost per request | USD 0.012 | USD 0.055 |
| Integration and acceptance estimate | 4 engineer-days | 4 engineer-days |
Your task: Recommend an initial deployment and state the business-priority assumption behind it. Explain the affected workflow, model-spend calculation, displaced work and evidence that would change your choice. A narrow release, planning-first release or approved later split can each be defensible; do not assume more integration capacity than the case gives you.
Updated
Official answer available
Read the opening below, then unlock the full answer and practical guidance.
For this fictional case, I'd initially prioritize urgent dispatch, assuming its timing has the larger immediate consequence. I'd confirm that with the customer owner. My first release would use A for urgent work, with planning kept in the agreed manual process...
Related Questions
-
easy
-
easy
-
easy
-
easy
-
medium
-
medium