Choose the next outreach experiment from mature cohorts

Instruction: Recommend whether to expand either message. Use downstream outcomes, acknowledge uncertainty, and outline the next test.

Context:

A fictional outreach pilot randomly assigned 100 eligible accounts to message A and 100 to message B. No account appears in both groups, each contributes at most one meeting, and outcomes matured at the same cutoff. Both groups used the same qualification gate, AE reviewers, and outreach budget. Neither had stop-contact policy violations.

Message Replies Booked Held AE accepted
A 12 8 6 4
B 18 9 6 3

Your manager says B won because it got more replies. Give your recommendation and next experiment. Do not claim statistical significance without a justified method.

Updated

Example Answer

B generated more replies, but that is not enough to make it our winner. Per 100 assigned accounts, A produced 12 replies, eight bookings, six held meetings, and four accepted; B produced 18, nine, six, and three. B's extra replies did not translate into more accepted meetings in this pilot. With only seven accepted outcomes, I would treat the choice as unresolved.

I would review reply content and AE rejection reasons to choose a hypothesis. For example, if B attracts interest in a feature we do not offer, I would test clearer wording about the actual use case. I would not assume that explanation from the counts alone.

Before the next test, I would agree on accepted meetings per assigned account as the primary outcome, then track reply quality and opt-outs as checks. We would change one message element, assign accounts exclusively, use the same follow-up window, and set the sample and stopping rule with someone qualified to assess uncertainty. I would retain all assigned accounts in the analysis.

For today's decision, I would recommend another controlled test rather than a broad rollout described as a proven improvement. If volume must continue, I would use an approved message while keeping any interim choice separate from the experiment's conclusion.

What the interviewer is assessing

Essential: Distinguishes reply rate from the 4% versus 3% observed accepted-meeting rates and avoids an unsupported winner claim.

Stronger: Connects qualitative review to one testable change and defines the outcome, assignment, maturity window, and analysis plan in advance.

Red flags: Claims significance from small counts, excludes unfavorable outcomes, or stops only when the preferred version leads.

Follow-up

Interviewer: Can you say A is one percentage point better?

Example response: I can say its observed accepted-meeting rate was one percentage point higher: 4% versus 3%. That is also about 33% higher relative to B’s observed rate, but neither description establishes a reliable future lift. I would lead with the counts and uncertainty.

Pressure test

Changed constraint: You discover B included twenty accounts already contacted by A before assignment.

Example response: I would preserve the assigned groups and flag prior exposure: this is no longer a clean comparison of first exposure to A versus B. Prior contact does not by itself prove assignment stopped being random, but it changes what we can infer about the messages. I would check contact history in both groups with the analyst, report the limitation, and fix eligibility checks before a fresh first-exposure test. I would not delete the twenty accounts after seeing results.

Related Questions