Product Data Science Case Interviews: How To Structure an Ambiguous Business Problem

By interviewDB Editorial Team

Published Updated

Quick summary

Summarize this blog with AI

A product data science case interview may begin with a question as loose as “Conversion fell. What would you investigate?” Knowing SQL and statistical definitions helps, but you still need to decide what to measure, which explanation to test first, and what action the evidence supports.

Start by making the question answerable. The method below is a practice framework, not an employer's scoring rubric. Both examples are original fictional exercises; the numbers are illustrative.

Your first five actions

  1. Name the decision: are we diagnosing a decline, evaluating a feature, or deciding what to build?
  2. Define the metric: numerator, denominator, population, time window, and unit of analysis.
  3. Validate the observation: check collection, completeness, and whether the comparison is fair.
  4. Prioritize explanations: investigate the highest-value distinction before listing every possible cause.
  5. Recommend a next step: explain what you know, what remains uncertain, and what would change your decision.

A useful opening is: “Before choosing an analysis, I want to confirm the decision and the conversion definition. Then I will check whether the decline is real, separate audience changes from behavior changes, and investigate the largest remaining contribution.” Ask questions whose answers change your approach. If details are unavailable, state an assumption and continue.

Worked case: purchases fell after a traffic campaign

Prompt: A shopping service reports that weekly conversion dropped after a campaign launched. Should the team roll back a checkout release shipped that same week?

Do not attribute the decline to either change yet. Agree on this measurement contract:

  • Population: signed-in shoppers with a qualifying storefront visit, excluding employees and known bots using the same predefined rules in both weeks.
  • Denominator: distinct eligible accounts visiting in each week. Each account contributes once per week, anchored to its first qualifying visit that week.
  • Numerator: those accounts with at least one completed paid purchase from that anchor time up to, but excluding, 24 hours later. Multiple purchases count as one converted account.
  • Windows: September 21, 2026, 00:00 to September 28, 00:00, and September 28, 00:00 to October 5, 00:00. Start boundaries are included; end boundaries are excluded. Both use America/Los_Angeles, which is PDT for these dates.
  • Completeness: wait until October 6, 00:00 PDT for the final outcome windows, plus the documented ingestion delay. Purchases may occur after the visit week ends.

This defines conversion per account, not per session or checkout attempt. An account can appear in both weeks, so observations across weeks are not necessarily independent. The calculations below describe the observed change; they are not a significance test.

Check whether the decline is measurable

Before slicing the dashboard, reconcile the visit and purchase counts with trusted event and transaction records. Check missing account IDs, duplicated events, payment-state changes, delayed purchases, timezone conversion, and joins that multiply rows. Confirm that the campaign and release did not change the definition or logging of a qualifying visit.

Ask whether revenue, distinct purchasers, and payment failures tell a compatible story. Stable payment records with falling client-side purchase events would shift attention toward instrumentation. Compare equivalent weekdays and note promotions, stock shortages, outages, and holidays. A consecutive-week comparison alone does not control for these influences.

For the implementation details, use SQL Interview Prep: Solve Unfamiliar Questions. That guide covers output grain, join multiplication, and small examples for checking a query. Test your understanding by producing one row per account per week, rejecting a purchase just outside its 24-hour window, and confirming that duplicate joined events cannot inflate the buyer count.

Separate audience mix from within-group change

Suppose the checks pass. Divide accounts into two exhaustive groups for each week: first-time visitors had no qualifying visit before that week's start; returning visitors did. Count each account in only one group per week. This classification uses prior history, not whether the account eventually purchases.

Fictional account-level conversion, with complete 24-hour outcomes
GroupPrior accountsPrior buyersPrior rateRecent accountsRecent buyersRecent rate
Returning8,0001,60020.0%4,00080020.0%
First-time2,0001005.0%6,0002404.0%
Total10,0001,70017.0%10,0001,04010.4%

The aggregate rate fell by 6.6 percentage points, or about 38.8% relative to 17%. Returning visitors still convert at 20%, but their share shrank from 80% to 40%. First-time visitors increased from 20% to 60% of traffic and their conversion fell from 5% to 4%.

Hold each group's prior conversion rate fixed and apply the recent audience shares:

(40% × 20%) + (60% × 5%) = 11.0%

Under this stated decomposition, changing the mix takes conversion from 17% to 11%: −6.0 percentage points. The remaining within-group change contributes 60% × (4% − 5%) = −0.6 percentage points. Together they explain the observed −6.6 points exactly.

This is an arithmetic comparison, not proof that the campaign caused the mix change or the release harmed first-time users. A different decomposition order can allocate interaction effects differently. State the method rather than treating the components as unique causal effects.

Choose the next investigation and give a decision

Investigate in an order that can change the decision quickly:

  1. Explain the missing returning traffic. Check acquisition sources, repeat-visit trends, and campaign targeting. More first-time traffic alone does not explain why returning accounts halved.
  2. Locate the first-time decline. Compare channel, device, geography, and landing page. Within each, examine counts and movement through product viewing, checkout, and payment. Distinguish a new low-intent audience from increased friction among comparable visitors.
  3. Check the release mechanism. Inspect errors, latency, payment failures, inventory, and changes that could affect first-time shoppers differently. Timing is a clue; a reproducible failure is stronger evidence.

Avoid searching dozens of slices until one looks dramatic. Start with plausible mechanisms, inspect sample sizes, and treat newly discovered patterns as hypotheses needing confirmation.

A defensible recommendation is: “The aggregate decline mainly reflects audience composition under this decomposition. I would not roll back solely on the headline rate. I would urgently investigate lost returning traffic and the one-point first-time decline. A confirmed checkout failure would justify mitigation; otherwise, I need more evidence before attributing the change to the release.”

Separate question: would a clearer delivery message help?

Now suppose research suggests first-time shoppers misunderstand delivery charges. A new message could help. Evaluating that proposal is a different task from explaining last week's decline.

Hypothesis: showing delivery charges earlier increases first-time shoppers' 24-hour purchase conversion without unacceptable harm to margin or reliability. Randomized controlled experiments can support causal attribution when assignment and measurement are trustworthy; see the original Practical Guide to Controlled Experiments on the Web.

  • Unit and eligibility: randomize eligible first-time signed-in accounts 50/50 at their first qualifying visit, before rendering the message. Keep assignment consistent across devices and later visits.
  • Primary outcome: the proportion of all assigned accounts purchasing within 24 hours of assignment. This is an intention-to-treat comparison: retain accounts even when rendering fails or they leave immediately.
  • Exposure: log whether the message rendered as a diagnostic. Do not redefine the primary analysis as “people who saw it”; treatment-dependent rendering or engagement can select different populations. Microsoft's post-experiment guidance explains why restricted exposure analyses need special care.
  • Guardrails: define allowable changes in contribution margin per assigned account, payment-error incidence, and page latency. Set thresholds before seeing results. If refunds matter, add a longer follow-up window and wait for it.

State the interference assumption: one account's treatment should not materially affect another's outcome. Shared inventory constraints or household behavior may challenge it; discuss a different assignment design if that matters.

Check allocation counts against the planned ratio, identity consistency, missing outcomes, and telemetry by arm. An unexplained sample ratio mismatch blocks a trustworthy conclusion; Microsoft's SRM guidance details why apparently favorable results can be misleading.

Explain uncertainty and the stopping rule

Before launch, choose the minimum worthwhile improvement, planned detectable effect, statistical error rates, sample size, and duration. Use expected traffic and outcome variance; “run for a week” is not a power calculation. Allow complete outcome windows and relevant weekday cycles.

As a separate numerical interpretation exercise, suppose 400 of 10,000 control accounts and 460 of 10,000 treatment accounts purchase. The difference is 0.6 percentage points: 4.6% minus 4.0%. With independent randomized accounts, a simple large-sample 95% interval for the difference is approximately 0.04 to 1.16 percentage points. These illustrative counts are not a sample-size recommendation.

If the business requires at least 0.5 points to justify the change, this interval still includes economically insufficient gains. A positive estimate does not establish that the worthwhile threshold was met. Explain the uncertainty alongside costs, guardrails, and the prespecified decision rule. Do not equate statistical significance with a sound launch decision.

Monitor failures while the test runs, but do not repeatedly apply an ordinary fixed-horizon significance test and stop at the first favorable result. Use the planned endpoint, or an appropriately designed sequential method. Safety stops and success declarations serve different purposes. Microsoft's during-experiment guidance discusses monitoring and accounting for repeated looks.

Practice the decision, not a memorized framework

Rehearse this case in 20 minutes: three minutes for clarification, four for data checks, six for decomposition, four for the next investigation, and three for your recommendation. Then change one assumption: anonymous visitors replace accounts, purchase completion takes seven days, or inventory is shared. Explain which definition or design must change.

Review your recording for three things: whether each question changed the analysis, whether you separated observation from causation, and whether your final answer supported a decision. A clear conditional recommendation is stronger than a list of techniques with no next step.

Public discussions and technical sources

An August 17, 2026 discussion raised difficulty preparing for contextual data science cases. An independent August 30 discussion described surprise at ambiguous questions after studying technical topics. These individual accounts motivated this guide; they do not establish how every employer interviews. The linked primary experimentation publications support the technical cautions. The practice cases, counts, and answer structure here are original.