Part II Bank 6 Experiments

Experimentation and Causal Inference

A model is not done when AUC goes up. It is done when an experiment proves it helped users. This bank covers how you run that proof, and what to do when you cannot run it.

Interviewers use these questions to sort scientists who ship from scientists who only train. A senior answer names the unit, the metric, the power, and the threats to validity before anyone asks. It also knows the fix for each threat. Every question below has worked numbers. Practice saying them out loud.

Contents

  1. Design an A/B test end to end
  2. Sample size and power
  3. North star, guardrail and proxy metrics
  4. Novelty and primacy effects
  5. Network effects and interference
  6. Peeking and multiple testing
  7. CUPED variance reduction
  8. Sample ratio mismatch
  9. Causal inference without a test
  10. Offline vs online disagreement
  11. Interleaving for ranking
Randomization is the whole trick. A coin flip makes the two groups the same in every way, seen and unseen. So any gap in the outcome must come from the treatment. Every topic on this page either protects that coin flip or tries to fake it when you cannot flip.

The questions

1. Design an A/B test end to end Medium

What they are testing

Can you run a full test without being led? They want a fixed order of steps. They want you to choose the unit and metrics on purpose. They want a decision rule written down before the data comes in.

Strong answer

Use one example the whole way through. Say we have a new ranking model for a shopping feed. Walk seven steps.

  1. Hypothesis. State it as a change, a metric and a direction. “The new ranker raises purchases per user by at least 1% with no drop in revenue per user.” Write the null too. H0: the true effect is zero.
  2. Unit of randomization. Pick the unit that matches the experience. For a feed, randomize by user, not by page view. Page view splits would show one person both rankers. That breaks the user experience and makes views within a user correlated. Use a stable hash: bucket = hash(user_id + salt) mod 1000. A new salt per test stops the same users from always landing in treatment.
  3. Metrics. Choose one primary metric (purchases per user). Add guardrails (revenue, latency, crash rate, unsubscribes). Add debug metrics (click rate, items seen) to explain the result later. Only the primary metric decides ship or no ship.
  4. Power and duration. Pick a minimum detectable effect (MDE). This is the smallest effect worth shipping. Compute the sample size per arm (see Q2). Divide by daily eligible users to get days. Round up to whole weeks. Weekday and weekend users behave differently, so a test of 9 days over-weights one weekday.
  5. Launch and checks. Start with a small ramp, like 1% then 5% then 50%. The ramp catches crashes and bugs early. Run an A/A check or look at pre-period metrics. Check sample ratio mismatch every day (see Q8).
  6. Analysis. Compare the means with a two-sample t-test or a z-test on proportions. Report the effect, a 95% confidence interval, and the p-value. Use CUPED to shrink variance (see Q7). If the metric is a ratio like clicks per view, use the delta method (see Q3). Slice by platform and new vs old users, but treat slices as hints, not proof.
  7. Decision. Apply the rule you wrote before the test. Ship if the primary metric is positive and significant and no guardrail fails. Do not ship if a guardrail drops past its limit. If the result is flat, the answer is usually “no ship” unless the change cuts cost or tech debt.

The analysis step in numbers. Control has 50,000 users with mean 0.200 purchases and standard deviation 0.80. Treatment has 50,000 users with mean 0.212.

import math
se = math.sqrt(0.80**2 / 50_000 + 0.80**2 / 50_000)   # standard error of the difference
diff = 0.212 - 0.200                                   # 0.012, a 6% lift
z = diff / se                                          # 0.012 / 0.00506 = 2.37
ci = (diff - 1.96 * se, diff + 1.96 * se)              # (0.0021, 0.0219)
# p is about 0.018. The interval excludes zero, so we reject H0.
Pre-registration. Write the hypothesis, primary metric, MDE, duration and decision rule in a doc before launch. This stops you from picking the best-looking metric after the fact.

Follow-ups they will ask

Common traps

Say this out loud: “I will walk seven steps. Hypothesis, unit, metrics, power, launch checks, analysis, and decision rule. I write the decision rule before launch so the data cannot talk me into it.”

2. Sample size and power Medium

What they are testing

Do you know where the sample size formula comes from? Can you compute it in your head for a conversion rate? Do you know which knobs change it, and by how much?

Strong answer

Four things set the sample size. They are the false positive rate α, the power 1−β, the metric variance σ², and the effect size δ. For a two-sided test with equal arms, the size per arm is:

n per arm = 2 (z_{1-α/2} + z_{1-β})² σ² / δ²

With α = 0.05 and 80% power, z0.975 = 1.96 and z0.80 = 0.84. So (1.96 + 0.84)² = 7.84. Times 2 gives 15.7, which rounds to 16. That gives the rule of thumb:

n per arm ≈ 16 σ² / δ²

Where it comes from. The difference in means has standard error √(2σ²/n). We reject when the z-score passes 1.96. We want that to happen 80% of the time when the true effect is δ. So δ must sit 1.96 + 0.84 standard errors above zero. Solve δ = 2.8 · √(2σ²/n) for n.

Worked example: conversion rate

Feel the knobs. Sample size scales with 1/δ². Halve the MDE to 2.5% relative and you need four times the users: 230,400 per arm. Go to 90% power and the constant rises from 16 to about 21. A lower baseline rate raises the relative variance. At p = 1%, the same 5% relative lift needs about 634,000 per arm.

from scipy.stats import norm

def n_per_arm(p_base: float, rel_mde: float,
              alpha: float = 0.05, power: float = 0.80) -> int:
    """Users per arm for a two-sided test on a conversion rate."""
    p_treat = p_base * (1 + rel_mde)
    delta = p_treat - p_base
    # Use the pooled variance under H1 (average of the two arms).
    var = (p_base * (1 - p_base) + p_treat * (1 - p_treat)) / 2
    z_a = norm.ppf(1 - alpha / 2)     # 1.96
    z_b = norm.ppf(power)             # 0.84
    return int(2 * (z_a + z_b) ** 2 * var / delta ** 2) + 1

print(n_per_arm(0.10, 0.05))   # about 57,800, close to the 57,600 rule of thumb
Unequal splits cost power. A 90/10 split needs more total users than 50/50. The variance of the difference is σ²(1/nA + 1/nB). That sum is smallest when the arms are equal. A 90/10 split needs about 2.8 times the total traffic of 50/50.

Follow-ups they will ask

Common traps

Say this out loud: “For 80% power at alpha 0.05, n per arm is 16 sigma squared over delta squared. At a 10% base rate and a 5% relative lift, that is 57,600 per arm. Halving the MDE quadruples it.”

3. North star, guardrail and proxy metrics Medium

What they are testing

Can you build a metric system, not just pick one number? Do you know the tension between a metric that moves fast and a metric that matters? Can you get the variance right for a ratio metric?

Strong answer

Metrics come in layers. Each layer has a job.

Sensitivity vs directionality

A good proxy has two traits. These traits often fight.

How to check a proxy. Collect past experiments, ideally 50 or more. For each one, record the proxy effect and the long-term north star effect. Look at how often the signs agree. Also look at the correlation of effects across tests. Do not use the correlation across users. A metric can correlate with retention across users and still point the wrong way in a test.

Ratio metrics and the delta method

Click-through rate is clicks over views. We randomize by user, but the ratio is over views. Views from one user are not independent. A naive binomial standard error treats each view as a coin flip. It is too small, and false positives follow.

Treat the metric as a ratio of two user-level means. Let Yi be clicks and Xi be views for user i. Then R = Ȳ / X̄. The delta method gives a first-order Taylor expansion:

Var(R) ≈ (1 / μ_X²) · [ Var(Ȳ) − 2R · Cov(Ȳ, X̄) + R² · Var(X̄) ]

where Var(Ȳ) = s_Y² / n,  Var(X̄) = s_X² / n,  Cov(Ȳ, X̄) = s_XY / n
import numpy as np

def ratio_se(clicks: np.ndarray, views: np.ndarray) -> tuple[float, float]:
    """Delta-method SE for sum(clicks)/sum(views), one row per user."""
    n = len(clicks)
    mx, my = views.mean(), clicks.mean()
    r = my / mx
    vx, vy = views.var(ddof=1), clicks.var(ddof=1)
    cxy = np.cov(clicks, views, ddof=1)[0, 1]
    var_r = (vy - 2 * r * cxy + r * r * vx) / (n * mx * mx)
    return r, float(np.sqrt(var_r))

Bootstrapping by user gives the same answer and is a fine cross-check. But the delta method is closed form and runs on billions of rows.

Follow-ups they will ask

Common traps

Say this out loud: “I want a proxy that is both sensitive and directional. I check directionality across past experiments, not across users. For CTR, the unit is the user, so I use the delta method for the variance.”

4. Novelty and primacy effects Medium

What they are testing

Do you know that the effect in week one is not the long-run effect? Can you detect a time-varying effect and correct for it?

Strong answer

Users react to change itself, not only to the new thing.

Both make a two-week test lie about the long-run effect. Novelty gives a false win. Primacy gives a false loss.

How to detect them

How to correct for them

Do not confuse novelty with a weekday cycle. A falling effect over five days might just be the week. Look across at least two full weeks before you call a trend.

Follow-ups they will ask

Common traps

Say this out loud: “I plot the daily effect, not the cumulative one. I split by new vs existing users. New users have no habit to break, so their effect is my best guess of the long run. I confirm with a holdout.”

5. Network effects and interference Hard

What they are testing

Do you know when a user-level A/B test breaks? Can you name the bias direction? Can you pick the right design for a marketplace, a social network, and a ride-share app?

Strong answer

A normal test assumes SUTVA. That is the stable unit treatment value assumption. It says one user’s outcome depends only on their own treatment. Interference breaks it. Treated users change the world for control users.

Two kinds of interference

Know the bias direction. Competition for a shared resource inflates the lift. Spillover of a good thing to friends shrinks it. Say which one applies before you pick a fix.

Designs that fix it

Follow-ups they will ask

Common traps

Say this out loud: “This is a two-sided market, so SUTVA fails. Treatment will steal drivers from control, which inflates the lift. I would run a switchback by city and time block, and analyze at the block level.”

6. Peeking and multiple testing Hard

What they are testing

Do you know why checking results daily breaks the p-value? Do you know the fixes, and when each fits? Can you tell FWER from FDR?

Strong answer

Why peeking inflates false positives

A p-value of 0.05 means a 5% false alarm rate for one look at a fixed sample size. Under the null, the running z-score wanders like a random walk. Say you look every day and stop the first time p < 0.05. Then you get many chances to cross the line. Some walk will cross by luck.

Looking is fine. Stopping early because of a look is the problem.

Fixes for peeking

Multiple testing

Test 20 metrics with no true effect at α = 0.05. The chance of at least one false win is 1 − 0.9520 = 64%. Same with 20 slices or 20 variants.

Worked example. Five metrics, q = 0.05. Sorted p-values are 0.001, 0.008, 0.020, 0.041, 0.30. The BH cutoffs are 0.01, 0.02, 0.03, 0.04, 0.05. The third passes (0.020 ≤ 0.03). The fourth fails (0.041 > 0.04). So BH rejects three. Bonferroni uses 0.01 for all and rejects only two.

import numpy as np

def benjamini_hochberg(pvals: list[float], q: float = 0.05) -> np.ndarray:
    """Return a boolean mask of rejected hypotheses, in the input order."""
    p = np.asarray(pvals)
    m = len(p)
    order = np.argsort(p)
    thresh = q * np.arange(1, m + 1) / m
    passed = p[order] <= thresh
    reject = np.zeros(m, dtype=bool)
    if passed.any():
        k = np.max(np.nonzero(passed)[0])   # largest k that passes
        reject[order[: k + 1]] = True       # reject all up to k
    return reject

print(benjamini_hochberg([0.001, 0.008, 0.020, 0.041, 0.30]))
# [ True  True  True False False]

Follow-ups they will ask

Common traps

Say this out loud: “Ten unplanned looks push false positives to about 19%. If we need to stop early, I use a group sequential plan or always-valid p-values. For many metrics, I use Bonferroni for guardrails and Benjamini-Hochberg for exploration.”

7. CUPED variance reduction Hard

What they are testing

Can you derive CUPED, not just name it? Do you know why it stays unbiased? Can you say how much it helps for a given correlation?

Strong answer

CUPED stands for Controlled-experiment Using Pre-Experiment Data. Much of a user’s metric is just who they are. Heavy users buy a lot in both arms. That user-to-user spread is noise for the test. CUPED removes the part you could predict from pre-test data.

Let Y be the metric during the test. Let X be a covariate measured before the test, often the same metric over the prior 2 to 4 weeks. Define:

Y_cuped = Y − θ · (X − E[X])

Why it stays unbiased

X is measured before randomization. So the treatment cannot change it, and E[X] is the same in both arms. The term θ(X − E[X]) has mean zero in each arm. The difference in means of Ycuped equals the difference in means of Y, in expectation. Only the variance changes.

Derivation of θ

Var(Y_cuped) = Var(Y) − 2θ · Cov(Y, X) + θ² · Var(X)

d/dθ = −2 Cov(Y, X) + 2θ Var(X) = 0
θ*   = Cov(Y, X) / Var(X)

Var(Y_cuped) = Var(Y) − Cov(Y, X)² / Var(X) = Var(Y) · (1 − ρ²)

θ* is the slope of an OLS regression of Y on X. ρ is the correlation between Y and X. Variance falls by a factor of (1 − ρ²). Sample size scales with variance, so the needed users fall by the same factor.

import numpy as np

def cuped(y: np.ndarray, x: np.ndarray) -> np.ndarray:
    """Return CUPED-adjusted y. Fit theta on BOTH arms pooled."""
    theta = np.cov(y, x, ddof=1)[0, 1] / np.var(x, ddof=1)
    return y - theta * (x - x.mean())

rng = np.random.default_rng(0)
n = 20_000
x = rng.gamma(2.0, 5.0, size=2 * n)                 # pre-period spend
y = 0.8 * x + rng.normal(0, 6, size=2 * n)          # in-test spend
arm = np.repeat([0, 1], n)
y = y + 0.3 * arm                                   # true effect = 0.3

y_adj = cuped(y, x)
for name, v in [("raw", y), ("cuped", y_adj)]:
    diff = v[arm == 1].mean() - v[arm == 0].mean()
    se = np.sqrt(v[arm == 1].var() / n + v[arm == 0].var() / n)
    print(f"{name:6s} diff={diff:.3f} se={se:.3f}")
# raw    diff about 0.3, se about 0.082
# cuped  diff about 0.3, se about 0.060  (rho about 0.69, so var about halves)
CUPED is regression adjustment. It is the same as running OLS of Y on treatment and centered X. You can add many pre-period covariates. Some platforms use an ML prediction of Y as the covariate. That method is called CUPAC. The rule is the same. The covariate must not be affected by the treatment.

Follow-ups they will ask

Common traps

Say this out loud: “CUPED subtracts theta times centered pre-period X. Theta is cov(Y, X) over var(X). X comes before randomization, so the estimate stays unbiased. Variance drops by rho squared. At rho 0.7, that halves the test length.”

8. Sample ratio mismatch Medium

What they are testing

Do you check data quality before you read results? Can you run the test by hand? Do you know the usual root causes?

Strong answer

Sample ratio mismatch (SRM) means the arms got a different share of users than planned. You set 50/50 but got 50.6/49.4. That sounds tiny. With big samples it is almost never chance. It means something non-random decided who got counted. Then the arms are no longer comparable, and the effect estimate is untrustworthy.

The chi-square test

Compare observed counts to expected counts:

χ² = Σ (observed − expected)² / expected,   with (arms − 1) degrees of freedom

Worked example. Total 100,000 users, planned 50/50. Observed 50,600 in control and 49,400 in treatment. Expected is 50,000 each.

from scipy.stats import chisquare

observed = [50_600, 49_400]
expected = [50_000, 50_000]
stat, p = chisquare(observed, f_exp=expected)
print(stat, p)   # 14.4, 0.000148  -> SRM, do not trust the results

Why use a strict threshold like 0.001? You run the check on every test, every day. A loose threshold would raise many false alarms.

Common causes

How to debug

Follow-ups they will ask

Common traps

Say this out loud: “Before any metric, I run a chi-square on the arm counts. 50,600 vs 49,400 gives chi-square 14.4 and p about 0.0002. That is SRM, so I find the cause before I trust anything.”

9. Causal inference without a test Hard

What they are testing

Can you estimate an effect when randomizing is not possible? Do you know the main methods and the assumption each one rests on? Can you say how you would test that assumption?

Strong answer

Sometimes you cannot run a test. The feature already launched. The change is a price or a policy. Or a test would be unethical. Every method below swaps randomization for an assumption. The senior skill is to name that assumption and stress-test it.

Difference-in-differences (DiD)

Propensity score matching and IPW

Instrumental variables (IV)

Regression discontinuity (RD)

Synthetic control

Pick the method from the data shape. Panel data with a clean launch date points to DiD. Rich covariates point to propensity methods. A natural nudge points to IV. A threshold rule points to RD. One treated region points to synthetic control.

Follow-ups they will ask

Common traps

Say this out loud: “Every quasi-experiment trades randomization for an assumption. For DiD, it is parallel trends, and I check pre-period plots and a placebo date. I state the assumption, test what I can, and give a sensitivity range for the rest.”

10. Offline vs online disagreement Hard

What they are testing

Have you shipped models and been burned? They want causes in a ranked list and a plan to close the gap. This question separates people who have run A/B tests from people who have read about them.

Strong answer

The setup: offline AUC rose from 0.780 to 0.795. The A/B test is flat. First rule out the boring causes. Then go through the real ones.

Rule out the boring causes

The real causes

What to do

  1. Check for skew and bugs first. Compare offline and online score distributions for the same requests.
  2. Check calibration as well as ranking. Plot predicted vs observed rate by bucket.
  3. Use an offline metric closer to the product. Per-user NDCG@k, or counterfactual estimates with IPS from randomized logs.
  4. Collect a small slice of randomized exploration traffic. It gives unbiased data for offline evaluation.
  5. Track a history of offline gain vs online gain across past launches. Fit the relation. Now you know how much offline gain you need for a real win.
  6. Use interleaving (see Q11) as a fast, sensitive step between offline and the A/B test.
Inverse propensity scoring (IPS). Estimate a new policy’s reward from old logs. Weight each logged reward by πnew(a|x) / πold(a|x). It is unbiased if the old policy had nonzero chance on every action. Variance can be large, so people clip the weights.

Follow-ups they will ask

Common traps

Say this out loud: “First I rule out skew, power and dilution. Then I look at proxy mismatch, feedback loops, position bias and shift. Long term, I calibrate offline gains against past A/B results so I know what offline lift predicts a win.”

11. Interleaving for ranking Hard

What they are testing

Do you know a faster way to compare rankers than an A/B test? Can you explain team-draft interleaving step by step? Do you know what it cannot tell you?

Strong answer

In an A/B test, each user sees one ranker. Users differ a lot, so the comparison is noisy. Interleaving shows each user a single list that mixes results from both rankers. The user’s clicks reveal which ranker they prefer. Each user is their own control. That removes the between-user noise.

Team-draft interleaving

Think of two captains picking teams in a schoolyard.

  1. Each round, flip a coin to see which ranker picks first.
  2. That ranker adds its highest-ranked result not already in the list. Tag the slot with that ranker’s team.
  3. The other ranker does the same.
  4. Repeat until the list is full.
  5. Show the list. Credit each click to the team that owns the slot.
  6. For each query, the ranker with more clicks wins. Ties count as ties.
  7. Across all queries, test whether A wins more often than B. Use a sign test or a binomial test on wins.
import random

def team_draft(rank_a: list[str], rank_b: list[str], k: int, rng=random):
    """Merge two rankings. Return the list and the team of each slot."""
    out, team, seen = [], [], set()
    ia = ib = 0
    while len(out) < k:
        order = ["A", "B"] if rng.random() < 0.5 else ["B", "A"]
        for who in order:
            src, i = (rank_a, ia) if who == "A" else (rank_b, ib)
            while i < len(src) and src[i] in seen:   # skip items already placed
                i += 1
            if i < len(src) and len(out) < k:
                out.append(src[i]); team.append(who); seen.add(src[i])
                i += 1
            if who == "A": ia = i
            else: ib = i
        if ia >= len(rank_a) and ib >= len(rank_b):
            break
    return out, team

def winner(team: list[str], clicked_slots: list[int]) -> str:
    a = sum(team[s] == "A" for s in clicked_slots)
    b = sum(team[s] == "B" for s in clicked_slots)
    return "A" if a > b else "B" if b > a else "tie"

Why it is so sensitive

Worked number. An A/B test needs 2 million users per arm to detect a ranker change. Interleaving with a 50x gain would need about 40,000 users total. That turns a two-week test into a day.

Limits

Use it as a funnel. Offline metrics screen many candidates. Interleaving picks the best two or three in days. A full A/B test confirms the winner on business metrics. Each stage is slower and more trusted than the last.

Follow-ups they will ask

Common traps

Say this out loud: “Team-draft interleaving merges both rankings into one list, with a coin flip per round. Clicks credit the ranker that owns the slot. It is 10 to 100 times more sensitive, but it measures preference, not revenue. So I use it to pick candidates and an A/B test to launch.”

Recap

← 5 — ML System Design 7 — Research Depth and Behavioral →