A model is not done when AUC goes up. It is done when an experiment proves it helped users. This bank covers how you run that proof, and what to do when you cannot run it.
Interviewers use these questions to sort scientists who ship from scientists who only train. A senior answer names the unit, the metric, the power, and the threats to validity before anyone asks. It also knows the fix for each threat. Every question below has worked numbers. Practice saying them out loud.
Randomization is the whole trick. A coin flip makes the two groups the same in every way, seen and unseen. So any gap in the outcome must come from the treatment. Every topic on this page either protects that coin flip or tries to fake it when you cannot flip.
The questions
1. Design an A/B test end to end Medium
What they are testing
Can you run a full test without being led? They want a fixed order of steps. They want you to choose the unit and metrics on purpose. They want a decision rule written down before the data comes in.
Strong answer
Use one example the whole way through. Say we have a new ranking model for a shopping feed. Walk seven steps.
Hypothesis. State it as a change, a metric and a direction. “The new ranker raises purchases per user by at least 1% with no drop in revenue per user.” Write the null too. H0: the true effect is zero.
Unit of randomization. Pick the unit that matches the experience. For a feed, randomize by user, not by page view. Page view splits would show one person both rankers. That breaks the user experience and makes views within a user correlated. Use a stable hash: bucket = hash(user_id + salt) mod 1000. A new salt per test stops the same users from always landing in treatment.
Metrics. Choose one primary metric (purchases per user). Add guardrails (revenue, latency, crash rate, unsubscribes). Add debug metrics (click rate, items seen) to explain the result later. Only the primary metric decides ship or no ship.
Power and duration. Pick a minimum detectable effect (MDE). This is the smallest effect worth shipping. Compute the sample size per arm (see Q2). Divide by daily eligible users to get days. Round up to whole weeks. Weekday and weekend users behave differently, so a test of 9 days over-weights one weekday.
Launch and checks. Start with a small ramp, like 1% then 5% then 50%. The ramp catches crashes and bugs early. Run an A/A check or look at pre-period metrics. Check sample ratio mismatch every day (see Q8).
Analysis. Compare the means with a two-sample t-test or a z-test on proportions. Report the effect, a 95% confidence interval, and the p-value. Use CUPED to shrink variance (see Q7). If the metric is a ratio like clicks per view, use the delta method (see Q3). Slice by platform and new vs old users, but treat slices as hints, not proof.
Decision. Apply the rule you wrote before the test. Ship if the primary metric is positive and significant and no guardrail fails. Do not ship if a guardrail drops past its limit. If the result is flat, the answer is usually “no ship” unless the change cuts cost or tech debt.
The analysis step in numbers. Control has 50,000 users with mean 0.200 purchases and standard deviation 0.80. Treatment has 50,000 users with mean 0.212.
import math
se = math.sqrt(0.80**2 / 50_000 + 0.80**2 / 50_000) # standard error of the difference
diff = 0.212 - 0.200 # 0.012, a 6% lift
z = diff / se # 0.012 / 0.00506 = 2.37
ci = (diff - 1.96 * se, diff + 1.96 * se) # (0.0021, 0.0219)
# p is about 0.018. The interval excludes zero, so we reject H0.
Pre-registration. Write the hypothesis, primary metric, MDE, duration and decision rule in a doc before launch. This stops you from picking the best-looking metric after the fact.
Follow-ups they will ask
Why not randomize by session? One user would see both versions. Learning carries over between sessions. Sessions from one user are also correlated, so the naive standard error is too small.
What if the result is significant but tiny? Compare it to the MDE and the cost. Statistical significance is not business significance. A 0.1% lift may not pay for a model that doubles serving cost.
What if the primary metric wins and a guardrail loses? Do not average them. Escalate the tradeoff with numbers. “We gain 1.2% purchases and lose 0.4% revenue. Here is the dollar value of each.”
How long is too long? Past four to six weeks, cookie churn and outside changes add noise. If you need longer, raise the MDE or cut variance.
Who is eligible? Only users who could see the change. Count users at the point of exposure, not all users. Diluting with unexposed users shrinks the measured effect.
Common traps
Picking the metric after looking at results. That is p-hacking with extra steps.
Running for a fixed number of days that is not a whole number of weeks.
Counting all users instead of triggered users. The effect gets diluted and power drops.
Say this out loud: “I will walk seven steps. Hypothesis, unit, metrics, power, launch checks, analysis, and decision rule. I write the decision rule before launch so the data cannot talk me into it.”
2. Sample size and power Medium
What they are testing
Do you know where the sample size formula comes from? Can you compute it in your head for a conversion rate? Do you know which knobs change it, and by how much?
Strong answer
Four things set the sample size. They are the false positive rate α, the power 1−β, the metric variance σ², and the effect size δ. For a two-sided test with equal arms, the size per arm is:
n per arm = 2 (z_{1-α/2} + z_{1-β})² σ² / δ²
With α = 0.05 and 80% power, z0.975 = 1.96 and z0.80 = 0.84. So (1.96 + 0.84)² = 7.84. Times 2 gives 15.7, which rounds to 16. That gives the rule of thumb:
n per arm ≈ 16 σ² / δ²
Where it comes from. The difference in means has standard error √(2σ²/n). We reject when the z-score passes 1.96. We want that to happen 80% of the time when the true effect is δ. So δ must sit 1.96 + 0.84 standard errors above zero. Solve δ = 2.8 · √(2σ²/n) for n.
Worked example: conversion rate
Baseline conversion is p = 10%. For a 0/1 metric, σ² = p(1−p) = 0.09.
We want to detect a 5% relative lift. That is 10.0% to 10.5%, so δ = 0.005.
n = 16 × 0.09 / 0.005² = 1.44 / 0.000025 = 57,600 users per arm.
With 20,000 eligible users a day split 50/50, each arm gets 10,000 a day. That is 5.8 days, so run 7 days.
Feel the knobs. Sample size scales with 1/δ². Halve the MDE to 2.5% relative and you need four times the users: 230,400 per arm. Go to 90% power and the constant rises from 16 to about 21. A lower baseline rate raises the relative variance. At p = 1%, the same 5% relative lift needs about 634,000 per arm.
from scipy.stats import norm
def n_per_arm(p_base: float, rel_mde: float,
alpha: float = 0.05, power: float = 0.80) -> int:
"""Users per arm for a two-sided test on a conversion rate."""
p_treat = p_base * (1 + rel_mde)
delta = p_treat - p_base
# Use the pooled variance under H1 (average of the two arms).
var = (p_base * (1 - p_base) + p_treat * (1 - p_treat)) / 2
z_a = norm.ppf(1 - alpha / 2) # 1.96
z_b = norm.ppf(power) # 0.84
return int(2 * (z_a + z_b) ** 2 * var / delta ** 2) + 1
print(n_per_arm(0.10, 0.05)) # about 57,800, close to the 57,600 rule of thumb
Unequal splits cost power. A 90/10 split needs more total users than 50/50. The variance of the difference is σ²(1/nA + 1/nB). That sum is smallest when the arms are equal. A 90/10 split needs about 2.8 times the total traffic of 50/50.
Follow-ups they will ask
What is power, in words? The chance the test detects an effect of size δ when it is really there. At 80%, one in five real wins look flat.
How do you pick the MDE? From the business, not from the data. Ask what lift pays for the build and the serving cost. Then check if the traffic can detect it in a few weeks.
What if you cannot get enough users? Cut variance (CUPED, trimming outliers). Use a more sensitive metric. Use interleaving for ranking. Or accept a larger MDE and say so.
What is the variance for a revenue metric? Estimate it from historical data. Revenue has a heavy tail, so cap it at the 99th or 99.9th percentile first. Capping can cut variance by half or more.
Why is an underpowered win a problem? When power is low, the wins that do pass are inflated. This is the “winner’s curse”. Expect the effect to shrink after launch.
Common traps
Mixing relative and absolute MDE. A 5% lift on 10% is 0.5 points, not 5 points.
Forgetting that n is per arm, not total.
Using the user count when the unit of analysis is page views. The variance must match the unit.
Say this out loud: “For 80% power at alpha 0.05, n per arm is 16 sigma squared over delta squared. At a 10% base rate and a 5% relative lift, that is 57,600 per arm. Halving the MDE quadruples it.”
3. North star, guardrail and proxy metrics Medium
What they are testing
Can you build a metric system, not just pick one number? Do you know the tension between a metric that moves fast and a metric that matters? Can you get the variance right for a ratio metric?
Strong answer
Metrics come in layers. Each layer has a job.
North star. The long-term goal of the product. Think 28-day retention, or sessions per user per week. It captures real value but moves slowly and is noisy.
Primary or proxy metric. A short-term metric that predicts the north star. Think clicks with a dwell of 30 seconds or more. It is what the test decides on.
Guardrails. Things that must not get worse. Some protect the business (revenue, ad load). Some protect trust (latency, crashes, reports, unsubscribes). Each one has a limit, like “no drop beyond 0.5%”.
Debug metrics. Help you explain the result. Not used to decide.
Sensitivity vs directionality
A good proxy has two traits. These traits often fight.
Sensitivity. It moves when the product changes. Clicks are very sensitive. Retention is not.
Directionality. When it moves up, the north star also moves up. Clickbait raises clicks and hurts retention. So raw clicks are sensitive but not directional.
How to check a proxy. Collect past experiments, ideally 50 or more. For each one, record the proxy effect and the long-term north star effect. Look at how often the signs agree. Also look at the correlation of effects across tests. Do not use the correlation across users. A metric can correlate with retention across users and still point the wrong way in a test.
Ratio metrics and the delta method
Click-through rate is clicks over views. We randomize by user, but the ratio is over views. Views from one user are not independent. A naive binomial standard error treats each view as a coin flip. It is too small, and false positives follow.
Treat the metric as a ratio of two user-level means. Let Yi be clicks and Xi be views for user i. Then R = Ȳ / X̄. The delta method gives a first-order Taylor expansion:
Var(R) ≈ (1 / μ_X²) · [ Var(Ȳ) − 2R · Cov(Ȳ, X̄) + R² · Var(X̄) ]
where Var(Ȳ) = s_Y² / n, Var(X̄) = s_X² / n, Cov(Ȳ, X̄) = s_XY / n
import numpy as np
def ratio_se(clicks: np.ndarray, views: np.ndarray) -> tuple[float, float]:
"""Delta-method SE for sum(clicks)/sum(views), one row per user."""
n = len(clicks)
mx, my = views.mean(), clicks.mean()
r = my / mx
vx, vy = views.var(ddof=1), clicks.var(ddof=1)
cxy = np.cov(clicks, views, ddof=1)[0, 1]
var_r = (vy - 2 * r * cxy + r * r * vx) / (n * mx * mx)
return r, float(np.sqrt(var_r))
Bootstrapping by user gives the same answer and is a fine cross-check. But the delta method is closed form and runs on billions of rows.
Follow-ups they will ask
How do you pick guardrail limits? From the cost of harm. Latency limits often come from past tests that tied latency to revenue. A 100 ms slowdown costing about 1% revenue is a common finding.
What is an overall evaluation criterion (OEC)? A single weighted score of several metrics. It forces the team to agree on tradeoffs in advance. The weights should come from long-term value models.
When would you use a mean per user vs a ratio? Use per-user means when you can. They are easier and robust. Use ratios when the business thinks in rates, like CTR. Then use the delta method.
What is a surrogate index? A model that predicts the long-term outcome from many short-term metrics. You fit it on past data, then use its prediction as the test metric.
Common traps
Using a binomial standard error on a per-view metric when the test splits by user.
Choosing a proxy because it correlates with the goal across users, not across experiments.
Having ten primary metrics. With many, one will win by chance.
Say this out loud: “I want a proxy that is both sensitive and directional. I check directionality across past experiments, not across users. For CTR, the unit is the user, so I use the delta method for the variance.”
4. Novelty and primacy effects Medium
What they are testing
Do you know that the effect in week one is not the long-run effect? Can you detect a time-varying effect and correct for it?
Strong answer
Users react to change itself, not only to the new thing.
Novelty effect. A new thing gets attention because it is new. Users click a new button out of curiosity. The lift starts high and decays.
Primacy effect. Also called change aversion. Users are used to the old way. A better design first looks worse. The effect starts low and rises as users learn.
Both make a two-week test lie about the long-run effect. Novelty gives a false win. Primacy gives a false loss.
How to detect them
Plot the daily treatment effect. Not the cumulative one. Cumulative curves smooth out the trend. A daily effect that slides from +4% to +1% is a warning.
Split by user tenure. Compare new users to existing users. New users have no old habit, so they show no novelty or primacy. If new users show +1% and old users show +4% that is fading, the true long-run effect is near +1%.
Cohort by first exposure day. Look at each cohort’s effect on its day 1, day 7, day 14. This removes the mix of old and new cohorts that confuses the daily plot.
How to correct for them
Run longer, until the daily effect is flat.
Read the decision off later days or off new users only.
Keep a long-term holdout. Hold 1% to 5% of users on the old version for months after launch. Compare them to everyone else. This gives the true long-run effect.
Fit a decay curve, like effect(t) = a + b·e−t/τ, and read off the asymptote a. Treat this as a rough guide.
Do not confuse novelty with a weekday cycle. A falling effect over five days might just be the week. Look across at least two full weeks before you call a trend.
Follow-ups they will ask
Which changes are prone to novelty? Visible UI changes. New badges, colors, layouts, notifications. Back-end ranking changes are less prone, since users cannot see them.
What about learning effects on the system side? A recommender that learns from treatment traffic gets better over time. This looks like primacy but comes from the model, not users.
How long should a holdout run? Long enough for the effect to settle, often one to three months. Weigh the cost of keeping users on a worse version.
Can you shorten the wait? Use a surrogate model that predicts long-term outcomes from early signals. Or test on new users first.
Common traps
Reading only the cumulative effect, which hides the decay.
Shipping a UI change on a strong first week without a holdout.
Killing a good redesign because week one looked negative.
Say this out loud: “I plot the daily effect, not the cumulative one. I split by new vs existing users. New users have no habit to break, so their effect is my best guess of the long run. I confirm with a holdout.”
5. Network effects and interference Hard
What they are testing
Do you know when a user-level A/B test breaks? Can you name the bias direction? Can you pick the right design for a marketplace, a social network, and a ride-share app?
Strong answer
A normal test assumes SUTVA. That is the stable unit treatment value assumption. It says one user’s outcome depends only on their own treatment. Interference breaks it. Treated users change the world for control users.
Two kinds of interference
Marketplaces share a supply. Say treatment riders get a lower price on a ride app. They book more rides. That uses up drivers, so control riders wait longer and book less. Treatment looks great and control looks bad. The measured lift is too big. At full launch, everyone competes for the same drivers, and the true gain is much smaller. The same thing happens with ad budgets, hotel rooms and delivery couriers.
Social networks share content. Say treatment users get a better share button. They share more. Their friends in control see more posts and engage more. Control rises too, so the measured lift is too small.
Know the bias direction. Competition for a shared resource inflates the lift. Spillover of a good thing to friends shrinks it. Say which one applies before you pick a fix.
Designs that fix it
Cluster randomization. Randomize groups that rarely interact with other groups. For social, use graph clusters from community detection. For marketplaces, use cities. Effects stay mostly inside the cluster. The cost is power. Units in a cluster are correlated. The design effect is 1 + (m−1)ρ, where m is cluster size and ρ is the intra-cluster correlation. With m = 100 and ρ = 0.02, the variance goes up 3 times.
Switchback tests. Flip the whole market between treatment and control over time. Use blocks like 30 minutes or 2 hours, per city. Each block is one unit. This suits ride-share and delivery, where supply resets fast. Watch for carryover. Drivers moved in one block are still out of place in the next. Drop a buffer at the start of each block, or model the carryover.
Geo experiments. Treat some regions and keep others as control. This is used for ads and TV spend. With few regions, pair similar ones, or use synthetic control (see Q9). Analyze at the region level. You may have only 20 to 200 units, so power is low.
Two-sided randomization. In a marketplace, split both buyers and sellers. Compare the cells. This lets you estimate how big the bias is.
Ego-cluster or budget-split designs. For ads, split the advertiser budget into two pools. Each pool serves only its arm. Then the arms cannot steal budget from each other.
Follow-ups they will ask
How do you know interference exists? Run the test at two levels at once. Use user-level in some clusters and cluster-level in others. If the effects differ, interference is present.
How do you analyze a cluster test? Aggregate to the cluster, or use cluster-robust standard errors. Never use the user-level t-test. It ignores the correlation and gives false wins.
Why not just make clusters huge? Fewer clusters means less power. There is a tradeoff between bias, which shrinks with size, and variance, which grows with size.
What switchback block length? Long enough for the system to settle and short enough to get many blocks. Base it on how fast supply rebalances, often 30 to 60 minutes for ride-share.
Common traps
Reporting a user-level marketplace lift as the launch impact.
Using cluster randomization but analyzing at the user level.
Ignoring carryover between switchback blocks.
Say this out loud: “This is a two-sided market, so SUTVA fails. Treatment will steal drivers from control, which inflates the lift. I would run a switchback by city and time block, and analyze at the block level.”
6. Peeking and multiple testing Hard
What they are testing
Do you know why checking results daily breaks the p-value? Do you know the fixes, and when each fits? Can you tell FWER from FDR?
Strong answer
Why peeking inflates false positives
A p-value of 0.05 means a 5% false alarm rate for one look at a fixed sample size. Under the null, the running z-score wanders like a random walk. Say you look every day and stop the first time p < 0.05. Then you get many chances to cross the line. Some walk will cross by luck.
Look 1 time: 5% false positives.
Look 5 times: about 14%.
Look 10 times: about 19%.
Look after every user, forever: it tends to 100%.
Looking is fine. Stopping early because of a look is the problem.
Fixes for peeking
Fixed horizon. Compute n up front. Look at the decision metric only at the end. Monitor guardrails daily for harm only.
Group sequential tests. Plan K looks in advance. Spend α across them with an alpha spending function. O’Brien-Fleming spending is strict early and loose late. For 5 looks, the first boundary is about z = 4.6 and the last is about z = 2.04. You can stop early for a huge effect and lose little power at the end.
Always-valid p-values. Methods like mSPRT (mixture sequential probability ratio test) give a p-value valid at any stopping time. You can look after every user. The cost is wider intervals at a fixed n. Many experiment platforms use this.
Bayesian with care. A posterior is not immune to peeking if you use a fixed threshold to stop. Calibrate the stopping rule by simulation.
Multiple testing
Test 20 metrics with no true effect at α = 0.05. The chance of at least one false win is 1 − 0.9520 = 64%. Same with 20 slices or 20 variants.
Bonferroni controls FWER. FWER is the family-wise error rate. That is the chance of even one false positive. Test each at α/m. With 20 metrics, use 0.0025. Simple and safe, but very strict. It loses a lot of power when m is large.
Benjamini-Hochberg controls FDR. FDR is the false discovery rate. That is the expected share of false wins among your wins. Sort the m p-values from small to large. Find the largest k with p(k) ≤ (k/m)·q. Reject the first k. It gives much more power when many effects are real.
Worked example. Five metrics, q = 0.05. Sorted p-values are 0.001, 0.008, 0.020, 0.041, 0.30. The BH cutoffs are 0.01, 0.02, 0.03, 0.04, 0.05. The third passes (0.020 ≤ 0.03). The fourth fails (0.041 > 0.04). So BH rejects three. Bonferroni uses 0.01 for all and rejects only two.
import numpy as np
def benjamini_hochberg(pvals: list[float], q: float = 0.05) -> np.ndarray:
"""Return a boolean mask of rejected hypotheses, in the input order."""
p = np.asarray(pvals)
m = len(p)
order = np.argsort(p)
thresh = q * np.arange(1, m + 1) / m
passed = p[order] <= thresh
reject = np.zeros(m, dtype=bool)
if passed.any():
k = np.max(np.nonzero(passed)[0]) # largest k that passes
reject[order[: k + 1]] = True # reject all up to k
return reject
print(benjamini_hochberg([0.001, 0.008, 0.020, 0.041, 0.30]))
# [ True True True False False]
Follow-ups they will ask
When use Bonferroni over BH? When one false positive is costly, like a safety guardrail. Use BH for exploring many metrics or slices.
Do guardrails need correction? Guardrails look for harm. Missing harm is the worse error. Many teams test them with less correction, or one-sided, to keep power.
Does one primary metric avoid all this? Mostly yes. That is the main reason to pick one. Everything else is labeled exploratory.
What about many variants? Compare each to control and apply a correction like Dunnett’s or Bonferroni. Or run a bandit if the goal is to pick a winner, not to measure.
Common traps
Saying “never look at the data”. You should look for harm. Just do not stop for a win without a sequential method.
Applying BH and then calling it FWER control.
Slicing by 30 segments and shipping for the one that won.
Say this out loud: “Ten unplanned looks push false positives to about 19%. If we need to stop early, I use a group sequential plan or always-valid p-values. For many metrics, I use Bonferroni for guardrails and Benjamini-Hochberg for exploration.”
7. CUPED variance reduction Hard
What they are testing
Can you derive CUPED, not just name it? Do you know why it stays unbiased? Can you say how much it helps for a given correlation?
Strong answer
CUPED stands for Controlled-experiment Using Pre-Experiment Data. Much of a user’s metric is just who they are. Heavy users buy a lot in both arms. That user-to-user spread is noise for the test. CUPED removes the part you could predict from pre-test data.
Let Y be the metric during the test. Let X be a covariate measured before the test, often the same metric over the prior 2 to 4 weeks. Define:
Y_cuped = Y − θ · (X − E[X])
Why it stays unbiased
X is measured before randomization. So the treatment cannot change it, and E[X] is the same in both arms. The term θ(X − E[X]) has mean zero in each arm. The difference in means of Ycuped equals the difference in means of Y, in expectation. Only the variance changes.
θ* is the slope of an OLS regression of Y on X. ρ is the correlation between Y and X. Variance falls by a factor of (1 − ρ²). Sample size scales with variance, so the needed users fall by the same factor.
ρ = 0.5: variance falls 25%.
ρ = 0.7: variance falls 51%. The test needs about half the users or half the time.
ρ = 0.9: variance falls 81%. Common for stable metrics like sessions per user.
import numpy as np
def cuped(y: np.ndarray, x: np.ndarray) -> np.ndarray:
"""Return CUPED-adjusted y. Fit theta on BOTH arms pooled."""
theta = np.cov(y, x, ddof=1)[0, 1] / np.var(x, ddof=1)
return y - theta * (x - x.mean())
rng = np.random.default_rng(0)
n = 20_000
x = rng.gamma(2.0, 5.0, size=2 * n) # pre-period spend
y = 0.8 * x + rng.normal(0, 6, size=2 * n) # in-test spend
arm = np.repeat([0, 1], n)
y = y + 0.3 * arm # true effect = 0.3
y_adj = cuped(y, x)
for name, v in [("raw", y), ("cuped", y_adj)]:
diff = v[arm == 1].mean() - v[arm == 0].mean()
se = np.sqrt(v[arm == 1].var() / n + v[arm == 0].var() / n)
print(f"{name:6s} diff={diff:.3f} se={se:.3f}")
# raw diff about 0.3, se about 0.082
# cuped diff about 0.3, se about 0.060 (rho about 0.69, so var about halves)
CUPED is regression adjustment. It is the same as running OLS of Y on treatment and centered X. You can add many pre-period covariates. Some platforms use an ML prediction of Y as the covariate. That method is called CUPAC. The rule is the same. The covariate must not be affected by the treatment.
Follow-ups they will ask
What about new users with no pre-period data? Fill X with a constant, often the mean, and add an indicator for “no history”. They get no variance reduction but stay unbiased.
Should θ be fit per arm? Use one pooled θ. Separate θs can add bias when the treatment changes the slope. Pooled is the standard choice.
Can X be measured during the test? Only if the treatment cannot affect it. Pre-period is the safe choice. A post-randomization covariate can absorb part of the effect and bias the answer.
Does CUPED fix bias from a bad split? It reduces chance imbalance on X. It does not fix SRM or a broken randomizer.
How does it combine with the delta method? Apply CUPED to both numerator and denominator, or linearize the ratio first. Then adjust the linearized metric.
Common traps
Using a covariate measured after the treatment starts.
Claiming CUPED changes the effect size. It only changes the variance.
Forgetting that a weak pre-period correlation gives almost no gain. At ρ = 0.2, variance falls only 4%.
Say this out loud: “CUPED subtracts theta times centered pre-period X. Theta is cov(Y, X) over var(X). X comes before randomization, so the estimate stays unbiased. Variance drops by rho squared. At rho 0.7, that halves the test length.”
8. Sample ratio mismatch Medium
What they are testing
Do you check data quality before you read results? Can you run the test by hand? Do you know the usual root causes?
Strong answer
Sample ratio mismatch (SRM) means the arms got a different share of users than planned. You set 50/50 but got 50.6/49.4. That sounds tiny. With big samples it is almost never chance. It means something non-random decided who got counted. Then the arms are no longer comparable, and the effect estimate is untrustworthy.
The chi-square test
Compare observed counts to expected counts:
χ² = Σ (observed − expected)² / expected, with (arms − 1) degrees of freedom
Worked example. Total 100,000 users, planned 50/50. Observed 50,600 in control and 49,400 in treatment. Expected is 50,000 each.
Most platforms flag SRM at p < 0.001. This test fails. Stop and debug before reading any metric.
from scipy.stats import chisquare
observed = [50_600, 49_400]
expected = [50_000, 50_000]
stat, p = chisquare(observed, f_exp=expected)
print(stat, p) # 14.4, 0.000148 -> SRM, do not trust the results
Why use a strict threshold like 0.001? You run the check on every test, every day. A loose threshold would raise many false alarms.
Common causes
Logging loss that differs by arm. Treatment is slower, so some users leave before the logging event fires. Those users vanish from treatment only.
Bots and filters. The treatment changes behavior so a bot filter removes more users from one arm.
Crashes. The new code crashes on some devices. Those users never log an exposure.
Redirects. The treatment page sits behind a redirect. Some browsers drop the cookie or the user bounces.
Trigger conditions that depend on treatment. You count a user only after an event the treatment can change. Also ramp changes mid-test, or a bad hash or salt that collides with another test.
How to debug
Slice the SRM by day, platform, browser, country and new vs old users. The cause usually lives in one slice.
Check whether the gap started on one date. That often matches a deploy or ramp change.
Compare assignment logs to exposure logs. If assignment is 50/50 but exposure is not, the loss is after assignment.
Follow-ups they will ask
Can you just reweight the arms? No. The missing users are not random. Reweighting assumes they are. Fix the cause and rerun.
What if SRM appears only in one slice? You may be able to drop that slice and analyze the rest. Only do this if you know the cause and it is unrelated to the treatment.
Does SRM apply to unequal splits? Yes. For a 90/10 test, expected counts are 90% and 10%. The test is the same.
Is a good-looking result safe if SRM is present? No. A slow treatment that drops impatient users can make the remaining users look better. That is survivor bias.
Common traps
Shrugging off 50.6/49.4 as “close enough” at a sample of 100,000.
Reading metrics first and checking SRM only when the result looks odd.
Patching SRM with weights instead of finding the cause.
Say this out loud: “Before any metric, I run a chi-square on the arm counts. 50,600 vs 49,400 gives chi-square 14.4 and p about 0.0002. That is SRM, so I find the cause before I trust anything.”
9. Causal inference without a test Hard
What they are testing
Can you estimate an effect when randomizing is not possible? Do you know the main methods and the assumption each one rests on? Can you say how you would test that assumption?
Strong answer
Sometimes you cannot run a test. The feature already launched. The change is a price or a policy. Or a test would be unethical. Every method below swaps randomization for an assumption. The senior skill is to name that assumption and stress-test it.
Difference-in-differences (DiD)
Idea. Compare the change over time in a treated group to the change in a control group. effect = (Tafter − Tbefore) − (Cafter − Cbefore).
Example. A feature launches in Canada only. Canada’s sessions go from 10.0 to 11.5. The US goes from 12.0 to 12.8. DiD = 1.5 − 0.8 = 0.7 sessions.
Key assumption. Parallel trends. Without the treatment, both groups would have moved the same way.
How to check. Plot several pre-periods. The lines should move together before launch. Run a placebo DiD on a fake launch date. It should show no effect.
Watch out. Staggered rollouts with a two-way fixed effects regression can give biased answers. Use newer estimators like Callaway-Sant’Anna.
Propensity score matching and IPW
Idea. Model the chance of treatment from covariates: e(x) = P(T=1 | X=x). Then compare treated and untreated users with similar scores. Matching pairs them up. Inverse propensity weighting (IPW) weights each user by 1/e(x) if treated and 1/(1−e(x)) if not.
Key assumptions. No unmeasured confounding. All causes of both treatment and outcome are in X. Also overlap. Every user has some chance of either arm, so 0 < e(x) < 1.
How to check. Check balance after weighting. Standardized mean differences should be under 0.1. Plot score overlap. Trim users with scores near 0 or 1, since their weights explode.
Better variant. Doubly robust estimators like AIPW combine an outcome model with IPW. They stay consistent if either model is right.
Weakness. You can never test for unmeasured confounders. Run a sensitivity analysis, like an E-value, to say how strong a hidden confounder must be to erase the effect.
Instrumental variables (IV)
Idea. Find a variable Z that pushes people into treatment but has no other path to the outcome. Use only the part of treatment that Z explains.
Example. Old A/B tests that nudged users to try a feature are a great instrument. Random assignment shifts usage, and it touches the outcome only through usage.
Key assumptions. Relevance: Z moves T. Check with a first-stage F-stat above 10. Exclusion: Z affects Y only through T. This cannot be tested, only argued. Also monotonicity: Z never pushes anyone away from treatment.
Estimate. Two-stage least squares. Wald form for a binary instrument: effect = (E[Y|Z=1] − E[Y|Z=0]) / (E[T|Z=1] − E[T|Z=0]). This is the local average treatment effect (LATE), the effect for compliers only.
Regression discontinuity (RD)
Idea. Treatment flips at a cutoff on a running variable. Users just above and just below are nearly the same. Compare them.
Example. Users with a spam score above 0.8 get a warning. Compare users at 0.79 to users at 0.81.
Key assumption. No one can precisely game the cutoff. Everything else is smooth through the cutoff.
How to check. A McCrary density test for bunching at the cutoff. Check that covariates do not jump. Try several bandwidths and see if the effect holds.
Weakness. The effect holds only near the cutoff. It may not generalize.
Synthetic control
Idea. One treated unit, like a city or country. Build a fake control as a weighted mix of untreated units. The weights make the mix match the treated unit’s pre-period path. The gap after launch is the effect.
Example. A new pricing policy in Seattle. Synthetic Seattle is 0.4 Portland + 0.35 Denver + 0.25 San Diego.
Key assumptions. The pre-period fit is good and long. The donors are not affected by the treatment. No other shock hits only the treated unit.
How to check. Placebo tests. Pretend each donor was treated and rerun. The real effect should be bigger than almost all placebo gaps.
Pick the method from the data shape. Panel data with a clean launch date points to DiD. Rich covariates point to propensity methods. A natural nudge points to IV. A threshold rule points to RD. One treated region points to synthetic control.
Follow-ups they will ask
Which assumption worries you most in practice? No unmeasured confounding. It is untestable and almost always a bit wrong. I trust designs with a source of quasi-random variation, like RD or IV, more than pure matching.
How do you validate any of these? Placebo outcomes that should not move. Placebo dates before launch. And if you can, compare to a real A/B test on a slice.
What is the difference between ATE, ATT and LATE? ATE averages over everyone. ATT averages over those who got treated. LATE averages over compliers moved by the instrument.
What is a DAG for? A causal graph tells you which variables to adjust for. Adjust for confounders. Never adjust for colliders or mediators, which can create bias.
Common traps
Naming a method without stating its assumption.
Controlling for a variable caused by the treatment, which blocks part of the effect.
Calling a weak instrument fine. A weak first stage makes IV biased toward OLS and very noisy.
Say this out loud: “Every quasi-experiment trades randomization for an assumption. For DiD, it is parallel trends, and I check pre-period plots and a placebo date. I state the assumption, test what I can, and give a sensitivity range for the rest.”
10. Offline vs online disagreement Hard
What they are testing
Have you shipped models and been burned? They want causes in a ranked list and a plan to close the gap. This question separates people who have run A/B tests from people who have read about them.
Strong answer
The setup: offline AUC rose from 0.780 to 0.795. The A/B test is flat. First rule out the boring causes. Then go through the real ones.
Rule out the boring causes
A bug. Is the new model really serving? Check that the features at serving match the features in training. Training-serving skew is the most common cause. Log the served features and score them offline.
No power. Was the test sized to detect the lift you expected? A small AUC gain may mean a 0.2% lift, below the MDE.
Dilution. Does the model touch only a few requests? Then the effect on all users is tiny.
The real causes
Proxy mismatch. AUC measures ranking of a label like click. The business metric is purchases or retention. A model can rank clicks better and pick more clickbait. AUC is also global. It counts pairs across users, but each user only sees their own top 10. A metric like NDCG@10 per user is closer to the product.
Feedback loops. The training data came from the old model’s choices. The new model is judged on items the old model chose to show. It never saw what happens with items it would show. Online, it shows new items, and the logged data cannot speak to them.
Position bias. Items at the top get clicks because they are at the top. Offline labels mix relevance with position. The model learns “was shown high” instead of “is good”. One fix is to train with position as a feature, then set it to a constant at serving. Another is to weight clicks by inverse position propensity.
Distribution shift. The offline set is last month. Users, items and seasons change. A model tuned to old data can lose its edge online.
System effects. The score feeds an auction, a blender or business rules. A better score can be cancelled by later stages. Calibration matters too. A model with better AUC but worse calibration can break an auction that uses predicted CTR times bid.
What to do
Check for skew and bugs first. Compare offline and online score distributions for the same requests.
Check calibration as well as ranking. Plot predicted vs observed rate by bucket.
Use an offline metric closer to the product. Per-user NDCG@k, or counterfactual estimates with IPS from randomized logs.
Collect a small slice of randomized exploration traffic. It gives unbiased data for offline evaluation.
Track a history of offline gain vs online gain across past launches. Fit the relation. Now you know how much offline gain you need for a real win.
Use interleaving (see Q11) as a fast, sensitive step between offline and the A/B test.
Inverse propensity scoring (IPS). Estimate a new policy’s reward from old logs. Weight each logged reward by πnew(a|x) / πold(a|x). It is unbiased if the old policy had nonzero chance on every action. Variance can be large, so people clip the weights.
Follow-ups they will ask
The test was flat. Do you trust the offline gain? No. The online test is the ground truth for the product. Offline gain is a hypothesis.
What if online is up but offline is down? Often the offline set is biased toward the old model. The new model wins on items the old one never showed.
How do you make offline eval more predictive? Use exploration data, per-user metrics, and a calibration check. Validate the offline metric against past A/B results.
Is AUC ever the right offline metric? For binary decisions with a threshold chosen later, it is reasonable. For top-k ranking, prefer per-query metrics.
Common traps
Arguing that the A/B test must be wrong because offline looked good.
Skipping the training-serving skew check.
Treating AUC gains as additive. An AUC jump of 0.015 has no fixed online meaning.
Say this out loud: “First I rule out skew, power and dilution. Then I look at proxy mismatch, feedback loops, position bias and shift. Long term, I calibrate offline gains against past A/B results so I know what offline lift predicts a win.”
11. Interleaving for ranking Hard
What they are testing
Do you know a faster way to compare rankers than an A/B test? Can you explain team-draft interleaving step by step? Do you know what it cannot tell you?
Strong answer
In an A/B test, each user sees one ranker. Users differ a lot, so the comparison is noisy. Interleaving shows each user a single list that mixes results from both rankers. The user’s clicks reveal which ranker they prefer. Each user is their own control. That removes the between-user noise.
Team-draft interleaving
Think of two captains picking teams in a schoolyard.
Each round, flip a coin to see which ranker picks first.
That ranker adds its highest-ranked result not already in the list. Tag the slot with that ranker’s team.
The other ranker does the same.
Repeat until the list is full.
Show the list. Credit each click to the team that owns the slot.
For each query, the ranker with more clicks wins. Ties count as ties.
Across all queries, test whether A wins more often than B. Use a sign test or a binomial test on wins.
import random
def team_draft(rank_a: list[str], rank_b: list[str], k: int, rng=random):
"""Merge two rankings. Return the list and the team of each slot."""
out, team, seen = [], [], set()
ia = ib = 0
while len(out) < k:
order = ["A", "B"] if rng.random() < 0.5 else ["B", "A"]
for who in order:
src, i = (rank_a, ia) if who == "A" else (rank_b, ib)
while i < len(src) and src[i] in seen: # skip items already placed
i += 1
if i < len(src) and len(out) < k:
out.append(src[i]); team.append(who); seen.add(src[i])
i += 1
if who == "A": ia = i
else: ib = i
if ia >= len(rank_a) and ib >= len(rank_b):
break
return out, team
def winner(team: list[str], clicked_slots: list[int]) -> str:
a = sum(team[s] == "A" for s in clicked_slots)
b = sum(team[s] == "B" for s in clicked_slots)
return "A" if a > b else "B" if b > a else "tie"
Why it is so sensitive
It is a within-user comparison. The big user-to-user spread cancels out.
The coin flip per round removes position bias between the two teams. On average, each team gets the top slot half the time.
Search and streaming teams report needing 10 to 100 times fewer users. The verdict matches the A/B test. Netflix reported over 100 times for some ranker comparisons.
Worked number. An A/B test needs 2 million users per arm to detect a ranker change. Interleaving with a 50x gain would need about 40,000 users total. That turns a two-week test into a day.
Limits
It measures preference, not business impact. It says which ranker gets more clicks. It does not give revenue, retention or session length. You still need an A/B test for the launch decision.
It needs two rankers over the same item pool. It does not fit page layout changes, UI changes or whole-page effects.
Set-level effects are lost. Diversity is a property of the full list. A mixed list hides how diverse each ranker would be alone.
Clicks are the signal. If clicks are a poor proxy, interleaving amplifies a poor proxy very efficiently.
Bias in edge cases. When rankers agree on most items, few slots differ, and the signal comes from a few slots. Variants like probabilistic or optimized interleaving fix some of this.
Use it as a funnel. Offline metrics screen many candidates. Interleaving picks the best two or three in days. A full A/B test confirms the winner on business metrics. Each stage is slower and more trusted than the last.
Follow-ups they will ask
Why not balanced interleaving? Balanced interleaving can favor one ranker when both rank many items the same. Team-draft credits each click to a clear owner, so it is fairer and simpler.
Does interleaving agree with A/B tests? Studies show high agreement on direction, often above 90%, for ranking changes. Check this on your own past tests first.
How do you handle a click on a shared item? In team-draft, each item has one owner, so there is no shared item. In other variants, shared items give no credit.
Can you interleave more than two rankers? Yes. Multileaving extends it to many rankers at once. It helps when tuning many candidates.
Common traps
Using an interleaving win as the launch decision.
Applying it to non-ranking changes like UI layout.
Forgetting to randomize who picks first each round. Without it, the first picker gets the top slot every time.
Say this out loud: “Team-draft interleaving merges both rankings into one list, with a coin flip per round. Clicks credit the ranker that owns the slot. It is 10 to 100 times more sensitive, but it measures preference, not revenue. So I use it to pick candidates and an A/B test to launch.”
Recap
Write the hypothesis, unit, primary metric, MDE and decision rule before launch.
n per arm ≈ 16σ²/δ². At a 10% base and 5% relative lift, that is 57,600 per arm.
Pick proxies that are sensitive and directional. Use the delta method for ratio metrics.
Plot daily effects and split by new users to catch novelty and primacy.
Interference breaks SUTVA. Use cluster, switchback or geo designs, and analyze at that level.
Unplanned peeking inflates false positives. Use sequential tests, Bonferroni or Benjamini-Hochberg.
CUPED cuts variance by ρ² and stays unbiased. Check SRM with a chi-square before reading any metric.
Every non-random method trades a coin flip for an assumption. Name it and test it.
Offline gains are hypotheses. Interleaving screens fast. The A/B test decides.