Every model is a guess about a distribution. Every experiment is a bet on the noise. This bank covers the math under both.
The breadth round almost always has a stats block. The questions look basic. The grade is not whether you know Bayes. It is whether you can work the numbers out loud, state the conditions, and say where the method breaks.
This page has ten questions. Each one has what they test, a strong answer, the follow-ups, the traps, and one line to say out loud. Work each one with a pen first. Then read the answer.
How to answer a stats question. Define the thing in one line. Write the formula. Plug in numbers and get a result. State the assumptions. Then name one case where it breaks and what you would do instead. Five steps, about two minutes. Numbers beat words every time.
Part A — Probability
1. A disease affects 1 in 1,000 people. A test catches 99% of cases and has a 5% false positive rate. You test positive. What is the chance you have it? Easy
What they are testing
Can you apply Bayes' rule without freezing?
Do you know that a low base rate swamps a good test?
Can you link it to precision under class imbalance in ML?
Strong answer
Start with the terms. Let D mean "has the disease" and + mean "tests positive".
So a positive test means about a 2% chance of disease. Most people guess 95% or 99%. That gap is the whole point of the question.
Say it with counts. Counts are easier to follow than fractions. Picture 100,000 people.
100 are sick. The test flags 99 of them.
99,900 are healthy. The test flags 5% of them, so 4,995.
That is 5,094 positives in total. Only 99 are real.
99 / 5,094 = 1.94%.
The yellow boxes are all the positives. The healthy group is so large that its 5% error wins.
The odds form is faster. Posterior odds equal prior odds times the likelihood ratio.
LR+ = P(+ | D) / P(+ | not D) = 0.99 / 0.05 = 19.8
prior odds = 0.001 / 0.999 = 0.001
after 1 test + = 0.001 * 19.8 = 0.0198 -> p = 1.9%
after 2 tests + = 0.0198 * 19.8 = 0.392 -> p = 28.2%
after 3 tests + = 0.392 * 19.8 = 7.77 -> p = 88.6%
This form makes repeat tests easy. Each positive multiplies the odds by 19.8. But that only holds if the test errors are independent given the true state.
The ML link. This is precision under class imbalance. Swap "sick" for "fraud" and "test" for "classifier".
Recall is P(+ | D). Precision is P(D | +). They differ by the base rate.
A fraud model with 99% recall and 5% FPR at 0.1% fraud has 2% precision. Reviewers would see 49 false alarms per real case.
ROC curves hide this. They do not depend on the base rate. Use precision-recall curves when positives are rare.
Follow-ups they will ask
How do you raise precision? Lower the FPR or raise the prior. Test only a high-risk group, so the prior is higher. Or add a second, independent test.
What FPR gives 50% precision? Set 0.99 * 0.001 = FPR * 0.999. So FPR is about 0.1%. The FPR must be near the base rate.
Why can the second test not be the same test run twice? Its errors are correlated with the first. A person who fooled it once will often fool it again. The 19.8 factor would be too big.
Train on balanced data, deploy on 0.1% positives. What happens to scores? The model's probabilities are too high. Correct the odds by the ratio of deploy prior odds to train prior odds. Or recalibrate on real traffic.
Common traps
Answering 99% or 95%. That confuses P(+ | D) with P(D | +). This is the base rate fallacy.
Forgetting the healthy positives in the denominator.
Multiplying by the likelihood ratio for repeat tests without checking independence.
Say this out loud: “Out of 100,000 people, 99 sick people test positive and about 5,000 healthy people do. So a positive means about 2%. The base rate dominates. In ML terms, high recall does not give high precision when positives are rare.”
2. State the central limit theorem. When does it hold, and why do ML and A/B tests rely on it? Medium
What they are testing
Can you state it precisely, with the scaling and the conditions?
Do you know it is about the sample mean, not the data?
Do you know when it fails in real product data?
Strong answer
Central limit theorem. Take X1, ..., Xn independent and identically distributed, with mean μ and finite variance σ2. Then the standardized sample mean tends to a standard normal as n grows:
√n (X̄ − μ) / σ → N(0, 1) in distribution.
Read it as: the sample mean is about N(μ, σ2/n) for large n. The data itself can be any shape. Clicks are 0 or 1. Revenue is skewed. The mean of many of them is still close to normal.
Three separate facts often get mixed up.
Law of large numbers.X̄ → μ. The mean lands on the truth.
Variance of the mean.Var(X̄) = σ2/n. This is exact for any iid data. No CLT needed.
CLT. The shape of the error around μ is normal. This is the part that lets us use z-scores.
The conditions.
Independence. Strict iid is not needed. Weak dependence is fine (mixing sequences). But strong correlation breaks the σ2/n rate.
Finite variance. This one is required. The mean of Cauchy draws is still Cauchy, for any n. It never gets tighter.
Not identical is fine. The Lindeberg condition covers that case. It says no single term may dominate the sum.
How fast? The Berry-Esseen bound gives the rate. The max gap between the true CDF and the normal CDF is at most C ρ / (σ3 √n). Here ρ = E|X − μ|3 and C < 0.48. So skewed, heavy-tailed data converges slowly. The "n > 30" rule is a myth for skewed data. Revenue per user may need tens of thousands of users before the mean looks normal.
Why A/B tests rely on it.
The z-test and the t-test assume the difference in means is normal. The CLT makes that true for large samples, even for 0/1 or skewed metrics.
Confidence intervals of the form mean ± 1.96 SE come straight from it.
Ratio metrics like CTR (clicks over views) use the delta method. It is a first-order Taylor expansion plus the CLT. It gives the variance of a ratio of two means.
Why ML relies on it.
Error bars on a test-set metric. Accuracy is a mean of 0/1 values. Its SE is √(p(1−p)/n).
Minibatch gradients. A batch gradient is a mean over examples. So its noise is roughly Gaussian with variance shrinking like 1/B. This is why the learning rate and batch size trade off.
Ensembles and averaging. Averaging many weakly correlated predictors shrinks variance.
Follow-ups they will ask
Where does it fail in product data? Heavy-tailed revenue with a few whales. Clustered data, where you randomize by user but analyze by session. Tiny samples of rare events.
Randomize by user, analyze CTR per page view. What breaks? Views from the same user are correlated. The naive SE is too small. Use the delta method at the user level or a cluster bootstrap.
What do you do with heavy tails? Cap (winsorize) the metric at the 99.9th percentile. Or log-transform. Or use a rank test. Or use CUPED to cut variance.
Is the t-test still valid if data is not normal? Yes, for large n, because of the CLT. For small n with skew it is not, so use a bootstrap or a permutation test.
Common traps
Saying "the data becomes normal". Only the sampling distribution of the mean does.
Quoting "n > 30" as a universal rule.
Forgetting the finite variance condition. Power-law data can break it.
Say this out loud: “For iid data with finite variance, the sample mean is about normal with variance sigma squared over n. That is why a z-test works on clicks or revenue. But heavy tails slow it down, and correlated units shrink the real sample size.”
3. What is a p-value, and what is it not? How do you read a confidence interval? Medium
What they are testing
Can you give the exact definition, with the conditioning in the right place?
Can you list the common wrong readings?
Do you know what a 95% confidence interval promises?
Strong answer
p-value. The probability of seeing a test statistic at least as extreme as the one you saw, if the null hypothesis were true. In symbols: p = P(T ≥ tobs | H0).
It measures how surprising the data is under the null. It conditions on H0. It says nothing direct about how likely H0 is.
What a p-value is not.
Not P(H0 | data). That needs a prior. This is the same flip as the disease test in question 1.
Not the chance the result is "due to luck". That is the same wrong reading in other words.
Not the effect size. A tiny, useless lift has a tiny p-value with enough users.
Not 1 − P(replication). A result at p = 0.05 replicates at p < 0.05 only about half the time, if the true effect equals the estimate.
p > 0.05 does not mean "no effect". It means the data cannot rule out zero. The test may simply lack power.
Why the flip matters. Work the numbers. Say 10% of the ideas you test truly work. Tests use α = 0.05 and 80% power. Run 1,000 tests.
100 real effects. 80 come out significant.
900 null effects. 45 come out significant by chance.
So 45 of the 125 wins are false. That is a 36% false discovery rate.
So "p < 0.05" does not mean "95% sure it works". With a low hit rate, a third of your wins can be noise.
95% confidence interval. A procedure that, over many repeated samples, gives an interval that covers the true value 95% of the time.
The 95% is a property of the method, not of one interval. Once you compute [1.2%, 3.4%], the true lift is either in it or not. In strict frequentist terms you cannot say "95% chance the truth is in here".
Correct readings.
"If we repeated this experiment many times, 95% of intervals built this way would contain the true lift."
"These are the values of the lift that a 5% test would not reject." This is the duality between tests and intervals. The interval excludes 0 exactly when p < 0.05 (two-sided).
The Bayesian credible interval does mean "95% probability the parameter is in here". But it needs a prior. With a flat prior and large n, the two often match in numbers.
Why the interval beats the p-value. It shows the size and the doubt together. "Lift is 0.1%, CI [0.05%, 0.15%]" is significant and maybe too small to matter. "Lift is 3%, CI [−1%, 7%]" is not significant but may be worth a bigger test.
Follow-ups they will ask
Two groups' CIs overlap. Is the difference not significant? Not necessarily. The SE of a difference is √(SE12 + SE22), which is less than SE1 + SE2. Intervals can overlap a bit while the difference is still significant. Build the CI of the difference.
What is p-hacking? Trying many metrics, segments or stop times and reporting the one that hit. Each extra look raises the real false positive rate.
What does peeking do? Checking daily and stopping at the first p < 0.05 can push the real Type I rate past 20%. Fix it with a fixed horizon, sequential tests, or always-valid p-values.
Under the null, what is the distribution of p? Uniform on [0, 1], for a continuous test statistic. A spike near 0 across many tests signals real effects. A spike right under 0.05 signals p-hacking.
Common traps
Saying "p = 0.03 means a 3% chance the null is true".
Saying "the CI has a 95% chance of containing the true value", then defending it as frequentist.
Treating "not significant" as "proved no effect".
Say this out loud: “A p-value is the chance of data this extreme if the null were true. It is not the chance the null is true. I would rather report the effect with its confidence interval, because that shows size and doubt together.”
4. Explain MLE vs MAP. Derive the MLE for a Bernoulli and a Gaussian. Show MAP with a Beta prior. How does MAP relate to L2 regularization? Medium
What they are testing
Can you write a likelihood, take logs, and solve for the max?
Do you see priors as regularizers?
Do you know where common losses come from?
Strong answer
MLE. Pick the parameter that makes the data most likely: θ̂MLE = argmaxθ P(D | θ). MAP. Pick the parameter that is most likely given the data: θ̂MAP = argmaxθ P(D | θ) P(θ). That is the mode of the posterior.
We always maximize the log. Logs turn products into sums and do not move the argmax.
MLE for a Bernoulli. We see n flips with k heads. The parameter is p.
L(p) = p^k (1 - p)^(n - k)
log L(p) = k log p + (n - k) log(1 - p)
d/dp = k / p - (n - k) / (1 - p) = 0
=> k (1 - p) = (n - k) p
=> p_hat = k / n
The second derivative is negative, so this is a max. The negative log-likelihood here is binary cross-entropy. So training a classifier with log loss is MLE.
MLE for a Gaussian. We see x1 ... xn from N(μ, σ2).
Two points to mention. First, the variance MLE divides by n, so it is biased low. Dividing by n − 1 fixes the bias. Second, the μ part is just least squares. So mean squared error is the Gaussian negative log-likelihood with fixed variance.
MAP with a Beta prior. Put a Beta(α, β) prior on p. Its density is proportional to pα−1 (1−p)β−1.
The Beta is conjugate to the Bernoulli (posterior stays in the same family). Updating is just adding counts.
Read α − 1 and β − 1 as fake heads and tails you saw before the data.
α = β = 2 gives (k + 1)/(n + 2). This is Laplace smoothing.
As n grows, the prior fades and MAP tends to MLE.
Product example. A new ad has 1 click in 3 views. MLE says 33% CTR, which is silly. With a prior Beta(2, 98) (about 2% CTR), the posterior mean is (1 + 2) / (3 + 100) = 2.9%. That shrinkage is how you rank new items fairly.
MAP and L2 regularization. Take linear regression with Gaussian noise: y = Xw + ε, ε ~ N(0, σ2). Put a prior w ~ N(0, τ2 I) on the weights.
That is ridge regression. A tight prior (small τ) means a big λ. More noise also means a big λ. A Laplace prior p(w) ∝ exp(−|w|/b) gives L1, the lasso. Its sharp peak at zero is why L1 sets weights to exactly zero.
Follow-ups they will ask
Is MAP Bayesian? Only partly. It uses a prior but keeps one point. Full Bayes keeps the whole posterior and averages predictions over it.
Why can MAP be odd? The mode is not invariant to reparameterization. Change p to log-odds and the MAP point moves. The MLE does not move.
When does MLE overfit? With few data points or many parameters. Perfectly separable data sends logistic regression weights to infinity. A prior or L2 stops that.
Is the MLE always unbiased? No. The Gaussian variance MLE is biased. But MLE is consistent and asymptotically efficient under regularity conditions.
What does weight decay in deep learning correspond to? A Gaussian prior on the weights, for plain SGD. With Adam, L2 in the loss and decoupled weight decay differ. That is why AdamW exists.
Common traps
Forgetting the −1 in the Beta mode, or mixing up the mode and the mean.
Saying the MLE variance divides by n − 1.
Getting the sign of λ backward. A stronger prior means more regularization.
Say this out loud: “MLE maximizes the likelihood. MAP adds a log prior. For coin flips, MLE is k over n, and a Beta prior just adds fake counts. A Gaussian prior on weights gives exactly L2, with lambda equal to sigma squared over tau squared.”
5. Expected value puzzles. How many fair coin flips until you see HH? Until HT? Then coupon collector, and one more. Hard
What they are testing
Can you set up states and solve a recurrence?
Do you use linearity of expectation instead of brute force?
Can you explain why the answer is not what intuition says?
Strong answer
Puzzle A: HH vs HT. Most people guess both are 4. They are not. HT takes 4 flips on average. HH takes 6.
Model it as a Markov chain. A state is how much of the target you have matched so far. Let Es be the expected flips left from state s.
Both chains need one H to leave start. They differ in what a miss costs.
HH. Every equation reads "spend one flip, then move".
E_start = 1 + 0.5 * E_H + 0.5 * E_start
E_H = 1 + 0.5 * 0 + 0.5 * E_H # an H keeps you in state H
From the second: E_H = 2
From the first: E_start = 2 + E_H = 4
The intuition. Both need an H first. After that, HT waits for a T, and an extra H costs nothing. HH waits for a second H. But a T throws away the H you had, and you start over. The per-position chance is the same, 1/4. But HH matches clump together (HHH holds two), so the gaps between them are longer.
Shortcut for any pattern. Sum 2j over every length j where the prefix of length j equals the suffix of length j. For HH, both j = 1 and j = 2 match, so 2 + 4 = 6. For HT, only j = 2 matches, so 4. For HHH it is 2 + 4 + 8 = 14. This comes from a martingale argument (the gambler's casino trick).
Puzzle B: coupon collector. There are n coupon types, each equally likely per draw. How many draws until you have all n?
Split the wait into stages. Once you hold i types, a new type appears with chance (n − i)/n. The wait for it is geometric, with mean n/(n − i). By linearity, add the stages:
E[T] = sum_{i=0}^{n-1} n / (n - i) = n * (1 + 1/2 + ... + 1/n) = n * H_n
~ n ln n + 0.577 n
n = 6 (die faces): 6 * 2.45 = 14.7 rolls
n = 50: 50 * 4.50 = 225 draws
Most of the time goes into the last few types. The final one alone takes n draws on average.
ML link. How many random samples until every class or every shard is seen? How many random negatives before you cover a catalog? It is about n ln n, not n.
Puzzle C: hat check.n people throw their hats in a pile. Each takes one at random. What is the expected number who get their own hat?
Let Ij = 1 if person j gets their own hat. Then P(Ij = 1) = 1/n. The total is ∑ Ij. By linearity, the mean is n · 1/n = 1. That holds for any n. The Ij are not independent, and it does not matter. Linearity never needs independence.
A bonus fact: the chance that nobody gets their own hat tends to 1/e ≈ 0.368.
Follow-ups they will ask
Expected rolls of a die until a 6? Geometric with p = 1/6, so 6. The general rule is 1/p.
Expected flips until two heads in a row with a biased coin, P(H) = p?(1 + p)/p2. For p = 1/2 that gives 6, which checks out.
Which pattern appears first in one sequence, HH or HT? Each is 50/50, since both need an H first and the next flip decides. The waiting times differ but the race is fair. For THH vs HHT, THH wins 3/4 of the time (Penney's game).
Variance of the coupon collector time? The stages are independent geometrics. The variance is below π2n2/6. So the SD is order n.
Expected number of distinct items in n draws from n? Linearity again. Each item is missed with chance (1 − 1/n)n ≈ 1/e. So about 0.632 n distinct items. This is why a bootstrap sample holds 63.2% of the data.
Common traps
Saying HH and HT both take 4 because each has chance 1/4.
Trying to enumerate sequences instead of writing state equations.
Thinking linearity of expectation needs independence.
Say this out loud: “I will define states by how much of the pattern I have matched. Then each state's expected time is one flip plus the average of where I go next. For HT a miss keeps my H, so it is 4. For HH a miss resets me, so it is 6.”
6. Sample k items uniformly from a stream of unknown length, in one pass and O(k) memory. Prove it works. Medium
What they are testing
Do you know reservoir sampling?
Can you prove uniformity by induction, not just assert it?
Can you extend it to weights and to many machines?
Strong answer
The algorithm (Algorithm R). Number the items from 1.
Put the first k items in the reservoir.
For item i > k, keep it with chance k/i.
If you keep it, replace a slot chosen uniformly at random.
At any point, the reservoir is a uniform sample of size k from the items seen so far. You never need to know the stream length.
import random
from typing import Iterable, TypeVar
T = TypeVar("T")
def reservoir_sample(stream: Iterable[T], k: int, seed: int | None = None) -> list[T]:
"""Return k items chosen uniformly at random from a stream.
Each item ends up in the result with probability k / n, where n is the
stream length. One pass, O(k) memory, O(1) work per item.
Args:
stream: Any iterable. Its length need not be known.
k: Sample size. Must be at least 1.
seed: Optional seed for reproducible tests.
Returns:
A list of min(k, n) items.
"""
if k < 1:
raise ValueError("k must be at least 1")
rng = random.Random(seed)
reservoir: list[T] = []
for i, item in enumerate(stream, start=1): # i is the 1-based position
if i <= k:
reservoir.append(item) # fill the first k slots
else:
j = rng.randint(1, i) # uniform on 1..i
if j <= k: # true with chance k / i
reservoir[j - 1] = item # j is also a uniform slot
return reservoir
A neat detail: one random draw does both jobs. j is uniform on 1..i. So j ≤ k has chance k/i. Given that, j is uniform over the k slots.
Proof by induction. Claim: after n ≥ k items, each item is in the reservoir with chance k/n.
Base. At n = k, every item is in. Chance k/k = 1.
Step. Assume it holds at n − 1. Item n enters with chance k/n. Good.
An old item is in at n − 1 with chance k/(n − 1). It gets kicked out only if item n enters and picks its slot. That has chance (k/n)(1/k) = 1/n.
So it survives with chance k/(n − 1) · (1 − 1/n) = k/(n − 1) · (n − 1)/n = k/n. Done.
You can also show it directly. Item i enters with chance k/i. It then survives each later step j with chance (j − 1)/j. The product telescopes:
(k / i) * (i / (i+1)) * ((i+1) / (i+2)) * ... * ((n-1) / n) = k / n
Per-item uniformity is not the full claim. A strong answer notes that every k-subset is equally likely, too. The same induction works on subsets.
Extensions.
Fewer random calls (Algorithm L). Draw how many items to skip, from a geometric-like law. Expected work is O(k (1 + log(n/k))) random draws instead of n. Good when the stream is huge.
Weighted sampling (A-Res). Give item i the key ui1/wi, with ui uniform on (0, 1). Keep the k largest keys in a min-heap. This samples without replacement, in proportion to weight.
Distributed. Give every item a random key. Each shard keeps its top k keys. Merge and keep the global top k. This works because "top k of random keys" is a uniform sample.
Follow-ups they will ask
What about k = 1? Replace the kept item with item i with chance 1/i. Each item ends with chance 1/n.
Why not just read the whole stream and shuffle? That needs O(n) memory and two passes. The stream may be endless, like a log feed.
How would you test it? Run it many times on a stream of 10 items with k = 3. Count how often each item appears. Each should be near 30%. Use a chi-square test on the counts.
Where is this used in ML? Keeping a uniform sample of production traffic for monitoring. Replay buffers in RL. Picking eval examples from logs. Online drift checks.
Common traps
Using chance 1/i instead of k/i when k > 1.
Off-by-one with 0-based indices. With 0-based i, the chance is k/(i + 1).
Drawing the keep decision and the slot with two separate calls, then biasing the slot choice.
Say this out loud: “Keep the first k. For item i, keep it with probability k over i and evict a random slot. Each old item survives step i with probability i minus 1 over i. The product telescopes to k over n, so every item is equally likely.”
Part B — Inference
7. Define Type I error, Type II error and power. How do you pick a sample size? Easy
What they are testing
Do you know the definitions cold?
Can you compute a sample size from the baseline and the effect?
Do you know what moves power and what quietly breaks it?
Strong answer
Type I error (false positive). You reject H0 when it is true. Its rate is α, often 0.05. You ship a change that does nothing.
Type II error (false negative). You fail to reject H0 when the effect is real. Its rate is β. You kill a good idea.
Power.1 − β. The chance you detect an effect of a given size, if it is real. Often set to 0.8.
Power is always tied to an effect size. "The test has 80% power" means nothing alone. Say "80% power to detect a 2% relative lift".
The sample size formula. For a two-sided test comparing two means, equal groups:
n per group = 2 * (z_{1 - alpha/2} + z_{1 - beta})^2 * sigma^2 / delta^2
alpha = 0.05, power = 0.8: (1.96 + 0.84)^2 = 7.84
=> n per group ~ 16 * sigma^2 / delta^2
The "16 rule" is worth memorizing. δ is the minimum detectable effect (MDE). σ2 is the metric's variance per unit.
Worked example. Baseline conversion is 10%. You want to detect a lift to 10.5%, a 5% relative lift.
sigma^2 = p (1 - p) = 0.1 * 0.9 = 0.09
delta = 0.005
n = 16 * 0.09 / 0.005^2 = 16 * 0.09 / 0.000025 = 57,600 per group
So you need about 115,000 users in total. If you get 20,000 eligible users a day, that is six days. Round up to a full week or two to cover weekday effects.
What moves power.
Effect size.n scales with 1/δ2. Halving the MDE needs four times the users.
Variance. Lower variance means fewer users. CUPED uses pre-period data as a covariate. It often cuts variance by 30% to 50%.
α. A stricter α needs more users.
Split. A 50/50 split is the most efficient. A 90/10 split needs far more total traffic.
One-sided tests gain power, but only if you truly would never act on a drop.
Move the line right and α shrinks but β grows. More data narrows both curves and shrinks both errors.
Follow-ups they will ask
You test 20 metrics. What happens? With α = 0.05 and no real effects, the chance of at least one false win is 1 − 0.9520 ≈ 64%. Pick one primary metric. Use Bonferroni (α/m) for strict control, or Benjamini-Hochberg to control the false discovery rate.
An underpowered test comes out significant. Should you trust the size? No. Only big noisy estimates clear the bar. So the reported effect is inflated. This is the winner's curse, or a Type M (magnitude) error. Low power can even give the wrong sign (Type S).
Why not run until it is significant? Peeking inflates α. Use a fixed horizon, a sequential test with alpha spending, or always-valid inference.
The randomization unit is the user, but the metric is per session. What changes? Sessions within a user are correlated. The real n is closer to the user count. Compute variance at the user level with the delta method.
Common traps
Computing power after the test with the observed effect. Post-hoc power is just a restated p-value.
Picking the MDE from what the traffic allows, not from what matters to the business.
Mixing up β and power.
Say this out loud: “Type I is shipping a dud, Type II is killing a winner. At alpha 0.05 and 80% power, n per group is about 16 sigma squared over delta squared. For a 10% baseline and a half-point lift, that is about 58,000 per group.”
8. What is Simpson's paradox? Give a numeric example and tie it to confounding. Medium
What they are testing
Can you build a real example with numbers?
Do you know when to trust the split view and when to trust the pooled view?
Can you spot it in product and A/B data?
Strong answer
Simpson's paradox. A trend that holds in every subgroup reverses, or vanishes, when you pool the groups. It happens when a third variable affects both the grouping and the outcome, and the groups have very different sizes.
The classic numbers: kidney stone treatments.
Small stones. Treatment A: 81 / 87 = 93%. Treatment B: 234 / 270 = 87%. A wins.
Large stones. Treatment A: 192 / 263 = 73%. Treatment B: 55 / 80 = 69%. A wins.
Why it flips. Doctors gave A to the hard cases. 263 of A's 350 patients had large stones. Only 80 of B's did. Large stones have lower success under any treatment. So A's pooled rate is dragged down by its hard case mix. The pooled rate is a weighted average. A and B use very different weights.
The link to confounding. Stone size is a confounder. It causes both the treatment choice and the outcome.
stone size
/ \
v v
treatment -> success
The pooled comparison mixes the treatment effect with the case mix effect. Stratifying by stone size blocks that back-door path. Then the within-group comparison estimates the causal effect. So here, A is the better treatment.
But do not always split. The data alone cannot tell you which view is right. The causal graph does.
If Z causes the treatment (a confounder), condition on it.
If the treatment causes Z (a mediator), do not. Splitting would hide part of the real effect.
If both cause Z (a collider), conditioning creates a fake link. This is selection bias.
Product examples.
Ramp-up in an A/B test. Day 1 at 10% treatment, day 2 at 50%. Weekend users convert less. If you pool across days with changing splits, the day mix confounds the result. Analyze per day, or only use the period with a fixed split.
Model launch by platform. A new ranker wins on iOS and on Android. But it was rolled out more on Android, where CTR is lower. Pooled CTR looks flat or worse.
Mix shift in a metric. Average order value falls while it rises in every country. A cheap new market grew fast. That is a mix change, not a drop.
Follow-ups they will ask
Does randomization fix it? Yes, in expectation. Random assignment makes treatment independent of every confounder. That is why we run A/B tests. But changing allocation over time can bring it back.
How do you adjust in observational data? Stratify, then reweight to a common mix. Or use regression with the confounder, propensity scores, or inverse probability weighting. All of these need you to measure the confounders.
What is the stratified estimate here? Weight each group's gap by the total group size. Small stones are 357 of 700 and large are 343. A beats B by about 6 points on small stones and 4 on large. Weighted, A wins by about 5 points.
Can it happen with continuous data? Yes. A regression slope can be positive inside each group and negative across groups. This is the same idea.
Common traps
Saying "always trust the subgroups". It depends on the causal role of the split variable.
Giving an example where the groups have equal sizes. Then the paradox cannot happen.
Forgetting that A/B ramp schedules can bring it back.
Say this out loud: “A wins in both subgroups but loses pooled, because A got most of the hard cases. Stone size is a confounder. Whether to split depends on the causal graph. Split on a confounder, never on a mediator or a collider.”
9. Derive the standard error of the mean. Then explain the bootstrap, code it, and say when it fails. Medium
What they are testing
Can you derive σ/√n from basic variance rules?
Do you know the bootstrap well enough to code it and know its limits?
Do you know the fixes for dependent data?
Strong answer
Standard error is the standard deviation of an estimator across repeated samples. It tells you how much the estimate would move if you re-ran the study.
Derivation for the mean. Take X1 ... Xn iid with variance σ2.
Var(X_bar) = Var((1/n) * sum_i X_i)
= (1/n^2) * Var(sum_i X_i) # Var(aY) = a^2 Var(Y)
= (1/n^2) * sum_i Var(X_i) # independence: covariances are 0
= (1/n^2) * n * sigma^2
= sigma^2 / n
SE(X_bar) = sigma / sqrt(n) ~ s / sqrt(n) (plug in the sample SD)
Two lessons. First, precision grows like √n. To halve the SE you need four times the data. Second, the independence step is the weak point. If the Xi share a correlation ρ inside clusters of size m, the variance grows by the design effect 1 + (m − 1)ρ.
The bootstrap. Often there is no formula. Think of the SE of a median, a ratio, a percentile, or an AUC. The bootstrap estimates it by simulation. It treats the sample as a stand-in for the population.
From your n data points, draw n with replacement.
Compute the statistic on that resample.
Repeat B times (1,000 to 10,000).
The SD of the B values estimates the SE.
For a 95% CI, take the 2.5th and 97.5th percentiles (the percentile method).
import numpy as np
from typing import Callable
def bootstrap(
data: np.ndarray,
stat: Callable[[np.ndarray], float],
n_boot: int = 5000,
alpha: float = 0.05,
seed: int = 0,
) -> tuple[float, float, tuple[float, float]]:
"""Bootstrap the standard error and a percentile CI of any statistic.
Args:
data: 1-D array of iid observations.
stat: Function that maps an array to a number, such as np.median.
n_boot: Number of resamples.
alpha: 1 - confidence level. 0.05 gives a 95% interval.
seed: Random seed.
Returns:
(point estimate, standard error, (ci_low, ci_high)).
"""
rng = np.random.default_rng(seed)
n = len(data)
# Each row is one resample: n indices drawn with replacement.
idx = rng.integers(0, n, size=(n_boot, n))
stats = np.array([stat(data[row]) for row in idx])
se = stats.std(ddof=1)
lo, hi = np.quantile(stats, [alpha / 2, 1 - alpha / 2])
return stat(data), se, (lo, hi)
# Example: SE of the median of skewed revenue data.
rng = np.random.default_rng(1)
revenue = rng.lognormal(mean=3.0, sigma=1.2, size=2000)
est, se, ci = bootstrap(revenue, np.median)
print(f"median={est:.2f} se={se:.2f} 95% CI=({ci[0]:.2f}, {ci[1]:.2f})")
# Sanity check: for the mean, the bootstrap SE should match s / sqrt(n).
_, se_mean, _ = bootstrap(revenue, np.mean)
print(se_mean, revenue.std(ddof=1) / np.sqrt(len(revenue)))
For an A/B test, resample each arm on its own, then take the difference of the statistics. For a ratio metric, resample users, not events. That is the cluster bootstrap.
When the bootstrap fails.
Extremes. The max or min of a sample. The resample max can never exceed the sample max. So the bootstrap distribution is badly wrong.
Dependent data. Plain resampling breaks time series and clustered data. Use the block bootstrap for time series. Use the cluster bootstrap for users with many events.
Infinite variance. Heavy power-law data. The bootstrap of the mean is not consistent.
Tiny samples. With n = 10, the sample is a poor stand-in for the population. Percentile intervals undercover.
Biased or non-random samples. The bootstrap measures sampling noise only. It cannot fix selection bias.
Non-smooth statistics at boundaries. A parameter on the edge of its range, or a statistic that jumps. BCa intervals help with skew and bias but not with these.
Follow-ups they will ask
What share of the data is in each resample? About 1 − 1/e ≈ 63.2% unique points. The other 36.8% are "out of bag". Random forests use them for free validation.
How do you bootstrap a billion rows? Use the Poisson bootstrap. Give each row a weight drawn from Poisson(1), once per replicate. It streams and splits across machines with no shuffle.
How do you pick B? 1,000 is fine for an SE. Percentile CIs need more, 5,000 to 10,000, because the tails are noisy.
Percentile vs BCa vs t intervals? Percentile is simple but biased for skewed statistics. BCa corrects for bias and skew. The bootstrap-t is most accurate when you can estimate the SE inside each resample.
Bootstrap vs permutation test? The bootstrap builds intervals and SEs. A permutation test shuffles labels to build the null and get a p-value.
Common traps
Resampling events when the unit of randomization is the user.
Bootstrapping a max, a min, or a tiny sample and trusting the result.
Thinking the bootstrap creates new information. It only measures the noise already there.
Say this out loud: “Variance of the mean is sigma squared over n, because the covariances vanish under independence. When there is no formula, I bootstrap: resample with replacement, recompute, take the spread. I resample at the randomization unit, and I do not trust it for extremes or tiny samples.”
10. Which distribution would you use, and when? Give real ML and product examples. Easy
What they are testing
Do you know the generating story behind each distribution?
Can you match a real metric to the right model?
Do you know the signs that a model is wrong, like overdispersion or heavy tails?
Strong answer
Pick a distribution by its story, not its shape. Ask what process made the data. Then check the fit.
Bernoulli(p). One trial, success or failure.
Mean p. Variance p(1 − p).
Examples: did the user click, convert, churn? The label in binary classification.
The loss: binary cross-entropy is its negative log-likelihood.
Binomial(n, p). Count of successes in n independent Bernoulli trials with the same p.
Mean np. Variance np(1 − p).
Examples: clicks out of 1,000 impressions. Correct answers on a 500-item eval set.
Watch for: users with different p. Then the variance is larger than np(1 − p). Use a beta-binomial.
Poisson(λ). Count of rare, independent events in a fixed window.
Mean and variance both equal λ.
Examples: requests per second to a server. Errors per hour. Purchases per user per day.
It is the limit of Binomial(n, p) with large n, small p, and np = λ.
Watch for: variance far above the mean (overdispersion). That is common in user activity, because users differ. Use a negative binomial (a Poisson with a Gamma-mixed rate).
Exponential(λ). Time between events in a Poisson process.
Mean 1/λ. Memoryless:P(T > s + t | T > s) = P(T > t).
Examples: time between arrivals in a queue. A simple baseline for time to churn in survival analysis.
Watch for: real hazards that change with time. Churn risk often falls as users settle in. Use a Weibull to allow that.
Normal(μ, σ2). Sums and averages of many small effects.
It comes from the CLT. It has the most entropy of any distribution with a fixed mean and variance.
Examples: the sampling distribution of an A/B lift. Measurement noise. Weight init in neural nets.
The loss: MSE is its negative log-likelihood.
Watch for: thin tails. It badly underrates extreme events in revenue or latency.
Beta(α, β). A distribution over a probability, on [0, 1].
Mean α / (α + β). Conjugate prior for Bernoulli and binomial.
Examples: your belief about an ad's CTR. Thompson sampling for bandits: draw from each arm's Beta, play the max. Shrinking the CTR of new items toward the global rate.
Log-normal. The log of the value is normal. It comes from many multiplicative effects.
Mean exp(μ + σ2/2), which sits above the median exp(μ).
Examples: revenue per purchaser. Session length. Latency. File sizes. Income.
Practice: model log(y). Report the median as well as the mean. Note that the mean of log y is not the log of the mean, so a lift in logs is not a lift in the mean.
Power law (Pareto, Zipf).P(X > x) ∝ x−a. A few items take most of the mass.
Examples: item popularity in recommendations. Followers per account. Word frequency (Zipf). City sizes.
Tail index matters. For the survival exponent a, the variance is infinite when a ≤ 2. The mean is infinite when a ≤ 1. Then the CLT and the bootstrap fail.
ML impact: popularity bias. A recommender trained on raw logs pushes the head and starves the tail. Fixes: sampled softmax with a log-Q correction, and down-weighting popular items as negatives.
Watch for: a straight line on a log-log plot is weak proof. Log-normals look the same over a few decades. Fit by MLE and compare models with a likelihood ratio test (the Clauset method).
How to check a fit.
Q-Q plot against the chosen law.
Compare the mean and variance. Count data with variance much larger than the mean is not Poisson.
Plot the survival function on log axes to see the tail.
Compare candidate models with log-likelihood or AIC on held-out data.
Follow-ups they will ask
How do you model "number of purchases per user" in a regression? Poisson regression with a log link. If the variance is too high, use a negative binomial. If most users have zero, use a zero-inflated model or a hurdle model.
Why is the geometric distribution related to the exponential? The geometric is the discrete version. It counts trials to the first success. Both are memoryless.
What is the sum of independent Poissons? Poisson, with the rates added. That is why you can merge traffic from many servers.
Revenue per user is mostly zero with a long tail. How do you A/B test it? Split it into conversion rate times revenue per converter. Cap outliers. Use CUPED. Check results with a bootstrap.
Common traps
Fitting a normal to skewed, positive data like latency or revenue.
Assuming Poisson without checking that the variance matches the mean.
Calling anything with a long tail a power law, based on a log-log plot.
Say this out loud: “I pick by the story. A yes or no is Bernoulli. Counts in a window are Poisson, unless the variance is too high, then negative binomial. Waits are exponential. Rates get a Beta prior. Multiplicative quantities like revenue are log-normal. Popularity is a power law, so I watch for infinite variance.”
Recap
Bayes: a low base rate beats a good test. Precision depends on the prior. Recall does not.
CLT: the mean is about normal with variance σ2/n. It needs finite variance and independent units.
A p-value is P(data | H0), not P(H0 | data). Report effects with intervals.
MLE gives the count ratio and least squares. MAP adds a prior. A Gaussian prior is L2.
HT takes 4 flips, HH takes 6. Coupon collector is n Hn. Linearity never needs independence.
Reservoir sampling keeps item i with chance k/i. The survival product telescopes to k/n.
Sample size is about 16 σ2/δ2 per group at 5% and 80% power.
Simpson's paradox comes from confounding and unequal mixes. The causal graph decides which view to trust.
SE of the mean is σ/√n. Bootstrap at the randomization unit. Never for extremes.
Pick a distribution by the process that made the data. Then check the fit.