Part I Bank 1 Classic ML

ML Fundamentals

The breadth round lives here. Every question looks easy. The grade comes from the second and third layer.

Interviewers open with these questions because they find gaps fast. Anyone can define precision. Few people can say when ROC-AUC misleads, or why log loss beats MSE for a classifier. This bank covers the twelve questions that come up most. Each answer goes one level deeper than a textbook.

Read each card in three passes. First, say the short answer out loud. Next, learn the strong answer with its math. Last, drill the follow-ups. The follow-ups are where the level gets set.

Contents

  1. Bias-variance trade-off
  2. Detecting and fixing overfitting
  3. L1 vs L2 regularization
  4. Precision, recall and when ROC-AUC lies
  5. Class imbalance
  6. Logistic regression loss and gradient
  7. Random forests vs gradient boosting
  8. Cross-validation and data leakage
  9. Feature engineering
  10. Calibration
  11. Generative vs discriminative
  12. Curse of dimensionality and PCA
How to answer a breadth question. Give a one-line definition first. Then give the math or the mechanism. Then give one real case where it matters. Stop there and let them pick the next layer. A 90-second answer with a clean structure beats a 5-minute tour.

1. Explain the bias-variance trade-off. Easy

What they are testing

Can you write the decomposition, not just name it? Do you know what each term means for a real model? Do you know where the classic story breaks down with deep nets?

Strong answer

Take a target y = f(x) + ε with noise of mean 0 and variance σ². We fit a model f̂ on a random training set D. At a fixed point x, the expected squared error over draws of D and noise splits into three parts.

E[(y - f̂(x))²] = (E[f̂(x)] - f(x))²  +  E[(f̂(x) - E[f̂(x)])²]  +  σ²
                  =        Bias²           +          Variance            + Irreducible noise

The trade-off comes from capacity. More capacity lowers bias but raises variance. The test error curve is U-shaped. The best model sits at the bottom of the U.

Model complexity Error Bias² Variance Total error sweet spot
Classic picture. Bias falls, variance rises, and total error has a minimum.

How to tell which one you have. Look at train and validation error side by side.

The modern twist. Very large networks can fit the training data to zero error and still generalize. Test error can rise near the point where the model just barely fits the data. Then it falls again as capacity keeps growing. This is called double descent. The decomposition still holds at every point. What changes is that variance does not grow forever with parameter count. Implicit regularization from SGD and the shape of the minimum it finds keep it in check.

Follow-ups they will ask

Common traps

Say this out loud: “Expected error splits into bias squared, variance and noise. Bias is the error of the average model. Variance is how much the model moves across training sets. I diagnose with the train and validation gap, then add capacity or add regularization based on which one dominates.”

2. How do you detect and fix overfitting? Easy

What they are testing

Do you have a real toolbox, ranked by cost? Do you check the data before you blame the model? Can you tell overfitting apart from a broken validation set?

Strong answer

Detect it. Overfitting means the model learned noise in the training set. The signal is a gap between training and held-out performance.

Fix it. Go from cheap and safe to expensive.

  1. Check the split first. Duplicates across train and test, time leakage, or user leakage can fake a gap or hide one.
  2. More data or augmentation. The most reliable fix. For images, use crops, flips and color jitter. For text, use back-translation or token dropout.
  3. Regularize the weights. Add L2 (weight decay) or L1. Tune the strength on validation.
  4. Cut capacity. Fewer layers, smaller trees, larger min_samples_leaf, fewer features.
  5. Early stopping. Stop at the best validation epoch. For gradient descent on a linear model, it acts much like L2.
  6. Dropout and noise. Dropout trains an implicit ensemble of thinned networks. Label smoothing stops the net from pushing logits to infinity.
  7. Ensembles. Bagging and averaging seeds lower variance directly.
Watch the validation set itself. If you tune hundreds of runs on one validation set, you overfit to it. Keep a final test set you touch once. Report that number.

Follow-ups they will ask

Common traps

Say this out loud: “I detect it with learning curves and the train and validation gap. First I rule out a leaky split. Then I add data, then regularize, then cut capacity, and I use early stopping throughout. I keep one clean test set for the final number.”

3. Compare L1 and L2 regularization. Medium

What they are testing

Three views of the same idea. Can you show why L1 gives sparse weights, with geometry and with the gradient? Do you know the Bayesian prior behind each one? Do you know when to use elastic net?

Strong answer

Both add a penalty on the weights to the loss.

L1 (Lasso):  J(w) = Loss(w) + λ ∑ |w_j|
L2 (Ridge):  J(w) = Loss(w) + λ ∑ w_j²

View 1: geometry. Write the penalty as a constraint. Minimize the loss with ∑|w_j| ≤ t or ∑w_j² ≤ t. The L1 ball is a diamond with corners on the axes. The L2 ball is a circle. The loss contours are ellipses that grow until they touch the ball. They usually touch the diamond at a corner. At a corner, some weights are exactly zero. A circle has no corners, so the touch point is almost never on an axis.

L1: hits a corner, w₁ = 0 L2: touches off axis, both small
Loss contours (red) grow until they touch the constraint set. The diamond’s corners sit on the axes.

View 2: the gradient. The L2 penalty has gradient 2λw. It pulls hard on big weights and gently on small ones. It shrinks weights toward zero but never quite gets there. The L1 penalty has a subgradient of λ·sign(w). It pulls with the same force no matter how small the weight is. If the loss gradient on a weight is smaller than λ, the weight goes to exactly zero and stays there. In one dimension with a squared loss, the closed forms show it.

Ridge:  w* = z / (1 + λ)                    # scale down, never zero
Lasso:  w* = sign(z) · max(|z| - λ, 0)        # soft threshold, exact zeros

Here z is the unpenalized least squares solution, with the penalty scaled to match. Solvers use coordinate descent or proximal gradient (ISTA) to apply this soft threshold.

View 3: Bayesian priors. Regularized least squares is MAP estimation.

One subtle point: the full Bayesian posterior under a Laplace prior is not sparse. Only the MAP point is. Draws from the posterior almost never land on exactly zero.

Practical differences.

Follow-ups they will ask

Common traps

Say this out loud: “L1 adds the absolute weights and L2 adds the squared weights. L1 gives exact zeros because its constraint set has corners on the axes, and its pull does not fade near zero. L1 is MAP with a Laplace prior and L2 is MAP with a Gaussian prior. With correlated features I reach for elastic net.”

4. Precision, recall, F1, and when does ROC-AUC lie? Medium

What they are testing

Do you know the formulas cold? Can you pick a metric from the business cost? Do you know why ROC-AUC looks great on a rare-event problem while the model is useless?

Strong answer

Start from the confusion matrix counts: TP, FP, TN, FN.

Precision = TP / (TP + FP)        # of the ones I flagged, how many were right
Recall    = TP / (TP + FN)        # of the real positives, how many I caught  (= TPR)
FPR       = FP / (FP + TN)        # of the real negatives, how many I flagged
F1        = 2 · P · R / (P + R)   # harmonic mean of P and R
Fβ        = (1 + β²) · P · R / (β² · P + R)   # β > 1 weights recall more

Pick by the cost of each error.

ROC-AUC. The ROC curve plots TPR against FPR as the threshold moves. AUC is the area under it. It has a clean meaning. AUC is the chance that a random positive scores higher than a random negative. It ignores the threshold and the class ratio.

When ROC-AUC lies. Its strength is also its flaw. FPR divides by the number of negatives. With heavy imbalance, negatives are huge. A large count of false positives is still a tiny FPR.

Worked example. 1,000,000 transactions, 1,000 are fraud. A model catches 800 of them (recall 0.8) at FPR 1%. That is 0.01 × 999,000 ≈ 9,990 false positives. ROC looks superb at that point. Precision is 800 / (800 + 9,990) ≈ 7.4%. Thirteen of every fourteen alerts are wrong. The review team drowns.

The PR curve plots precision against recall. Precision divides by what you flagged, not by all negatives. So it shows the false-positive cost directly. Report PR-AUC, also called average precision, for rare positive classes.

Follow-ups they will ask

Common traps

Say this out loud: “Precision is how many flags were right. Recall is how many positives I caught. ROC-AUC uses false positive rate, which divides by all negatives. Under heavy imbalance a flood of false positives still looks like a tiny rate. So for rare events I report PR-AUC with its baseline and precision at the operating recall.”

5. How do you handle class imbalance? Medium

What they are testing

Do you know that imbalance is often not the problem? Can you list fixes at the data, loss and decision levels? Do you know what each fix does to calibration?

Strong answer

First, ask if imbalance is really hurting you. A well-specified model trained with log loss handles a 1% positive rate fine. It learns the right probabilities. The real problems are three. The metric hides poor performance. The default 0.5 threshold is wrong. And there are too few positives to learn from. Fix those directly.

Level 1: the metric. Drop accuracy. Use PR-AUC, recall at a fixed precision, or expected cost. Use stratified splits so every fold has positives.

Level 2: the decision threshold. Often this is the only fix you need. Train normally, then choose the threshold on validation. With costs C_FP and C_FN and calibrated probabilities, the cost-optimal rule is simple.

Predict positive when  p(y=1|x) > C_FP / (C_FP + C_FN)

If a miss costs 9 times a false alarm, the threshold is 0.1, not 0.5.

Level 3: the loss.

FL(p_t) = -α_t (1 - p_t)^γ log(p_t)       # p_t = prob of the true class, γ ~ 2

When p_t is 0.9, the factor (1-p_t)² is 0.01. So easy examples barely count. γ = 0 gives plain weighted cross-entropy.

Level 4: the data.

The calibration cost. Resampling and class weights shift the predicted probabilities. If you undersample negatives at rate s (keep fraction s), you can correct the output.

p_true = p / (p + (1 - p) / s)

Equivalently, subtract log(1/s) from the logit. Anything downstream that needs real probabilities, like bidding or expected value, needs this fix.

import numpy as np

def focal_loss(p, y, gamma=2.0, alpha=0.25, eps=1e-7):
    """Binary focal loss. p: predicted P(y=1). y: 0/1 labels."""
    p = np.clip(p, eps, 1 - eps)
    p_t = np.where(y == 1, p, 1 - p)              # prob given to the true class
    a_t = np.where(y == 1, alpha, 1 - alpha)      # class balance weight
    return np.mean(-a_t * (1 - p_t) ** gamma * np.log(p_t))

def best_threshold(p, y, c_fp=1.0, c_fn=9.0):
    """Pick the threshold that minimises total cost on a validation set."""
    grid = np.unique(p)
    costs = [c_fp * np.sum((p >= t) & (y == 0)) + c_fn * np.sum((p < t) & (y == 1))
             for t in grid]
    return grid[int(np.argmin(costs))]

Follow-ups they will ask

Common traps

Say this out loud: “First I fix the metric and the threshold, because a log-loss model often learns fine on imbalanced data. If the model still ignores the rare class, I add class weights or focal loss, or undersample negatives. Resampling shifts the probabilities, so I correct the logit before anything uses them.”

6. Derive the logistic regression loss and gradient. Why not MSE? Medium

What they are testing

Can you derive log loss from maximum likelihood on a whiteboard? Can you get the clean gradient Xᵀ(p - y)? Can you give two separate reasons MSE is a bad fit?

Strong answer

The model. Logistic regression models the log odds as linear in x.

z = wᵀx + b
p = σ(z) = 1 / (1 + e^(-z))          # P(y = 1 | x)
log(p / (1 - p)) = z                    # log odds are linear

The loss from likelihood. Each label is a Bernoulli draw with success chance p. The chance of one label is p^y (1-p)^(1-y). Take the product over the data, then the negative log. That gives binary cross-entropy, also called log loss.

L(w) = -(1/n) ∑_i [ y_i log p_i + (1 - y_i) log(1 - p_i) ]

The gradient. One fact makes this short. The sigmoid has derivative σ'(z) = σ(z)(1 - σ(z)) = p(1-p). For one example:

∂L/∂p = -y/p + (1-y)/(1-p) = (p - y) / (p(1-p))
∂p/∂z = p(1-p)
∂L/∂z = (p - y)                    # the p(1-p) terms cancel
∂L/∂w = (p - y) x

Batch:  ∇_w L = (1/n) Xᵀ(p - y)      ∇_b L = (1/n) ∑(p_i - y_i)

The gradient is the error times the input. It has the same form as linear regression. That is no accident. Both are generalized linear models with a canonical link.

The Hessian. H = (1/n) Xᵀ S X, where S is diagonal with entries p_i(1-p_i). All entries of S are positive, so H is positive semi-definite. The loss is convex. Gradient descent finds the global minimum. Newton’s method on this loss is IRLS (iteratively reweighted least squares).

import numpy as np

def sigmoid(z):
    # stable form: avoid overflow in exp for large |z|
    return np.where(z >= 0, 1 / (1 + np.exp(-z)), np.exp(z) / (1 + np.exp(z)))

def fit_logreg(X, y, lr=0.1, epochs=1000, l2=0.0):
    """Batch gradient descent on L2-regularized log loss."""
    n, d = X.shape
    w, b = np.zeros(d), 0.0
    for _ in range(epochs):
        p = sigmoid(X @ w + b)
        err = p - y                          # dL/dz for every row
        w -= lr * (X.T @ err / n + l2 * w)   # do not regularize b
        b -= lr * err.mean()
    return w, b

def log_loss(y, z):
    """Log loss from logits, stable: log(1 + e^z) - y*z."""
    return np.mean(np.logaddexp(0, z) - y * z)

Why not MSE? Give both reasons.

  1. Non-convex. MSE on top of a sigmoid, (σ(wᵀx) - y)², is not convex in w. Gradient descent can stall in flat regions.
  2. Vanishing gradient when badly wrong. The MSE gradient with respect to z is 2(p - y) · p(1-p). Say y = 1 and the model says p = 0.001. Then p(1-p) is about 0.001, so the gradient is tiny. The model is confidently wrong and learns slowly. Log loss cancels that factor. Its gradient is p - y ≈ -1, full strength.
  3. Wrong likelihood. MSE is the MLE under Gaussian noise. Labels in {0, 1} are Bernoulli, not Gaussian. Log loss is the matching likelihood, and it is a proper scoring rule, so it rewards honest probabilities.

Follow-ups they will ask

Common traps

Say this out loud: “Labels are Bernoulli with p equal to sigmoid of w-transpose-x. The negative log likelihood is cross-entropy. The sigmoid derivative cancels, so the gradient is X-transpose times p minus y. The loss is convex. MSE would be non-convex and would give tiny gradients exactly when the model is confidently wrong.”

7. Random forests vs gradient boosting. Medium

What they are testing

Do you know bagging and boosting as mechanisms, not just names? Can you say which error each one attacks? Do you know what makes XGBoost and LightGBM fast and well-regularized?

Strong answer

Random forest: bagging plus feature sampling

Why it works. Deep trees have low bias and high variance. Averaging B trees with pairwise correlation ρ gives variance ρσ² + (1-ρ)σ²/B. More trees kill the second term. Feature sampling lowers ρ, which shrinks the first term. Adding more trees never overfits. It only stops helping.

Bonus. Each row is left out of about 37% of bootstraps, since (1 - 1/n)^n ≈ 1/e. Score each row with the trees that did not see it. That gives a free out-of-bag error estimate.

Gradient boosting: sequential fits to the gradient

Why it works. It is gradient descent in function space. Shallow trees have high bias and low variance. Each round removes bias. Too many rounds will overfit, so use early stopping on validation.

XGBoost details

LightGBM details

When to pick which

Follow-ups they will ask

Common traps

Say this out loud: “A random forest averages deep, decorrelated trees, so it cuts variance. Boosting adds shallow trees in sequence, each fit to the loss gradient, so it cuts bias. XGBoost adds a second-order step and a penalty on leaves. LightGBM adds histograms and leaf-wise growth for speed. I start with a forest as a baseline and ship tuned boosting.”

8. How do you set up cross-validation and avoid data leakage? Hard

What they are testing

This is the most common real-world failure. Can you match the split to how the model will be used? Can you name concrete leakage cases from experience? Do you know that preprocessing must sit inside the fold?

Strong answer

The rule. The validation split must copy the production setting. In production, the model predicts on future data, for new users, with only the features known at that moment. Every gap between your split and that reality is a leak.

Pick the right split

The leakage catalog

from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.impute import SimpleImputer
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import GroupKFold, TimeSeriesSplit, cross_val_score

# Every fitted step lives inside the pipeline, so it is refit on each training fold.
pipe = make_pipeline(SimpleImputer(strategy="median"), StandardScaler(),
                     LogisticRegression(max_iter=1000))

# One user never appears in both train and validation.
scores = cross_val_score(pipe, X, y, groups=user_id, cv=GroupKFold(n_splits=5),
                         scoring="average_precision")

# Rows sorted by time. gap= leaves an embargo between train and validation.
ts_scores = cross_val_score(pipe, X, y, cv=TimeSeriesSplit(n_splits=5, gap=7))

How to catch leaks.

Follow-ups they will ask

Common traps

Say this out loud: “My split copies production. Time splits with a gap for temporal data. Group splits when entities repeat. All fitted preprocessing lives inside the fold. Then I audit each feature for when its value becomes known, and I check that offline gains show up online.”

9. Walk me through feature engineering for a tabular model. Medium

What they are testing

Practical judgment. Do you know which models need scaling? Can you handle a category with a million values? Do you know how target encoding leaks? Do you treat missing values as signal?

Strong answer

Numeric features

Categorical features

Target encoding and its pitfalls

Target encoding replaces a category with the mean label for that category. It is powerful for high-cardinality data like zip code or merchant ID. It is also the easiest way to leak the label.

enc(c) = (n_c · mean_c + m · global_mean) / (n_c + m)      # m = prior strength
import numpy as np
import pandas as pd
from sklearn.model_selection import KFold

def target_encode_oof(cat: pd.Series, y: pd.Series, m: float = 20.0, k: int = 5, seed: int = 0):
    """Out-of-fold, smoothed target encoding. Each row never sees its own label."""
    out = np.zeros(len(cat))
    prior = y.mean()
    for tr, va in KFold(k, shuffle=True, random_state=seed).split(cat):
        stats = y.iloc[tr].groupby(cat.iloc[tr]).agg(["sum", "count"])
        enc = (stats["sum"] + m * prior) / (stats["count"] + m)
        out[va] = cat.iloc[va].map(enc).fillna(prior).to_numpy()   # unseen -> prior
    return out

Missing values

Follow-ups they will ask

Common traps

Say this out loud: “I scale for distance and gradient models but not for trees. Low-cardinality categories get one-hot. High-cardinality ones get hashing, embeddings, or out-of-fold smoothed target encoding. Missing values get an imputed value plus a missing flag, since missingness is often signal. Every fitted transform lives inside the fold.”

10. What is calibration and how do you fix it? Hard

What they are testing

Do you know the difference between ranking and calibrated probabilities? Can you read a reliability diagram? Do you know Platt, isotonic and temperature scaling, and when each fits? Do you know where calibration matters in a product?

Strong answer

Definition. A model is calibrated when its scores match real frequencies. Of all cases where it says 0.7, about 70% should be positive. Formally, P(Y = 1 | p̂ = p) = p for all p.

Why it matters. Ranking only needs order. Many decisions need the actual number.

Reliability diagram. Bin predictions by score, often 10 or 15 bins. For each bin, plot mean predicted score against observed positive rate. A calibrated model sits on the diagonal. Points below the diagonal mean overconfidence. Points above mean underconfidence.

mean predicted probability observed positive rate perfect overconfident
Red points below the diagonal at high scores: the model claims more than it delivers.

Measure it.

ECE = ∑_b (n_b / n) · | acc(b) - conf(b) |        # weighted gap across bins
MCE = max_b | acc(b) - conf(b) |                  # worst bin

ECE depends on the binning. Equal-mass bins are more stable than equal-width ones. Also report a proper scoring rule like log loss or the Brier score, mean((p - y)²). Those catch miscalibration and poor ranking together.

Who is miscalibrated, and which way.

Fix it. Fit a small map from score to probability on a held-out calibration set. Never on the training set.

import numpy as np

def ece(probs, labels, n_bins=15):
    """Expected calibration error for binary probabilities, equal-width bins."""
    probs, labels = np.asarray(probs), np.asarray(labels)
    edges = np.linspace(0, 1, n_bins + 1)
    idx = np.clip(np.digitize(probs, edges[1:-1]), 0, n_bins - 1)
    total = 0.0
    for b in range(n_bins):
        mask = idx == b
        if mask.any():
            total += mask.mean() * abs(labels[mask].mean() - probs[mask].mean())
    return total

def fit_temperature(logits, labels, grid=np.linspace(0.5, 5, 91)):
    """Pick T that minimises multiclass NLL on a validation set."""
    def nll(T):
        z = logits / T
        z = z - z.max(axis=1, keepdims=True)
        logp = z - np.log(np.exp(z).sum(axis=1, keepdims=True))
        return -logp[np.arange(len(labels)), labels].mean()
    return min(grid, key=nll)

Follow-ups they will ask

Common traps

Say this out loud: “Calibration means a score of 0.7 comes true 70% of the time. I check it with a reliability diagram, ECE, and log loss on held-out data. To fix it I fit Platt or isotonic on a separate calibration set, or temperature scaling for neural nets. It matters any time the number feeds a decision, like a bid or an expected cost.”

11. Generative vs discriminative models. Medium

What they are testing

Can you state what each one models? Do you know the classic Naive Bayes vs logistic regression result? Can you say when each one wins?

Strong answer

The Naive Bayes and logistic regression pair

Naive Bayes assumes features are independent given the class.

p(y | x) ∝ p(y) ∏_j p(x_j | y)

Take the log odds for binary y. With Gaussian features that share a variance per feature, or with Bernoulli features, the log odds are linear in x.

log p(y=1|x)/p(y=0|x) = log p(y=1)/p(y=0) + ∑_j log p(x_j|y=1)/p(x_j|y=0) = wᵀx + b

So both models have the same functional form for the posterior. They differ in how they fit w.

Ng and Jordan (2001). Naive Bayes reaches its (higher) best error with O(log d) examples. Logistic regression needs O(d) examples but reaches a lower best error when the independence assumption is wrong. So Naive Bayes often wins with little data. Logistic regression wins as data grows.

When to pick which

Follow-ups they will ask

Common traps

Say this out loud: “Generative models learn p of x and y and use Bayes’ rule. Discriminative models learn p of y given x directly. Naive Bayes and logistic regression give the same linear log odds, but Naive Bayes fits by counting under an independence assumption. So it converges faster with little data, while logistic regression wins with more data and correlated features.”

12. Explain the curse of dimensionality. Derive PCA. Hard

What they are testing

Can you make the curse concrete with numbers? Can you derive PCA from variance maximization to the eigenvectors of the covariance? Do you know the SVD link and the cases where PCA fails?

Strong answer

The curse of dimensionality

Why ML still works. Real data is not uniform. It lies near a low-dimensional manifold with structure. Models that exploit smoothness, sparsity or invariance beat the curse. Regularization, feature selection, and dimension reduction all help.

PCA derivation

Center the data so each column has mean 0. Call it X, with shape n by d. The sample covariance is Σ = XᵀX / (n-1).

Goal: find a unit vector u so that the projected data Xu has the most variance.

maximize   Var(Xu) = uᵀ Σ u
subject to uᵀu = 1

Lagrangian:  ℒ(u, λ) = uᵀΣu - λ(uᵀu - 1)
∂ℒ/∂u = 2Σu - 2λu = 0    ⇒    Σu = λu

So u must be an eigenvector of Σ. Plug back in: the variance is uᵀΣu = λuᵀu = λ. The best u is the eigenvector with the largest eigenvalue. The next component maximizes variance while staying orthogonal to the first. That gives the second eigenvector, and so on. Σ is symmetric, so its eigenvectors are orthogonal and the eigenvalues are real and non-negative.

The same answer comes from a second view. PCA finds the k-dimensional subspace with the smallest squared reconstruction error. Max variance and min reconstruction error are the same problem, since total variance is fixed.

The SVD link. Write X = U S Vᵀ. Then XᵀX = V S² Vᵀ. So the columns of V are the principal directions. The eigenvalues are s_i² / (n-1). The projected data is XV_k = U_k S_k. In practice use SVD on X, not an eigendecomposition of XᵀX. Forming XᵀX squares the condition number and loses precision. For huge data, use randomized SVD.

import numpy as np

def pca(X, k):
    """PCA by SVD. Returns projected data, components, and explained variance ratio."""
    Xc = X - X.mean(axis=0)                       # centering is required
    U, S, Vt = np.linalg.svd(Xc, full_matrices=False)
    var = S**2 / (len(X) - 1)                     # eigenvalues of the covariance
    components = Vt[:k]                           # rows are principal directions
    Z = Xc @ components.T                         # same as U[:, :k] * S[:k]
    return Z, components, var[:k] / var.sum()

Picking k. Keep enough components to explain 90 to 99% of variance. Or look for an elbow in the scree plot. Or pick k by the downstream model’s validation score, which is what matters.

When PCA fails

Follow-ups they will ask

Common traps

Say this out loud: “In high dimensions, space is empty and distances concentrate, so local methods break down. PCA maximizes projected variance subject to a unit norm. The Lagrangian gives Sigma u equals lambda u, so the top components are the top eigenvectors of the covariance. I compute it with SVD of the centered data. It fails on unscaled, non-linear, or label-irrelevant variance.”

Recap

← Guide 1 — The Applied Science Loop 2 — Statistics and Probability →