Everything an Applied Scientist should know that is not a model. This is the craft that turns a good model into shipped impact.
Interviewers rarely fail people for not knowing a loss function. They fail people who pick the wrong metric. They fail people who trust noisy labels or leak the future into training. This page covers the work around the model. It spans framing, data, evaluation, operations, ethics, and how you tell the story. Each topic gives a plain definition, the key moves, and why it matters. Each one ends with a short interview check.
The ML system design page shows how to design one full system in a round. This page is different. It is the toolbox you reach into for every system. Read it as a checklist of habits, not a list of facts.
Problem framing is the step that turns a business goal into something a model can learn. It names the decision, the prediction target, the metric, and the guardrails. Most failed ML projects fail here, long before training.
From business goal to ML objective
A business goal is fuzzy. “Reduce churn” does not tell you what to predict. You work down a ladder of four rungs. Each rung is more concrete than the one above.
Business goal. The outcome leaders care about. Example: keep more paying users after 90 days.
Decision. The action the system will take. Example: send a discount offer to some users.
Prediction target. What the model outputs to drive that decision. Example: the lift in 90-day retention if this user gets the offer.
Training objective and metric. The loss you optimize and the offline metric you track. Example: an uplift model scored by Qini curve.
Notice the trap in that example. A churn risk model predicts who will leave. But the decision needs who will change because of the offer. Users who will leave anyway waste the discount. Users who would stay anyway also waste it. Framing caught that the target is causal, not just predictive.
Always ask: what decision does this score drive? If no decision changes, the model has no value. If the decision is an intervention, you likely need a causal or uplift target.
A framing checklist
Who uses the output? A person, a ranking system, or a downstream model. Each needs a different output shape.
What is the label, exactly? Write the SQL in your head. Define the time window and the unit.
When is the prediction made? Only features known at that moment are legal.
What does an error cost? A false positive and a false negative rarely cost the same.
What is the online metric? Name the A/B test metric and the guardrails before you build.
What is the latency and cost budget? A 10 ms budget rules out many models.
Proxy metrics and their gaps
You can rarely train on the true goal. Long-term value is slow and noisy. So you train on a proxy like clicks or watch time. Every proxy has a gap. Clicks reward clickbait. Watch time rewards long, dull videos. Write the gap down. Then add a guardrail metric that catches it.
When NOT to use ML
Good scientists say no to ML often. Use a rule or a lookup when any of these hold.
The logic is known and stable. Tax rules and eligibility checks are code, not models.
There is little data. A few hundred labels will not beat an expert rule.
Errors are very costly and must be explained. Rules are easy to audit.
The value is small. A model adds upkeep forever. A 0.1% gain may not pay for it.
No feedback loop exists. If you can never see outcomes, you cannot tell if the model works.
Baselines and heuristics first
Build the dumbest thing that could work before anything clever. A baseline tells you how hard the problem is. It also gives you a number to beat in the launch review.
Constant baseline. Predict the mean or the majority class. Any model must beat this.
Popularity baseline. For ranking, sort by global popularity. It is often shockingly strong.
Simple rule. “Flag if amount is over three times the user's average.” Domain experts can give you five of these.
Linear or tree model. Logistic regression or gradient boosting on a few features. This is the real bar for a deep model.
If a deep model beats logistic regression by only a small margin, ask whether the extra cost is worth it. Report both numbers. Reviewers trust a scientist who shows the baseline.
Why it matters in practice
A wrong frame wastes months. Teams build a great churn predictor and then find it does not change retention. A one-page framing doc with the decision, label, and metric saves that. It also gets partners to agree on success before work starts.
Interview check
Q: A PM asks you to “use AI to reduce support tickets.” What do you do first? Ask what decision changes. Options include routing tickets, auto-answering, or fixing root causes. Then pick a measurable target and check if a rule already works.
Q: When would you refuse to use ML? When the rules are known, data is thin, value is small, or no feedback exists. Say a rule ships faster and is easy to audit.
Q: Why is a churn model the wrong tool for targeting offers? It predicts risk, not response to the offer. You need uplift, the change caused by treatment.
Data quality and labeling
Data quality is how well your data matches the world you want to model. Labeling is how you get the target values. Labels set the ceiling on model quality. No model can be more right than its labels.
Kinds of label noise
Random noise. Labels flip at random with some rate. Models are fairly robust to small amounts.
Class-dependent noise. One class gets mislabeled more often. Rare fraud labeled as normal is the classic case.
Instance-dependent noise. Hard or ambiguous items get mislabeled more. This is the most common and the most harmful.
Systematic bias. Labels reflect a process, not the truth. Arrest data reflects where police patrol.
Delayed labels. Chargebacks arrive weeks later. Recent data looks cleaner than it is.
Finding and fixing noisy labels
Train and inspect high-loss items. Items the model is confident about but labeled the other way are often label errors.
Confident learning. Use out-of-fold predicted probabilities to estimate which labels are likely wrong. The cleanlab library does this.
Re-label a sample. Send 500 items to expert raters. Measure the error rate of your source labels.
Robust losses. Label smoothing and noise-robust losses reduce the pull of wrong labels.
Keep a gold set. A small, expert-checked test set is worth more than a large noisy one.
Inter-annotator agreement and Cohen's kappa
Raw agreement is misleading. If 90% of items are “safe”, two lazy raters agree 82% of the time by chance alone. Cohen's kappa corrects for chance.
Cohen's kappa for two raters is κ = (p_o − p_e) / (1 − p_e). Here p_o is observed agreement and p_e is the agreement expected by chance from each rater's label rates. A kappa of 1 is perfect. A kappa of 0 is no better than chance.
import numpy as np
def cohens_kappa(a, b):
"""a, b: label arrays from two raters on the same items."""
a, b = np.asarray(a), np.asarray(b)
labels = np.union1d(a, b)
p_o = np.mean(a == b)
# chance agreement: sum over classes of P(rater A says k) * P(rater B says k)
p_e = sum(np.mean(a == k) * np.mean(b == k) for k in labels)
return (p_o - p_e) / (1 - p_e)
a = [1, 1, 0, 1, 0, 0, 1, 1, 0, 1]
b = [1, 0, 0, 1, 0, 1, 1, 1, 0, 1]
print(round(cohens_kappa(a, b), 3)) # 0.583
Rough reading. Below 0.4 is weak. From 0.4 to 0.6 is moderate. From 0.6 to 0.8 is good. Above 0.8 is strong.
More than two raters. Use Fleiss' kappa or Krippendorff's alpha. Alpha also handles missing ratings and ordinal scales.
Low kappa is a signal. It often means the guidelines are unclear, not that raters are bad.
Agreement caps your metric. If humans agree 85% of the time, a model at 84% is near the ceiling.
Writing labeling guidelines
Guidelines are a product spec for raters. Treat them like code. Version them, test them, and fix bugs.
State the task in one sentence. Then define every label with a short rule.
Give three clear positive examples and three clear negatives per label.
Spend most of the page on edge cases. These are where raters disagree.
Add a “cannot tell” option. Forced guesses add noise.
Run a pilot of 100 items. Measure kappa. Rewrite the rules where raters split. Repeat.
Seed known gold items into the queue to track rater quality over time.
Active learning
Active learning lets the model pick which unlabeled items to send for labeling next. The goal is the most model gain per labeling dollar.
Uncertainty sampling. Pick items where the model is least sure. Use margin, entropy, or probability near 0.5.
Query by committee. Train several models. Pick items where they disagree most.
Diversity sampling. Pick items that cover the feature space, like cluster centers. This avoids many near-duplicates.
Hybrid. Take the top uncertain items, then pick a diverse subset. This is the usual practical choice.
Active learning biases your data. The labeled set is no longer a random sample. Never use it as your test set. Keep a separate random sample for evaluation.
Why it matters in practice
Fixing labels often beats changing the model. A team that cleans 5% of wrong labels can gain more than a month of tuning. Clear guidelines also make labels stable across vendors and quarters. Without that, metric changes reflect rater drift, not model change.
Interview check
Q: Two raters agree 90% of the time. Is that good? Not alone. Compute kappa. With a skewed label mix, 90% can be close to chance.
Q: You have a budget for 10,000 labels. How do you spend it? A random gold test set first. Then a pilot to fix guidelines. Then active learning batches for training.
Q: How do you find mislabeled training items? Use out-of-fold predictions. Flag confident disagreements. Have experts review a sample to confirm.
Evaluation done right
Offline evaluation measures a model on held-out data before launch. Done right, it predicts the online result. Done wrong, it gives you false confidence and a failed A/B test.
Pick the metric by task. The metric must match how the output gets used.
Ranking metrics
Ranking metrics reward putting good items near the top. They differ in how they treat grades and position.
MRR (mean reciprocal rank). Score is 1 over the rank of the first relevant item. Use it when the user needs one right answer, like search for a known page.
MAP (mean average precision). Averages precision at each relevant item's rank. Use it for binary relevance with several good items.
NDCG@k. Handles graded relevance and discounts lower positions by a log factor. It is the default for feeds and search with ratings.
Recall@k. Fraction of relevant items in the top k. Use it for candidate generation, where later stages re-rank.
ROC AUC. Chance a random positive scores above a random negative. It ignores the class mix. It can look great on rare-event problems that are still unusable.
PR AUC. Area under precision and recall. Use it when positives are rare, like fraud or abuse.
Precision and recall at an operating point. The business runs at one threshold. Report the numbers there.
Log loss. Scores the probabilities, not just the order. Use it when downstream code consumes probabilities.
F1. Balances precision and recall. Use it only when you truly weight them equally.
Regression metrics
RMSE. Punishes large errors hard. Matches a squared-error cost.
MAE. Robust to outliers. Its best constant is the median, not the mean.
MAPE. Percent error. It blows up near zero and favors under-forecasting. Prefer WAPE or sMAPE for sales data.
Quantile (pinball) loss. Use it when you forecast a percentile, like stock levels at P90.
Calibration
Calibration means predicted probabilities match real frequencies. Among items scored 0.3, about 30% should be positive.
Why care. Ads auctions, expected value math, and thresholds all use the raw probability. A ranking metric will not catch a miscalibrated model.
How to check. Plot a reliability diagram. Bin predictions, then compare mean prediction to mean outcome per bin.
Expected calibration error (ECE). The weighted average gap across bins. Simple but sensitive to bin choice.
How to fix. Fit Platt scaling or isotonic regression on a held-out set. Temperature scaling works well for neural nets.
Common cause. Negative downsampling during training. Correct it with the known sampling rate.
Slices and fairness metrics
An average hides failures. A model can gain 1% overall and lose 10% for new users. Always report key slices.
Standard slices: new versus returning users, country, device, language, traffic source, and item age.
Sensitive slices: protected groups where you have consent and legal basis to measure.
Fairness gaps: compare positive rates, true positive rates, and false positive rates across groups.
Watch sample sizes. A small slice has wide error bars. Do not chase noise.
Confidence intervals on metrics
A metric from a test set is an estimate. A 0.3 point gain in AUC may be noise. Put an interval on it. The bootstrap is the general tool.
import numpy as np
from sklearn.metrics import roc_auc_score
def paired_bootstrap_delta(y, p_a, p_b, metric=roc_auc_score, n=2000, seed=0):
"""95% CI for metric(B) - metric(A) on the SAME test items."""
rng = np.random.default_rng(seed)
y, p_a, p_b = map(np.asarray, (y, p_a, p_b))
deltas = []
for _ in range(n):
idx = rng.integers(0, len(y), len(y)) # resample items with replacement
if y[idx].min() == y[idx].max():
continue # AUC needs both classes
deltas.append(metric(y[idx], p_b[idx]) - metric(y[idx], p_a[idx]))
lo, hi = np.percentile(deltas, [2.5, 97.5])
return float(np.mean(deltas)), (float(lo), float(hi))
Pair the comparison. Resample the same items for both models. This cancels item difficulty and gives tighter intervals.
Resample the right unit. For ranking, resample queries, not documents. For users, resample users. Items in a cluster are not independent.
Closed forms exist. Use the Wilson interval for a proportion. Use DeLong's test for comparing two AUCs.
Holdout hygiene
Split by time when the model will predict the future. A random split leaks future patterns.
Split by entity when the same user or item appears many times. Otherwise the model memorizes the user.
Remove near-duplicates across train and test. Copied text and re-uploads inflate scores.
Lock the test set. Tune on validation only. Touch the test set once per major decision.
Refresh it. Repeated peeking overfits the test set over many months. Rotate in a fresh one.
Check feature timing. Each feature must be computed as of the prediction time. This is point-in-time correctness.
Offline and online can disagree. Offline data comes from the old model's choices. Your new model shows items the old one never showed. So offline gains are a filter, not a verdict. The A/B test is the verdict.
Why it matters in practice
The wrong metric picks the wrong model. A model with higher AUC can lose money if it is miscalibrated in an auction. A gain without an interval may vanish online. Clean holdouts and slice reports are what make launch reviews fast and trusted.
Interview check
Q: Fraud is 0.1% of transactions. Which metric do you report? PR AUC and recall at a fixed precision or alert budget. ROC AUC looks high even for weak models here.
Q: Model B beats A by 0.4 points of NDCG. Ship it? First get a paired bootstrap interval over queries. Check slices. Then run an A/B test, since offline data is biased by the old ranker.
Q: When is NDCG better than MRR? When relevance is graded and many items matter. MRR fits tasks with a single right answer.
Error analysis
Error analysis is the habit of looking at what the model gets wrong and why. It turns a single metric into a ranked list of fixes.
Look at 100 examples
This is the most useful hour in any project. Pull 100 errors at random. Read each one. Write a short note on why it failed. Then count the notes.
Sample errors at random, not the worst ones. You want the true mix.
Look at the raw input, the label, the prediction, and the score.
Tag each error with one or more causes. Make up tags as you go.
Merge similar tags after about 30 items. Finish the rest with the merged list.
Count the tags. The biggest bucket is your next project.
Often 20% of errors turn out to be label errors. That changes your plan. It also means your true accuracy is higher than reported.
Confusion slices
Confusion matrix by class. Shows which classes get swapped. Two swapped classes may need merging or better guidelines.
Error rate by slice. Group by metadata like country, length, or source. Sort slices by error count times error rate.
Automatic slice finding. Fit a shallow decision tree to predict “model was wrong”. Its leaves are slices with high error.
Score bands. Errors near the threshold need different fixes than confident errors.
Build a failure taxonomy
A taxonomy is a stable list of failure types. It lets you track progress across model versions. A good one has under ten buckets and covers 90% of errors.
Label error. The model was right. Fix the data.
Missing signal. The input lacks the needed info. Add a feature or data source.
Rare pattern. Too few training examples. Collect or augment more.
Ambiguous case. Humans also disagree. Accept it or refine the task.
Pipeline bug. Wrong preprocessing, truncation, or stale features. Fix the code.
Distribution shift. New patterns since training. Retrain or add fresher data.
Figure 1. A failure taxonomy turns a vague error rate into a ranked list of projects. The bar length is the ceiling on what each fix can gain.
The ceiling rule. If a bucket holds 10% of errors, fixing it fully cuts errors by at most 10%. Use that to rank projects before you start any of them.
Why it matters in practice
Teams without error analysis guess at fixes. They try a bigger model when the real issue is a truncation bug. A failure taxonomy also makes progress visible to leaders. You can say the “missing signal” bucket shrank from 32 to 12.
Interview check
Q: Your model's accuracy plateaued. What next? Sample 100 errors and tag causes. Rank buckets by size and cost to fix. Pick the top one.
Q: Why sample errors at random, not the worst? The worst cases are often outliers or label errors. Random sampling shows the true mix you need to fix.
Q: How do you find a weak slice you did not think of? Fit a shallow tree on metadata to predict errors. Inspect the leaves with high error rates.
Reproducibility
Reproducibility means anyone can rerun your experiment and get the same result. It needs the same code, data, config, and environment. Without it, you cannot trust or debug a gain.
The four things to pin
Code. Record the git commit hash for every run. Refuse to run with uncommitted changes in serious work.
Data. Record a version or snapshot ID for every dataset. A table that changes daily is not a dataset.
Config. Store every hyperparameter in a config file, not in code or the command line alone.
Environment. Pin library versions with a lock file or a container image.
Seeds and determinism
import os, random
import numpy as np
import torch
def set_seed(seed: int = 42):
random.seed(seed)
np.random.seed(seed)
torch.manual_seed(seed)
torch.cuda.manual_seed_all(seed)
os.environ["PYTHONHASHSEED"] = str(seed)
# deterministic kernels: slower, but bit-for-bit repeatable on the same hardware
torch.backends.cudnn.deterministic = True
torch.backends.cudnn.benchmark = False
torch.use_deterministic_algorithms(True, warn_only=True)
Seed data loader workers too. Each worker needs its own derived seed.
Some GPU ops are not deterministic. Exact repeats may need the same hardware and drivers.
Seeds are not a fix for variance. Run three to five seeds and report the mean and spread. A gain smaller than seed noise is not a gain.
Data versioning
Snapshot training tables with a date partition and never overwrite it.
Use tools like DVC, lakeFS, or Delta Lake time travel for files and tables.
Store a hash and row count with each version. Check both at load time.
Log the exact SQL or pipeline version that built each dataset.
Experiment tracking
A tracker logs every run with its inputs and outputs. MLflow and Weights & Biases are common choices. Log these for each run.
Git hash, data version, full config, and seed.
Training curves, final metrics, and slice metrics.
Hardware, wall time, and cost.
The model artifact and a short note on the hypothesis.
Configs done well
# config/ranker_v7.yaml (read by the training script, logged with the run)
# data:
# train_table: events.ranker_train
# snapshot: 2026-09-30
# label: long_click_30s
# model:
# type: gbdt
# num_trees: 800
# learning_rate: 0.05
# eval:
# slices: [country, device, new_user]
# seed: 13
from dataclasses import dataclass
import yaml
@dataclass(frozen=True)
class Config:
data: dict
model: dict
eval: dict
seed: int
def load_config(path: str) -> Config:
with open(path) as f:
return Config(**yaml.safe_load(f))
One config file fully defines one run. Tools like Hydra add overrides and sweeps.
Freeze the config object so code cannot change it mid-run.
Diff two configs to explain why two runs differ.
Why it matters in practice
Unrepeatable gains are a common reason launches stall. A reviewer asks you to rerun with one change and you cannot match the old number. Good tracking also saves weeks when a model breaks in production. You can rebuild the last good version in hours.
Interview check
Q: You cannot reproduce last month's best model. Where do you look? Data version first, since tables change. Then code commit, config, library versions, and seed.
Q: A new idea gains 0.2% AUC. Is it real? Run several seeds for both arms. Compare the gain to the seed spread and a bootstrap interval.
Q: What do you log for every run? Commit, data version, config, seed, metrics, slices, cost, and the artifact.
The ML lifecycle and MLOps
MLOps is the set of tools and habits that take a model from notebook to production. It also keeps the model healthy after launch. It covers features, training pipelines, registries, deployment, and monitoring.
Figure 2. The ML lifecycle. One feature definition feeds both training and serving. Each new model passes the registry, shadow, and canary before full traffic. Monitoring closes the loop.
Feature stores
A feature store keeps one definition per feature and serves it two ways. An offline store gives history for training. An online store gives low-latency reads for serving.
Point-in-time joins. For each training row, it returns the feature value as of that row's timestamp. This prevents leakage.
Reuse. Teams share features instead of rebuilding them. This cuts cost and bugs.
Freshness. Batch features update daily. Streaming features update in seconds. Know which you have.
Training-serving skew
Training-serving skew is any difference between what the model saw in training and what it sees in production. It quietly drops online quality while offline metrics look fine.
Code skew. Python computes a feature in training. Java recomputes it in serving with a small bug.
Data skew. Training used a cleaned log. Serving gets raw, late, or missing values.
Time skew. Training used a feature computed at end of day. Serving sees it hours stale.
Feedback skew. The model changes which items users see. So the input mix shifts after launch.
How to catch it. Log the exact features at serving time. Train on those logs when you can. Compare training and serving feature distributions daily. Alert on any feature whose mean or null rate moves past a threshold.
Model registry
Stores each model version with its metrics, data version, config, and owner.
Tracks the stage: candidate, staging, production, or retired.
Makes rollback one click. Point production back at the last good version.
Holds model cards that note intended use and known limits.
CI for models
Software CI checks that code works. Model CI also checks that the model is good enough to ship.
Unit tests for feature code and transforms. Test edge cases like nulls and empty lists.
Data tests on each new training set. Check schema, row counts, null rates, and label rate. Great Expectations is one tool.
Training smoke test. Overfit a tiny batch. Loss should go near zero. If not, something is broken.
Quality gates. The new model must match or beat production on the gold set and key slices.
Behavioral tests. Invariance checks, like a name swap should not change a sentiment score.
Serving tests. Latency, memory, and output parity between offline and online scoring.
Shadow and canary deployment
Shadow. The new model scores live traffic but its output is thrown away. You compare its scores, latency, and errors to production. No user risk.
Canary. The new model serves a small slice of real users, like 1%. Watch health and guardrail metrics. Then ramp to 5%, 25%, and 50%.
A/B test. The canary becomes a proper experiment with a holdout. This measures the real effect.
Auto rollback. Define the alarms in advance. If error rate or a guardrail breaks, traffic goes back without a meeting.
Monitoring after launch
System health. Latency, error rate, throughput, and timeouts.
Input drift. Feature distributions against the training set. Use population stability index or KL divergence.
Output drift. Score distribution and predicted positive rate.
Outcome metrics. Real accuracy once labels arrive, plus the business metric.
Retraining policy. Retrain on a schedule or when drift crosses a line. Always gate the new model with the same CI.
Why it matters in practice
Most production ML incidents are not model bugs. They are stale features, schema changes, or skew. MLOps turns these silent failures into alerts. It also makes it safe to ship often, which compounds gains over a year.
Interview check
Q: Offline AUC rose but the A/B test is flat. What could cause it? Training-serving skew, stale features, or a feedback loop. Also a proxy metric that does not track the online goal. Start by comparing logged serving features to training features.
Q: Shadow or canary first? Shadow first, since it has no user risk. It catches crashes, latency, and score drift. Then canary to see real user response.
Q: What belongs in a model CI gate? Data checks, a smoke test, quality on gold set and slices, behavioral tests, and latency limits.
Model efficiency
Model efficiency is getting the same quality for less compute, memory, latency, or money. At scale, it decides whether a model ships at all.
Cost per prediction
Start every efficiency talk with this number. It ties model choices to the budget.
def cost_per_1k_predictions(gpu_hour_usd, qps_per_gpu, utilization=0.6):
"""Rough serving cost. utilization covers headroom for peaks."""
preds_per_hour = qps_per_gpu * 3600 * utilization
return 1000 * gpu_hour_usd / preds_per_hour
# A big model: 200 QPS per GPU at $2.50/hr
print(round(cost_per_1k_predictions(2.50, 200), 4)) # 0.0058 dollars
# Distilled model: 1500 QPS per GPU
print(round(cost_per_1k_predictions(2.50, 1500), 4)) # 0.0008 dollars
At one billion predictions a day, that gap is about $5,000 a day. A small quality loss may be well worth it. Put both numbers in the launch doc.
Distillation
Train a small student model to match a large teacher's outputs.
Soft targets carry more information than hard labels. They show which wrong answers are close.
Use a temperature to soften the teacher's softmax. Mix the distillation loss with the true label loss.
The teacher can label large unlabeled sets. This often matters more than the loss form.
Typical result: a much smaller model that keeps most of the teacher's quality.
Quantization
Store weights, and often activations, in fewer bits. Move from FP32 to FP16, INT8, or INT4.
Post-training quantization. Quick and needs only a small calibration set. Works well at INT8 for most models.
Quantization-aware training. Simulates low precision during training. Needed for aggressive cuts like INT4 on small models.
Gains: two to four times less memory and faster math on hardware that supports it.
Risk: outlier weights and activations lose precision. Always check slice metrics, not just the average.
Pruning
Unstructured pruning. Zero out small weights. Saves size but rarely speed without sparse kernels.
Structured pruning. Remove whole neurons, heads, channels, or layers. Gives real speedups on normal hardware.
Prune gradually and fine-tune after each step to recover quality.
Systems tricks
Caching. Cache scores or embeddings for repeat inputs. Item embeddings can be computed once a day.
Batching. Group requests to fill the GPU. Dynamic batching waits a few ms to form a batch. Trade a little latency for big throughput.
Cascades. A cheap model filters most items. The expensive model scores only the top few hundred.
Early exit. Return when the model is already confident. Send only hard inputs through the full model.
Feature pruning. Drop features with low importance and high fetch cost. Feature fetch often dominates latency.
Measure tail latency, like p99, not just the mean. Users and timeouts feel the tail. A batch that waits too long can blow the p99 budget.
Why it matters in practice
Compute is often the largest line in an ML team's budget. A model that wins by 0.5% but costs four times more may be rejected. Showing cost per prediction and a cheaper variant makes your launch easy to approve.
Interview check
Q: Your model is too slow for a 20 ms budget. What do you try? Profile first. Then cut feature fetch, cache embeddings, add a cascade, and quantize or distill the model.
Q: Why does unstructured pruning often not speed things up? Dense hardware still does the multiply on zeros. You need structured pruning or sparse kernels.
Q: Why do soft targets help distillation? They encode how similar the classes are. The student learns more per example than from one-hot labels.
Responsible AI
Responsible AI is building models that are fair, private, safe, and understandable enough for their use. It is a set of measurable checks, not a slogan.
Fairness definitions
Let A be a group attribute, Y the true label, and Ŷ the prediction.
Demographic parity. Positive rates are equal across groups. P(Ŷ=1 | A=a) is the same for all a.
Equal opportunity. True positive rates are equal. Qualified people get approved at the same rate.
Equalized odds. Both true positive and false positive rates are equal across groups.
Predictive parity. Precision is equal. A positive prediction means the same thing for each group.
Calibration within groups. A score of 0.7 means 70% for every group.
Individual fairness. Similar people get similar predictions.
Why the definitions conflict
You cannot have them all at once. This is a theorem, not a tooling gap.
If base rates differ across groups, calibration and equalized odds cannot both hold. Only a perfect model escapes this.
Demographic parity forces equal rates even when true rates differ. That breaks calibration or equal opportunity.
So fairness is a choice. Pick the definition that matches the harm. Write down why.
Lending or hiring. Missing a qualified person is the harm. Equal opportunity is a common fit.
Risk scores read by people. A score must mean the same for everyone. Calibration within groups fits.
Ad delivery for jobs or housing. Exposure itself matters. Some form of parity may be required by law.
Mitigation options
Pre-processing. Fix the data. Reweight, resample, or relabel biased examples.
In-processing. Add a fairness constraint or penalty to the loss.
Post-processing. Use group-specific thresholds. Simple, but may raise legal questions.
Removing the sensitive feature is not enough. Zip code and name act as proxies.
Privacy
Collect less. Only collect and keep data the task needs. Set retention limits.
Anonymization is weak. Removing names does not stop re-identification from other fields.
Models can leak. Membership inference can tell if a person was in the training set. Large models can recite training text.
Federated learning. Train on devices and send only updates. Data stays local, but updates can still leak without more protection.
Differential privacy intuition
Differential privacy promises the output barely changes if any one person's data is added or removed. So nobody can learn much about one person from the output. The privacy budget ε sets how much it may change. Smaller ε means more privacy and more noise.
For counts. Add Laplace noise scaled to how much one person can change the count, divided by ε.
For training. DP-SGD clips each example's gradient, then adds Gaussian noise to the sum.
The budget adds up. Each query spends some ε. Many queries mean weak privacy overall.
The cost. Accuracy drops, most for small groups and rare patterns. DP can hurt fairness for minorities.
import numpy as np
def dp_count(true_count, epsilon, sensitivity=1.0, rng=np.random.default_rng()):
"""Laplace mechanism: one person changes a count by at most 1."""
return true_count + rng.laplace(0.0, sensitivity / epsilon)
# epsilon=1: noise std about 1.4. epsilon=0.1: about 14.
print(round(dp_count(1200, epsilon=1.0)), round(dp_count(1200, epsilon=0.1)))
Explainability and its limits
SHAP. Splits a prediction into feature contributions using Shapley values. TreeSHAP is exact and fast for tree models.
LIME. Fits a simple local model around one prediction using perturbed samples.
Global views. Permutation importance and partial dependence plots show average effects.
Know the limits. Interviewers like this part.
Not causal. SHAP shows what the model used, not what drives the outcome in the world.
Correlated features. Credit splits in odd ways between correlated features. Perturbations create impossible inputs.
Unstable. LIME results change with the random sample and kernel width. Two runs can disagree.
Baseline choice. SHAP values depend on the background data you pick.
Can be gamed. A model can be built to look fair under LIME or SHAP while still being biased.
Why it matters in practice
Fairness and privacy failures cause legal risk, press, and lost trust. Regulators now ask for documented checks in many domains. A scientist who can name the right fairness metric and its trade-off earns trust in reviews. Knowing explainability limits stops teams from over-claiming in front of regulators.
Interview check
Q: Can a model be calibrated and have equal false positive rates across groups? Not when base rates differ, unless it is perfect. You must choose based on the harm.
Q: We dropped race from the features. Is the model fair now? No. Proxies like zip code carry the signal. Measure outcomes by group directly.
Q: What does epsilon mean in differential privacy? It bounds how much one person can change the output. Smaller is more private and noisier.
Recommender and ranking essentials
A recommender picks and orders items for a user. Most learn from implicit feedback, like clicks and watch time, which is plentiful but biased.
Implicit feedback
Positives are noisy. A click may be a mis-tap or clickbait. Use stronger signals like dwell time or long watches.
Negatives are missing. A non-click may mean the user never saw the item. Only treat shown and skipped items as negatives.
Confidence weighting. Weight positives by strength, like more watch time means more confidence.
Negative sampling. Sample random items as negatives for retrieval. Correct for popularity, since popular items get sampled more.
Multi-task. Predict several signals at once. Combine them with tuned weights into one score.
Position bias
Users click top items more, whatever their quality. Training on raw clicks teaches the model to copy the old ranking.
Measure it. Randomize the top few positions for a small slice of traffic. Click rate by position gives the propensity.
Inverse propensity weighting. Weight each click by one over the chance its position was seen.
Position as a feature. Train with position as an input. At serving time, set it to a fixed value.
A shallow bias tower. Learn position effects in a separate small network and add it to the logit. Drop it at serving.
Exploration with bandits
If you only show what the model likes, you never learn about new items. This is the explore and exploit trade-off. Multi-armed bandits handle it.
Epsilon-greedy. Show the best arm most of the time. With chance epsilon, show a random arm. Simple, but explores blindly and forever.
UCB. Pick the arm with the highest upper confidence bound. Arms with few trials get a bonus. Exploration shrinks as data grows.
Thompson sampling. Keep a belief over each arm's reward. Sample once from each belief. Pick the arm with the highest sample. It often wins in practice and handles delayed feedback well.
Contextual bandits. Condition on user and item features. LinUCB and neural Thompson sampling are common.
import numpy as np
class ThompsonBernoulli:
"""Beta-Bernoulli Thompson sampling for click/no-click arms."""
def __init__(self, n_arms, seed=0):
self.a = np.ones(n_arms) # successes + 1
self.b = np.ones(n_arms) # failures + 1
self.rng = np.random.default_rng(seed)
def choose(self):
return int(np.argmax(self.rng.beta(self.a, self.b)))
def update(self, arm, reward):
self.a[arm] += reward
self.b[arm] += 1 - reward
def ucb1_choose(counts, sums, t):
"""counts, sums: arrays per arm. t: total pulls so far."""
if (counts == 0).any():
return int(np.argmin(counts)) # try every arm once
means = sums / counts
bonus = np.sqrt(2 * np.log(t) / counts)
return int(np.argmax(means + bonus))
def eps_greedy_choose(means, eps, rng):
return int(rng.integers(len(means))) if rng.random() < eps else int(np.argmax(means))
Figure 3. Each arm has a Beta belief. The new item's wide belief sometimes draws above the old item's narrow one. That gives it traffic in proportion to how likely it is to be the best.
Offline evaluation of policies
Logged data came from the old policy. You cannot just replay it for a new one.
Inverse propensity scoring. Reweight logged rewards by new policy probability over logging probability. It needs logged propensities.
Doubly robust. Combine a reward model with IPS. Lower variance, still unbiased if either part is right.
So log the probability of every action you take. Without it, you cannot evaluate new policies offline.
Other must-knows
Cold start. Use content features for new items and context for new users. Reserve exploration traffic.
Feedback loops. The model shapes what users see, which shapes training data. Popular items get more popular.
Diversity. Re-rank to avoid ten near-identical items. Use MMR or a determinantal point process.
Long-term value. Clicks today can hurt retention later. Track a long-term holdout.
Why it matters in practice
Ignoring position bias makes a ranker learn its own past. Ignoring exploration starves new items and creators. Logging propensities is cheap now and expensive to add later. These details often separate a flat A/B test from a winning one.
Interview check
Q: How do you correct for position bias in click training data? Estimate propensity by position with a small randomization. Then use IPS weights or a position feature fixed at serving.
Q: Epsilon-greedy versus Thompson sampling? Epsilon-greedy explores at a fixed random rate. Thompson explores where it is uncertain and tapers off. Thompson usually has lower regret.
Q: Why log action probabilities? Off-policy evaluation needs them. Without them, you can only judge new policies by online tests.
Time series essentials
Time series data has an order in time, and the order matters. Nearby points are related. The future must never leak into the past.
Stationarity
Stationary means the mean, variance, and autocorrelation do not change over time.
Many classic models, like ARIMA, assume it after transforms.
Make it stationary. Difference the series. Take logs to steady the variance. Remove seasonality.
Test it. The ADF test has a null of a unit root. The KPSS test has a null of stationarity. Use both, since they can disagree.
Look first. Plot the series, its rolling mean, and its ACF before any test.
Leakage traps
Random splits. They put tomorrow in training and today in test. Always split by time.
Scaling on all data. Fit scalers and encoders on the training window only.
Centered windows. A rolling mean centered on t uses future values. Use trailing windows.
Revised data. Sales and economic figures get revised later. Train on what was known at the time.
Target-based features. Lags must be at least the forecast horizon. A 7-day forecast cannot use yesterday's value.
Backtesting
Backtesting simulates forecasting in the past, many times. It is cross-validation for time series.
import numpy as np
def rolling_origin_splits(n, initial, horizon, step):
"""Yield (train_idx, test_idx). Train always ends before test starts."""
end = initial
while end + horizon <= n:
yield np.arange(0, end), np.arange(end, end + horizon)
end += step
def backtest(y, fit_predict, initial=365, horizon=28, step=28):
errs = []
for tr, te in rolling_origin_splits(len(y), initial, horizon, step):
pred = fit_predict(y[tr], horizon)
errs.append(np.sum(np.abs(y[te] - pred)) / np.sum(np.abs(y[te]))) # WAPE
return float(np.mean(errs)), float(np.std(errs))
def seasonal_naive(train, horizon, season=7):
last = train[-season:]
return np.resize(last, horizon) # repeat last week
Expanding window. Train on all history up to each cut. Good when old data still helps.
Sliding window. Train on a fixed recent span. Good when patterns drift.
Gap. Leave a gap between train and test if features have reporting delays.
Report error by horizon. Day-1 and day-28 accuracy differ a lot.
Forecasting baselines
These are hard to beat. Always report them.
Naive. Tomorrow equals today.
Seasonal naive. Next Monday equals last Monday.
Moving average. The mean of the last k periods.
Exponential smoothing. ETS and Holt-Winters handle trend and season with few parameters.
Gradient boosting on lags. Lag, rolling, and calendar features in LightGBM. It is the workhorse for many related series.
For many related series, like sales per store, one global model often beats one model per series. It shares strength across series and handles new ones better.
Why it matters in practice
Leaky time series models look brilliant offline and fail on day one. Forecasts drive inventory, staffing, and capacity, so errors cost real money. A backtest with honest baselines is what makes a forecast credible to finance and ops partners.
Interview check
Q: Why not use k-fold cross-validation on time series? It trains on the future to predict the past. Use rolling-origin backtests instead.
Q: Your forecast model barely beats seasonal naive. Ship it? Check the gain across backtest folds and horizons. If it is small and unstable, the simpler model is safer.
Q: Name a subtle leak in time series features. A centered rolling mean, or lags shorter than the forecast horizon.
Writing and communication
Scientific communication is turning your work into a decision for someone else. A result that nobody understands does not ship. Writing is part of the job, not an extra.
The one-page doc
Leaders read the top and decide. Put the answer first. Use this order.
The ask or the answer. One or two sentences. “We should launch ranker v7 to 100%.”
Why it matters. The business problem and the size of the prize.
What we did. The approach in plain words. No equations unless needed.
Results. The key numbers with intervals. One chart at most.
Risks and trade-offs. Cost, guardrail moves, and known failure modes.
Next steps. Who does what, by when.
The experiment write-up
Hypothesis. What you expected and why, written before the test.
Setup. Arms, traffic split, duration, unit of randomization, and primary metric.
Results. Primary metric with confidence interval. Guardrails. Key slices.
Checks. Sample ratio mismatch, novelty effects, and pre-period balance.
Decision. Ship, iterate, or stop. Say why in one line.
Learnings. What this tells you about users or the model, even if it failed.
Write up failed experiments too. A clear null result saves the next team a quarter. It also shows judgment, which matters as much as wins for promotion.
Telling the impact story with numbers
Translate the metric. “+0.8% sessions” becomes “about 2 million more sessions a day.”
Use relative and absolute. “A 12% relative cut in fraud loss, or $3.1M a year.”
Show the interval. “+0.8%, 95% CI +0.5% to +1.1%.” It builds trust.
Compare to a baseline. “Twice the gain of the last three launches.”
Name the cost. “Adds 3 ms of latency and $40K a year in compute.”
Say what you did. In promotion packets and interviews, separate your part from the team's.
Example impact line. “I replaced the rule-based fraud filter with a gradient boosted model. At the same review budget, it caught 18% more fraud. That saved about $3M a year. False positives stayed flat.”
Habits that help
Short sentences and plain words. Define each acronym once.
Chart titles that state the finding, like “New users gain most”.
Pre-read the doc with one partner before the big review.
Match depth to the reader. Leaders get the decision. Peers get an appendix.
Why it matters in practice
At senior levels, influence comes through writing. Your doc reaches rooms you are not in. A clear one-pager can win headcount or kill a bad project. Interviewers also judge the research talk and behavioral rounds on how clearly you tell impact.
Interview check
Q: How do you present a result to a VP in two minutes? Lead with the decision and the business number. Then one sentence on how, and one on the main risk.
Q: Your experiment was flat. How do you write it up? State the hypothesis, the power, the interval, and what it rules out. Then say what you will try next.
Q: How do you state impact in a resume line? Action, metric moved, size in business terms, and your role.
Reading papers and staying current
Staying current means knowing which new ideas matter for your problems, without reading everything. It is a filtering skill, not a volume skill.
The three-pass read
Pass one, five minutes. Title, abstract, figures, and conclusion. Decide if it is worth more time.
Pass two, one hour. Read the method and main results. Skip proofs. Note the baselines and datasets.
Pass three, a few hours. Only for papers you will build on. Rederive key steps. Run the code if it exists.
Questions to ask every paper
What problem does it solve, and do I have that problem?
Are the baselines strong and well tuned, or straw men?
How big is the gain, and is there an error bar or multiple seeds?
What does it cost in compute, latency, and data?
Does the ablation show which part matters?
Would it work at my scale, with my noisy, shifting data?
Benchmark gains rarely transfer one to one. A 2% gain on an academic set can vanish on production data. Treat papers as a list of ideas to test, not as promises.
An efficient system
Curate sources. Follow a few strong labs and authors. Use a paper digest or a reading group.
Read by need. Go deep on what you are building now. Skim the rest.
Keep notes. One paragraph per paper: idea, result, and when you would use it.
Reproduce one. Re-implementing a key paper each quarter builds real depth.
Read industry papers. Papers from large companies show what works under real constraints.
Budget time. Two to three hours a week is enough if it is focused.
Turning papers into impact
Map each idea to a known failure bucket from your error analysis.
Estimate the ceiling gain and the cost before trying it.
Try the cheapest version first. Often 20% of the method gives 80% of the gain.
Share a short summary with the team. This spreads value and builds your reputation.
Why it matters in practice
The field moves fast. A scientist who ignores it falls behind. A scientist who chases every paper ships nothing. The skill is picking the few ideas that fit your failure modes and testing them cheaply.
Interview check
Q: Tell me about a recent paper you liked. State the problem, the key idea, and the result. Then give one weakness and how you would apply it.
Q: How do you judge if a paper's gain is real? Check baseline strength, error bars, ablations, and compute parity.
Q: How do you stay current with limited time? A few curated sources, three-pass reading, short notes, and depth only where it maps to current work.
Recap
Frame the decision first. Build baselines and rules before models.
Labels set the ceiling. Measure agreement with kappa and fix the guidelines.
Pick metrics by task. Add calibration, slices, and confidence intervals. Keep holdouts clean.
Read 100 errors. Build a failure taxonomy and fix the biggest bucket.
Pin code, data, config, and environment. Run several seeds.
Use one feature definition for training and serving. Ship through shadow, canary, and A/B.
Know cost per prediction. Distill, quantize, cache, and batch.
Fairness definitions conflict. Choose by the harm. Know the limits of SHAP and LIME.
Correct for position bias. Explore with bandits. Log propensities.
Split time series by time. Backtest against seasonal naive.
Write the answer first. Tell impact in business numbers with intervals.
Read papers in three passes. Test the cheap version against your failure buckets.