Part IV Theory 6 Craft

The Practitioner's Toolkit

Everything an Applied Scientist should know that is not a model. This is the craft that turns a good model into shipped impact.

Interviewers rarely fail people for not knowing a loss function. They fail people who pick the wrong metric. They fail people who trust noisy labels or leak the future into training. This page covers the work around the model. It spans framing, data, evaluation, operations, ethics, and how you tell the story. Each topic gives a plain definition, the key moves, and why it matters. Each one ends with a short interview check.

The ML system design page shows how to design one full system in a round. This page is different. It is the toolbox you reach into for every system. Read it as a checklist of habits, not a list of facts.

Contents

  1. Problem framing
  2. Data quality and labeling
  3. Evaluation done right
  4. Error analysis
  5. Reproducibility
  6. The ML lifecycle and MLOps
  7. Model efficiency
  8. Responsible AI
  9. Recommender and ranking essentials
  10. Time series essentials
  11. Writing and communication
  12. Reading papers and staying current

Problem framing

Problem framing is the step that turns a business goal into something a model can learn. It names the decision, the prediction target, the metric, and the guardrails. Most failed ML projects fail here, long before training.

From business goal to ML objective

A business goal is fuzzy. “Reduce churn” does not tell you what to predict. You work down a ladder of four rungs. Each rung is more concrete than the one above.

  1. Business goal. The outcome leaders care about. Example: keep more paying users after 90 days.
  2. Decision. The action the system will take. Example: send a discount offer to some users.
  3. Prediction target. What the model outputs to drive that decision. Example: the lift in 90-day retention if this user gets the offer.
  4. Training objective and metric. The loss you optimize and the offline metric you track. Example: an uplift model scored by Qini curve.

Notice the trap in that example. A churn risk model predicts who will leave. But the decision needs who will change because of the offer. Users who will leave anyway waste the discount. Users who would stay anyway also waste it. Framing caught that the target is causal, not just predictive.

Always ask: what decision does this score drive? If no decision changes, the model has no value. If the decision is an intervention, you likely need a causal or uplift target.

A framing checklist

Proxy metrics and their gaps

You can rarely train on the true goal. Long-term value is slow and noisy. So you train on a proxy like clicks or watch time. Every proxy has a gap. Clicks reward clickbait. Watch time rewards long, dull videos. Write the gap down. Then add a guardrail metric that catches it.

When NOT to use ML

Good scientists say no to ML often. Use a rule or a lookup when any of these hold.

Baselines and heuristics first

Build the dumbest thing that could work before anything clever. A baseline tells you how hard the problem is. It also gives you a number to beat in the launch review.

If a deep model beats logistic regression by only a small margin, ask whether the extra cost is worth it. Report both numbers. Reviewers trust a scientist who shows the baseline.

Why it matters in practice

A wrong frame wastes months. Teams build a great churn predictor and then find it does not change retention. A one-page framing doc with the decision, label, and metric saves that. It also gets partners to agree on success before work starts.

Interview check

Data quality and labeling

Data quality is how well your data matches the world you want to model. Labeling is how you get the target values. Labels set the ceiling on model quality. No model can be more right than its labels.

Kinds of label noise

Finding and fixing noisy labels

Inter-annotator agreement and Cohen's kappa

Raw agreement is misleading. If 90% of items are “safe”, two lazy raters agree 82% of the time by chance alone. Cohen's kappa corrects for chance.

Cohen's kappa for two raters is κ = (p_o − p_e) / (1 − p_e). Here p_o is observed agreement and p_e is the agreement expected by chance from each rater's label rates. A kappa of 1 is perfect. A kappa of 0 is no better than chance.
import numpy as np

def cohens_kappa(a, b):
    """a, b: label arrays from two raters on the same items."""
    a, b = np.asarray(a), np.asarray(b)
    labels = np.union1d(a, b)
    p_o = np.mean(a == b)
    # chance agreement: sum over classes of P(rater A says k) * P(rater B says k)
    p_e = sum(np.mean(a == k) * np.mean(b == k) for k in labels)
    return (p_o - p_e) / (1 - p_e)

a = [1, 1, 0, 1, 0, 0, 1, 1, 0, 1]
b = [1, 0, 0, 1, 0, 1, 1, 1, 0, 1]
print(round(cohens_kappa(a, b), 3))   # 0.583

Writing labeling guidelines

Guidelines are a product spec for raters. Treat them like code. Version them, test them, and fix bugs.

Active learning

Active learning lets the model pick which unlabeled items to send for labeling next. The goal is the most model gain per labeling dollar.
Active learning biases your data. The labeled set is no longer a random sample. Never use it as your test set. Keep a separate random sample for evaluation.

Why it matters in practice

Fixing labels often beats changing the model. A team that cleans 5% of wrong labels can gain more than a month of tuning. Clear guidelines also make labels stable across vendors and quarters. Without that, metric changes reflect rater drift, not model change.

Interview check

Evaluation done right

Offline evaluation measures a model on held-out data before launch. Done right, it predicts the online result. Done wrong, it gives you false confidence and a failed A/B test.

Pick the metric by task. The metric must match how the output gets used.

Ranking metrics

Ranking metrics reward putting good items near the top. They differ in how they treat grades and position.

import numpy as np

def dcg_at_k(rels, k):
    rels = np.asarray(rels, dtype=float)[:k]
    # gain 2^rel - 1 rewards high grades more; log2(rank+1) discount
    return np.sum((2 ** rels - 1) / np.log2(np.arange(2, len(rels) + 2)))

def ndcg_at_k(rels_in_model_order, k):
    ideal = sorted(rels_in_model_order, reverse=True)
    idcg = dcg_at_k(ideal, k)
    return dcg_at_k(rels_in_model_order, k) / idcg if idcg > 0 else 0.0

def mrr(list_of_binary_rels):
    out = []
    for rels in list_of_binary_rels:
        hits = np.flatnonzero(rels)
        out.append(1.0 / (hits[0] + 1) if len(hits) else 0.0)
    return float(np.mean(out))

print(round(ndcg_at_k([3, 2, 0, 1], k=4), 3))   # 0.993
print(mrr([[0, 1, 0], [1, 0, 0], [0, 0, 0]]))   # 0.5

Classification metrics

Regression metrics

Calibration

Calibration means predicted probabilities match real frequencies. Among items scored 0.3, about 30% should be positive.

Slices and fairness metrics

An average hides failures. A model can gain 1% overall and lose 10% for new users. Always report key slices.

Confidence intervals on metrics

A metric from a test set is an estimate. A 0.3 point gain in AUC may be noise. Put an interval on it. The bootstrap is the general tool.

import numpy as np
from sklearn.metrics import roc_auc_score

def paired_bootstrap_delta(y, p_a, p_b, metric=roc_auc_score, n=2000, seed=0):
    """95% CI for metric(B) - metric(A) on the SAME test items."""
    rng = np.random.default_rng(seed)
    y, p_a, p_b = map(np.asarray, (y, p_a, p_b))
    deltas = []
    for _ in range(n):
        idx = rng.integers(0, len(y), len(y))      # resample items with replacement
        if y[idx].min() == y[idx].max():
            continue                               # AUC needs both classes
        deltas.append(metric(y[idx], p_b[idx]) - metric(y[idx], p_a[idx]))
    lo, hi = np.percentile(deltas, [2.5, 97.5])
    return float(np.mean(deltas)), (float(lo), float(hi))

Holdout hygiene

Offline and online can disagree. Offline data comes from the old model's choices. Your new model shows items the old one never showed. So offline gains are a filter, not a verdict. The A/B test is the verdict.

Why it matters in practice

The wrong metric picks the wrong model. A model with higher AUC can lose money if it is miscalibrated in an auction. A gain without an interval may vanish online. Clean holdouts and slice reports are what make launch reviews fast and trusted.

Interview check

Error analysis

Error analysis is the habit of looking at what the model gets wrong and why. It turns a single metric into a ranked list of fixes.

Look at 100 examples

This is the most useful hour in any project. Pull 100 errors at random. Read each one. Write a short note on why it failed. Then count the notes.

  1. Sample errors at random, not the worst ones. You want the true mix.
  2. Look at the raw input, the label, the prediction, and the score.
  3. Tag each error with one or more causes. Make up tags as you go.
  4. Merge similar tags after about 30 items. Finish the rest with the merged list.
  5. Count the tags. The biggest bucket is your next project.

Often 20% of errors turn out to be label errors. That changes your plan. It also means your true accuracy is higher than reported.

Confusion slices

Build a failure taxonomy

A taxonomy is a stable list of failure types. It lets you track progress across model versions. A good one has under ten buckets and covers 90% of errors.

Share of 100 sampled errors by cause Missing signal 32 Label error 22 Rare pattern 17 Ambiguous 13 Pipeline bug 10 Shift 6 Read it as a plan 1. New feature: up to 32% 2. Relabel: 22% are not real model errors 3. Bugs: cheap, fix today
Figure 1. A failure taxonomy turns a vague error rate into a ranked list of projects. The bar length is the ceiling on what each fix can gain.
The ceiling rule. If a bucket holds 10% of errors, fixing it fully cuts errors by at most 10%. Use that to rank projects before you start any of them.

Why it matters in practice

Teams without error analysis guess at fixes. They try a bigger model when the real issue is a truncation bug. A failure taxonomy also makes progress visible to leaders. You can say the “missing signal” bucket shrank from 32 to 12.

Interview check

Reproducibility

Reproducibility means anyone can rerun your experiment and get the same result. It needs the same code, data, config, and environment. Without it, you cannot trust or debug a gain.

The four things to pin

Seeds and determinism

import os, random
import numpy as np
import torch

def set_seed(seed: int = 42):
    random.seed(seed)
    np.random.seed(seed)
    torch.manual_seed(seed)
    torch.cuda.manual_seed_all(seed)
    os.environ["PYTHONHASHSEED"] = str(seed)
    # deterministic kernels: slower, but bit-for-bit repeatable on the same hardware
    torch.backends.cudnn.deterministic = True
    torch.backends.cudnn.benchmark = False
    torch.use_deterministic_algorithms(True, warn_only=True)

Data versioning

Experiment tracking

A tracker logs every run with its inputs and outputs. MLflow and Weights & Biases are common choices. Log these for each run.

Configs done well

# config/ranker_v7.yaml  (read by the training script, logged with the run)
# data:
#   train_table: events.ranker_train
#   snapshot: 2026-09-30
#   label: long_click_30s
# model:
#   type: gbdt
#   num_trees: 800
#   learning_rate: 0.05
# eval:
#   slices: [country, device, new_user]
# seed: 13

from dataclasses import dataclass
import yaml

@dataclass(frozen=True)
class Config:
    data: dict
    model: dict
    eval: dict
    seed: int

def load_config(path: str) -> Config:
    with open(path) as f:
        return Config(**yaml.safe_load(f))

Why it matters in practice

Unrepeatable gains are a common reason launches stall. A reviewer asks you to rerun with one change and you cannot match the old number. Good tracking also saves weeks when a model breaks in production. You can rebuild the last good version in hours.

Interview check

The ML lifecycle and MLOps

MLOps is the set of tools and habits that take a model from notebook to production. It also keeps the model healthy after launch. It covers features, training pipelines, registries, deployment, and monitoring.
Raw datalogs, tables Feature storeone definition Trainingpipeline + CI Registryversioned model Shadowno user impact Canary1% then ramp Online servingsame features, same code Monitoringdrift, skew, metrics logged outcomes become new labels served features
Figure 2. The ML lifecycle. One feature definition feeds both training and serving. Each new model passes the registry, shadow, and canary before full traffic. Monitoring closes the loop.

Feature stores

A feature store keeps one definition per feature and serves it two ways. An offline store gives history for training. An online store gives low-latency reads for serving.

Training-serving skew

Training-serving skew is any difference between what the model saw in training and what it sees in production. It quietly drops online quality while offline metrics look fine.

How to catch it. Log the exact features at serving time. Train on those logs when you can. Compare training and serving feature distributions daily. Alert on any feature whose mean or null rate moves past a threshold.

Model registry

CI for models

Software CI checks that code works. Model CI also checks that the model is good enough to ship.

Shadow and canary deployment

Monitoring after launch

Why it matters in practice

Most production ML incidents are not model bugs. They are stale features, schema changes, or skew. MLOps turns these silent failures into alerts. It also makes it safe to ship often, which compounds gains over a year.

Interview check

Model efficiency

Model efficiency is getting the same quality for less compute, memory, latency, or money. At scale, it decides whether a model ships at all.

Cost per prediction

Start every efficiency talk with this number. It ties model choices to the budget.

def cost_per_1k_predictions(gpu_hour_usd, qps_per_gpu, utilization=0.6):
    """Rough serving cost. utilization covers headroom for peaks."""
    preds_per_hour = qps_per_gpu * 3600 * utilization
    return 1000 * gpu_hour_usd / preds_per_hour

# A big model: 200 QPS per GPU at $2.50/hr
print(round(cost_per_1k_predictions(2.50, 200), 4))   # 0.0058 dollars
# Distilled model: 1500 QPS per GPU
print(round(cost_per_1k_predictions(2.50, 1500), 4))  # 0.0008 dollars

At one billion predictions a day, that gap is about $5,000 a day. A small quality loss may be well worth it. Put both numbers in the launch doc.

Distillation

Quantization

Pruning

Systems tricks

Measure tail latency, like p99, not just the mean. Users and timeouts feel the tail. A batch that waits too long can blow the p99 budget.

Why it matters in practice

Compute is often the largest line in an ML team's budget. A model that wins by 0.5% but costs four times more may be rejected. Showing cost per prediction and a cheaper variant makes your launch easy to approve.

Interview check

Responsible AI

Responsible AI is building models that are fair, private, safe, and understandable enough for their use. It is a set of measurable checks, not a slogan.

Fairness definitions

Let A be a group attribute, Y the true label, and Ŷ the prediction.

Why the definitions conflict

You cannot have them all at once. This is a theorem, not a tooling gap.

Mitigation options

Privacy

Differential privacy intuition

Differential privacy promises the output barely changes if any one person's data is added or removed. So nobody can learn much about one person from the output. The privacy budget ε sets how much it may change. Smaller ε means more privacy and more noise.
import numpy as np

def dp_count(true_count, epsilon, sensitivity=1.0, rng=np.random.default_rng()):
    """Laplace mechanism: one person changes a count by at most 1."""
    return true_count + rng.laplace(0.0, sensitivity / epsilon)

# epsilon=1: noise std about 1.4. epsilon=0.1: about 14.
print(round(dp_count(1200, epsilon=1.0)), round(dp_count(1200, epsilon=0.1)))

Explainability and its limits

Know the limits. Interviewers like this part.

Why it matters in practice

Fairness and privacy failures cause legal risk, press, and lost trust. Regulators now ask for documented checks in many domains. A scientist who can name the right fairness metric and its trade-off earns trust in reviews. Knowing explainability limits stops teams from over-claiming in front of regulators.

Interview check

Recommender and ranking essentials

A recommender picks and orders items for a user. Most learn from implicit feedback, like clicks and watch time, which is plentiful but biased.

Implicit feedback

Position bias

Users click top items more, whatever their quality. Training on raw clicks teaches the model to copy the old ranking.

Exploration with bandits

If you only show what the model likes, you never learn about new items. This is the explore and exploit trade-off. Multi-armed bandits handle it.

import numpy as np

class ThompsonBernoulli:
    """Beta-Bernoulli Thompson sampling for click/no-click arms."""
    def __init__(self, n_arms, seed=0):
        self.a = np.ones(n_arms)   # successes + 1
        self.b = np.ones(n_arms)   # failures + 1
        self.rng = np.random.default_rng(seed)

    def choose(self):
        return int(np.argmax(self.rng.beta(self.a, self.b)))

    def update(self, arm, reward):
        self.a[arm] += reward
        self.b[arm] += 1 - reward

def ucb1_choose(counts, sums, t):
    """counts, sums: arrays per arm. t: total pulls so far."""
    if (counts == 0).any():
        return int(np.argmin(counts))           # try every arm once
    means = sums / counts
    bonus = np.sqrt(2 * np.log(t) / counts)
    return int(np.argmax(means + bonus))

def eps_greedy_choose(means, eps, rng):
    return int(rng.integers(len(means))) if rng.random() < eps else int(np.argmax(means))
Thompson sampling: wide beliefs get explored, narrow ones get exploited 0% click rate New item: 3 clicks in 10 views Old item: 2,000 in 10,000 This round's samples: new item draws high, wins Next round it may lose. Its share tracks its chance to be best.
Figure 3. Each arm has a Beta belief. The new item's wide belief sometimes draws above the old item's narrow one. That gives it traffic in proportion to how likely it is to be the best.

Offline evaluation of policies

Other must-knows

Why it matters in practice

Ignoring position bias makes a ranker learn its own past. Ignoring exploration starves new items and creators. Logging propensities is cheap now and expensive to add later. These details often separate a flat A/B test from a winning one.

Interview check

Time series essentials

Time series data has an order in time, and the order matters. Nearby points are related. The future must never leak into the past.

Stationarity

Leakage traps

Backtesting

Backtesting simulates forecasting in the past, many times. It is cross-validation for time series.

import numpy as np

def rolling_origin_splits(n, initial, horizon, step):
    """Yield (train_idx, test_idx). Train always ends before test starts."""
    end = initial
    while end + horizon <= n:
        yield np.arange(0, end), np.arange(end, end + horizon)
        end += step

def backtest(y, fit_predict, initial=365, horizon=28, step=28):
    errs = []
    for tr, te in rolling_origin_splits(len(y), initial, horizon, step):
        pred = fit_predict(y[tr], horizon)
        errs.append(np.sum(np.abs(y[te] - pred)) / np.sum(np.abs(y[te])))  # WAPE
    return float(np.mean(errs)), float(np.std(errs))

def seasonal_naive(train, horizon, season=7):
    last = train[-season:]
    return np.resize(last, horizon)   # repeat last week

Forecasting baselines

These are hard to beat. Always report them.

For many related series, like sales per store, one global model often beats one model per series. It shares strength across series and handles new ones better.

Why it matters in practice

Leaky time series models look brilliant offline and fail on day one. Forecasts drive inventory, staffing, and capacity, so errors cost real money. A backtest with honest baselines is what makes a forecast credible to finance and ops partners.

Interview check

Writing and communication

Scientific communication is turning your work into a decision for someone else. A result that nobody understands does not ship. Writing is part of the job, not an extra.

The one-page doc

Leaders read the top and decide. Put the answer first. Use this order.

  1. The ask or the answer. One or two sentences. “We should launch ranker v7 to 100%.”
  2. Why it matters. The business problem and the size of the prize.
  3. What we did. The approach in plain words. No equations unless needed.
  4. Results. The key numbers with intervals. One chart at most.
  5. Risks and trade-offs. Cost, guardrail moves, and known failure modes.
  6. Next steps. Who does what, by when.

The experiment write-up

Write up failed experiments too. A clear null result saves the next team a quarter. It also shows judgment, which matters as much as wins for promotion.

Telling the impact story with numbers

Example impact line. “I replaced the rule-based fraud filter with a gradient boosted model. At the same review budget, it caught 18% more fraud. That saved about $3M a year. False positives stayed flat.”

Habits that help

Why it matters in practice

At senior levels, influence comes through writing. Your doc reaches rooms you are not in. A clear one-pager can win headcount or kill a bad project. Interviewers also judge the research talk and behavioral rounds on how clearly you tell impact.

Interview check

Reading papers and staying current

Staying current means knowing which new ideas matter for your problems, without reading everything. It is a filtering skill, not a volume skill.

The three-pass read

  1. Pass one, five minutes. Title, abstract, figures, and conclusion. Decide if it is worth more time.
  2. Pass two, one hour. Read the method and main results. Skip proofs. Note the baselines and datasets.
  3. Pass three, a few hours. Only for papers you will build on. Rederive key steps. Run the code if it exists.

Questions to ask every paper

Benchmark gains rarely transfer one to one. A 2% gain on an academic set can vanish on production data. Treat papers as a list of ideas to test, not as promises.

An efficient system

Turning papers into impact

Why it matters in practice

The field moves fast. A scientist who ignores it falls behind. A scientist who chases every paper ships nothing. The skill is picking the few ideas that fit your failure modes and testing them cheaply.

Interview check

Recap

← T5 — Linear Algebra and Matrix Calculus M1 — Being a Successful Applied Science Manager →