Part II Bank 5 Design

ML System Design

One framework, nine designs. Frame the problem and the metric first. The model comes later than you think.

The design round sets your level more than any other round. The interviewer gives you a vague product goal. You have 45 minutes to turn it into a system that could ship. This page gives you a seven-step framework. Then it walks through the nine designs that come up most. Each design follows the same seven steps, so you can practice the habit, not just the answers.

Contents

  1. The seven-step framework
  2. News feed ranking
  3. Video recommendations
  4. Ads click prediction
  5. Fraud and spam detection
  6. Search ranking
  7. Harmful content moderation
  8. People You May Know
  9. Monitoring, drift and retraining
  10. An LLM assistant with RAG

The seven-step framework

Every ML design answer has the same bones. Interviewers carry a rubric that maps to these steps. If you skip one, you lose its points, no matter how good the rest is. Say the steps out loud at the start. Then the interviewer knows where you are going and can steer you.

1. The seven-step framework Core

Step 1. Clarify the goal and constraints

Step 2. Frame it as an ML task and pick metrics

The three metric layers must line up. Offline gains that do not move the online metric are a red flag. Say how you will check that offline and online agree.

Step 3. Data and labels

Step 4. Features

Step 5. Model

Step 6. Serving and scale

Step 7. Evaluate and iterate

The 45-minute time budget

The most common failure: spending 25 minutes on the model. The interviewer then has no time to test serving, metrics or experiments. Those are where senior points live.
Say this out loud: “I will go in seven steps. Clarify, frame and metrics, data, features, model, serving, then evaluation. Stop me if you want to go deeper on any one.”

The designs

Each design below follows the seven steps. The numbers are realistic for a large consumer app. Use them as anchors, but always ask the interviewer for their numbers first.

2. Design news feed ranking Hard

Clarify

Frame and metrics

Value model. A weighted sum of predicted action rates. The weights are a product decision, not a learned parameter. Comments may weigh 5 to 15 times a like, because they signal more value and create more content.

Data and labels

Features

Model

Inventory ~2,000 posts Light ranker keep ~500 Multi-task ranker P(like), P(comment) P(share), P(hide) Value model weighted sum Re-rank diversity, policy Feature store ~20 ms ~80 ms ~10 ms
Feed ranking funnel. A light ranker trims the pool. A multi-task model predicts each action. A value model turns them into one score.

Serving

Evaluation and iteration

Follow-ups they will ask

What separates senior answers

Say this out loud: “I will predict several actions with one multi-task model. Then I combine them with a value model whose weights the product team controls. Hides and reports get negative weight.”

3. Design video recommendations Hard

Clarify

Frame and metrics

Watch time alone is a trap. It rewards long videos and clickbait. Weight it with completion rate and satisfaction surveys.

Data and labels

Features

Model

User tower Item tower ANN index 1B vectors Retrieval ~2,000 merged Ranker score ~500 Re-rank diversity, explore built offline query vector ~10 ms ~60 ms top 20 shown
Two-tower retrieval feeds a heavy ranker, then a re-ranker. Item vectors are built offline. The user vector is built per request.

Serving

Evaluation and iteration

Follow-ups they will ask

What separates senior answers

Say this out loud: “I will use a two-tower model with an ANN index to pull about a thousand videos fast. A heavier ranker with cross features orders them. A re-ranker adds diversity and a few exploration slots.”

4. Design ads click prediction Hard

Clarify

Frame and metrics

Normalized cross entropy. Log loss divided by the log loss of always predicting the average CTR. Lower is better. A value of 0.80 means 20% better than the naive guess.
Calibration ratio. Sum of predicted clicks over sum of real clicks. Aim for 1.00 overall and in every slice, like each country and placement.

Data and labels

Features

Model

Serving

Evaluation and iteration

Follow-ups they will ask

What separates senior answers

Say this out loud: “The auction multiplies bid by P(click), so I need calibrated probabilities, not just a good ranking. I will track normalized cross entropy and calibration per slice, and retrain every few hours.”

5. Design fraud and spam detection Hard

Clarify

Frame and metrics

Accuracy and ROC AUC lie under extreme imbalance. If fraud is 0.1%, a model that says “never fraud” is 99.9% accurate. Report precision at the operating point and recall at a low FPR.

Data and labels

Features

Model

Event Rules GBDT model Graph score Decision layer Allow Challenge / review Block New labels
Rules and models run side by side. A decision layer maps scores to actions. Review outcomes become new labels.

Serving

Evaluation and iteration

Follow-ups they will ask

What separates senior answers

Say this out loud: “Fraud is rare and adversarial. I will combine fast rules with a tree model and graph features. A decision layer maps the score to allow, challenge or block. I will track recall at a fixed low false positive rate.”

6. Design search ranking Hard

Clarify

Frame and metrics

NDCG@10. Gain from each result, discounted by log of its rank, summed over the top 10. Then divided by the best possible score. 1.0 means a perfect order.

Data and labels

Features

Model

Serving

Evaluation and iteration

Follow-ups they will ask

What separates senior answers

Say this out loud: “I will clean and classify the query first. Retrieval is hybrid, BM25 plus dense vectors. Then a LambdaMART ranker and a cross-encoder on the top 50. I will train on debiased clicks and judge on NDCG@10.”

7. Design harmful content moderation Hard

Clarify

Frame and metrics

Severity sets the trade-off. For child safety or terror content, favor recall and accept more review load. For borderline speech, favor precision, because false removals hurt trust and free expression.

Data and labels

Features

Model

Serving

Evaluation and iteration

Follow-ups they will ask

What separates senior answers

Say this out loud: “I will run hash matching, then a multimodal classifier with one head per policy. Each policy gets its own thresholds by severity. The middle band goes to human review, ranked by expected harm. I will track prevalence from a random sample.”

8. Design People You May Know Hard

Clarify

Frame and metrics

Data and labels

Features

Adamic-Adar. Sum over mutual friends of 1 / log(their friend count). A mutual friend with 50 friends counts more than one with 5,000.

Model

Serving

Evaluation and iteration

Follow-ups they will ask

What separates senior answers

Say this out loud: “Most candidates come from friends of friends, capped to a couple of thousand. A ranker predicts that the request is sent and accepted. I will filter for privacy and safety first, and use cluster randomization in the test.”

9. Design monitoring, drift detection and retraining Medium

Clarify

Frame and metrics

Data and labels

Features

PSI formula. Sum over bins of (live% − train%) × ln(live% / train%). It is a symmetric form of KL divergence on binned data.

Model

Serving

Evaluation and iteration

Follow-ups they will ask

What separates senior answers

Say this out loud: “I will monitor three layers. Data quality, input and prediction drift with PSI, and online model metrics when labels arrive. New models go through shadow, then canary, then ramp, with auto rollback. Retrain cadence comes from a measured decay curve.”

10. Design an LLM assistant with RAG Hard

Clarify

Frame and metrics

Groundedness. The share of claims in the answer that the retrieved text supports. Also called faithfulness. An answer can be true but ungrounded, and that is still a risk.

Data and labels

Features

Model

Documents Chunk + embed offline Vector + BM25 index Question + rewrite Hybrid retrieve top 100 Reranker keep top 5-8 LLM prompt + chunks Guardrails grounding, citations ~50 ms ~100 ms ~1-3 s ~100 ms
RAG pipeline. Documents are chunked and indexed offline. Each question is rewritten, retrieved, reranked, answered, then checked.

Serving

Evaluation and iteration

Follow-ups they will ask

What separates senior answers

Say this out loud: “I will build a gold eval set first. Retrieval is hybrid with a cross-encoder reranker. The prompt forces citations and allows ‘I don’t know’. A groundedness check runs before the answer ships. I will measure retrieval recall and answer groundedness apart.”

Recap

← 4 — ML Coding from Scratch 6 — Experimentation and Causal Inference →