One framework, nine designs. Frame the problem and the metric first. The model comes later than you think.
The design round sets your level more than any other round. The interviewer gives you a vague product goal. You have 45 minutes to turn it into a system that could ship. This page gives you a seven-step framework. Then it walks through the nine designs that come up most. Each design follows the same seven steps, so you can practice the habit, not just the answers.
Every ML design answer has the same bones. Interviewers carry a rubric that maps to these steps. If you skip one, you lose its points, no matter how good the rest is. Say the steps out loud at the start. Then the interviewer knows where you are going and can steer you.
1. The seven-step framework Core
Step 1. Clarify the goal and constraints
Ask what the business wants. More time spent? More revenue? Fewer bad actors? The answer picks your metric.
Ask about scale. How many users, items and requests per second? 1,000 QPS and 1,000,000 QPS lead to very different designs.
Ask about latency. A feed must load in about 200 ms end to end. A nightly batch job can take hours.
Ask what exists today. Is there a heuristic or an old model? That is your baseline and your source of labels.
Write the assumptions down. Say them out loud. “I will assume 500M daily users and a 300 ms budget.”
Step 2. Frame it as an ML task and pick metrics
Name the ML task. Binary classification, ranking, regression, retrieval, or generation. Say what the input and the output are.
Business metric. The thing the company cares about. Revenue, daily active users, fraud loss in dollars.
Online metric. What you measure in an A/B test. Click-through rate, watch time, report rate.
Offline metric. What you measure on held-out data before launch. AUC, log loss, NDCG@10, recall@K.
Guardrail metrics. Things that must not get worse. Latency, crash rate, creator diversity, user reports.
The three metric layers must line up. Offline gains that do not move the online metric are a red flag. Say how you will check that offline and online agree.
Step 3. Data and labels
Where do labels come from? Implicit feedback like clicks, explicit feedback like ratings, or human raters.
Name the bias in each source. Clicks have position bias. Ratings are sparse. Raters are slow and costly.
Define negatives with care. An item shown and skipped is a real negative. An item never shown is not.
Split by time, not at random. Train on days 1 to 28 and test on day 29. A random split leaks the future.
Handle imbalance. Downsample negatives, then correct the predicted probability afterwards.
Step 4. Features
User features. Age bucket, country, device, long-term interests, recent actions.
Item features. Type, age, creator, text and image embeddings, past engagement rates.
Context features. Time of day, network speed, surface, session depth.
Cross features. User by item signals. How often has this user engaged with this creator?
Freshness matters. Real-time counters, like clicks in the last hour, are often the strongest features.
Step 5. Model
Start with a baseline. Popularity, a rule, or logistic regression. It gives you a floor and a sanity check.
Then upgrade. Gradient boosted trees for tabular data. Deep models for sparse IDs, text and images.
Use stages when the item pool is large. You cannot score a billion items with a heavy model.
Retrieval. Cheap models pull thousands of candidates from millions or billions.
Ranking and re-ranking. A heavy model scores hundreds. Then business rules fix diversity and policy.
Step 6. Serving and scale
Split the latency budget. Give each stage a number in milliseconds. They must add up.
Cache what you can. Item embeddings and slow user features can be precomputed.
Use a feature store. One place serves the same feature values to training and serving.
Watch for training-serving skew. If features are computed by different code online and offline, the model sees different inputs. Log features at serving time and train on those logs.
Plan for failure. If the ranker times out, fall back to a cached or simpler result.
Step 7. Evaluate and iterate
Offline first. Compare to the baseline on a time-based holdout. Slice by country, new users and device.
Then an A/B test. Pick the unit, the length and the minimum effect you care about. Run at least one full week.
Monitor after launch. Track feature drift, prediction drift, latency and the online metric.
Retrain on a schedule. Daily or hourly for fast-moving products. Name the trigger for an early retrain.
Close the loop. Say what you would try next and why.
The 45-minute time budget
0 to 5 min. Clarify the goal, scale, latency and what exists.
5 to 10 min. Frame the ML task. Name the business, online and offline metrics.
10 to 17 min. Data, labels and features.
17 to 27 min. Model. Baseline, then the real design with stages.
27 to 35 min. Serving, latency budget, feature store and skew.
35 to 42 min. Evaluation, A/B test, monitoring and retraining.
42 to 45 min. Sum up, name the risks, take follow-ups.
The most common failure: spending 25 minutes on the model. The interviewer then has no time to test serving, metrics or experiments. Those are where senior points live.
Say this out loud: “I will go in seven steps. Clarify, frame and metrics, data, features, model, serving, then evaluation. Stop me if you want to go deeper on any one.”
The designs
Each design below follows the seven steps. The numbers are realistic for a large consumer app. Use them as anchors, but always ask the interviewer for their numbers first.
2. Design news feed ranking Hard
Clarify
Goal. Show each user the posts they find most valuable, so they come back. Not just the posts they click.
Scale. 1B daily users. About 100K feed loads per second at peak.
Inventory. Posts from friends, groups and pages the user follows. Maybe 1,000 to 5,000 new eligible posts per user per day.
Latency. About 200 ms server time for the whole ranking call.
Ask. Do we include unconnected content, like suggested posts? That changes retrieval a lot.
Frame and metrics
ML task. Multi-task prediction. For each (user, post) pair, predict the chance of several actions. Then combine them into one score.
The actions. Like, comment, share, click, dwell over 10 s, hide, and report.
Business metric. Daily active users and sessions per user over weeks.
Online metrics. Meaningful interactions per user, time spent, hide rate, report rate.
Offline metrics. Per-task AUC and normalized log loss. Also NDCG on the final ranked list.
Value model. A weighted sum of predicted action rates. The weights are a product decision, not a learned parameter. Comments may weigh 5 to 15 times a like, because they signal more value and create more content.
Data and labels
Training rows. Every impression is a row. Labels are the actions that followed within a window, like 24 hours.
Volume. Tens of billions of impressions a day. Downsample negatives for the common tasks.
Rare tasks. Hide and report happen on well under 1% of impressions. Keep all positives and weight them.
Position bias. Top slots get more actions just because they are on top. Add position as a training feature, then set it to a fixed value at serving.
Time split. Train on the last 2 to 4 weeks. Evaluate on the next day.
Features
Viewer. Country, device, how active they are, topics they engage with.
Author. Page or friend, follower count, past post quality.
Viewer-author affinity. Past likes, comments and profile visits between them. Often the top feature group.
Post. Type (photo, video, link, text), age, text and image embeddings, early engagement counts.
Real-time counters. Likes and comments in the last 15 minutes. Served from a streaming store.
Model
Baseline. Reverse time order. Then logistic regression on P(like).
Main model. One shared network with one head per task. Sparse ID embeddings feed a shared trunk, then task towers.
Why multi-task. Rare tasks like share borrow signal from common tasks like click. One model is cheaper to serve than six.
Task conflict. If tasks fight, use MMoE (multi-gate mixture of experts) so each task picks its own mix of experts.
Stages. Candidate gathering pulls about 2,000 posts. A light model cuts to about 500. The heavy multi-task model scores those 500. A re-ranker applies diversity and policy rules.
Feed ranking funnel. A light ranker trims the pool. A multi-task model predicts each action. A value model turns them into one score.
Serving
Budget. 30 ms to gather candidates, 20 ms for the light ranker, 80 ms for the heavy ranker, 10 ms to re-rank. The rest is network and feature fetch.
Batching. Score all 500 posts in one batched call on GPU or a tuned CPU fleet.
Caching. Cache post embeddings when the post is created. Cache viewer embeddings for a few minutes.
Skew. Log the exact features used at serving time. Train on those logs, not on a nightly recompute.
Fallback. If the heavy ranker times out, serve the light ranker order.
Evaluation and iteration
Offline. Per-head AUC and calibration. Check that predicted P(comment) matches the real rate in each bucket.
A/B test. Randomize by user. Run 2 weeks to catch novelty effects. Watch sessions, interactions and hide rate.
Tuning weights. Run a small grid of weight sets as parallel test arms. Pick the one that moves long-term retention best.
Long-term holdout. Keep 1% of users on the old system for months. This catches slow harms like clickbait fatigue.
Follow-ups they will ask
“How do you pick the weights?” Start from how each action predicts long-term retention. Then tune with online tests. Weights change as product goals change.
“How do you stop clickbait?” Put negative weight on hide and report. Add a quality classifier. Value dwell time over clicks.
“A new post has no engagement. What then?” Use content embeddings and author history. Add a small exploration bonus for fresh posts.
“Why not one model for the final score?” Separate heads let product change weights without retraining. Each head is also easier to debug.
What separates senior answers
They name the value model and say the weights are a product lever.
They raise negative feedback signals without being asked.
They talk about long-term holdouts and the gap between short-term clicks and retention.
They mention calibration, because a weighted sum of badly calibrated probabilities is meaningless.
Say this out loud: “I will predict several actions with one multi-task model. Then I combine them with a value model whose weights the product team controls. Hides and reports get negative weight.”
3. Design video recommendations Hard
Clarify
Goal. Recommend videos on the home page and the up-next slot. Grow satisfied watch time.
Scale. 500M daily users. 1B videos in the corpus. About 50K requests per second.
Latency. 150 ms server time.
Ask. Short-form or long-form? Short-form gets feedback every few seconds, so the model can learn fast within a session.
Frame and metrics
ML task. Retrieval plus ranking. Retrieval finds a few thousand relevant videos. Ranking orders them.
Ranking target. Expected watch time, plus P(like), P(share) and P(not interested).
Business metric. Daily active users and total satisfied watch time.
Online metrics. Watch time per user, completion rate, survey satisfaction, “not interested” rate.
Offline metrics. Recall@K for retrieval, like recall@500. AUC and NDCG for ranking.
Watch time alone is a trap. It rewards long videos and clickbait. Weight it with completion rate and satisfaction surveys.
Data and labels
Positives. Watched more than 50% or more than 30 s, liked, shared, or added to a playlist.
Negatives. Shown and skipped within 2 s, or marked “not interested”.
Retrieval negatives. Use in-batch negatives. Other users’ positives in the same batch act as negatives for you.
Popularity correction. In-batch negatives oversample popular videos. Subtract log of item frequency from the logit, called logQ correction.
Features
User tower. Watch history as a sequence of video IDs, search queries, country, language, device.
Item tower. Video ID, creator ID, title and thumbnail embeddings, length, upload age, language.
Ranker extras. User-creator affinity, the video’s recent completion rate, time of day, and the previous video watched.
Model
Baseline. Trending in the user’s country. Then item-to-item co-watch, “people who watched X watched Y”.
Two-tower retrieval. One tower embeds the user. One tower embeds the video. Score is their dot product. Train with sampled softmax.
ANN index. Precompute all 1B video vectors. Use approximate nearest neighbor search, like HNSW or IVF-PQ, to find the top 1,000 in a few ms.
Several retrieval sources. Two-tower, co-watch, subscriptions, and fresh uploads. Merge to about 2,000 candidates.
Ranker. A deep multi-task model scores about 500. It uses cross features the two towers cannot see.
Re-rank. Limit videos per creator. Mix topics. Insert a few exploration slots.
Two-tower retrieval feeds a heavy ranker, then a re-ranker. Item vectors are built offline. The user vector is built per request.
Serving
Budget. 10 ms for the user tower, 10 ms for ANN, 60 ms for the ranker, 10 ms to re-rank.
Index refresh. Rebuild item vectors daily. Add new uploads to the index every few minutes.
Version lock. User and item towers must come from the same training run. Mixed versions give garbage dot products.
Session signals. Stream the last few watches into the user tower input, so recommendations react within a session.
Evaluation and iteration
Offline. Recall@500 on next-day watches for retrieval. NDCG and calibration for the ranker.
A/B test. User-level split for 2 to 4 weeks. Watch time, satisfaction surveys, creator diversity.
Creator side. Track how many creators get their first 1,000 views. A healthy platform needs new supply.
Follow-ups they will ask
“How do you handle cold-start users?” Use country, language and device. Ask for topic picks at signup. Lean on popular and diverse videos, then adapt fast from the first few watches.
“How do you handle cold-start videos?” Embed the title, thumbnail and audio, so the item tower works with no ID history. Give new uploads a small guaranteed exposure budget.
“How do you explore?” Reserve 5% of slots for epsilon-greedy or Thompson sampling over uncertain items. Log the propensity, so you can debias later.
“Why not just use the ranker on everything?” Scoring 1B videos at 50K QPS is impossible. The two-tower model allows precomputed item vectors and fast ANN.
“What about filter bubbles?” Add topic diversity in re-ranking. Measure the spread of topics each user sees over a month.
What separates senior answers
They explain why two towers cannot use cross features, and why that is the price of ANN.
They mention logQ correction for in-batch negatives.
They treat exploration as a must, not a nice-to-have, and log propensities.
They think about creators and supply, not only viewers.
Say this out loud: “I will use a two-tower model with an ANN index to pull about a thousand videos fast. A heavier ranker with cross features orders them. A re-ranker adds diversity and a few exploration slots.”
4. Design ads click prediction Hard
Clarify
Goal. Predict P(click) for each eligible ad. The auction uses it to pick ads and set prices.
Scale. 1M ad requests per second. Each request has about 1,000 eligible ads after targeting.
Latency. The whole ad call gets about 100 ms. The CTR model gets maybe 20 to 30 ms of that.
Ask. Do we also predict conversions? Many advertisers bid on purchases, not clicks.
Frame and metrics
ML task. Binary classification that outputs a calibrated probability.
Why calibration matters. The auction ranks by bid × P(click). If P(click) is 2× too high, the advertiser pays 2× too much. Ranking alone is not enough.
Business metric. Revenue, plus advertiser return on ad spend.
Online metrics. Revenue per thousand requests, CTR, advertiser cost per click, user ad hide rate.
Offline metrics. Normalized cross entropy (NE) and calibration ratio. AUC is a secondary check.
Normalized cross entropy. Log loss divided by the log loss of always predicting the average CTR. Lower is better. A value of 0.80 means 20% better than the naive guess.
Calibration ratio. Sum of predicted clicks over sum of real clicks. Aim for 1.00 overall and in every slice, like each country and placement.
Data and labels
Rows. Every ad impression. Label is 1 if clicked within a window.
Imbalance. CTR is often 0.5% to 2%. Downsample negatives at rate w. Then fix each prediction with p′ = p / (p + (1 − p) / w).
Delayed feedback. Conversions can arrive days after the click. If you train too soon, you label future converters as negatives. Use a wait window, or a model that estimates the delay, or importance weights.
Position bias. Top slots get clicks regardless of the ad. Train with position as a feature, then serve with a fixed position. Or train a separate shallow position tower that you drop at serving.
Features
Sparse IDs. User ID, ad ID, advertiser ID, campaign ID, page ID. Billions of values, so use hashed embedding tables.
User. Age bucket, country, device, interests, past ad clicks by category.
Ad. Creative embeddings, format, landing page category, historical CTR.
Context. Placement, time of day, position, app version.
Feature crosses. User country by ad language. User interest by ad category. Crosses carry much of the signal.
Model
Baseline. Logistic regression with hashed feature crosses. Strong, fast and easy to calibrate.
Trees plus LR. An older strong design. Gradient boosted trees make leaf features, then LR learns on them.
DLRM. Embedding tables for sparse IDs, an MLP for dense features, then pairwise dot products between all embeddings. Then a top MLP.
DCN. Deep and Cross Network. Cross layers learn explicit feature crosses of bounded degree. A deep branch learns the rest.
Stages. A light model narrows 1,000 ads to about 100. The heavy model scores 100. The auction picks the winners.
Calibration layer. Add isotonic regression or Platt scaling per slice after training. Recheck daily.
Serving
Embedding tables. Can reach terabytes. Shard them across many hosts. Cache hot IDs in memory.
Online training. Ads change fast. Update the model every few hours, or train continuously on a stream.
Freshness check. A model one day stale can lose 1% or more in NE. Measure staleness cost and set the cadence by it.
Skew. Log features at request time. Join the click later by impression ID.
Evaluation and iteration
Offline. NE on the next day of data. Calibration per country, placement and advertiser size.
A/B test. Split by user. Advertisers share budgets across arms, so watch for budget leakage between arms.
Guardrails. Advertiser cost per conversion, user ad hides, ad load.
Follow-ups they will ask
“AUC went up but revenue fell. Why?” Most likely calibration broke. AUC ignores the scale of the scores. The auction does not.
“How do you handle a brand-new ad?” Back off to advertiser and campaign averages. Use creative embeddings. Give a small exploration budget.
“Why not just use the raw downsampled output?” It is biased upward by the sample rate. The auction would overcharge everyone.
“How do you deal with conversion delay?” Wait a fixed window for most labels. Or model the delay and weight early negatives as partial negatives.
What separates senior answers
They lead with calibration and say why the auction needs it.
They use NE, not only AUC.
They explain negative downsampling and the correction formula.
They raise delayed feedback for conversions and training freshness.
Say this out loud: “The auction multiplies bid by P(click), so I need calibrated probabilities, not just a good ranking. I will track normalized cross entropy and calibration per slice, and retrain every few hours.”
5. Design fraud and spam detection Hard
Clarify
Goal. Catch fake accounts, spam posts or fraudulent payments. Block them before they cause harm.
Scale. 10K payments per second, or 50K new posts per second.
Latency. Payments need a decision in under 100 ms. Account checks can run in batch every hour.
Cost of errors. A missed fraud costs money. A false block angers a real user. Ask which hurts more.
Frame and metrics
ML task. Binary classification with a risk score. Then a policy maps the score to an action: allow, challenge, review or block.
Business metric. Fraud loss in dollars. Spam prevalence, the share of views that hit spam.
Offline metrics. Precision-recall AUC, and recall at a fixed low false positive rate, like recall at 0.1% FPR.
Accuracy and ROC AUC lie under extreme imbalance. If fraud is 0.1%, a model that says “never fraud” is 99.9% accurate. Report precision at the operating point and recall at a low FPR.
Data and labels
Label sources. Chargebacks, user reports, manual review, and accounts later banned.
Label delay. A chargeback can arrive 30 to 90 days later. Recent data looks cleaner than it is. Train on data old enough to be mature, and add fast proxy labels.
Selection bias. You only see labels for what you let through. Blocked items have no outcome. Let a tiny random slice through, or send it to review, to get unbiased labels.
Imbalance. Positives can be 1 in 1,000 or rarer. Downsample negatives, use class weights, or focal loss.
Features
Velocity. Actions in the last minute, hour and day. Fraud often comes in bursts.
Account. Age, verified email or phone, profile completeness, past flags.
Device and network. Device fingerprint, IP reputation, how many accounts share this device.
Graph features. Shared devices, cards or IPs link accounts. Rings show up as dense clusters. Use connected components, PageRank on a risk graph, or a GNN.
Content. Text and URL embeddings for spam. Known bad link domains.
Model
Rules first. Hard rules catch known patterns at once. They are clear and fast to change. Keep them.
Baseline model. Gradient boosted trees on tabular features. Strong for this data.
Graph model. A GNN or label propagation over the account graph. It catches rings that look fine one by one.
Anomaly detection. Isolation forest or an autoencoder flags new patterns with no labels yet.
Combine. Rules and model scores feed a decision layer. High score blocks. Middle score goes to review or a challenge like 2FA.
Rules and models run side by side. A decision layer maps scores to actions. Review outcomes become new labels.
Serving
Real-time path. Velocity counters in a low-latency store like Redis. Trees score in about 5 ms.
Batch path. Graph features and GNN scores refresh every hour or day. Serve them from a cache.
Fail mode. Decide in advance. For payments, fail closed above some amount and fail open below it.
Secrecy. Do not tell attackers which feature caught them. Vague errors slow their learning.
Evaluation and iteration
Offline. Recall at 0.1% FPR on a mature, time-split holdout. Report dollars caught, not just counts.
Online. A/B tests are hard because attackers adapt. Use a small random holdout and shadow scoring.
Adversarial drift. Attackers probe and change tactics within days. Retrain weekly or faster. Watch for sudden drops in score distribution.
Red team. Run planned attacks to test the system.
Follow-ups they will ask
“How do you set the block threshold?” Pick the FPR the business accepts, like 1 false block per 1,000 good users. Use cost per error to pick the point. Use different thresholds per action severity.
“How do you catch a new attack with no labels?” Anomaly detection and graph clusters. Send the top anomalies to analysts. Their labels seed the next model.
“Why keep rules if you have a model?” Rules react in minutes to a live attack. Retraining takes hours or days. Rules are also easy to explain to regulators.
“How do you measure recall if you never see missed fraud?” Sample random traffic for review. Track chargebacks on allowed traffic.
What separates senior answers
They frame the output as an action policy with tiers, not one yes or no.
They raise label delay and selection bias on their own.
They treat the attacker as an opponent who adapts.
They pick the metric as recall at a fixed low FPR, tied to dollar cost.
Say this out loud: “Fraud is rare and adversarial. I will combine fast rules with a tree model and graph features. A decision layer maps the score to allow, challenge or block. I will track recall at a fixed low false positive rate.”
6. Design search ranking Hard
Clarify
Goal. Return the most relevant results for a text query. Products, documents or posts.
Scale. 100M documents. 20K queries per second.
Latency. 200 ms end to end. Users notice anything slower.
Ask. Is this web search, product search or in-app search? Product search adds price, stock and purchase signals.
Frame and metrics
ML task. Learning to rank. Given a query and candidate documents, output an ordered list.
Business metric. Successful searches per user. For shopping, revenue from search.
Online metrics. Click-through rate, time to first click, reformulation rate, abandonment rate.
Offline metrics. NDCG@10 on human-rated sets. MRR for navigational queries. Recall@1000 for retrieval.
NDCG@10. Gain from each result, discounted by log of its rank, summed over the top 10. Then divided by the best possible score. 1.0 means a perfect order.
Data and labels
Human ratings. Raters grade query-document pairs on a 0 to 4 scale. Clean but costly. Use for evaluation and some training.
Click logs. Huge and cheap, but biased. Top results get clicked because they are on top.
Click models. Models like the cascade model or DBN separate “examined” from “relevant”. They turn raw clicks into better relevance labels.
Unbiased learning to rank. Estimate the examine chance per position with small result swaps. Then weight clicks by inverse propensity.
Good signals. A long click with dwell over 30 s is a strong positive. A quick bounce back is a negative.
Features
Query understanding. Spell fix, query expansion with synonyms, intent class (navigational, info, shopping), entity tags.
Text match. BM25 score per field, like title and body. Exact phrase match.
Semantic match. Cosine between query and document embeddings.
Document quality. PageRank or popularity, freshness, spam score.
Behavior. Past CTR for this query-document pair. Personal history for logged-in users.
Model
Retrieval, hybrid. BM25 over an inverted index for exact terms. A dense bi-encoder with ANN for meaning. Merge to about 1,000 candidates.
First ranker. GBDT with LambdaMART on about 100 features. Cuts to about 100.
Second ranker. A cross-encoder reads query and document together. Most accurate but slow, so only the top 50 to 100.
Pointwise. Predict a relevance score per document. Simple, but ignores order.
Pairwise. Learn which of two documents is better, like RankNet. Fits ranking better.
Listwise. Optimize the whole list, like LambdaMART or ListNet. LambdaMART weights each pair by its NDCG change. It is the standard strong choice.
Serving
Budget. 10 ms query understanding, 30 ms retrieval, 30 ms GBDT, 60 ms cross-encoder, the rest for network and snippets.
Caching. Head queries repeat a lot. Cache results for the top 10% of queries for a few minutes.
Index updates. Add new docs to the inverted index in near real time. Rebuild dense vectors in batch.
Distill. If the cross-encoder is too slow, distill it into a smaller model.
Evaluation and iteration
Offline. NDCG@10 on a fixed rated set. Slice by head, torso and tail queries.
Interleaving. Mix results from two rankers in one list. See which ranker’s results get clicked. It needs 10 to 100 times less traffic than an A/B test.
A/B test. Confirm the winner on CTR, reformulation and abandonment.
Follow-ups they will ask
“Why keep BM25 if you have embeddings?” Dense models miss exact matches on rare terms, part numbers and names. BM25 nails those. Hybrid beats either alone.
“How do you handle tail queries?” They have no click history. Rely on text and semantic match. Query rewriting helps.
“Pointwise or listwise?” Listwise usually wins on NDCG. Pointwise is fine for a first stage or when you need calibrated scores.
“How do you fix position bias in clicks?” Use a click model or inverse propensity weights from small randomized swaps.
What separates senior answers
They put query understanding first, before any ranking.
They use hybrid retrieval and explain why each half is needed.
They know clicks are biased and name a click model or IPS fix.
They suggest interleaving for faster ranker comparisons.
Say this out loud: “I will clean and classify the query first. Retrieval is hybrid, BM25 plus dense vectors. Then a LambdaMART ranker and a cross-encoder on the top 50. I will train on debiased clicks and judge on NDCG@10.”
7. Design harmful content moderation Hard
Clarify
Goal. Find and act on posts that break policy. Hate speech, violence, nudity, self-harm, scams.
Scale. 500M new posts a day, about 6K per second. Text, images, video and links.
Latency. Check before or just after posting, within seconds. Deeper checks can run later.
Ask. Which policies, and in which languages? Each policy can need its own model and threshold.
Frame and metrics
ML task. Multi-label classification. One score per policy. A post can break several at once.
Actions. Remove, reduce reach, add a warning screen, send to human review, or allow.
Business metric. Prevalence. The share of views that land on violating content.
Online metrics. Proactive rate (caught before any report), appeal overturn rate, reviewer agreement.
Offline metrics. Precision and recall per policy and per severity.
Severity sets the trade-off. For child safety or terror content, favor recall and accept more review load. For borderline speech, favor precision, because false removals hurt trust and free expression.
Data and labels
Labels. Human reviewers apply a written policy. Track agreement between reviewers. Low agreement means a vague policy, not a bad model.
Sampling. Violations are rare. Use active learning to pick posts near the decision boundary for labeling.
Prevalence sample. Label a random sample of views each week. This gives an unbiased prevalence estimate.
Languages. Low-resource languages have few labels. Use multilingual encoders and translate data.
Reviewer welfare. Blur images by default and limit exposure time. This is a real design need.
Features
Text. Multilingual transformer embeddings of the post, comments and OCR text from images.
Image and video. Vision encoder embeddings. Sample key frames from video. Audio transcripts.
Hash match. Perceptual hashes like PDQ match known bad media at once.
Author and context. Past strikes, account age, group type, how fast the post is spreading.
Model
Stage 1. Hash match against known bad content. Exact, fast and very precise.
Stage 2. A multimodal classifier. It fuses text, image and context embeddings with one head per policy.
Stage 3. LLM-based review for hard cases. Give it the policy text and ask for a label with a reason.
Thresholds per policy. Above a high bar, auto-remove. In the middle band, send to humans. Below, allow.
Review queue ranking. Order the queue by expected harm. That is P(violation) × severity × expected views.
Serving
At upload. Hash match and a fast text model in under 100 ms.
Async. Heavy image and video models run within a minute. Rescore when the post goes viral.
Human capacity. Reviewers are the bottleneck. Set the review band width by how many items they can handle per day.
Appeals. Every removal can be appealed. Appeals go to a different reviewer. Overturned cases feed back as labels.
Evaluation and iteration
Offline. Precision and recall per policy, per language and per severity.
Online. Prevalence trends from the weekly random sample. Appeal overturn rate as a precision signal.
Fairness. Check false positive rates across dialects and groups. Some models flag certain dialects as toxic more often.
Policy changes. When policy changes, old labels can become wrong. Relabel a sample and retrain.
Follow-ups they will ask
“How do you pick thresholds?” Per policy, set the auto-remove bar at the score where precision is 95% or more. Set the review band by reviewer capacity.
“Attackers misspell words to evade. What then?” Character-level and subword models, OCR on images, and fast retraining on new evasions.
“How do you measure recall?” You cannot from flagged items alone. Use the random prevalence sample.
“Why not let an LLM do everything?” Cost and latency at 6K posts per second are too high. Use it for the hard middle band only.
What separates senior answers
They use prevalence as the north star and know how to measure it.
They set different precision and recall targets per severity.
They design the human review queue and appeals as part of the system.
They raise fairness across languages and dialects.
Say this out loud: “I will run hash matching, then a multimodal classifier with one head per policy. Each policy gets its own thresholds by severity. The middle band goes to human review, ranked by expected harm. I will track prevalence from a random sample.”
8. Design People You May Know Hard
Clarify
Goal. Suggest people a user knows in real life and will connect with. More good connections make the network more useful.
Scale. 2B users. Average 300 friends each. The graph has hundreds of billions of edges.
Latency. Most work can be precomputed daily. The page itself loads in 100 ms.
Ask. Can we use contact uploads? That is a privacy question as much as a modeling one.
Frame and metrics
ML task. Link prediction. Predict P(request sent and accepted) for a user pair.
Why accepted. Predicting only “request sent” rewards spammy suggestions. The receiver must want it too.
Business metric. New accepted connections per user. Downstream engagement from those connections.
Offline metrics. Recall@K on edges formed in the next week. AUC on accepted versus ignored requests.
Data and labels
Positives. Suggestions that led to an accepted request.
Negatives. Suggestions shown and ignored, removed, or requests declined.
Time split. Build the graph as of day T. Predict edges formed after T. Never let future edges leak into features.
Features
Graph. Mutual friend count, Adamic-Adar score, Jaccard of friend sets.
Shared context. Same school, workplace, city, or groups.
Interaction. Profile views, tags in the same photo, comments on the same posts.
Node embeddings. From node2vec or a GNN. Close vectors mean close in the graph.
Adamic-Adar. Sum over mutual friends of 1 / log(their friend count). A mutual friend with 50 friends counts more than one with 5,000.
Model
Candidate generation. Friends of friends is the main source. It has the highest hit rate. Add shared groups, schools and contacts.
The blowup. 300 friends × 300 friends is 90K friends of friends. Cap it. Take top friends by interaction, then top candidates by mutual count. Keep about 1,000 to 2,000.
Baseline. Rank by mutual friend count. Strong and simple.
Ranker. GBDT or a deep model on all features. Predicts P(accept).
GNN. GraphSAGE learns node embeddings from neighbor features. It works for new users who have few edges but some profile data.
Serving
Batch. Run friends-of-friends and scoring as a daily distributed graph job. Store the top 100 per user.
Real-time refresh. After a user adds a friend, add that friend’s friends as fresh candidates within minutes.
Filters. Remove blocked users, people already declined, and anyone who opted out of suggestions.
Evaluation and iteration
Offline. Recall@20 on next-week edges. Acceptance AUC.
A/B test. Network effects break the usual setup. Two users in different arms can still connect. Use graph cluster randomization to keep friends in the same arm.
Long-term. Track whether new connections keep interacting after 30 days.
Follow-ups they will ask
“What are the privacy risks?” Suggestions can reveal secrets. A therapist’s clients could see each other. Avoid using sensitive signals like location co-visits or health contacts. Let users opt out.
“What about fairness?” Popular users get suggested more and grow faster. Cap how often one person is suggested. Check that new and minority users get fair exposure.
“How do you help a brand-new user?” Contacts, if allowed. School and work fields. A GNN that uses profile features. Ask who they know.
“What about safety?” Do not suggest adults to minors who are strangers. Apply hard filters before ranking.
What separates senior answers
They predict acceptance, not just sending, and care about the receiver.
They size the friends-of-friends blowup and show how to cap it.
They raise network effects in A/B tests and propose cluster randomization.
They treat privacy and minor safety as hard limits, not features.
Say this out loud: “Most candidates come from friends of friends, capped to a couple of thousand. A ranker predicts that the request is sent and accepted. I will filter for privacy and safety first, and use cluster randomization in the test.”
9. Design monitoring, drift detection and retraining Medium
Clarify
Goal. Keep a live model healthy. Catch problems before they hurt users or revenue.
Scope. Ask how many models. A platform for 500 models needs automation. One model can use dashboards.
Ask. How fast do labels arrive? Clicks arrive in seconds. Fraud labels take weeks. That shapes what you can monitor.
Frame and metrics
Task. This is a system, not one model. It detects drift, triggers alerts, and safely ships new model versions.
Business metric. Time to detect and time to fix a model problem. Revenue lost to stale models.
System metrics. Latency p50 and p99, error rate, timeouts, feature null rate.
Model metrics. Online accuracy or NE when labels arrive. Prediction mean and spread when they do not.
Data and labels
Data drift. The input distribution P(x) changes. New country launch, app update, or a broken upstream pipeline.
Concept drift. The link P(y | x) changes. Same inputs, new meaning. A holiday changes what people buy.
Label drift. The base rate P(y) changes. Fraud rate doubles during an attack.
Log everything. Features, predictions, model version and request ID. Join labels later.
Features
PSI. Population Stability Index. Bin the feature, compare training and live shares. Rule of thumb: under 0.1 is stable, 0.1 to 0.25 is a moderate shift, over 0.25 needs action.
KL divergence. Measures how one distribution differs from another. Not symmetric. JS divergence is a symmetric, bounded version.
KS test. For continuous features. But with millions of rows, tiny harmless shifts become “significant”. Use effect size, not p-values.
Data quality checks. Null rate, out-of-range values, schema changes, and row counts. Most incidents are broken pipelines, not real drift.
PSI formula. Sum over bins of (live% − train%) × ln(live% / train%). It is a symmetric form of KL divergence on binned data.
Model
Detection. Compute PSI per feature and on the prediction score every hour. Compare to a rolling baseline.
Alerting. Rank alerts by feature importance. A shift in a top-10 feature matters. A shift in a minor feature often does not.
Retrain cadence. Set by how fast performance decays. Measure it. Train a model, then score it on each later day. Plot the loss growth.
Typical cadence. Ads and feeds: hourly to daily. Fraud: daily to weekly. Slow domains like credit: monthly.
Triggered retrain. Also retrain early when PSI on key features passes a bar or online NE drops past a limit.
Serving
Model registry. Every version is stored with its data snapshot, code version and offline metrics.
Shadow mode. The new model scores live traffic but its output is not used. Compare its predictions and latency to the live model.
Canary. Send 1% of traffic to the new model. Watch errors, latency and key metrics for hours. Then ramp to 5%, 25% and 100%.
Auto rollback. If guardrails break during the canary, roll back with no human in the loop.
Validation gate. Block deploys that fail offline checks. NE must not regress more than a set amount, and calibration must stay near 1.0.
Evaluation and iteration
Postmortems. For each incident, ask why monitoring did not catch it sooner. Add the missing check.
Alert quality. Track the false alarm rate. Noisy alerts get ignored.
Feedback loops. The model shapes the data it learns from. A recommender only sees clicks on what it showed. Keep a small random exploration slice to break the loop.
Follow-ups they will ask
“Labels take 60 days. How do you monitor?” Watch input drift and prediction drift as early warning. Use fast proxy labels. Backfill true metrics when labels land.
“PSI alert fired but metrics look fine. What do you do?” Check if the feature matters to the model. Check for a pipeline bug. Not every drift needs a retrain.
“Shadow or canary?” Shadow first, since it has zero user risk. Canary next, because only real exposure shows the effect on user behavior.
“Why not retrain every hour by default?” Cost, and the risk of learning from a bad data hour. Retrain as often as decay justifies, with validation gates.
What separates senior answers
They separate data drift, concept drift and pipeline bugs, and say bugs are most common.
They pick retrain cadence from a measured decay curve.
They describe shadow, canary, ramp and auto rollback as one release path.
They name feedback loops and keep an exploration slice.
Say this out loud: “I will monitor three layers. Data quality, input and prediction drift with PSI, and online model metrics when labels arrive. New models go through shadow, then canary, then ramp, with auto rollback. Retrain cadence comes from a measured decay curve.”
10. Design an LLM assistant with RAG Hard
Clarify
Goal. Answer employee or customer questions from a private document set. Answers must be correct and cite sources.
Scale. 5M documents, about 50M chunks. 200 queries per second at peak.
Latency. First token in under 1.5 s. Full answer in under 5 s.
Ask. Do users have different access rights? Then retrieval must filter by permission. Is wrong worse than no answer? Usually yes.
Frame and metrics
ML task. Retrieval plus conditional generation. Retrieve relevant chunks, then generate an answer grounded in them.
Why RAG. The base LLM does not know your private data. Fine-tuning is slow to update and still hallucinates. RAG updates when the documents do.
Business metric. Support tickets deflected, or time saved per employee.
Online metrics. Thumbs up rate, follow-up rate, escalation to a human, citation click rate.
Offline metrics. Retrieval recall@K and MRR. Answer groundedness, correctness and relevance.
Groundedness. The share of claims in the answer that the retrieved text supports. Also called faithfulness. An answer can be true but ungrounded, and that is still a risk.
Data and labels
Ingest. Parse PDFs, HTML and wikis. Keep titles, headings, tables as text, and permissions.
Chunking. Split into chunks of about 300 to 800 tokens with 10% to 20% overlap. Split on headings and paragraphs, not mid-sentence.
Chunk context. Prepend the doc title and section heading to each chunk. It helps both retrieval and the LLM.
Eval set. Build 500 to 2,000 real questions with gold answers and gold source chunks. Mine them from support tickets and logs.
Features
Dense embeddings. A text embedding model maps each chunk to a vector, about 768 to 1,536 dims.
Sparse signals. BM25 over the same chunks. It catches exact terms like error codes and product names.
Metadata. Doc date, source, product and access list. Use it to filter and to boost fresh docs.
Query rewrite. Turn a follow-up like “what about for Mac?” into a full standalone query using chat history.
Model
Baseline. Dense top-5 retrieval, stuff into the prompt, generate. Measure it before adding parts.
Hybrid retrieval. BM25 top 50 plus dense top 50. Merge with reciprocal rank fusion.
Reranker. A cross-encoder scores the 100 merged chunks against the query. Keep the top 5 to 8.
Prompt. System rules, then the chunks with IDs, then the question. Tell the model to answer only from the chunks, cite chunk IDs, and say “I don’t know” if the answer is missing.
Generator. A strong hosted LLM, or a smaller fine-tuned model if cost or privacy demands it.
RAG pipeline. Documents are chunked and indexed offline. Each question is rewritten, retrieved, reranked, answered, then checked.
Serving
Budget. 50 ms retrieval, 100 ms rerank, then the LLM. Stream tokens so the first one shows fast.
Cost. Cost scales with prompt tokens. 8 chunks of 500 tokens is 4K tokens per call. Fewer, better chunks save money and often improve answers.
Caching. Cache embeddings of common queries. Cache full answers for repeated questions. Use prompt caching for the fixed system prompt.
Routing. Send easy questions to a small, cheap model. Send hard ones to the large model.
Access control. Filter chunks by the user’s permissions at retrieval time. Never rely on the LLM to hide data.
Freshness. Re-embed changed docs within minutes. Delete removed docs from the index at once.
Evaluation and iteration
Evaluate retrieval alone. Recall@5 against gold chunks. If the right chunk is not retrieved, no prompt can save the answer.
Evaluate generation. Groundedness, correctness against the gold answer, and citation accuracy.
LLM as judge. Use a strong model with a rubric to grade thousands of answers. Check it against human grades on a sample first.
Online. A/B test prompt and retrieval changes on thumbs up, escalation rate and cost per answer.
Error buckets. Sort failures into retrieval miss, bad chunking, ignored context, or hallucination. Fix the biggest bucket first.
Follow-ups they will ask
“How do you stop hallucinations?” Tell the model to answer only from context and cite. Run a groundedness check that verifies each claim against the chunks. Refuse or hedge when support is weak.
“RAG or fine-tuning?” RAG for facts that change. Fine-tuning for style, format or domain language. They combine well.
“How do you pick chunk size?” Sweep it on the eval set. Small chunks are precise but lose context. Large chunks keep context but add noise and cost.
“What about prompt injection?” A document could contain hidden instructions. Treat chunks as data, not commands. Filter known patterns. Limit what tools the model can call.
“Long context windows are huge now. Why retrieve?” Cost and latency grow with tokens. Models also miss facts buried in the middle. Retrieval stays cheaper and more precise.
What separates senior answers
They evaluate retrieval and generation apart, and build a gold eval set first.
They use hybrid retrieval and a reranker, and can say why.
They do access control in retrieval, not in the prompt.
They give token math for cost and latency.
Say this out loud: “I will build a gold eval set first. Retrieval is hybrid with a cross-encoder reranker. The prompt forces citations and allows ‘I don’t know’. A groundedness check runs before the answer ships. I will measure retrieval recall and answer groundedness apart.”
Recap
Use seven steps every time. Clarify, frame and metrics, data, features, model, serving, evaluation.
Name three metric layers. Business, online and offline. Show they line up.
Large item pools need stages. Cheap retrieval, heavy ranking, then rule-based re-ranking.
Give each stage a latency budget in milliseconds.
Calibration matters when scores feed an auction or a weighted sum.
Rare and adversarial problems need recall at a fixed low FPR, not accuracy.
Ship through shadow, canary and ramp. Monitor drift and retrain on a measured cadence.