← Back to CatZep
Applied Science Interview Questions and Preparation Guide
One guide and seven question banks. Each question has what the interviewer is testing, a strong answer, the follow-ups they will ask, and one line to say out loud.
An Applied Scientist sits between research and engineering. You are hired to turn models into products. So the loop tests three things at once. Do you know the theory? Can you write the code? Can you ship a model that moves a business metric?
This guide covers all three. Start with the loop guide. It tells you which rounds to expect and how to spend eight weeks. Then work through the banks in order.
How to use this. Read a question. Close the page. Answer it out loud in two minutes. Then open the strong answer and compare. The gap between your answer and the strong one is your study list.
Breadth first, then depth. Most loops ask easy questions in many areas, then drill hard into one. Know every bank at the first level. Know your own research area at the deepest level.
1 guide · 7 question banks · 6 theory pages · 2 manager pages · 90+ worked questions.
Start here
What the role is, every round in the loop, what each round grades, how levels differ, and a week-by-week study plan.
Part I — Foundations
The classic theory round. Bias and variance, regularization, metrics, imbalance, trees, leakage and calibration.
Bias-variance trade-off
Detecting and fixing overfitting
L1 vs L2 regularization
Precision, recall and when ROC-AUC lies
Class imbalance
Logistic regression loss and gradient
Random forests vs gradient boosting
Cross-validation and data leakage
Feature engineering
Calibration
Generative vs discriminative
Curse of dimensionality and PCA
The math under every model and every experiment. Bayes, the CLT, p-values, MLE and MAP, and puzzles.
Bayes and the rare disease test
Central limit theorem
What a p-value is not
MLE vs MAP
Expected value puzzles
Reservoir sampling
Type I, Type II and power
Simpson's paradox
Standard error and the bootstrap
Which distribution, when
From backprop to the transformer. Optimizers, normalization, attention, fine-tuning, and how LLM inference works.
Backpropagation
Vanishing and exploding gradients
SGD, momentum, Adam, AdamW
BatchNorm vs LayerNorm
Dropout and other regularizers
Convolutions and receptive field
Attention and the transformer block
Positional encodings and RoPE
Fine-tuning, LoRA, RLHF and DPO
KV cache and decoding
Embeddings and contrastive learning
Scaling laws and mixed precision
Part II — Building
Implement the classics in NumPy with no libraries. Every problem has full code, the trap, and the complexity.
Linear regression by gradient descent
Logistic regression
K-means
K-nearest neighbours
Stable softmax and cross-entropy
Best decision-tree split
ROC-AUC from scores
Self-attention forward pass
Two-layer MLP with backprop
K-fold cross-validation
A seven-step framework, then the designs that come up most. Feeds, recommendations, ads, fraud, search and RAG.
The seven-step framework
News feed ranking
Video recommendations
Ads click prediction
Fraud and spam detection
Search ranking
Harmful content moderation
People You May Know
Monitoring, drift and retraining
An LLM assistant with RAG
How you prove a model helped. A/B test design, power, metrics, interference, CUPED, and what to do when you cannot randomize.
Design an A/B test end to end
Sample size and power
North star, guardrail and proxy metrics
Novelty and primacy effects
Network effects and interference
Peeking and multiple testing
CUPED variance reduction
Sample ratio mismatch
Causal inference without a test
Offline vs online disagreement
Interleaving for ranking
Part III — You
The paper deep dive, the research talk, and the stories. How to show impact, judgment and ownership.
Walk me through your best project
The research presentation
Picking problems that matter
A failed experiment
Disagreeing with a PM or engineer
Research to production
Critique a paper on the spot
Mentoring and influence
A problem with no labels
Questions to ask them
Part IV — Theory
The ideas under every answer. Read these when a question bank sends you deeper.
Why models generalize. Risk, bias and variance, PAC bounds, VC dimension, double descent and distribution shift.
Estimators, Fisher information, tests, intervals, the delta method, regression theory, GLMs and Bayesian inference.
Inequalities, the Gaussian, Markov chains, entropy, KL, mutual information and sampling methods.
Convexity, convergence rates, SGD theory, KKT and duality, non-convex landscapes, EM and hyperparameter search.
Projections, eigenvectors, SVD, stable solvers, matrix gradients with shapes, einsum and low-rank structure.
Everything that is not a model. Framing, labels, evaluation, error analysis, MLOps, bandits, fairness and writing.
Part V — Leading Applied Science
For managers, and for scientists who want to know how their manager thinks.
The four hats, the move from scientist to manager, hiring, growing and judging scientists, and your first 90 days.
Portfolio bets, kill criteria, experiment reviews, measuring impact, compute budgets, and ten manager interview questions.
An original study companion for Applied Science interviews. Part of CatZep.