Know the rounds before you study for them. Each round grades a different thing. Prepare for the grade, not the topic.
Most people fail the Applied Science loop in a round they did not prepare for. A strong researcher fails the coding round. A strong engineer fails the ML depth round. This page maps every round, says what it grades, and gives you a plan to cover all of them.
Companies use the title in different ways. But the core is the same. An Applied Scientist owns the model and its science. An ML Engineer owns the system around it. A Research Scientist owns new methods and papers.
Research Scientist. Judged on new ideas and papers. Loop leans on research depth and the job talk.
Applied Scientist. Judged on models that ship and move a metric. Loop tests theory, coding, design and experiments.
ML Engineer. Judged on systems that train and serve at scale. Loop leans on coding and infrastructure design.
Data Scientist. Judged on decisions and insight. Loop leans on statistics, SQL and product sense.
The one-line pitch. “I take a fuzzy business problem, frame it as an ML problem, build the model, and prove with an experiment that it helped.” Every round checks one part of that sentence.
The rounds
A typical onsite has five or six rounds of 45 to 60 minutes. Names vary by company. The content does not.
Phone screen. One coding problem plus ML questions. Gate to the onsite.
Coding. One or two algorithm problems. Sometimes ML coding in NumPy.
ML breadth. Rapid questions across the whole field.
ML depth. A deep drill into your area or your past project.
ML system design. Design a full ML product from data to experiment.
Research talk. Present your work for 30 to 45 minutes, then take questions.
Behavioral. Stories about impact, conflict, failure and ownership. Often mixed into other rounds.
Behavioral is never only one round. Many companies score leadership principles in every round. The last five minutes of a design round are often a behavioral question. Have stories ready all day.
The phone screen
What it grades: can you code, and do you know the basics? The bar is “no red flags”, not “brilliant”.
One medium coding problem in 25 to 30 minutes.
Five to ten quick ML questions. Bias and variance. Overfitting. Precision and recall. How logistic regression works.
Sometimes a short “tell me about your work”.
Do not ramble on easy questions. “What is overfitting?” wants three sentences, not five minutes. Give the short answer, then offer to go deeper. Long answers eat time for the coding problem.
Coding
What it grades: clean, correct code under time pressure, and clear talk while you write it.
There are two kinds of coding round. Ask your recruiter which one you have.
Algorithm coding. The same as a software engineer round. Arrays, hash maps, trees, graphs, DP. The Python Coding Interview pages cover this.
ML coding. Implement a model from scratch in NumPy. K-means, logistic regression, attention. See ML Coding from Scratch.
Say this out loud: “Before I code, let me confirm the input shapes. X is n by d, y is length n, and labels are zero or one. Is that right?”
ML breadth
What it grades: do you know the whole field well enough to pick the right tool?
Expect 15 to 25 questions in an hour. They jump between topics. Each one is easy alone. The test is that you never freeze.
Use a fixed shape for every answer. Define it in one line. Give the intuition. Say when you would use it. Name one trap. That takes 45 seconds and sounds senior.
ML depth
What it grades: do you truly understand one area, all the way down?
The interviewer picks your strongest area, often from your resume. Then they keep asking “why?” until you reach the edge of what you know. Reaching the edge is expected. How you behave there is the grade.
Know the math of your main method. Write the loss. Derive the gradient.
Know why you chose it over the two obvious alternatives.
Know what broke, and how you found out.
Know the recent papers in your area from the last two years.
At the edge, say: “I have not worked through that case. My guess is X, because of Y. I would check it by Z.” That is a pass. Bluffing is a fail.
ML system design
What it grades: can you own an ML product end to end? This round sets your level more than any other.
You get a vague prompt like “design a feed ranker”. You must frame the problem, pick labels and features, choose a model, plan serving, and design the experiment. The framework is in ML System Design. The experiment half is in Experimentation.
Senior candidates drive. A junior candidate waits for the next question. A senior candidate states the plan, keeps time, and says what they would cut. Lead the interviewer through your seven steps.
Research talk
What it grades: can you explain hard work clearly, and does it show real impact?
30 to 45 minutes of slides, then 15 to 30 minutes of questions.
The audience is mixed. Some are experts, some are not.
Pick one or two projects. Depth beats a tour of everything.
What it grades: will you be a good teammate, and do you own outcomes?
Prepare six to eight stories. Each one should fit many questions. Use STAR: Situation, Task, Action, Result. Spend most of the time on Action. End with a number.
Write your six to eight STAR stories. Say each one out loud in two minutes.
Build your research talk slides.
Week 8 — Rehearse
Give the research talk to a friend twice. Take their hardest questions.
Two full mock loops.
Redo every question you marked red in week 1.
Mocks are the highest-return hour. Reading feels like progress. Speaking under a clock is the skill being graded. Do at least four mocks across the eight weeks.
The last week and the day
Stop learning new topics three days out. Review only.
Reread your resume. Every line on it is fair game.
Prepare two questions to ask in each round.
Sleep. A tired brain loses more points than any topic gap.
On the day, reset between rounds. A bad round does not sink a loop. Carrying it into the next one does.
Common mistakes
Jumping to the model. In design, frame the problem and the metric first. The model comes fourth.
Only offline metrics. AUC went up is not impact. Say what the A/B test showed.
“We” for everything. In stories, say what you did. The interviewer is hiring you, not your team.
Bluffing. One confident wrong answer costs more than three honest “I do not know”s.
Skipping coding prep. Scientists often fail here. It is the easiest round to fix with practice.
Recap
Six or so rounds. Each grades a different thing.
System design sets your level more than any other round.
Have six to eight stories ready for every round.
Eight weeks, ten hours a week, at least four mocks.