The deep dive, the talk, and the stories. This is where they decide if they trust you with a real problem.
Every Applied Science loop has a part about you. Sometimes it is a full round. Often it is the last ten minutes of every round. The questions sound soft. The grading is not. Interviewers write down what you did, what it changed, and how you think when things go wrong.
This page covers the ten questions you will hear most. Each one has what they are testing, a shape for your answer, a full example, the follow-ups, and the traps. The examples are fictional. Build your own from your real work, then practice them out loud.
Most story questions want the same four parts. Use them in order. Keep the setup short and the action long.
Situation. Where you were and what was going on. Two or three sentences.
Task. What you owned and why it was hard. One or two sentences.
Action. What you did, step by step, and why. This is the core.
Result. What changed, with a number. Then one line on what you learned.
The time split
A good story runs two to three minutes. Spend about 10% on Situation, 10% on Task, 60% on Action and 20% on Result. Most people flip this. They spend a minute on context and rush the part that gets scored.
Say “I”, not “we”
The interviewer is hiring you, not your team. Say “we” once to set the scene. Then say “I” for what you did. If a teammate did a key part, name it. Then say what you did next to it.
End with a number
“It went well” scores nothing. “Click rate rose 2.1% in a two-week A/B test” scores well. If you have no business number, use a model number, a time saved, or a cost cut. If the result was bad, say the number anyway.
One story fits many prompts
You do not need forty stories. You need six to eight strong ones. Each can answer several prompts if you change the angle. Here is one set mapped to common prompts.
Story A: the ranking model you shipped. Fits “best project”, “research to production”, “biggest impact” and “hard technical call”.
Story B: the launch that failed its A/B test. Fits “a failure”, “a mistake”, “what you learned” and “bad news to a leader”.
Story C: pushing back on a PM about the metric. Fits “disagreement”, “influence without authority” and “saying no”.
Story D: the fraud model built with no labels. Fits “ambiguity”, “no labels”, “scrappy start” and “creative solution”.
Story E: the intern you guided to a launch. Fits “mentoring”, “growing others” and “delegation”.
Story F: killing your own project. Fits “picking problems”, “priorities” and “a hard decision”.
Write a story bank. For each story, write one line per STAR part and the one number. Then list which prompts it fits. Read the bank the night before. Do not memorize scripts. Memorize the beats.
1. Walk me through your best project Medium
What they are testing
Depth. Do you understand your own work down to the details?
Judgment. Did you choose the approach for good reasons?
Ownership. What part was yours, and what part was the team's?
Impact. Did it change something real, both offline and online?
How to structure it
Problem. One sentence on the user or business problem. Not the model.
Why it mattered. The metric it moved and how big the prize was.
Approach. The framing, the data, the model, and the loss. Keep it to the key choices.
Alternatives you rejected. Two options and why you said no. This shows judgment.
Results. Offline metrics first, then the online A/B test. Give both numbers.
What you would do differently. One honest change. It shows you still think about it.
Aim for four to five minutes on the first pass. Then let them drill. They will pick one part and go deep. Be ready to go three levels down on any choice.
A strong example answer
“At a food delivery app, our search ranking was a gradient boosted tree from three years back. Users who searched and did not order were 41% of search sessions. Each point of search conversion was worth about $6M a year. I owned the new ranker end to end.
I framed it as learning to rank with order as the label. Clicks were noisy, so I used them only as a second task. I built a two-tower model for candidate scoring, then a small cross-attention model for the top 200. The towers used query text, past orders, time of day and store load.
I rejected two ideas. A pure LLM reranker scored well offline but cost 90 ms at p99. Our budget was 30 ms. I also rejected training on clicks alone. In a pilot it pushed cheap fast food that people clicked but rarely ordered.
Offline, NDCG@10 rose from 0.412 to 0.447. Online, a three-week test on 10% of traffic lifted search conversion 1.8%. Order value stayed flat. Latency rose 4 ms at p99. We shipped to everyone.
If I did it again, I would fix position bias earlier. I added inverse propensity weights late. That alone added 0.4% in a later test. I should have checked for it in week one.”
Follow-ups they will ask
“Why that loss?” Explain pairwise or listwise versus pointwise. Tie it to what the metric rewards.
“How did you know the offline gain was real?” Mention a time-based split, no leakage, and confidence intervals.
“Which feature mattered most?” Name it and say how you measured it, like an ablation.
“What did you personally build?” Be exact. Name the parts others built.
“What broke after launch?” Have a real answer. Every launch has one.
Red flags to avoid
Spending three minutes on the model and none on why it mattered.
Only offline numbers. “AUC went up” is not impact.
Not knowing a detail of your own project, like the size of the dataset.
Saying “we” for every step so they cannot tell what you did.
Say this out loud: “The problem was lost orders from search. It was worth about $6M a year per point. I owned the ranker. I picked a two-tower model plus a small reranker and rejected an LLM reranker on latency. Offline NDCG rose 8%. Online conversion rose 1.8%. I would fix position bias sooner next time.”
2. The research presentation Hard
What they are testing
Can you explain deep work to a mixed room? Some listeners are experts. Some are not.
Do you know why your work matters, not only what it is?
Can you defend your choices under hard questions without getting defensive?
Do you know the limits of your own work?
How to structure it
Pick one project, two at most. Depth beats breadth. Pick work that shipped or that you can tie to a real problem. A talk on three papers in 40 minutes feels like a list. A talk on one problem feels like a story.
Use this rule for the room. The first ten minutes are for everyone. The middle is for the experts. The last five minutes are for everyone again.
Slide-by-slide outline for a 40-minute talk
This plan leaves 15 minutes for questions in a 55-minute slot. Aim for about 20 slides. Timings are cumulative.
Title slide (0:00 to 0:30). Your name, the talk title, and one line on the result. Do not read it.
About me (0:30 to 1:30). Three bullets on your path. Skip a long bio.
The problem in plain words (1:30 to 4:00). One picture of a user hitting the problem. No math yet.
Why it matters (4:00 to 6:00). The business metric and the size of the prize. One number in big type.
Why it is hard (6:00 to 8:30). Two or three real obstacles. Scale, noise, latency, no labels.
What others tried (8:30 to 10:30). Prior work and past team attempts. Say where each falls short.
The key idea (10:30 to 13:00). Your main insight in one sentence and one diagram. This is the slide they remember.
Problem framing (13:00 to 15:00). Inputs, outputs, label, and loss. Now the math starts.
Data (15:00 to 17:00). Sources, size, how you cleaned it, and the train and test split.
Method part 1 (17:00 to 20:00). The model design with one clean figure.
Method part 2 (20:00 to 22:30). The trick that made it work. Show the one equation that matters.
Offline results (22:30 to 25:00). Your method against strong baselines. Error bars on every number.
Ablations (25:00 to 27:00). Remove each part and show the drop. This proves each piece earns its place.
Online results (27:00 to 29:30). The A/B test. Primary metric, guardrails, duration, and traffic share.
Error analysis (29:30 to 31:30). Two real failure cases with examples. Say why they fail.
Production story (31:30 to 33:30). Latency, cost, monitoring, and what you had to simplify.
Limits (33:30 to 35:00). What it does not do. Say it before they ask.
What I would do next (35:00 to 37:00). Two or three concrete next steps.
Other work (37:00 to 38:30). One slide with two short blurbs. Shows range without stealing time.
Takeaways (38:30 to 40:00). Three bullets. Problem, idea, impact. Then stop and invite questions.
Practice plan. Run the talk out loud at least three times. Do one run for a non-expert friend and one for a peer who will attack it. Time every run. If you run long, cut slides, do not talk faster.
The 8 hardest audience questions
“Is your baseline strong enough?” Name the strongest baseline you ran and how you tuned it. Say you gave it the same tuning budget as your method. If you skipped a known baseline, admit it and say what you expect.
“How do I know this is not just more parameters?” Point to a size-matched baseline. Or show the ablation that removes your idea but keeps the size.
“The online gain is small. Is it real?” Give the confidence interval and test length. Mention a pre-test A/A check or a holdback that confirmed it.
“Could there be leakage in your split?” Explain the split by time or by user. Name one leak you found and fixed. That builds trust.
“Why not just use an LLM?” Give the cost and latency numbers. Say if you tested one. If not, say how you would compare.
“Does it work for new users or rare items?” Show a sliced result. If it is weaker there, say so and say what you would try.
“What would you do with ten times the data?” Say which part of the method would gain most. Mention a scaling curve if you have one.
“What is the weakest part of this work?” Pick a real weakness. Never say “nothing”. Then say how you would fix it.
When you do not know, say so. Then reason out loud for 20 seconds. “I did not test that. My guess is X, because of Y. I would check it with Z.” That scores better than a bluff.
A strong example answer
Here is how one candidate opened a talk on fraud detection, in the first two minutes.
“Last year, fake sellers cost our marketplace about $14M in refunds. Our rules caught 60% of them, but only after the first bad order. I want to show you how I caught 85% of them before their first sale.
The hard part was labels. We only learn a seller is fake weeks later, when refunds pile up. And fakers change tactics every month. So a model trained on last quarter's fraud goes stale fast.
My key idea was to stop modeling the seller alone. I built a graph of sellers, devices, bank accounts and addresses. Fake sellers share these things far more than real ones do. A graph neural network learned those patterns from 2.3M sellers.
Offline, recall at 1% false positive rate rose from 0.58 to 0.81. In a six-week shadow run and then a live test, we blocked 85% of fake sellers before their first sale. False blocks on real sellers stayed under 0.3%. Refund losses fell by about $9M a year.
Next, I will show how I built the graph. Then why the obvious model failed, and what still breaks. Please stop me any time.”
Follow-ups they will ask
“Go back to slide 11.” Know every slide cold. Have backup slides for details you cut.
“Why this project and not another?” Say it best shows how you work from problem to impact.
“What did your co-authors do?” Give clear credit. Then say what was yours.
“How would this apply to our product?” Prepare one slide or one answer that maps your idea to their domain.
Red flags to avoid
Going over time. It signals you cannot prioritize.
Thirty slides of math with no picture of the problem.
Arguing with a questioner. Say “good point”, answer, and move on.
No limits slide. They will find the limits for you, and it looks worse.
Say this out loud: “I will spend ten minutes on the problem and why it matters. Then twenty on the method and results. The rest goes to limits and next steps. Please stop me any time.”
3. Picking problems that matter Medium
What they are testing
Do you tie science work to a business metric, or do you chase cool ideas?
Can you size impact before you spend months on it?
Can you say no to work that does not matter, even when someone senior asks?
How to structure it
The choice. Name the options you had. Show you did not just take the first one.
The metric. Say which top-line metric each option could move.
The sizing. Show the rough math. Reach times lift times value, then divided by effort.
The no. Say what you turned down and how you said it.
The result. Show the bet paid off, with a number.
Quick sizing formula. Impact ≈ users touched × expected lift × value per unit × chance it works. Divide by weeks of effort. Compare options on this one number. It does not need to be exact. It needs to be roughly right and written down.
A strong example answer
“I joined a streaming app's personalization team as the second scientist. My manager handed me three ideas from the roadmap. One was a new video embedding model. One was better ranking for the home page top row. One was a churn model for the retention team.
The embedding model was the most fun. But I sized all three in a doc first. The top row touched 100% of daily users. A 1% gain in plays from that row was worth about 0.3% in watch time. The embedding model only fed a side shelf that 12% of users scrolled to. The churn model was useful, but the retention team had no lever to act on its scores yet.
My math said the top row was worth about 4x the embedding work for similar effort. I shared the doc with my manager and the PM. I proposed we do the top row first and park the embedding model.
The PM had promised the embedding work to a director. So I offered a two-week spike to test the embedding idea on the side shelf. The spike showed a 0.2% gain. That settled it.
The top row work shipped in one quarter. It lifted watch time 0.6% in a four-week test. That was the team's biggest win that half.”
Follow-ups they will ask
“What if your sizing was wrong?” Say you set a checkpoint. If the first test showed under half the expected lift, you would stop.
“How do you pick when nothing ties to revenue?” Use a proxy that leaders already track. Engagement, retention, or cost.
“How do you balance long research bets?” Say you keep about 20% of time for bets with a clear test point.
“Tell me about a time you said no to a leader.” Use data, offer an option, and let them decide.
Red flags to avoid
Choosing work because it was interesting or would make a good paper.
No numbers in the sizing. “It felt bigger” is not a reason.
Saying no with no data and no other option offered.
Say this out loud: “Before I build, I size it. Users touched, times expected lift, times value, divided by effort. I write it down and share it. Then I say no to the smaller bets. But I offer a cheap test so the other side sees the data.”
4. A failed experiment Medium
What they are testing
Honesty. Do you own the failure, or blame others?
Rigor. Did you find out why it failed, or just move on?
Learning. What do you do differently now?
Judgment. Did you stop at the right time?
How to structure it
The bet. What you expected and why it was a reasonable bet.
The failure. What happened, with numbers. Do not soften it.
The diagnosis. How you found the root cause. This is the most scored part.
The response. What you did next. Fix, pivot, or stop.
The lesson. One habit you changed, with proof you still use it.
Pick a real failure that cost something. A fake failure, like “I work too hard”, loses trust fast. But pick one where you were not reckless.
A strong example answer
“I built a new ad click model for a news site. It had a transformer over each user's last 100 page views. Offline, log loss fell 3.2%. That is big for ads. I was sure it would win.
The A/B test ran two weeks on 5% of traffic. Revenue per thousand views dropped 1.4%. Click rate was flat. I was stunned.
I dug in for three days. First I checked the logging. It was fine. Then I sliced by user type. Logged-out users made up 55% of traffic. They had almost no page history. For them the new model was worse than the old one. In the offline data, logged-out users were only 30% of rows. The training log had dropped many of their events.
So my offline test set did not match real traffic. I fixed the data pipeline to keep logged-out events. Then I added a fallback path for users with short history. The second test lifted revenue 0.9%.
The lesson changed how I work. I now check that the offline test set matches live traffic on the top three slices before any A/B test. I added that check to our launch template. It caught a similar gap on a teammate's model two months later.”
Follow-ups they will ask
“Could you have caught it sooner?” Yes. Say how. A slice check before launch.
“How did you tell your manager?” Fast and plain. Results, what you know, what you will check next.
“When do you stop a failing project?” When the best case no longer beats the cost. Set that bar up front.
“Tell me about a failure you could not fix.” Have a second story where you stopped the work. That shows judgment too.
Red flags to avoid
Blaming the data team, the PM, or bad luck.
A failure with no root cause. “It just did not work” shows no rigor.
A lesson that is vague, like “communicate more”.
A failure so big it raises doubts, like a privacy breach you caused.
Say this out loud: “Offline it won by 3%. Online it lost 1.4% of revenue. I sliced the results and found the test set did not match live traffic. I fixed the pipeline and the next test won. Now I check slice match before every launch.”
5. Disagreeing with a PM or engineer Medium
What they are testing
Can you push back with data, not with rank or volume?
Do you listen? Do you change your mind when you should?
Can you disagree and still keep a good working bond?
Do you commit once a decision is made?
How to structure it
The stakes. What the disagreement was about and why it mattered.
Their view. State it fairly. Show you understood why they held it.
Your view. State it with the evidence you had.
How you resolved it. Data, a test, a third person, or a compromise.
The outcome. Who was right, the number, and how the bond held up.
A strong example answer
“I worked on a shopping app's recommendations. Our PM wanted to optimize the feed for clicks. Clicks were the team's goal for the quarter. I thought clicks would push clickbait items and hurt purchases.
I first asked the PM why clicks. The goal came from her director. Clicks were fast to measure, and purchases took a week to show up. That was fair. A purchase goal would slow every test.
So I did not argue in the meeting. I pulled data from our last six tests. In two of them, clicks rose but purchases fell. One had cost about $400K in a month before anyone noticed. I shared a one-page note with the PM first, not with the whole team.
I proposed a middle path. Optimize a blend of clicks and add-to-cart, which shows up within a day. Keep purchases as a guardrail that can block a launch. She agreed to try it on the next test.
The blended model lifted clicks 1.1%, less than the pure click model's 2.4%. But it lifted purchases 0.8%. The click model had dropped them 0.3%. The PM took the blend to her director, who made it the new team goal. She and I still work together on goal design for each quarter.”
Follow-ups they will ask
“What if the data had sided with the PM?” Say you would have backed the click goal fully.
“Tell me about a time you lost the argument.” Have one. Show you committed and helped it succeed.
“What about an engineer who said your model was too slow?” Say you treat latency as a hard limit and look for a cheaper model.
“How do you handle a disagreement that gets personal?” Move to one-on-one, focus on the goal, and bring in a manager only if stuck.
Red flags to avoid
Painting the other person as dumb or lazy.
Winning by going over their head first.
A story where you were right and nothing else happened. Show the bond survived.
Say this out loud: “I asked why she wanted clicks before I argued. Then I brought data from six past tests in a private note. I offered a blend with purchases as a guardrail. It won on both metrics and became the team goal.”
6. Research to production Hard
What they are testing
Do you know that a model is not done until it serves real users?
Can you trade accuracy for latency, cost and safety with good sense?
Do you own the launch, or throw the model over the wall?
Do you plan for drift and failure after launch?
How to structure it
Walk through the gap between the lab and the product. Hit these five points. Not every story uses all five. Pick the three that were hardest.
Latency. The budget, your first number, and how you closed the gap.
Simplification. What you cut from the research model and what it cost in accuracy.
Distillation. If you trained a small student model from a big teacher, say how much quality it kept.
Monitoring. What you watch after launch. Input drift, score drift, and the business metric.
Owning the launch. Ramp plan, rollback plan, on-call, and the first bad day.
Distillation (a small model copies a big one). You train the big teacher model first. Then you train a small student model to match the teacher's scores, not only the true labels. The student is fast enough to serve and often keeps most of the quality.
A strong example answer
“I built a model to flag toxic comments for a social app. My research model was a 350M parameter text model. It hit 0.93 F1 on our test set. The old keyword filter was at 0.71. But it took 120 ms per comment on a GPU. The budget was 15 ms on CPU, because comments post in real time.
I tried three things. First I cut the input to 128 tokens. Ninety-six percent of comments fit, and F1 dropped only 0.004. Second, I distilled the big model into a 6-layer student with 22M parameters. I trained it on 40M unlabeled comments scored by the teacher. The student hit 0.91 F1. Third, I quantized it to 8-bit. That got p99 latency to 11 ms on CPU.
I owned the launch. I ran a two-week shadow mode with no action taken. Then I ramped 1%, 10%, 50%, 100% over three weeks. I set alerts on the share of comments flagged, the score mix, and appeal rate. I wrote the rollback runbook and took on-call for the first month.
In week five, the flag rate jumped 30% overnight. A new slang term had gone viral. My alert caught it in two hours. I added 2,000 labeled examples and retrained in four days. Reports of toxic comments fell 38% after the full launch. Cost was $4K a month, not the $60K a GPU model would have needed.”
Follow-ups they will ask
“How did you pick the student size?” Say you swept three sizes and plotted F1 against latency. Pick the knee of the curve.
“What else could you have done for latency?” Caching, early exit, pruning, batching, or a cheap first-stage filter.
“How do you retrain?” Give a cadence, a trigger, and how you validate before swap.
“Who else was involved?” Name the infra engineers. Say what you did and what they did.
“What would make you roll back?” Name the guardrail and the threshold.
Red flags to avoid
“Then the engineers deployed it.” That says you do not own impact.
No monitoring plan. Models break silently.
Not knowing your latency or cost numbers.
Say this out loud: “The research model was 120 ms on GPU. The budget was 15 ms on CPU. I cut input length, distilled to a small student, and quantized it. It kept 0.91 of 0.93 F1 at 11 ms. I owned the ramp, the alerts and on-call. An alert caught a drift problem in two hours.”
7. Critique a paper on the spot Hard
What they are testing
Can you read work critically and fast?
Do you know what makes a result believable?
Are you fair? Do you see the strengths, not only the flaws?
Can you turn a critique into a next step?
How to structure it
They may hand you a paper, show you an abstract, or ask about a famous paper. Read the abstract and the main results first. Then walk this checklist out loud. Start with one honest strength.
The critique checklist
Claim. What exactly do they claim? Is it about accuracy, speed, or a new ability? Is the claim bigger than the evidence?
Baselines. Are the baselines strong and current? Were they tuned as hard as the new method? A weak baseline makes any idea look good.
Ablations. Did they remove each new part to show it helps? Without ablations, you cannot tell which idea did the work.
Data leakage. Could test data have leaked into training? Look at the split. Random splits on time series or users leak. Pretraining data may contain the test set.
Metrics. Do the metrics match the real goal? Is accuracy hiding class imbalance? Did they report only the metric that looks best?
Statistical significance. Do they report error bars, seeds, or a test? A 0.3 point gain from one run may be noise.
Compute fairness. Did the new method get more parameters, more data, or more training time? Compare at equal compute.
Reproducibility. Is code or data released? Are the hyperparameters and training details enough to rebuild it?
End with a fix. After the flaws, say what experiment would settle the biggest doubt. “I would want a size-matched baseline and three seeds.” That turns a critic into a collaborator.
A strong example answer
The interviewer showed an abstract. It claimed a new attention variant beat a standard transformer by 2.1 points on a reading benchmark.
“First, the idea is clean. It cuts attention cost from quadratic to linear. If it holds, that matters a lot for long documents.
Now the doubts. The claim is about accuracy, but the selling point is cost. I want to see a speed and memory chart, not only accuracy.
On baselines, they compare to a base transformer from 2019. I would want a modern baseline with the same training recipe. Their model has 180M parameters and the baseline has 110M. That gap alone could explain two points. So compute fairness is my top worry.
There is one ablation table, and it removes only the gating part. I want an ablation that keeps the size but drops the new attention.
They report one run per model. On this benchmark, seeds can swing a full point. So I want five seeds and a confidence interval.
On leakage, they pretrained on a web crawl. The benchmark's passages come from the web. They should check for overlap.
Code is promised but not out, so I cannot check the details.
Overall, the idea is promising, but the evidence is weak. The one experiment I would run first is a size-matched comparison with five seeds. If the gain holds, I would believe it.”
Follow-ups they will ask
“Would you accept it as a reviewer?” Give a clear verdict. Weak accept or weak reject, plus the one change that flips it.
“How would you use this at our company?” Name a product problem it could help and the first test you would run.
“What is the best paper you read this year?” Have one ready. Say what it showed and one doubt you have.
“Critique your own paper.” Run the same checklist on your work. Be honest.
Red flags to avoid
Only praise, or only attacks. Both show weak judgment.
Vague doubts like “I am not sure it generalizes” with no reason.
Bluffing about a paper you have not read. Say so and reason from the abstract.
Say this out loud: “The idea is promising. My top doubt is compute fairness, since their model is 60% larger. I also want ablations, five seeds, and a leakage check. The first thing I would run is a size-matched baseline.”
8. Mentoring and influence Medium
What they are testing
Do you make the people around you better?
Can you lead without a manager title?
Do you spread good science habits beyond your own work?
For senior roles, can you set direction for a group?
How to structure it
The person or group. Who they were and where they were stuck.
Your approach. How you helped. Teaching, pairing, reviews, or giving them room.
Their growth. What they could do after that they could not before.
The wider effect. A habit, tool, or process that spread past one person.
Two good stories cover most prompts. One about growing a single person. One about changing how a team works.
A strong example answer
“I mentored a summer intern named Priya. She was a strong PhD student in vision but new to industry. Her project was to improve product image search on our shopping site.
In week two, she was stuck. She wanted to train a big model from scratch and had spent a week on data loading. I did not take over. I asked her what result she needed by week six to call the summer a win. She said a launched A/B test. Then I asked what the fastest path to a test was.
She landed on fine-tuning an existing image model. I set up a weekly one-hour review where she showed one chart and one question. I paired with her for two days on our serving stack. I also had her present her plan to the PM herself. That built her confidence and got the PM's support early.
She ran her A/B test in week nine. Image search clicks rose 3.4%, and it shipped. She got a return offer and joined full time.
The wider effect came after. Her weekly one-chart review worked so well that I wrote it up as a two-page guide. Our team of 14 now uses it for every intern and new hire. Average time to first A/B test for new hires fell from 11 weeks to 7.”
Follow-ups they will ask
“What if they had kept struggling?” Say you would narrow the scope, then talk to their manager early.
“How did you influence a team you did not lead?” Use a proof point, a small pilot, then a doc others could adopt.
“Tell me about mentoring someone more senior.” Reverse mentoring counts. A staff engineer learning causal methods from you, for one.
“How do you give hard feedback?” Private, specific, soon, and tied to their goal.
Red flags to avoid
Taking over the work and calling it mentoring.
No outcome for the mentee. Growth needs a before and after.
Influence stories that are really about being loud in meetings.
Say this out loud: “I asked her what win she needed by week six, then helped her find the fastest path there. I did weekly one-chart reviews and paired on serving. She shipped a 3.4% lift. I turned the review format into a team guide. Time to first test fell from 11 weeks to 7.”
9. A problem with no labels Hard
What they are testing
Can you start when the data is not ready? Real problems rarely come with clean labels.
Do you know the toolkit, and when each tool fits?
Do you spend a labeling budget wisely?
Do you check that a proxy label tracks the real goal?
How to structure it
Walk through the options in order of cost. Then say which you chose and why.
Proxies. Find a signal you already log that tracks the true label. Check how well it agrees on a small hand-labeled set.
Weak supervision (many noisy rules, combined into labels). Write rules from domain experts. A label model learns how much to trust each rule and combines them.
Active learning. Train a first model, then send humans the examples it is least sure about. Each label teaches the model more.
Human labeling budget. Decide how many labels you can afford. Spend some on a clean test set first. Spend the rest where the model is weakest.
Self-supervised or pretrained models. Learn from raw data or start from a big model, then fine-tune on a few labels.
Always build a gold test set first. Even with no training labels, hand-label 300 to 1,000 examples. Without it, you cannot tell if any method works. Have two people label each one and measure how often they agree.
A strong example answer
“Our support team wanted to route customer emails to the right team automatically. There were 22 teams. We had no labels. Emails were forwarded by hand, and nobody logged where they went.
First, I built a gold set. I worked with two support leads to label 800 emails. They agreed on 87% of them. That told me the best a model could hope for was near 87%.
Then I looked for a proxy. The ticket system logged which team closed each ticket. On my gold set, the closing team matched the right team only 64% of the time. Tickets bounced a lot. So I could not trust it alone.
I used weak supervision. The support leads helped me write 45 rules, like keywords, sender domains, and product names. I added the closing team as one more noisy source. A label model combined them. It labeled 210,000 emails with about 78% accuracy on the gold set.
I trained a text classifier on those labels. Then I ran active learning. Each week, agents labeled the 500 emails the model was least sure about. I had budget for 5,000 labels in total. After four rounds, accuracy on the gold set reached 84%.
We launched it with a confidence cutoff. Emails under the cutoff still went to a human. The model routed 71% of emails on its own. Average first response time fell from 9 hours to 3.”
Follow-ups they will ask
“How did you pick which rules to write?” Start with the biggest classes. Look at errors on the gold set and add rules there.
“Why not just label 50,000 emails?” Give the cost. At $0.40 per label with two labelers, that is $40K and six weeks.
“What about an LLM to label?” Say you would test it as another weak source. Check it against the gold set before trusting it.
“How do you know the proxy does not drift?” Relabel a small sample each month and track agreement.
Red flags to avoid
Jumping straight to a fancy model with no plan for labels.
Trusting a proxy with no check on how well it agrees with truth.
No gold test set. Then every number is a guess.
Say this out loud: “First I hand-label a gold set to measure anything at all. Then I check cheap proxies against it. Next I combine expert rules with weak supervision. Then I spend the labeling budget on active learning, where the model is least sure.”
10. Questions to ask them Easy
What they are testing
Are you curious about the work, or only about the offer?
Do you think like a scientist who will own impact here?
Did you do your homework on the team and product?
How to structure it
Have two questions ready for each round. Match the question to the person. Never say “No, I think you covered it.” That reads as low interest. Listen to the answer and ask one follow-up. That turns it into a real talk.
Hiring manager
What would success look like for this role in six months? In a year?
What is the team's top metric this half, and how does science move it?
How do projects get picked? Who decides what scientists work on?
What is the hardest open problem the team has not cracked yet?
What did the last person who grew fast on your team do differently?
Peer scientist
What does a normal week look like for you? How much is research versus shipping?
How long does it take to go from an idea to an A/B test?
What tools and compute do you have? Is getting GPUs a fight?
How are offline and online results reviewed? Is there a launch review?
What is a project you are proud of here, and what made it hard?
Engineer
How do scientists and engineers split work on a model launch?
What does the path to production look like? Who owns serving and on-call?
What are the latency and cost limits on models in this product?
What is one thing scientists do that makes your job harder?
Skip-level or director
Where do you see this team in two years? What bets are you making?
How does leadership judge the value of science work here?
What is the biggest risk to the team's plans right now?
How do scientists grow to senior and principal levels here?
A strong example answer
Here is how a good exchange with a hiring manager might sound.
“Thanks, I do have a couple. I read your team's blog post on the new feed ranker. You mentioned position bias was still an open issue. Is that the hardest unsolved problem on the team, or is there a bigger one?
[The manager says the bigger issue is cold start for new creators.]
That is interesting. In my last role, I worked on cold start for new stores. We used content features plus an explore budget of about 2% of traffic. How do you handle the explore side today? Is there a set budget, or does it vary by test?
[The manager explains they have no explore budget yet.]
Got it. My last question is about success. If I joined, what would you want to see from me at six months that would make you glad you hired me?”
This works for three reasons. It shows homework. It links to past work without a speech. And the last question tells the candidate exactly what the bar is.
Follow-ups they will ask
“Why our team?” Tie it to the product, the problem, and how you like to work. Avoid pay and perks.
“What are you looking for in your next role?” Name two things this role offers. Ownership of a model in production, for one.
“Do you have other offers?” Be honest and brief. Save the details for the recruiter.
Red flags to avoid
Asking about pay, vacation, or remote rules in a technical round. Save those for the recruiter.
Asking things the job post or website already answers.
The same question to every interviewer. They compare notes.
Say this out loud: “What is the hardest open problem on the team right now? And if I joined, what would you want to see from me at six months?”
Recap
Use STAR. Short setup, long action, and end with a number.
Say “I”. Name what others did, then say what you did.
Build six to eight stories and map each one to several prompts.
For a deep dive, cover problem, why it mattered, choices you rejected, and online results.
For the talk, one project, 40 minutes, a limits slide, and practice out loud three times.
To critique a paper, start with a strength, run the checklist, and end with the test you would run.
With no labels, build a gold set first. Then try proxies, weak supervision and active learning.
Ask two real questions in every round, matched to who you are talking to.