Part III Bank 7 Stories

Research Depth and Behavioral

The deep dive, the talk, and the stories. This is where they decide if they trust you with a real problem.

Every Applied Science loop has a part about you. Sometimes it is a full round. Often it is the last ten minutes of every round. The questions sound soft. The grading is not. Interviewers write down what you did, what it changed, and how you think when things go wrong.

This page covers the ten questions you will hear most. Each one has what they are testing, a shape for your answer, a full example, the follow-ups, and the traps. The examples are fictional. Build your own from your real work, then practice them out loud.

Contents

  1. The STAR shape
  2. Walk me through your best project
  3. The research presentation
  4. Picking problems that matter
  5. A failed experiment
  6. Disagreeing with a PM or engineer
  7. Research to production
  8. Critique a paper on the spot
  9. Mentoring and influence
  10. A problem with no labels
  11. Questions to ask them

The STAR shape

Most story questions want the same four parts. Use them in order. Keep the setup short and the action long.

The time split

A good story runs two to three minutes. Spend about 10% on Situation, 10% on Task, 60% on Action and 20% on Result. Most people flip this. They spend a minute on context and rush the part that gets scored.

Say “I”, not “we”

The interviewer is hiring you, not your team. Say “we” once to set the scene. Then say “I” for what you did. If a teammate did a key part, name it. Then say what you did next to it.

End with a number

“It went well” scores nothing. “Click rate rose 2.1% in a two-week A/B test” scores well. If you have no business number, use a model number, a time saved, or a cost cut. If the result was bad, say the number anyway.

One story fits many prompts

You do not need forty stories. You need six to eight strong ones. Each can answer several prompts if you change the angle. Here is one set mapped to common prompts.

Write a story bank. For each story, write one line per STAR part and the one number. Then list which prompts it fits. Read the bank the night before. Do not memorize scripts. Memorize the beats.

1. Walk me through your best project Medium

What they are testing

How to structure it

  1. Problem. One sentence on the user or business problem. Not the model.
  2. Why it mattered. The metric it moved and how big the prize was.
  3. Approach. The framing, the data, the model, and the loss. Keep it to the key choices.
  4. Alternatives you rejected. Two options and why you said no. This shows judgment.
  5. Results. Offline metrics first, then the online A/B test. Give both numbers.
  6. What you would do differently. One honest change. It shows you still think about it.

Aim for four to five minutes on the first pass. Then let them drill. They will pick one part and go deep. Be ready to go three levels down on any choice.

A strong example answer

“At a food delivery app, our search ranking was a gradient boosted tree from three years back. Users who searched and did not order were 41% of search sessions. Each point of search conversion was worth about $6M a year. I owned the new ranker end to end.

I framed it as learning to rank with order as the label. Clicks were noisy, so I used them only as a second task. I built a two-tower model for candidate scoring, then a small cross-attention model for the top 200. The towers used query text, past orders, time of day and store load.

I rejected two ideas. A pure LLM reranker scored well offline but cost 90 ms at p99. Our budget was 30 ms. I also rejected training on clicks alone. In a pilot it pushed cheap fast food that people clicked but rarely ordered.

Offline, NDCG@10 rose from 0.412 to 0.447. Online, a three-week test on 10% of traffic lifted search conversion 1.8%. Order value stayed flat. Latency rose 4 ms at p99. We shipped to everyone.

If I did it again, I would fix position bias earlier. I added inverse propensity weights late. That alone added 0.4% in a later test. I should have checked for it in week one.”

Follow-ups they will ask

Red flags to avoid

Say this out loud: “The problem was lost orders from search. It was worth about $6M a year per point. I owned the ranker. I picked a two-tower model plus a small reranker and rejected an LLM reranker on latency. Offline NDCG rose 8%. Online conversion rose 1.8%. I would fix position bias sooner next time.”

2. The research presentation Hard

What they are testing

How to structure it

Pick one project, two at most. Depth beats breadth. Pick work that shipped or that you can tie to a real problem. A talk on three papers in 40 minutes feels like a list. A talk on one problem feels like a story.

Use this rule for the room. The first ten minutes are for everyone. The middle is for the experts. The last five minutes are for everyone again.

Slide-by-slide outline for a 40-minute talk

This plan leaves 15 minutes for questions in a 55-minute slot. Aim for about 20 slides. Timings are cumulative.

  1. Title slide (0:00 to 0:30). Your name, the talk title, and one line on the result. Do not read it.
  2. About me (0:30 to 1:30). Three bullets on your path. Skip a long bio.
  3. The problem in plain words (1:30 to 4:00). One picture of a user hitting the problem. No math yet.
  4. Why it matters (4:00 to 6:00). The business metric and the size of the prize. One number in big type.
  5. Why it is hard (6:00 to 8:30). Two or three real obstacles. Scale, noise, latency, no labels.
  6. What others tried (8:30 to 10:30). Prior work and past team attempts. Say where each falls short.
  7. The key idea (10:30 to 13:00). Your main insight in one sentence and one diagram. This is the slide they remember.
  8. Problem framing (13:00 to 15:00). Inputs, outputs, label, and loss. Now the math starts.
  9. Data (15:00 to 17:00). Sources, size, how you cleaned it, and the train and test split.
  10. Method part 1 (17:00 to 20:00). The model design with one clean figure.
  11. Method part 2 (20:00 to 22:30). The trick that made it work. Show the one equation that matters.
  12. Offline results (22:30 to 25:00). Your method against strong baselines. Error bars on every number.
  13. Ablations (25:00 to 27:00). Remove each part and show the drop. This proves each piece earns its place.
  14. Online results (27:00 to 29:30). The A/B test. Primary metric, guardrails, duration, and traffic share.
  15. Error analysis (29:30 to 31:30). Two real failure cases with examples. Say why they fail.
  16. Production story (31:30 to 33:30). Latency, cost, monitoring, and what you had to simplify.
  17. Limits (33:30 to 35:00). What it does not do. Say it before they ask.
  18. What I would do next (35:00 to 37:00). Two or three concrete next steps.
  19. Other work (37:00 to 38:30). One slide with two short blurbs. Shows range without stealing time.
  20. Takeaways (38:30 to 40:00). Three bullets. Problem, idea, impact. Then stop and invite questions.
Practice plan. Run the talk out loud at least three times. Do one run for a non-expert friend and one for a peer who will attack it. Time every run. If you run long, cut slides, do not talk faster.

The 8 hardest audience questions

  1. “Is your baseline strong enough?” Name the strongest baseline you ran and how you tuned it. Say you gave it the same tuning budget as your method. If you skipped a known baseline, admit it and say what you expect.
  2. “How do I know this is not just more parameters?” Point to a size-matched baseline. Or show the ablation that removes your idea but keeps the size.
  3. “The online gain is small. Is it real?” Give the confidence interval and test length. Mention a pre-test A/A check or a holdback that confirmed it.
  4. “Could there be leakage in your split?” Explain the split by time or by user. Name one leak you found and fixed. That builds trust.
  5. “Why not just use an LLM?” Give the cost and latency numbers. Say if you tested one. If not, say how you would compare.
  6. “Does it work for new users or rare items?” Show a sliced result. If it is weaker there, say so and say what you would try.
  7. “What would you do with ten times the data?” Say which part of the method would gain most. Mention a scaling curve if you have one.
  8. “What is the weakest part of this work?” Pick a real weakness. Never say “nothing”. Then say how you would fix it.

When you do not know, say so. Then reason out loud for 20 seconds. “I did not test that. My guess is X, because of Y. I would check it with Z.” That scores better than a bluff.

A strong example answer

Here is how one candidate opened a talk on fraud detection, in the first two minutes.

“Last year, fake sellers cost our marketplace about $14M in refunds. Our rules caught 60% of them, but only after the first bad order. I want to show you how I caught 85% of them before their first sale.

The hard part was labels. We only learn a seller is fake weeks later, when refunds pile up. And fakers change tactics every month. So a model trained on last quarter's fraud goes stale fast.

My key idea was to stop modeling the seller alone. I built a graph of sellers, devices, bank accounts and addresses. Fake sellers share these things far more than real ones do. A graph neural network learned those patterns from 2.3M sellers.

Offline, recall at 1% false positive rate rose from 0.58 to 0.81. In a six-week shadow run and then a live test, we blocked 85% of fake sellers before their first sale. False blocks on real sellers stayed under 0.3%. Refund losses fell by about $9M a year.

Next, I will show how I built the graph. Then why the obvious model failed, and what still breaks. Please stop me any time.”

Follow-ups they will ask

Red flags to avoid

Say this out loud: “I will spend ten minutes on the problem and why it matters. Then twenty on the method and results. The rest goes to limits and next steps. Please stop me any time.”

3. Picking problems that matter Medium

What they are testing

How to structure it

  1. The choice. Name the options you had. Show you did not just take the first one.
  2. The metric. Say which top-line metric each option could move.
  3. The sizing. Show the rough math. Reach times lift times value, then divided by effort.
  4. The no. Say what you turned down and how you said it.
  5. The result. Show the bet paid off, with a number.
Quick sizing formula. Impact ≈ users touched × expected lift × value per unit × chance it works. Divide by weeks of effort. Compare options on this one number. It does not need to be exact. It needs to be roughly right and written down.

A strong example answer

“I joined a streaming app's personalization team as the second scientist. My manager handed me three ideas from the roadmap. One was a new video embedding model. One was better ranking for the home page top row. One was a churn model for the retention team.

The embedding model was the most fun. But I sized all three in a doc first. The top row touched 100% of daily users. A 1% gain in plays from that row was worth about 0.3% in watch time. The embedding model only fed a side shelf that 12% of users scrolled to. The churn model was useful, but the retention team had no lever to act on its scores yet.

My math said the top row was worth about 4x the embedding work for similar effort. I shared the doc with my manager and the PM. I proposed we do the top row first and park the embedding model.

The PM had promised the embedding work to a director. So I offered a two-week spike to test the embedding idea on the side shelf. The spike showed a 0.2% gain. That settled it.

The top row work shipped in one quarter. It lifted watch time 0.6% in a four-week test. That was the team's biggest win that half.”

Follow-ups they will ask

Red flags to avoid

Say this out loud: “Before I build, I size it. Users touched, times expected lift, times value, divided by effort. I write it down and share it. Then I say no to the smaller bets. But I offer a cheap test so the other side sees the data.”

4. A failed experiment Medium

What they are testing

How to structure it

  1. The bet. What you expected and why it was a reasonable bet.
  2. The failure. What happened, with numbers. Do not soften it.
  3. The diagnosis. How you found the root cause. This is the most scored part.
  4. The response. What you did next. Fix, pivot, or stop.
  5. The lesson. One habit you changed, with proof you still use it.

Pick a real failure that cost something. A fake failure, like “I work too hard”, loses trust fast. But pick one where you were not reckless.

A strong example answer

“I built a new ad click model for a news site. It had a transformer over each user's last 100 page views. Offline, log loss fell 3.2%. That is big for ads. I was sure it would win.

The A/B test ran two weeks on 5% of traffic. Revenue per thousand views dropped 1.4%. Click rate was flat. I was stunned.

I dug in for three days. First I checked the logging. It was fine. Then I sliced by user type. Logged-out users made up 55% of traffic. They had almost no page history. For them the new model was worse than the old one. In the offline data, logged-out users were only 30% of rows. The training log had dropped many of their events.

So my offline test set did not match real traffic. I fixed the data pipeline to keep logged-out events. Then I added a fallback path for users with short history. The second test lifted revenue 0.9%.

The lesson changed how I work. I now check that the offline test set matches live traffic on the top three slices before any A/B test. I added that check to our launch template. It caught a similar gap on a teammate's model two months later.”

Follow-ups they will ask

Red flags to avoid

Say this out loud: “Offline it won by 3%. Online it lost 1.4% of revenue. I sliced the results and found the test set did not match live traffic. I fixed the pipeline and the next test won. Now I check slice match before every launch.”

5. Disagreeing with a PM or engineer Medium

What they are testing

How to structure it

  1. The stakes. What the disagreement was about and why it mattered.
  2. Their view. State it fairly. Show you understood why they held it.
  3. Your view. State it with the evidence you had.
  4. How you resolved it. Data, a test, a third person, or a compromise.
  5. The outcome. Who was right, the number, and how the bond held up.

A strong example answer

“I worked on a shopping app's recommendations. Our PM wanted to optimize the feed for clicks. Clicks were the team's goal for the quarter. I thought clicks would push clickbait items and hurt purchases.

I first asked the PM why clicks. The goal came from her director. Clicks were fast to measure, and purchases took a week to show up. That was fair. A purchase goal would slow every test.

So I did not argue in the meeting. I pulled data from our last six tests. In two of them, clicks rose but purchases fell. One had cost about $400K in a month before anyone noticed. I shared a one-page note with the PM first, not with the whole team.

I proposed a middle path. Optimize a blend of clicks and add-to-cart, which shows up within a day. Keep purchases as a guardrail that can block a launch. She agreed to try it on the next test.

The blended model lifted clicks 1.1%, less than the pure click model's 2.4%. But it lifted purchases 0.8%. The click model had dropped them 0.3%. The PM took the blend to her director, who made it the new team goal. She and I still work together on goal design for each quarter.”

Follow-ups they will ask

Red flags to avoid

Say this out loud: “I asked why she wanted clicks before I argued. Then I brought data from six past tests in a private note. I offered a blend with purchases as a guardrail. It won on both metrics and became the team goal.”

6. Research to production Hard

What they are testing

How to structure it

Walk through the gap between the lab and the product. Hit these five points. Not every story uses all five. Pick the three that were hardest.

Distillation (a small model copies a big one). You train the big teacher model first. Then you train a small student model to match the teacher's scores, not only the true labels. The student is fast enough to serve and often keeps most of the quality.

A strong example answer

“I built a model to flag toxic comments for a social app. My research model was a 350M parameter text model. It hit 0.93 F1 on our test set. The old keyword filter was at 0.71. But it took 120 ms per comment on a GPU. The budget was 15 ms on CPU, because comments post in real time.

I tried three things. First I cut the input to 128 tokens. Ninety-six percent of comments fit, and F1 dropped only 0.004. Second, I distilled the big model into a 6-layer student with 22M parameters. I trained it on 40M unlabeled comments scored by the teacher. The student hit 0.91 F1. Third, I quantized it to 8-bit. That got p99 latency to 11 ms on CPU.

I owned the launch. I ran a two-week shadow mode with no action taken. Then I ramped 1%, 10%, 50%, 100% over three weeks. I set alerts on the share of comments flagged, the score mix, and appeal rate. I wrote the rollback runbook and took on-call for the first month.

In week five, the flag rate jumped 30% overnight. A new slang term had gone viral. My alert caught it in two hours. I added 2,000 labeled examples and retrained in four days. Reports of toxic comments fell 38% after the full launch. Cost was $4K a month, not the $60K a GPU model would have needed.”

Follow-ups they will ask

Red flags to avoid

Say this out loud: “The research model was 120 ms on GPU. The budget was 15 ms on CPU. I cut input length, distilled to a small student, and quantized it. It kept 0.91 of 0.93 F1 at 11 ms. I owned the ramp, the alerts and on-call. An alert caught a drift problem in two hours.”

7. Critique a paper on the spot Hard

What they are testing

How to structure it

They may hand you a paper, show you an abstract, or ask about a famous paper. Read the abstract and the main results first. Then walk this checklist out loud. Start with one honest strength.

The critique checklist

  1. Claim. What exactly do they claim? Is it about accuracy, speed, or a new ability? Is the claim bigger than the evidence?
  2. Baselines. Are the baselines strong and current? Were they tuned as hard as the new method? A weak baseline makes any idea look good.
  3. Ablations. Did they remove each new part to show it helps? Without ablations, you cannot tell which idea did the work.
  4. Data leakage. Could test data have leaked into training? Look at the split. Random splits on time series or users leak. Pretraining data may contain the test set.
  5. Metrics. Do the metrics match the real goal? Is accuracy hiding class imbalance? Did they report only the metric that looks best?
  6. Statistical significance. Do they report error bars, seeds, or a test? A 0.3 point gain from one run may be noise.
  7. Compute fairness. Did the new method get more parameters, more data, or more training time? Compare at equal compute.
  8. Reproducibility. Is code or data released? Are the hyperparameters and training details enough to rebuild it?
End with a fix. After the flaws, say what experiment would settle the biggest doubt. “I would want a size-matched baseline and three seeds.” That turns a critic into a collaborator.

A strong example answer

The interviewer showed an abstract. It claimed a new attention variant beat a standard transformer by 2.1 points on a reading benchmark.

“First, the idea is clean. It cuts attention cost from quadratic to linear. If it holds, that matters a lot for long documents.

Now the doubts. The claim is about accuracy, but the selling point is cost. I want to see a speed and memory chart, not only accuracy.

On baselines, they compare to a base transformer from 2019. I would want a modern baseline with the same training recipe. Their model has 180M parameters and the baseline has 110M. That gap alone could explain two points. So compute fairness is my top worry.

There is one ablation table, and it removes only the gating part. I want an ablation that keeps the size but drops the new attention.

They report one run per model. On this benchmark, seeds can swing a full point. So I want five seeds and a confidence interval.

On leakage, they pretrained on a web crawl. The benchmark's passages come from the web. They should check for overlap.

Code is promised but not out, so I cannot check the details.

Overall, the idea is promising, but the evidence is weak. The one experiment I would run first is a size-matched comparison with five seeds. If the gain holds, I would believe it.”

Follow-ups they will ask

Red flags to avoid

Say this out loud: “The idea is promising. My top doubt is compute fairness, since their model is 60% larger. I also want ablations, five seeds, and a leakage check. The first thing I would run is a size-matched baseline.”

8. Mentoring and influence Medium

What they are testing

How to structure it

  1. The person or group. Who they were and where they were stuck.
  2. Your approach. How you helped. Teaching, pairing, reviews, or giving them room.
  3. Their growth. What they could do after that they could not before.
  4. The wider effect. A habit, tool, or process that spread past one person.

Two good stories cover most prompts. One about growing a single person. One about changing how a team works.

A strong example answer

“I mentored a summer intern named Priya. She was a strong PhD student in vision but new to industry. Her project was to improve product image search on our shopping site.

In week two, she was stuck. She wanted to train a big model from scratch and had spent a week on data loading. I did not take over. I asked her what result she needed by week six to call the summer a win. She said a launched A/B test. Then I asked what the fastest path to a test was.

She landed on fine-tuning an existing image model. I set up a weekly one-hour review where she showed one chart and one question. I paired with her for two days on our serving stack. I also had her present her plan to the PM herself. That built her confidence and got the PM's support early.

She ran her A/B test in week nine. Image search clicks rose 3.4%, and it shipped. She got a return offer and joined full time.

The wider effect came after. Her weekly one-chart review worked so well that I wrote it up as a two-page guide. Our team of 14 now uses it for every intern and new hire. Average time to first A/B test for new hires fell from 11 weeks to 7.”

Follow-ups they will ask

Red flags to avoid

Say this out loud: “I asked her what win she needed by week six, then helped her find the fastest path there. I did weekly one-chart reviews and paired on serving. She shipped a 3.4% lift. I turned the review format into a team guide. Time to first test fell from 11 weeks to 7.”

9. A problem with no labels Hard

What they are testing

How to structure it

Walk through the options in order of cost. Then say which you chose and why.

Always build a gold test set first. Even with no training labels, hand-label 300 to 1,000 examples. Without it, you cannot tell if any method works. Have two people label each one and measure how often they agree.

A strong example answer

“Our support team wanted to route customer emails to the right team automatically. There were 22 teams. We had no labels. Emails were forwarded by hand, and nobody logged where they went.

First, I built a gold set. I worked with two support leads to label 800 emails. They agreed on 87% of them. That told me the best a model could hope for was near 87%.

Then I looked for a proxy. The ticket system logged which team closed each ticket. On my gold set, the closing team matched the right team only 64% of the time. Tickets bounced a lot. So I could not trust it alone.

I used weak supervision. The support leads helped me write 45 rules, like keywords, sender domains, and product names. I added the closing team as one more noisy source. A label model combined them. It labeled 210,000 emails with about 78% accuracy on the gold set.

I trained a text classifier on those labels. Then I ran active learning. Each week, agents labeled the 500 emails the model was least sure about. I had budget for 5,000 labels in total. After four rounds, accuracy on the gold set reached 84%.

We launched it with a confidence cutoff. Emails under the cutoff still went to a human. The model routed 71% of emails on its own. Average first response time fell from 9 hours to 3.”

Follow-ups they will ask

Red flags to avoid

Say this out loud: “First I hand-label a gold set to measure anything at all. Then I check cheap proxies against it. Next I combine expert rules with weak supervision. Then I spend the labeling budget on active learning, where the model is least sure.”

10. Questions to ask them Easy

What they are testing

How to structure it

Have two questions ready for each round. Match the question to the person. Never say “No, I think you covered it.” That reads as low interest. Listen to the answer and ask one follow-up. That turns it into a real talk.

Hiring manager

Peer scientist

Engineer

Skip-level or director

A strong example answer

Here is how a good exchange with a hiring manager might sound.

“Thanks, I do have a couple. I read your team's blog post on the new feed ranker. You mentioned position bias was still an open issue. Is that the hardest unsolved problem on the team, or is there a bigger one?

[The manager says the bigger issue is cold start for new creators.]

That is interesting. In my last role, I worked on cold start for new stores. We used content features plus an explore budget of about 2% of traffic. How do you handle the explore side today? Is there a set budget, or does it vary by test?

[The manager explains they have no explore budget yet.]

Got it. My last question is about success. If I joined, what would you want to see from me at six months that would make you glad you hired me?”

This works for three reasons. It shows homework. It links to past work without a speech. And the last question tells the candidate exactly what the bar is.

Follow-ups they will ask

Red flags to avoid

Say this out loud: “What is the hardest open problem on the team right now? And if I joined, what would you want to see from me at six months?”

Recap

← 6 — Experimentation and Causal Inference T1 — Learning Theory →