A science manager runs a portfolio of bets. The job is to pick good bets, kill bad ones early, and prove the winners paid off.
The first manager page covered people. This page covers the machine. It explains how a strong manager plans, reviews, measures and scales science work. Each section gives principles, a template you can reuse, a short story with numbers, and the ways it goes wrong. The last section has ten interview questions with strong answers.
The viewpoint is a manager of 8 to 40 scientists at a large tech company. The team ships models into real products. Leaders judge it on business results, not papers.
A science team is not a feature team. Most of its ideas will fail. That is normal. The manager's job is to make the whole set of bets pay off, even when single bets lose.
Research portfolio. The full set of projects a team funds at once. Each one has a size, a time horizon, a chance of success and a payoff. You manage the mix, not just the parts.
The 70/20/10 split
Most mature teams split effort into three buckets. The numbers are a starting point, not a law.
70% core. Improve models that already ship. Known metric, known path, high odds. Payoff lands this quarter or next.
20% adjacent. New features or new surfaces for proven methods. Medium odds. Payoff in two to four quarters.
10% frontier. New methods that could change the product. Low odds, large payoff. Horizon of a year or more.
The right split depends on the team's stage. A new team with no shipped model may run 90/10/0. It must earn trust first. A team that owns a mature ranker may run 60/25/15. Its core gains are getting smaller each half.
A 70/20/10 portfolio. Odds fall and horizons grow as you move right.
Horizons
Tag every project with a horizon. It sets how you review it and when you expect a result.
H1, this half. Review on online metrics. Weekly check-ins.
H2, next year. Review on milestones. Each milestone answers a question.
H3, two years or more. Review on learning. Ask what you now know that you did not know last quarter.
Expected value of a project
You cannot rank projects without a number. A rough one is fine. Use this simple form.
EV = P(success) × payoff − cost. Then divide by scientist-months to get value per month. Rank on that. A good guess beats no guess. The act of writing the number exposes weak plans.
P(success). Base it on history. How often did similar bets on this team ship? Most core bets land 50 to 70% of the time.
Payoff. Annual business value if it ships. Use the same metric-to-dollar rate for every project.
Cost. Scientist-months, engineer-months, and compute. Include the cost to keep it running.
Option value. Some bets unlock later bets. A new embedding service may enable five projects. Add that as a note, not a fake number.
Bet sizing
Fund a frontier idea small first. One scientist for six weeks is a cheap option.
Double down only after a gate passes. Never staff a bet to full size on a slide.
Keep no more than one large, risky bet per 10 scientists. Two at once can sink a half.
Name an owner for every bet. A bet with three part-time owners has none.
Scenario: the ads ranking team
Priya runs 14 scientists on ads ranking. Her list has nine candidate projects for the half. She scores each one.
Feature refresh on the main model. P = 0.7, payoff $12M a year, cost 6 scientist-months. EV per month is about $1.4M.
Multi-task model for clicks and conversions. P = 0.4, payoff $40M, cost 18 months. EV per month is about $0.9M.
LLM-based ad text features. P = 0.15, payoff $80M, cost 6 months. EV per month is about $2M, with wide error bars.
She funds the feature refresh fully. She funds the multi-task model with a gate at week 8. She funds the LLM bet with one scientist for six weeks. The LLM bet fails its first gate. She shifts that scientist to the multi-task work, which passes its gate. The half ends with $14M in shipped gains and one clean failure.
Failure modes
The 100% core trap. Leaders push for this quarter's number. The frontier slice slowly goes to zero. In two years the team has nothing new.
Pet projects. A senior person's idea never gets a real EV estimate. It eats 20% of the team for a year.
Fake precision. EV numbers are treated as facts. Report ranges, not single values.
Too many small bets. Twenty bets for 12 people means nothing gets enough depth to work.
Say this out loud: “I run the team as a portfolio. About 70% goes to core models, 20% to adjacent bets, and 10% to frontier work. I rank projects by expected value per scientist-month. I fund risky bets small and grow them only after a gate.”
Planning under uncertainty
Science plans fail when they look like engineering plans. You cannot promise “model done by June” when nobody knows if the model will work. You can promise to answer key questions by set dates.
Milestones as questions answered
Write each milestone as a question with a yes or no answer. The answer decides the next step.
Weak milestone. “Train the graph model.”
Strong milestone. “Does the graph model beat the baseline by 1% NDCG offline on last month's data? Answer by March 15.”
A question milestone can fail and still be a success. A clear “no” saves months. Celebrate it.
Time-boxed spikes
Spike. A short, fixed-length test of one risky idea. Usually one or two scientists for two to six weeks. The goal is learning, not shipping.
Fix the end date before you start. Do not extend it more than once.
Write the question and the bar for success on day one.
Allow rough code. A spike is throwaway work.
End with a one-page memo. It says what you learned and what you recommend.
Kill criteria up front
Decide how you will stop before you start. Once people have spent months, they always find a reason to keep going.
Kill criteria template. “We will stop this project if any of these are true by DATE. (1) Offline gain is below X. (2) Latency at p99 exceeds Y ms. (3) Online test shows no lift above Z with 80% power. (4) Data we need is not approved by legal.” Sign it with the product partner.
Stage gates
Large bets move through gates. Each gate has a clear bar and a small review. Staffing grows only after a gate passes.
Stage gates. Staffing grows only after a bet passes a gate. Most kills should happen at the first two gates.
Gate 1, idea to spike. Is the problem worth solving? Is there a plausible path?
Gate 2, spike to prototype. Does a rough version show signal offline?
Gate 3, prototype to online test. Does it meet latency, cost and privacy limits? Is the offline gain big enough to detect online?
Gate 4, test to launch. Did the pre-registered metric move? Are guardrails clean?
Scenario: the fraud model rewrite
Marcus leads 6 scientists on payment fraud. His director wants a new sequence model by Q3. Marcus does not promise a date for the model. He promises three answers instead.
By week 4: at the same false positive rate, does it catch 5% more fraud?
By week 10: can it score in under 20 ms at p99?
By week 16: does a 5% shadow test confirm the offline gain?
Week 4 says yes, with a 7% gain. Week 10 says no, at 45 ms. The team spends three weeks on distillation and gets to 18 ms with a 6% gain. Week 16 confirms it. They launch in week 20. The director gets a later date than hoped but no surprises.
Failure modes
Zombie projects. No kill criteria, so a project lives on at 30% effort for a year.
Endless spikes. A four-week spike becomes twelve weeks. The time box meant nothing.
Gate theater. Every gate passes because the review is a formality.
Dates on unknowns. The manager promises a launch date before any signal. Trust breaks when it slips.
Say this out loud: “I plan science as questions with dates. Each milestone has a yes or no answer. We write kill criteria before we start and sign them with product. Staffing grows only when a bet passes a gate.”
The experiment review culture
A team's results are only as good as its review habits. One bad launch based on a false win can cost more than a year of real gains. Rigor is a team habit, not a personal trait.
Pre-registration of metrics
Before an A/B test starts, the owner writes down what will count as success. This stops people from picking the best-looking metric after the fact.
Pre-registration template.
Hypothesis in one sentence.
Primary metric, with the minimum effect worth shipping.
Two to four secondary metrics.
Guardrail metrics and their limits. Latency, revenue, complaints.
Sample size, power, and planned run length.
Segments you will look at, named now.
The launch decision rule.
The review checklist
Every launch review runs the same list. A senior scientist who did not work on the project runs it.
Sample ratio. Is the split what you planned? A mismatch means a bug.
Pre-period check. Did the groups look the same before the test? An A/A test helps.
Novelty. Does the effect fade over the weeks? Plot it by day.
Interference. Can treated users affect control users? Common in markets and social apps.
Multiple tests. How many metrics and segments were checked? Adjust for it.
Offline to online. Does the online gain match the offline gain in sign and rough size?
Guardrails. Any guardrail outside its limit blocks the launch.
Long-term risk. Could this help now and hurt in six months? Plan a holdout.
Launch criteria
Primary metric up with a 95% interval above zero.
No guardrail worse than its limit.
Cost to serve within budget, or a business case for the extra cost.
A long-term holdout of 1 to 5% for big changes.
A rollback plan that someone has tested.
Avoiding p-hacking in teams
P-hacking is rarely fraud. It is usually tired people under pressure trying one more cut. Build a system that makes it hard.
Lock the analysis plan before you look at results.
Show every metric in the review, not just the ones that moved.
Do not reward people for wins alone. Reward clean tests, even null ones.
Track the team's “win rate”. If 80% of tests win, your bar is too low.
Rerun surprising wins. A second test costs less than a bad launch.
The peeking trap. A scientist checks the test daily and stops it the first day p < 0.05. This can push the false win rate from 5% to over 25%. Use a fixed run length or a sequential test built for peeking.
Scenario: the search launch that was not
A search team ran a new ranker for two weeks. The owner reported a 1.2% lift in clicks. The reviewer checked the pre-registration. The primary metric was purchases, not clicks. Purchases were flat at +0.1%, with an interval of −0.4% to +0.6%. The reviewer also found 14 segments in the deck. Only the three best were shown on the first slide.
The manager did not blame the owner. She changed the template. Every deck now opens with the pre-registered primary metric. The team reran the test with a fix for a known position bias. Purchases rose 0.5% and the model shipped.
Failure modes
Review by status. Junior reviewers do not challenge a senior owner. Rotate reviewers and make it their job to push.
Metric switching. The primary metric is quietly changed after results come in.
Too slow. Reviews take three weeks. People start to route around them. Keep a light path for small changes.
No memory. Nobody logs past results. The team repeats the same failed idea every year.
Say this out loud: “Every test is pre-registered with one primary metric and guardrails. A scientist outside the project runs a fixed review checklist. We reward clean tests, including nulls. If our win rate gets too high, I treat that as a warning sign.”
Measuring and communicating impact
Science that leaders cannot see does not get funded. You must turn model gains into business terms. You must do it in a way that holds up when finance checks it.
From offline gains to business dollars
Build a chain from the model metric to money. Each link should rest on a past test, not a guess.
Offline metric. AUC up 0.8%, or NDCG up 2%.
Online metric. A/B test shows conversions up 0.6%.
Business metric. 0.6% of conversions is 120,000 more orders a month.
Dollars. At $8 margin per order, that is about $11.5M a year.
Agree on the conversion rates with finance once. Then use the same rates in every report. This stops each team from inflating its own math.
Attribution
Many teams touch the same metric. If every team claims the full gain, the total is three times the real number. Leaders notice.
Use holdouts. A long-term holdout from all launches gives the true combined gain.
Claim the test result only. Your share is the lift your test measured, not the whole metric trend.
Split shared wins. If science and engineering both made a launch work, say so. Credit grows when you share it.
Discount for decay. Many gains shrink over time. Report the holdout number after three months.
The impact doc
Impact doc template (one page).
What shipped. One sentence a VP can read.
Result. Primary metric lift with interval. Dollar value with the rate used.
How we know. Test design, length, traffic, and holdout.
Cost. People, compute, and serving cost.
What we learned. Including what did not work.
Who did it. Name every person, across all teams.
Executive updates
Lead with the outcome. “We added $11M a year in margin.” Then the how.
Show the trend over quarters, not just this month.
Include one bet that failed and what it taught you. It builds trust.
Ask for one thing. Headcount, compute, or a decision.
Keep it to five bullets or one slide. Put details in an appendix.
Scenario: the recommendations team
Lena's team shipped four model changes in a half. Each test showed lift. Added up, they claimed 6.2% more watch time. The long-term holdout showed 3.9%. Some gains overlapped and one faded.
Lena reported 3.9% to her VP. She explained the gap in one line. Finance trusted her numbers from then on. The next planning cycle, her team got two of the three extra heads it asked for. A peer team that reported raw sums got none.
Failure modes
Offline wins as impact. “AUC up 2%” is not a business result.
Double counting. Summing test lifts that overlap.
Hiding cost. A $5M gain that costs $4M in GPUs is a $1M gain.
Invisible work. Platform and data work gets no credit. Measure it by what it enabled.
Say this out loud: “I tie every launch to dollars using rates we agreed with finance. I report holdout numbers, not summed test lifts. Every impact doc includes cost and names everyone who helped.”
Roadmaps and OKRs for science
A science roadmap must be honest about risk. It must also show leaders what they get. OKRs (goals plus measurable key results) help when written well.
Outcome vs output
Output. What you build. “Ship a transformer ranker.”
Outcome. What changes for users or the business. “Raise search purchase rate 1%.”
Write key results as outcomes when you can. That leaves room to find the best method. Use output key results for platform work or learning goals. Even then, say what the output enables.
How to write science OKRs
OKR template.
Objective. A short, plain goal. “Make search results feel personal.”
KR1, outcome. “Search purchase rate +1.0% in A/B tests by end of Q2.”
KR2, learning. “Decide yes or no on LLM query rewriting by May 1, with a written memo.”
KR3, health. “Cut ranker training cost per run by 30%.”
Set committed KRs at 90% odds. Set stretch KRs at 50% odds. Label them.
Use a learning KR for frontier bets. A clear “no” counts as done.
Keep three to five KRs per team. More means none matter.
Name one owner per KR.
The science roadmap
Now. Committed work for this quarter, with dates.
Next. Likely work for next quarter, with gates.
Later. Bets and themes, with no dates.
Show confidence on each item. A simple high, medium or low works. Update the roadmap every month. Leaders forgive change. They do not forgive surprise.
Scenario: the trust and safety team
A safety team had an OKR that read “Ship v3 classifier.” They shipped it on time. Harmful content views did not drop. The new manager rewrote the OKR as “Cut harmful views per million by 15%.” The team found that half the gap came from new content types the model never saw. Better labeling did more than a new model. Views fell 18% that half.
Failure modes
Output theater. All KRs are “ship X”. The team ships and nothing changes.
Sandbagging. Every KR is hit at 100%. The bar was too low.
Dated frontier work. Putting launch dates on H3 bets. They always slip.
Set and forget. OKRs are written in January and read again in June.
Say this out loud: “I write science OKRs as outcomes, plus one learning goal for risky bets. Committed goals should land 90% of the time. Stretch goals 50%. The roadmap shows now, next and later, with confidence on each.”
Operating rhythm
A good rhythm makes work visible without slowing it down. It gives people regular places to get feedback, share ideas and make decisions.
The core meetings
Weekly science review (60 min). Two or three projects present results. Focus on methods and decisions, not status. Status lives in a written tracker.
Reading group (biweekly, 45 min). One paper, one presenter. Rotate. End with “could we use this?”
Experiment review (weekly, 30 min). Launch decisions using the checklist. Product joins.
Demo day (monthly or quarterly). Short live demos open to product and engineering. Builds energy and cross-team trust.
Quarterly planning. Review the portfolio, rescore EVs, kill or grow bets, set OKRs.
Weekly science review format
Science review template.
A short doc sent 24 hours ahead. Two pages max.
First 10 minutes: silent reading.
The question being answered and the result.
What the presenter plans next and what help they need.
End with a decision or a named next step.
Protect deep work
Keep two or three no-meeting afternoons a week.
Cap recurring meetings at about four hours a week for an IC scientist.
Use written updates for status. Use meetings for debate.
Scenario: a team drowning in syncs
A new manager joined a team of 11. Scientists sat in 9 hours of recurring meetings a week. Three meetings covered the same project status. She replaced them with one written weekly tracker and one science review. Meeting time fell to 4 hours. Experiments launched per scientist rose from 1.1 to 1.7 a month over the next quarter.
Failure modes
Status meetings in disguise. The science review becomes a round of updates.
Reading group decay. It drifts to papers nobody can use. Attendance falls.
Demo day for the boss. Only polished wins are shown. People stop learning from failures.
Planning without data. Quarterly planning ignores last quarter's results.
Say this out loud: “I run a weekly science review and a weekly launch review. I add a reading group, a demo day and quarterly planning. Status goes in writing. Meetings are for decisions.”
Working with engineering and product
Science alone ships nothing. A model reaches users through engineering and product. Most failed science projects fail at this seam, not in the math.
RACI for a model launch
RACI names who is Responsible, Accountable, Consulted and Informed for each task. Write it at the start of any project that crosses teams.
Problem and success metric. Product is accountable. Science is consulted and must agree.
Model design and offline results. Science is accountable and responsible.
Serving, latency and reliability. Engineering is accountable. Science is consulted.
Experiment design and analysis. Science is accountable. Data science or product is consulted.
Launch decision. Product is accountable. Science and engineering must sign off.
Post-launch monitoring. Engineering runs alerts. Science owns model quality and drift.
Handoff vs embedded
Handoff model. Science builds a model and hands it to engineering. Fast for science. Often slow to ship, since engineering must learn the model.
Embedded model. Scientists sit in product teams. Fast to ship. Risk of losing depth and drifting into ad hoc work.
Hybrid. Scientists report to a science manager but sit with a product team. Most large companies land here.
Own the launch, not just the model. A scientist who stays with a model through launch ships twice as often. Make launch ownership part of your team's job, and part of how you rate people.
Infra asks
Science often needs engineering help. A feature store, a new logging field, faster training. These asks compete with product features.
Write every ask as a one-pager with value in dollars.
Bundle asks. One large, clear ask beats ten small ones.
Offer to do part of the work. Scientists can write the first pipeline.
Bring asks to planning, not mid-quarter.
Scenario: the stalled pricing model
A pricing science team built a model with a 4% margin gain offline. It sat for two quarters. Engineering had not planned for it. The new manager wrote a RACI with the engineering lead. She moved one scientist to sit with the serving team for eight weeks. The scientist wrote the feature pipeline. Engineering owned serving. The model launched in nine weeks and delivered a 2.7% margin gain online.
Failure modes
Throw it over the wall. Science hands off a notebook and moves on.
No shared metric. Product wants engagement, science tunes for clicks. Nobody notices until launch.
Infra asks with no value. “We need a feature store” with no business case gets ignored.
Embedded drift. Embedded scientists become analysts who pull dashboards all day.
Say this out loud: “I write a RACI at the start of every cross-team project. Product owns the metric, engineering owns serving, and science owns model quality. My scientists own the launch, not just the model. Infra asks come with a dollar value.”
Build vs buy, and foundation models
Many problems no longer need a custom model. A large language model or a vendor API may get you 80% of the way in a week. The manager must decide when that is enough.
Questions to ask
Is this core to our edge? If the model is the product, build. If it is plumbing, buy or reuse.
Do we have unique data? Unique labels or behavior data favor a custom or fine-tuned model.
What are the latency and cost limits? A large LLM call may cost 100 times a small model per request.
How much does quality matter at the margin? In ads ranking, 0.1% is millions. In internal tagging, good enough is fine.
What are the privacy and legal limits? Some data cannot leave your systems.
The usual path
Prompt a foundation model. Fast baseline in days. Good for labeling, prototypes and low-volume tasks.
Add retrieval or few-shot examples. Often closes much of the gap.
Fine-tune a mid-size model. When volume is high or quality must rise.
Distill to a small custom model. When latency and cost per call matter at scale.
Train from scratch. Rare. Only for core models with huge data and a clear edge.
A cost check
Do the math early. Say a task gets 50 million calls a day. An LLM API at $0.002 per call costs $100,000 a day. That is $36M a year. A distilled model on owned hardware might cost $2M a year. If the quality gap is small, the custom model wins. At 50,000 calls a day, the API costs $36,000 a year. Buy.
Scenario: support ticket routing
A support team wanted a new ticket classifier. The old custom model hit 82% accuracy. A scientist tried a prompted LLM in three days. It hit 88%. Volume was 200,000 tickets a day. The API cost was about $150,000 a year. The team shipped the LLM version in four weeks. They used its labels to train a small model six months later. That model hit 89% at a tenth of the cost.
Failure modes
Not invented here. The team builds a custom model for a solved problem because it is more fun.
LLM for everything. A ranking task that needs 5 ms latency gets an LLM. It never ships.
Ignoring lock-in. A vendor changes price or quality. The team has no fallback.
No eval set. The team cannot compare options because it has no shared test data.
Say this out loud: “I start with the cheapest thing that might work, often a prompted foundation model. I move to fine-tuning or a custom model when volume, latency, cost or edge demand it. I do the cost math in week one.”
Managing compute and data budgets
Compute is now one of the biggest costs on a science team. GPUs can cost more than salaries. A manager who cannot explain the compute bill will lose it.
GPU allocation
Reserve for production. Retraining and serving jobs get guaranteed capacity first.
Quota for research. Give each sub-team a fixed share. Let them trade.
Pool for big bets. Hold 10 to 20% for large runs that pass a gate.
Use it or lose it. Reclaim idle quota each month.
Cost per experiment
Track what each experiment costs. Make it visible. People spend less when they can see the bill.
Tag every job with a project and owner.
Report GPU-hours and dollars per project each week.
Set a cost cap for any single run. Above it needs a lead's sign-off.
Prove ideas at small scale first. Most ideas that fail at 1% scale also fail at full scale.
Data access and privacy
Know which data needs approval before you plan. Approvals can take weeks.
Use the least data you need. Fewer fields means faster review and less risk.
Keep a log of datasets, owners and approved uses.
Delete data on schedule. Retention limits are rules, not hints.
Train scientists on the privacy rules every year.
Scenario: the runaway training bill
A recommendations team spent $1.8M on GPUs in one quarter. The budget was $1.1M. The manager tagged jobs and found 40% of spend came from hyperparameter sweeps at full scale. Two sweeps had run for a week after the owner left on vacation. He set a rule. Sweeps run at 10% data first. Any run over $20,000 needs a lead's approval. Next quarter's spend was $1.05M. The number of launches did not change.
Failure modes
GPU hoarding. Teams hold idle capacity in case they need it.
Invisible cost. Nobody knows what a project costs until finance asks.
Privacy as an afterthought. A model is built on data that legal later blocks.
Starving research. All compute goes to production. No new ideas get tested.
Say this out loud: “I treat compute like headcount. Production gets reserved capacity, research gets quotas, and big bets draw from a gated pool. Every job is tagged, so I know the cost per project. Data approvals go into the plan from day one.”
Publishing, patents and external presence
External work can help hiring, retention and the company's name. It can also leak secrets and distract from the product. A manager needs a clear policy.
When publishing helps the business
Hiring. Strong papers draw strong candidates. This is often the biggest return.
Retention. Many scientists want to publish. A path to do so keeps them.
Quality. Peer review is a hard check on your methods.
Standards. Shaping open benchmarks or tools can pull the field toward your needs.
When it does not
The method is a direct edge over competitors.
The paper would reveal user data or business numbers.
The work is weeks of extra effort with no product link.
A simple policy
Product work comes first. Papers come from work you were doing anyway.
Every paper goes through legal and comms review before submission.
Remove business numbers. Use public datasets for headline results when you can.
File patents before you publish, if the method is worth protecting.
Count papers in reviews only when they tie to product or hiring value.
Patents
File when a method is new, useful, and likely to be copied.
Do not file on every small trick. It costs legal time and money.
Make the invention disclosure easy. A one-page form and a short call.
Scenario: the conference paper debate
A scientist wanted to publish a new ranking loss. It had added 0.7% to revenue. The manager asked legal and the product lead. They agreed the loss was general but the features were not. The team filed a patent on the loss. They wrote the paper using two public datasets. Revenue numbers were left out. The paper was accepted. Two strong candidates later named it as why they applied.
Failure modes
Paper-first culture. Scientists chase venues, not product impact.
Leaks. A paper reveals a secret feature or a revenue number.
No path at all. A team that never publishes loses top hires to one that does.
Patent spam. Filing on everything wastes legal time.
Say this out loud: “We publish when it helps hiring or quality and does not give away our edge. Papers come from product work, not side projects. Legal reviews every paper, and we file patents first when the method is worth protecting.”
Managing risk and responsible AI reviews
Models can harm users, break laws, or embarrass the company. A science manager owns these risks for the team's models. Good process catches them early, when fixes are cheap.
The main risk types
Fairness. The model works worse for some groups of users.
Privacy. The model leaks or misuses personal data.
Safety. The model shows harmful content or gives harmful advice.
Robustness. The model fails on new inputs or under attack.
Legal and policy. The model breaks a rule in some market.
The responsible AI review
Review checklist.
What decisions does the model make about people?
What data trains it? Does it include sensitive fields, or proxies for them?
How does it perform across key groups? Show the numbers.
What happens when it is wrong? Who is hurt, and how badly?
Can a human override it? Can users appeal?
For generative models, what red-team tests ran? What did they find?
What monitoring runs after launch? Who gets the alert?
Make it part of the flow
Run a short risk check at gate 1. Most issues are cheap to fix then.
Run the full review at gate 3, before the online test.
Scale the review to the risk. A tagging model needs one page. A credit model needs a deep review.
Keep a risk register. Log each issue, its owner, and its status.
Scenario: the hiring screen model
A team built a model to rank job applicants for recruiters. Offline accuracy was strong. The fairness check showed women were ranked 12% lower on average at equal outcomes. The cause was a feature tied to past job titles. The team removed it and added a group-wise check to the launch criteria. Accuracy fell 1%. The gap fell to under 2%. The launch went ahead six weeks late but cleared legal review.
Failure modes
Review at the end. Risk checks happen the week before launch. Fixes mean months of delay.
Checkbox reviews. Forms get filled but no one tests anything.
Average-only metrics. The overall number looks fine and hides a group that does badly.
No monitoring. The model drifts after launch and nobody looks.
Say this out loud: “We run a light risk check at the first gate and a full responsible AI review before any online test. We report results by group, not just averages. Every risk has an owner in a register, and monitoring is part of launch.”
Scaling the org
What works for 5 scientists breaks at 15. What works at 15 breaks at 50. The manager must change the structure before it breaks, not after.
Stages of growth
5 to 10 scientists
One manager. Everyone knows every project.
The manager still reviews code and models.
Process is light. A weekly review and a shared tracker.
10 to 25 scientists
Split into two or three sub-teams by problem area.
Name tech leads for each. They guide methods but may not manage people.
Add the first line manager if the team passes about 12 direct reports.
Write things down. Pre-registration, review checklists, impact docs.
25 to 50 scientists
You become a manager of managers. Three to five managers report to you.
Add staff scientists who set technical direction across sub-teams.
Your job shifts to strategy, hiring, and the portfolio.
Key roles
Line manager. Owns people, growth, ratings and hiring for 6 to 10 scientists.
Tech lead. Owns the technical plan for one area. Often a senior scientist.
Staff scientist. Owns direction across teams. Solves the hardest problems. Mentors tech leads.
Science program manager. Runs planning, tracking and cross-team logistics. Worth hiring at about 30 people.
Two ladders, equal respect. Strong scientists must be able to grow without managing. Make the staff and principal path as visible and rewarded as the manager path. Otherwise your best ICs become unhappy managers.
Picking your first managers
Pick people who already grow others. Look for those whom juniors seek out.
Do not promote your best scientist by default. You lose a scientist and may gain a weak manager.
Offer a trial. Six months leading a small group, with a clear path back.
Coach new managers weekly for the first quarter.
Scenario: from 8 to 35 in two years
Ravi started with 8 scientists on a growing marketplace. At 14 he split into ranking and pricing, each with a tech lead. At 18 he hired an outside manager for pricing and promoted a senior scientist to manage ranking. At 28 he added a staff scientist for shared modeling and a program manager. At 35 he had four managers, two staff scientists, and a shared eval platform. Attrition stayed under 8% a year. Launches per quarter rose from 3 to 11.
Failure modes
The hero manager. One person has 20 reports and reviews every model. Everything waits on them.
Too many layers too soon. Managers with three reports each. Slow decisions and bored managers.
Siloed sub-teams. Each team builds its own tools. Five eval pipelines for one product.
Culture loss. Rigor fades as new people join without the habits.
Say this out loud: “I change structure before it breaks. Past about 12 people I split into sub-teams with tech leads. Past 25 I hire managers and add staff scientists. I keep the IC ladder as strong as the manager ladder.”
Manager interview questions
These are common questions for science manager roles. Each strong answer uses a first-person story with numbers. Adapt it to your own experience. Never memorize someone else's story.
1. How do you prioritize a science roadmap? Medium
What they are testing
Do you have a repeatable method, or do you go with gut feel?
Do you balance short-term wins with long-term bets?
Can you say no, and explain why?
A strong answer
“I use a portfolio approach. Each half, I gather every candidate project from the team, product and leadership. Last cycle that was 22 ideas for 16 scientists.
I score each one on four things. Chance of success, annual payoff in dollars, cost in scientist-months, and time to payoff. I use dollar rates we agreed with finance, so the numbers compare fairly. Then I rank by expected value per scientist-month.
I do not just take the top of the list. I aim for about 70% core work, 20% adjacent and 10% frontier. Last half, the top eight items were all core. I pulled in one frontier bet on LLM features, funded at one person for six weeks.
I review the draft with my product partners and two senior scientists. They catch bad estimates. Then I publish the list with what we are not doing and why. That list matters as much as the yes list.
The result was 11 funded projects. Seven shipped, three were killed at gates, and one carried over. Shipped gains were about $19M a year.”
Follow-ups
“What if your VP's pet project scores low?” Show the math. Offer a small spike to test it.
“How do you value platform work?” By the projects it unlocks and the time it saves.
“How often do you re-rank?” Quarterly, plus any time a gate fails.
Say this out loud: “I rank by expected value per scientist-month, then shape the mix to 70/20/10. I publish what we will not do. Last half, 7 of 11 projects shipped for about $19M a year.”
2. Tell me about a project you killed Medium
What they are testing
Can you stop work that is not working, even when people care about it?
Did you set kill criteria up front?
How did you handle the people involved?
A strong answer
“We had a project to replace our search ranker with a graph neural network. It had three scientists and strong support from a senior leader. I set kill criteria at the start. We needed a 1.5% offline NDCG gain by week 10 and p99 latency under 25 ms.
At week 10 we had a 0.9% gain. Latency was 60 ms. The team asked for six more weeks. I asked what would change. Their plan was more tuning, with no new idea for the latency gap.
I killed it. First I told the three scientists in person. I explained the criteria we had all signed. I made clear this was a project decision, not a judgment of their work. Then I told the senior leader with a one-page memo. It showed the numbers, the cost of six more weeks, and what we learned.
We moved two of the scientists to a feature project. One graph feature from the dead project became an input there. It added 0.4% to purchases.
The lead scientist got credit in her review for the clean test and the memo. She told me later that the quick decision helped her. She did not lose a year to a dead end.”
Follow-ups
“What if the leader pushed back?” Offer a small, time-boxed test of their best counter-idea.
“Did you ever kill something too early?” Have an honest example and what you changed.
“How do you keep morale up?” Credit the learning. Move people to strong work fast.
Say this out loud: “I set kill criteria up front and signed them with the team. When we missed them, I stopped the project, told the team first, and gave leaders a memo. We reused one piece and credited the scientists for a clean result.”
3. A scientist whose work is brilliant but never ships Hard
What they are testing
Can you coach a strong person without crushing what makes them good?
Do you find the real cause, or just push harder?
Do you set clear expectations about impact?
A strong answer
“I had a scientist, call him Dev, with two strong papers from his first year. Neither idea had reached production. His peers saw him as the smartest person on the team. His rating was at risk because the role expects shipped impact.
I first looked for the cause. I read his last three project docs and talked to his engineering partners. The pattern was clear. He stopped at the offline result. Then he moved to the next interesting problem. Nobody owned the path to launch.
I had a direct talk with him. I said his ideas were the best on the team. I also said our job is to change the product, and his work was not doing that yet. I showed him the rating guide.
We agreed on a plan. He would take his best idea all the way to an A/B test in one quarter. I paired him with a strong ML engineer. I set check-ins on launch steps, not on model gains.
It took 14 weeks. The model lifted engagement 1.1% and shipped. Dev found he liked seeing users react. Over the next year he shipped three more models. He still wrote one paper a year, now based on shipped work.”
Follow-ups
“What if he did not change?” Consider a research role elsewhere. Rate him honestly meanwhile.
“Is it ever right to let someone stay pure research?” Only if the team has a funded research charter.
“How do you avoid killing creativity?” Keep a 10% frontier slice he can lead.
Say this out loud: “I find the real cause first. Here it was no ownership past the offline result. I set a goal to take one idea to an A/B test. I paired him with an engineer and tracked launch steps. He shipped in 14 weeks and kept shipping.”
4. How do you measure your team's impact? Medium
What they are testing
Do you think in business outcomes, not model metrics?
Is your attribution honest?
Do you count work that is hard to measure?
A strong answer
“I measure impact at three levels.
The first is shipped business value. Every launch has an A/B test with a pre-registered primary metric. We convert lifts to dollars using rates set with finance. For example, one point of conversion is worth $7M a year on our surface.
The second is honest attribution. Summed test lifts overstate the truth. So we keep a 2% long-term holdout from all our launches. Last year our tests summed to 5.8% engagement. The holdout showed 4.1%. I report 4.1%.
The third is enabling work. Platform work does not move a metric by itself. I measure it by what it unlocked. Our new eval pipeline cut the time from idea to test from 19 days to 6. That doubled the number of tests we could run.
I also track health signals. Test velocity, launch rate, cost per test, and the share of tests that win. A win rate over 60% tells me our bar is too low.
Each quarter I write a one-page impact summary. It names every person who contributed, including partners in engineering.”
Follow-ups
“How do you credit shared wins?” Name all teams. Do not split the dollars to make yours look bigger.
“What about long-term effects?” Use holdouts and check at three and six months.
“How do you include cost?” Report net value after compute and serving cost.
Say this out loud: “I measure shipped value in dollars and report holdout numbers, not summed lifts. I credit platform work by what it unlocks. Last year our tests summed to 5.8%, but I reported the 4.1% holdout.”
5. Conflict between research and product deadlines Hard
What they are testing
Can you find a path that serves both sides?
Do you protect quality without blocking the business?
Can you manage up and across with data?
A strong answer
“Our product team had promised a new personalized home feed for a holiday launch. The date was fixed by marketing. My team's new model was eight weeks from ready. The launch was in five.
I met with the product lead to understand what was truly fixed. The date was. The full model was not. What mattered was that the feed felt personal for returning users.
I proposed two tracks. Track one was a simple version for the launch. We would use our existing model with three new features. That gave about 60% of the expected gain. I was confident it could pass review in four weeks. Track two kept the full model on its path, with a test planned for January.
I wrote this up with risks and gains for each option. I shared it with product, engineering and my director in one meeting. We agreed on the plan in 30 minutes.
The simple version launched on time and lifted engagement 2.1%. The full model shipped in January and added another 1.6%. I did not cut the experiment review on either one. That was my line, and I said so early.”
Follow-ups
“What if product wanted to skip the A/B test?” Offer a shorter test or a staged rollout with guardrails. Never no test.
“What if the simple version did not exist?” Be clear about the risk and let the accountable owner decide.
“How do you stop this from repeating?” Join product planning earlier.
Say this out loud: “I find what is truly fixed. Then I offer a smaller version that meets the date and keep the full work on its own track. I wrote both options up and got a decision in one meeting. I never cut the review step.”
6. Hiring a senior scientist: what do you look for? Medium
What they are testing
Do you know what senior means beyond technical skill?
Do you have a structured process?
Do you hire for the team's gaps?
A strong answer
“For a senior scientist, I look for four things.
First, depth. They must know one area far better than anyone on my team. I test this with a deep dive on their own past work. I keep asking why until they reach the edge.
Second, problem framing. Given a vague business problem, can they turn it into a clear ML problem? I give a real but simplified problem from our team. Strong people ask about the metric and the data before the model.
Third, shipped impact. They should have taken models to production and measured the result online. I ask for numbers and what broke after launch.
Fourth, multiplier behavior. Seniors make others better. I ask for times they mentored someone, set a technical direction, or changed a team's methods.
I also hire for gaps. Last year my team was strong in ranking and weak in causal inference. I wrote the role around that. The person I hired built our uplift modeling. Her first project improved coupon spend efficiency by 22%.
I use a written rubric for every round and hold a debrief with evidence, not feelings.”
Follow-ups
“Biggest red flag?” Cannot say what they did versus the team. Or no curiosity about the business.
“Papers or shipped work?” Shipped work for applied roles. Papers are a plus.
“How do you close a candidate?” Show them the problems and the people. Be honest about the hard parts.
Say this out loud: “I look for depth, problem framing, shipped impact, and multiplier behavior. I hire for the team's gaps. I use a written rubric and evidence-based debriefs. My last senior hire filled our causal inference gap and cut coupon waste by 22%.”
7. A failed launch you owned Hard
What they are testing
Do you take ownership, or blame others?
Did you respond well in the moment?
What did you change so it would not happen again?
A strong answer
“We launched a new delivery time model at a food app. The A/B test showed a 0.8% lift in orders. We shipped to everyone. Within a week, support tickets about late orders rose 30%.
The model was more optimistic in busy hours. That pulled orders, but many arrived late. Our test ran for two weeks during a quiet season. Our guardrails tracked cancellations but not lateness.
I owned it. I rolled back within 24 hours of finding the cause. I sent a short note to product and support leaders the same day. It said what happened, what we did, and what came next.
Then I ran a blameless review with the team. We found three gaps. First, no lateness guardrail. Second, a test window that missed peak load. Third, no check on calibration by time of day.
We fixed all three. Lateness became a standard guardrail for every logistics model. Tests now must cover at least one peak period. Calibration by segment joined our review checklist. We relaunched six weeks later. Orders rose 0.6% and late deliveries fell 4%.
The lesson for me was that the review checklist is a living thing. I now update it after every incident.”
Follow-ups
“Who was at fault?” The system. I owned the system.
“How did you rebuild trust?” Fast rollback, honest notes, and a clean relaunch.
“What would you do differently?” Ask support for their guardrails before the test.
Say this out loud: “I owned it, rolled back in a day, and told leaders the same day. A blameless review found three gaps in our process. We fixed all three and relaunched with a better result. I now update our checklist after every incident.”
8. How do you raise scientific rigor? Medium
What they are testing
Do you build systems, or rely on heroes?
Can you add rigor without slowing the team down?
Do you have evidence it worked?
A strong answer
“When I joined my current team, launch reviews were informal. Each owner chose their own metrics. In my first month I re-analyzed ten past launches. Three of the claimed wins did not hold up. One had a sample ratio mismatch. Two had picked the best metric after the test.
I made four changes. First, a one-page pre-registration for every test, filed before launch. Second, a fixed review checklist run by a scientist outside the project. Third, a shared results log, so we could see all past tests. Fourth, I changed how we rate people. Clean tests count, even nulls.
I knew process could slow us down. So I added a light path for small changes. It uses a checklist of five items and needs no meeting.
After two quarters, our win rate fell from 70% to 45%. That sounds bad but it meant the bar was real. Our holdout gain rose from 2.1% to 3.4% a year. Fewer fake wins meant more real ones. Review time per launch stayed under three days.”
Follow-ups
“How did senior people react?” Some pushed back. I asked two of them to co-write the checklist.
“How do you handle peeking?” Fixed run lengths or sequential tests.
“What about offline rigor?” Shared eval sets, time-based splits, and baselines in every report.
Say this out loud: “I audited past launches to show the gap. Then I added pre-registration, an outside reviewer, a results log and credit for nulls. Our win rate fell, but holdout gains rose from 2.1% to 3.4%. Real rigor means more real wins.”
9. How do you handle a low performer? Hard
What they are testing
Do you act early and clearly?
Are you fair and humane?
Do you follow through, either way?
A strong answer
“I had a mid-level scientist, call her Ana, who missed three milestones in a row. Her peers were starting to carry her work.
First I looked for causes. I asked her directly how things were going. It turned out she was struggling with a new area. She came from NLP and was now on tabular ranking. She was also hesitant to ask for help.
I was clear with her. I said her work was below the bar for her level, and gave three concrete examples. I also said I wanted her to succeed and believed she could.
We wrote a plan together with three goals over eight weeks. Each was concrete. Ship one feature test. Write a design doc reviewed by a senior. Present at science review. I paired her with a senior scientist for two hours a week. I met her weekly to check progress.
She met two of three goals in eight weeks, and the third two weeks later. Her next rating was at the bar. In other cases, the plan has not worked. Then I move to a formal process with HR. I am honest with the person about where it is heading. Waiting too long is unfair to them and to the team.”
Follow-ups
“How early do you act?” At the second missed signal, not the review cycle.
“What if it is a skill mismatch?” Consider a move to a team that fits better.
“How do you protect the team?” Rebalance work so peers are not silently carrying it.
Say this out loud: “I act early, find the cause, and give clear feedback with examples. We set concrete goals with support and weekly check-ins. If it works, great. If not, I move to a formal process and stay honest with the person.”
10. Why should a scientist join your team? Medium
What they are testing
Do you know what scientists value?
Can you describe a real team culture, not slogans?
Are you honest about the hard parts?
A strong answer
“I give candidates four reasons, and one honest warning.
First, the problems are real and big. Our models touch 200 million users a day. A good idea can show up in a test within three weeks.
Second, the work ships. Last year, 70% of our funded projects reached an online test. Scientists own their models through launch, so they see the results.
Third, the bar is high and fair. Every test is pre-registered and reviewed. People get credit for clean nulls, not just wins. That makes it a safe place to try hard ideas.
Fourth, people grow. I hold a career talk with each person every quarter. In the last two years, six of my scientists were promoted. Two moved to staff. We also keep a 10% frontier slice, and we publish when it does not give away our edge.
The warning is that we are not a pure research lab. If you want to write papers with no product link, this is not the right fit. I would rather say that now.
The best way to judge is to meet the team. I always offer a call with two scientists who are not in the loop.”
Follow-ups
“What is the hardest part of your team?” Name it honestly, like on-call for model quality.
“How do you keep people?” Growth, ownership, and fair credit.
“What is your management style?” Clear goals, high trust, and direct feedback.
Say this out loud: “Big real problems, work that ships, a high and fair bar, and real growth. Six promotions in two years. I am honest that we are not a pure research lab. Candidates can talk to my team directly.”
Recap
Run science as a portfolio. Rank by expected value per scientist-month and shape the mix to about 70/20/10.
Plan with questions and dates. Set kill criteria up front and grow staffing only after gates.
Pre-register every test, use an outside reviewer, and reward clean nulls.
Report impact in dollars using holdouts, and include cost.
Write outcome OKRs, keep a simple rhythm, and own the launch with engineering and product.
Start with the cheapest model that might work. Track compute like headcount.
Build risk reviews into the gates. Change the org shape before it breaks.