Science managers bet on ideas that may not work. Your job is to place good bets, grow the people who make them, and turn the winners into product impact.
Most new science managers were strong scientists first. That helps, but it is not the job. The job is to build a team that does better science than you could alone. It also has to ship that science into products that matter. This page covers the full role. It is written from the seat of a manager who has run science teams at a large tech company. Each topic gives principles, practices, a short scenario, common failures, and a line you can use in a manager interview.
The loop for an Applied Science manager tests two things at once. It checks that you can still judge science. It also checks that you can lead people, plans and partners. Expect behavioral rounds on hiring, performance and conflict. Expect a technical round where you review a design or an experiment. Expect a round on strategy and roadmap.
An engineering manager mostly runs work that is known to be possible. The open questions are cost and time. A science manager runs work where the answer itself is unknown. Many projects will fail, and that is normal. This one fact changes how you plan, judge and report.
Core principles
Uncertainty is the product. If you knew the answer, it would be engineering. Your team exists to reduce uncertainty for the business.
Feedback loops are long. A model idea can take six weeks to test offline. An online A/B test can take another four. You may learn the truth a quarter later.
Success is non-linear. Ten projects may give zero wins, then one win pays for all ten. Progress does not look like a burn-down chart.
Effort and outcome are loosely tied. A great scientist can run a clean study and get a null result. A weak one can get lucky. You must judge the process as well as the result.
Knowledge is an output. A well-run failed experiment still teaches the team what not to try. That has value if you write it down.
The research-to-product spectrum. Applied Science sits near product, but a team with no research bets runs dry within a year.
Concrete practices
Plan in bets, not tasks. Each bet gets a hypothesis, a cost cap, and a kill date.
Report confidence, not just status. “60% likely to beat baseline by June” is more useful than “on track”.
Build short loops on purpose. Invest in offline metrics that predict online wins. Invest in fast experiment tools.
Keep a decision log. Write down why each bet started and why it stopped.
Run a portfolio, not a project. Mix safe bets with a few long shots so the team never has a quarter with nothing to show.
Scenario: the ranking team with a quiet quarter
Priya runs a team of seven scientists on feed ranking. In Q2 three of four bets fail offline. Her VP asks why the team “did nothing”. Priya does not argue effort. She shows the decision log. Each failed bet had a cost cap of three scientist-weeks and stopped on time. The team spent nine weeks total on failures, not thirty. One failure showed that a popular feature leaked label data. Fixing that leak gave a 0.4% engagement gain on its own. The fourth bet is now in an A/B test at 1.1% gain. The VP leaves with a clear story: cheap failures, one real win, one bug fixed.
Common failure modes
Running science like a sprint board. Story points on research make people hide risk and pad estimates.
No kill dates. Bets drift for months because no one wants to call them dead.
All long shots. The team looks busy and ships nothing. Trust with product drops.
All safe bets. The team becomes a tuning shop. Top scientists leave.
Say this in a manager interview: “In science, many good bets fail. So I manage the portfolio, not single projects. Every bet has a hypothesis, a cost cap and a kill date. I report confidence levels to partners, so a failed bet is a planned outcome and not a surprise.”
Your job in one sentence and the four hats
The job in one sentence. Turn the right scientific bets into durable business impact, through a team that keeps getting stronger.
Each part of that sentence matters. “Right bets” means you choose the problems. “Durable impact” means it ships and stays shipped. “Keeps getting stronger” means people grow and stay. You wear four hats to do this. On any day one hat matters most. Over a quarter you need all four.
The four hats
People. Hire, grow, coach, judge and keep scientists. This is the hat only you can wear. Expect it to take 35 to 45% of your time.
Portfolio. Choose which problems the team works on and in what mix. Kill bets that stall. Fund bets that show signal.
Partnerships. Earn trust with product, engineering, data and legal. Turn science into roadmaps they can plan around.
Technical judgment. Review designs and experiments. Set the rigor bar. Spot the flawed result before it ships.
Concrete practices
Audit your calendar each month. Tag each meeting with one hat. Look for a hat with less than 10% of your time.
Write a one-page team charter. Name the mission, the metrics you own, and what you will not do.
Pick one owner per hat on your bench. A senior scientist can own much of the technical review. A tech lead can own partner syncs for one product.
Review the portfolio every quarter with your manager. Bring the bet list, the mix, and what you killed.
Scenario: the manager stuck in one hat
Marco manages nine scientists on ads quality. His calendar audit shows 60% technical reviews, 25% 1:1s, 10% partner meetings, and 5% planning. Two partner teams have stopped inviting him to roadmap talks. He hands the weekly model review to his strongest senior scientist, Lena. That frees six hours a week. He uses four for partner syncs and two for a monthly portfolio review. In one quarter, his team gets two new asks from product. Lena gets a clear growth story for her promotion case.
Common failure modes
Super-IC manager. Lives in the technical hat. The team waits on every review. People and partners drift.
Pure people manager. Loses the technical hat. Can no longer tell a sound result from a broken one.
Order taker. Lets partners set the whole portfolio. The team never builds its own long bets.
Say this in a manager interview: “I think of the job as four hats: people, portfolio, partnerships and technical judgment. I track my time against them each month. When one gets thin, I delegate from another rather than drop it.”
The scientist-to-manager transition
The hardest part of the switch is not new skills. It is letting go of the old ones. You were rewarded for your own models. Now you are rewarded for other people's models. Your output is the team's output.
Core principles
Let go of the keyboard. If you own a critical path task, you block the team. Your calendar will break it.
Stay technical enough. You must still read a paper, follow a loss curve, and spot a leak. You do not need to write the code.
Your leverage is other people. One hour spent unblocking three scientists beats one hour of your own work.
Feedback slows down. You lose the daily win of a passing test. Learn to find wins in people growing.
How much to code
Team of 3 to 5. Up to 20% hands-on is fine. Take work that is off the critical path.
Team of 6 to 10. About 5 to 10%. Notebooks, data checks, small analyses. Nothing anyone waits on.
Manager of managers. Close to zero production code. Stay sharp by reading code and results.
Ways to stay technical without blocking
Read every experiment write-up your team produces. Ask one hard question on each.
Reproduce one key result per quarter from the team's notebook.
Read two papers a month in your area. Share one with the team.
Join design reviews as a reviewer, not as the author.
Do small data dives that answer a partner question fast.
Scenario: the new manager who kept the model
Wei was promoted to manage the five-person search relevance team. He kept ownership of the main ranking model. After two months, the model launch slipped three weeks. Wei had spent his time in hiring loops and planning. Two juniors had waited on his code reviews for days. He handed the model to Aisha, a senior scientist, with a clear launch goal. He took a side analysis on query intent drift instead. The model shipped four weeks later with a 2.3% gain in click-through. Aisha used the launch as the core of her promotion packet.
Common failure modes
Owning the critical path. Your meetings become the team's bottleneck.
Rewriting people's work. It teaches them to wait for you. It also kills ownership.
Going fully non-technical. In a year you cannot tell good science from bad. The team notices.
Hiding in the work. Coding feels safe. Hard people talks do not. Many new managers retreat to the code.
Say this in a manager interview: “When I moved into management, I gave away the main model. I kept small analyses that nobody waited on. I stay technical by reading every experiment write-up and reproducing one key result each quarter.”
Hiring scientists
Hiring is the highest-leverage thing you do. One great scientist can define a team for years. One poor hire can cost a year of your time. Science hiring is also hard to get right. Papers and degrees are easy to see. Taste and shipping ability are not.
Core principles
Define the role before the loop. Know what this person must do in their first year.
Test for research taste. Can they pick problems worth solving and drop ones that are not?
Test for shipping. Have they taken a model all the way to production and measured impact?
Judge the work, not the brand. A top school or lab is a signal, not proof.
Selling starts at the first call. Strong scientists have options. They choose teams, not just companies.
Role definition checklist
What problem will they own in month six?
Where do they sit on the spectrum? Research-heavy, applied, or close to production?
What level is it really? Scope of work, not years of experience.
Which skills are must-have? Which can they learn on the job?
Who will they partner with most? Engineers, PMs, other scientists?
Loop design
Coding. One round. Real data handling, not just puzzles. Scientists must write code that others can run.
ML breadth. Covers the field. Screens for gaps that would hurt in your domain.
ML depth on their own work. Drill into one past project. Ask why each choice was made. Ask what failed.
Applied problem. Give a fuzzy business problem from your domain. Watch them frame it, pick a metric, and plan an experiment.
Research talk. For senior roles. Grades clarity and how they handle hard questions.
Behavioral. Conflict with a partner, a failed project, and a time they changed course.
Assessing research taste
Ask: “What is a problem in your field you think is overrated? Why?”
Ask: “Tell me about a project you killed. How did you know?”
Ask: “If you had three months and one engineer on our problem, what would you try first?”
Strong answers rank ideas by value and cost. Weak answers list the newest methods.
Assessing shipping ability
Ask for the online result, not the offline metric. What did the A/B test show?
Ask what broke after launch. People who shipped have scars.
Ask who the engineering partner was and how they split the work.
Ask how they handled latency or memory limits.
Avoiding pedigree bias
Write a rubric before the loop. Score each round against it before the debrief.
Have interviewers submit scores before they see other scores.
Source from more than top labs. Strong scientists come from industry, smaller schools, and other fields like physics.
Judge the papers by their content. A first-author workshop paper with a clear idea can beat a tenth-author oral.
Watch for “not a culture fit” feedback with no evidence. Ask the interviewer for the specific moment.
Closing candidates
Sell the problem first. Scientists join for interesting work and good data.
Name the first project and who they will work with.
Be honest about publication policy and conference time.
Connect them with a scientist on the team, not just you.
Move fast. Each extra week loses candidates.
Scenario: the hire that the brand almost hid
Dana has one senior scientist slot on a fraud team. Two finalists remain. Candidate A has a top PhD and six papers at major venues, but no shipped models. Candidate B has a master's degree from a smaller school and four years at a fintech firm. B shipped a fraud model that cut chargebacks by 18%. In the applied round, A proposes a graph neural net and cannot say how to label the data. B asks about label delay first, then proposes a simple model with a plan to improve it. Dana's rubric weights framing and shipping for this role. She hires B. A year later B leads the team's main model and mentors two juniors.
Common failure modes
Hiring for the paper count. You get great researchers who never ship.
Copying the engineering loop. You test LeetCode and miss research taste.
No rubric. The debrief becomes a vote on who people liked.
Slow loops. Your best candidate signs elsewhere in week three.
Overselling. You promise research freedom you cannot give. They leave in a year.
Say this in a manager interview: “I define the role by what the person must own in month six. My loop tests research taste and shipping, not just theory. I use a written rubric and blind scoring so the debrief is about evidence, not pedigree.”
Growing scientists
Your team gets stronger only if each person grows. Science growth is less obvious than engineering growth. It is not just bigger systems. It is better problem choice, sharper rigor, and wider influence.
Core principles
Use the ladder as a map, not a checklist. Show people the next level in plain terms.
Separate scope, complexity and impact. They grow on different axes. Name which one is the gap.
Taste is learned. People learn problem choice by watching it done and getting feedback on their own picks.
Sponsor, do not just mentor. Mentors give advice. Sponsors put people forward for big work.
Scope vs complexity vs impact
Scope. How wide is the problem? One model, one product area, or a whole org? Seniors own wider scope.
Complexity. How hard is it? Novel methods, messy data, hard constraints. A junior can do complex work in narrow scope.
Impact. What changed for the business? Metrics moved, costs cut, decisions made. This is what calibration weighs most.
A common stuck case is high complexity, low scope. The scientist does hard work on a narrow problem. They need a broader problem, not a harder one.
Concrete practices
Write a growth plan with each person twice a year. Name one ladder gap and one project that closes it.
Pair each junior with a senior on a real project. Make the senior's mentoring part of their review.
Let juniors write the first draft of the problem framing. Review it with them, not for them.
Give seniors a cross-team problem. Scope grows when they must influence people they do not manage.
Name people in rooms they are not in. Put their names on launch posts and in staff meetings.
Publication policy
Set a clear team policy and share it in hiring. Ambiguity breeds resentment.
A common split: publish methods, protect product data and business metrics.
Budget the time. A paper costs about two to four weeks of extra work after the result.
Tie papers to work the team already did. Avoid side projects that only exist for a paper.
Get legal and privacy review early, not the week before the deadline.
Scenario: the strong scientist stuck at one level
Omar has been at the same level for two years. His models are some of the best on the team. His manager, Grace, reads the ladder with him. The gap is scope, not skill. Omar only works on the one model his tech lead hands him. Grace gives him a new problem: unify three teams' separate churn models into one shared model. It needs buy-in from two other managers. Omar struggles at first with the meetings. Grace coaches him on writing a one-page proposal. Nine months later the shared model saves 40% of training cost and lifts recall by 3 points. Omar is promoted in the next cycle.
Common failure modes
Harder, not wider. You give a stuck senior a harder problem when they need broader scope.
Mentoring with no sponsorship. Lots of advice. No high-visibility work.
Hoarding the best people. You block a transfer that would grow them. They leave the company instead.
Vague ladder talk. “Show more leadership” with no example of what that means.
Say this in a manager interview: “I separate scope, complexity and impact when I coach. Most stuck seniors do hard work in narrow scope. So I find them a cross-team problem, sponsor them into it, and coach them on influence.”
Performance management for science
Judging science work is hard because outcomes are noisy. A good scientist can have a bad year of results. A weak one can ride a lucky launch. Your job is to judge the quality of the bets and the work, and then weigh the impact fairly.
Core principles
Judge decisions, not just outcomes. A good bet that failed for a reason no one could know is still a good bet.
But impact still counts. Over time, good process must produce some wins. Two years of clean failures is a pattern.
Evidence beats impressions. Write things down all year. Memory favors the last month.
No surprises. Nobody should learn their rating for the first time in the review.
Judging work when experiments fail
Was the hypothesis clear and worth testing?
Was the cost capped? Did they stop on time?
Was the experiment sound? Right metric, enough power, no leakage?
Did they share the learning so others did not repeat it?
Did they change course well after the result?
Rewarding good bets, not lucky outcomes
Ask for a pre-registration note before each big experiment. Hypothesis, metric, expected effect, stop rule.
Credit the person who killed a bad project early. That saved real money.
Question big wins too. Was the gain real? Did it hold for four weeks? Was the baseline weak?
Look at the full year. One lucky launch should not outweigh poor rigor all year.
Writing calibration packets
Lead with impact. First line: what changed for the business, with a number.
Show scope. How wide was the problem? Who else depended on the work?
Show how. One or two examples of rigor, taste or leadership.
Translate science. Calibration rooms have engineering managers. Say “cut fraud loss by $2M a year”, not “improved AUC by 0.03”.
Handle failed bets directly. Name the bet, why it was sound, and what the team learned.
Compare to the ladder. Quote the level text the work meets.
Handling underperformance
Find the cause first. Skill gap, wrong role, unclear goals, or something personal?
Give clear, written feedback early. Name the gap and what good looks like.
Set a short plan with check-ins every two weeks. Pick goals they can hit in six to eight weeks.
Offer support. A senior pair, a smaller problem, or training.
If there is no progress, move to the formal process. Work with HR. Do not drag it out for a year.
Scenario: two scientists, two very different years
Sam and Jo both work on recommendations. Sam shipped one model with a 2% gain. But the review showed his baseline had a bug. The fair gain was closer to 0.6%. He also skipped the holdout test on a second launch. Jo ran three bets. Two failed offline within their caps. Her write-ups stopped two other teams from trying the same ideas. Her third bet is in an A/B test showing 1.4%. Their manager, Ken, rates Jo above Sam. In calibration he explains the baseline bug and Jo's saved effort in dollar terms. He also gives Sam clear feedback on rigor with a plan for next half.
Common failure modes
Outcome bias. The lucky launcher gets the top rating. The rigorous scientist with bad luck gets “meets”.
Science jargon in calibration. The room does not understand the impact, so it discounts it.
Waiting too long on underperformance. The team sees it and loses faith in you.
Endless benefit of the doubt. “Research takes time” becomes a cover for no output at all.
Say this in a manager interview: “I judge science on the quality of the bet and the work, then weigh impact over the full year. I ask for a short pre-registration note before big experiments. That lets me reward a sound bet that failed, and question a lucky win.”
1:1s and coaching scientists
The 1:1 is the main tool you have for each person. It belongs to them, not to you. Done well, it catches problems early and grows people. Done badly, it is a status meeting that wastes 30 minutes.
Core principles
Their agenda first. Ask what they want to talk about before you raise your items.
Not a status update. Status goes in docs and standups. 1:1s are for blockers, growth and trust.
Split technical and career talks. If you mix them, the technical talk always wins.
Feedback on research quality is a gift. Give it often, kindly and with specifics.
A simple 1:1 agenda
Their topics. What is on your mind?
Blockers. What is slowing you down?
Feedback both ways. One thing that went well, one thing to change.
Your topics. Context they need, decisions coming.
Once a month: career and growth only.
Technical vs career 1:1s
Technical 1:1. Walk through a result, a design, or a stuck problem. Ask questions rather than give answers. Good questions: “What would convince you this is wrong?” and “What is the simplest baseline?”
Career 1:1. Monthly. Talk about goals, ladder gaps, and what kind of work energizes them. Ask: “Where do you want to be in two years?”
Giving feedback on research quality
Be specific. “Your holdout set overlaps the training window by two weeks” beats “check your data”.
Separate the person from the work. Talk about the analysis, not their ability.
Ask first. “How confident are you in this result? What would you check next?” Often they find the flaw themselves.
Name what was good too. Clear write-up, smart baseline, honest error bars.
Follow up. Check the fix in the next review.
Scenario: catching a flaw without crushing a junior
Nina is six months into her first job. She shows her manager, Tom, a model with a 9% offline lift. Tom suspects leakage, since gains on this task are usually 1 to 2%. He does not say that. He asks, “What is the biggest single feature by importance?” It is a timestamp field. He asks how that field is set. Nina realizes it is filled after the label event. She fixes it, and the lift falls to 1.3%. Tom praises her for tracing it fast. In the next team review, Nina presents the leak as a lesson. Two other scientists find the same field in their models.
Common failure modes
Status-only 1:1s. Thirty minutes of task updates. Nothing about growth or trust.
Cancelling often. It tells people they are not a priority.
Telling instead of asking. You fix the bug. They learn nothing.
Vague praise and vague critique. “Nice work” and “tighten this up” teach nothing.
Say this in a manager interview: “My 1:1s start with the scientist's agenda. I keep career talks separate and monthly, so technical topics do not crowd them out. When I see a flaw in an analysis, I ask questions that lead them to it. They learn more that way.”
Motivation and retention
Scientists are hard to hire and easy to lose. Most leave for one of three reasons. They lost interest in the work. They stopped growing. Or they felt like a service desk. Pay matters, but it is rarely the only reason.
Core principles
Autonomy. Give people a problem, not a recipe. Let them choose the method.
Mastery. Make room to learn deeply. Reading groups, conferences, hard problems.
Purpose. Tie each project to a user or business outcome they care about.
Recognition. Make good work visible. Credit the scientist by name.
Concrete practices
Fund conference time. One conference a year per scientist is a common norm.
Run a weekly reading group. Rotate who leads it.
Protect 10 to 20% time for exploratory ideas tied to the team's mission.
Support publishing within the team policy.
Ask in career 1:1s: “What would make you leave?” Act on the answer before they do.
Avoiding the ticket factory
A ticket factory is a team that only takes requests. Product sends tickets. Scientists tune models. No one owns a problem end to end. Top people leave first.
Own problems, not requests. Agree on a metric your team owns.
Cap request work. Keep at least 30% of capacity for the team's own roadmap.
Push back on vague asks. Turn “make the model better” into a clear goal with a metric.
Rotate the boring work. Nobody should be the permanent retraining person.
Burnout
Watch for signs. Late-night commits, short replies, missed 1:1s, cynical jokes.
Long failure streaks burn people out too. A scientist with four failed bets in a row needs a sure win next.
Do not reward heroics. A launch at 2 a.m. is a planning failure.
Model good behavior. Take your own time off and do not send weekend messages.
Scenario: turning around a team that felt like a service desk
Elena takes over a six-person team that supports four product groups. Attrition was 40% last year. Exit interviews say “no ownership”. She agrees with her director that the team will own one metric: recommendation diversity. She caps support requests at 40% of capacity, enforced through a simple weekly intake. Each scientist gets one problem they own from start to launch. She funds two conference trips and starts a reading group. In the next year attrition drops to one person. The team ships a diversity model that lifts long-term retention by 0.8%.
Common failure modes
Saying yes to every request. Partners love it. Your scientists leave.
Freedom with no direction. Autonomy without a mission gives scattered work and no impact.
Ignoring failure streaks. Repeated null results wear people down even when the work is good.
Retention by counter-offer only. Money keeps people for six months, not two years.
Say this in a manager interview: “Scientists stay for ownership, growth and recognition. So I make my team own a metric, not a ticket queue. I cap request work at a set share of capacity. I watch for failure streaks and make sure the next project has a good chance to land.”
Team structure
Where scientists sit shapes what they work on. There is no single right model. Each one trades depth for closeness to product. Good managers know the trade-offs and change the model as the org grows.
The three common models
Embedded
What it is. Scientists sit inside product teams and report to product or engineering leads.
Strengths. Close to the problem. Fast to ship. Deep domain context.
Weaknesses. Isolated scientists. Uneven rigor. Weak career paths. Often judged by managers who are not scientists.
Best for. Small orgs or mature products with stable problems.
Centralized
What it is. All scientists sit in one science org and take projects from product teams.
Strengths. High rigor. Shared tools. Strong peer learning. Clear career ladder.
Weaknesses. Far from product. Risk of building things no one ships. Can become a ticket factory.
Best for. Early science functions or deep, cross-product problems.
Hub-and-spoke
What it is. Scientists report to a central science leader but sit with product teams day to day.
Strengths. Product closeness plus shared standards and career growth.
Weaknesses. Two bosses in practice. Needs clear rules on who sets priorities.
Best for. Most large tech companies once they have 30 or more scientists.
Team size
Six to eight direct reports is a sweet spot for a science manager who still reviews work.
Above ten, your technical review quality drops. Add a tech lead or split the team.
Below four, you are likely a tech lead with a title. That is fine for a first team.
Avoid lone scientists. Pairs at least, so each has a peer to check their work.
Ratio of scientists to engineers
A common healthy range is one scientist to one to three engineers on product-facing teams.
Too few engineers and scientists spend half their time on pipelines. Launches slow down.
Too many engineers and science becomes the bottleneck. Engineers build ahead of proven ideas.
Research-heavy teams can run closer to two scientists per engineer.
Track one signal: how many weeks a proven model waits before it ships. Over four weeks means you need more engineering.
Scenario: moving from embedded to hub-and-spoke
A growing marketplace has 22 scientists spread across nine product teams. Each team uses its own A/B test method. Two launches were rolled back after false wins. Science promotion rates lag engineering by half. The new science director, Ravi, moves to hub-and-spoke. Scientists now report into three science managers but keep their desks and standups with product. He sets one shared experiment standard. Within a year, false-win rollbacks drop to zero. Promotion rates match engineering. Product leads say they still get the same daily support.
Common failure modes
Picking a model by fashion. Copying another company's org chart without its size or problems.
Lone embedded scientists. No peer review. Rigor slides. They leave.
Hub with no spokes. A central team that writes papers no product uses.
Ignoring the engineer ratio. Models sit in notebooks for months.
Say this in a manager interview: “At scale I prefer hub-and-spoke. Scientists sit with product for context and report into science for rigor and growth. I watch how long proven models wait to ship. That tells me if my engineer ratio is right.”
Building trust with product and engineering
Science only creates value when it ships. That needs partners who trust you. Product and engineering leads plan in quarters with fixed dates. Science gives uncertain answers on uncertain timelines. Your job is to bridge that gap.
Core principles
Speak business. Talk in revenue, cost, user growth and risk. Not in AUC.
Make uncertainty plannable. Give ranges, odds and decision dates. Partners can plan around those.
Under-promise, over-deliver. Commit to what you are sure of. Treat the rest as upside.
Share credit. A launch is a joint win. Say so in public.
Roadmaps with uncertainty
Use tiers. Committed (90% sure), likely (60%), and exploratory (under 30%). Partners plan only on committed.
Use decision gates. “By March 15 we will know if this works offline. If yes, we need two engineers in Q2.”
Give a fallback. If the new model fails, what ships instead? Often a tuned baseline.
Update often. Send a short monthly note on each bet's odds. No surprises at quarter end.
Concrete practices
Meet each partner lead every two weeks. Ask what keeps them up at night.
Translate each result into one business line. “This cuts false declines by 12%, worth about $3M a year.”
Write a one-page launch plan with engineering. Name owners for data, model, serving and monitoring.
Say no with a reason and an option. “We cannot do both. If we drop X, we can do Y by May.”
Show up for incidents. If your model causes a problem, own it first.
The portfolio quadrant. Quick wins buy partner trust. Moonshots keep the team ahead. Core bets pay the bills. See M2 for how to run the portfolio week to week.
Scenario: the roadmap that survived a failed bet
Lucas runs science for a payments product. His PM partner wants a new risk model in Q3 for a market launch. Lucas splits his plan into tiers. Committed: retune the current model on new market data for a 5% cut in fraud loss. Exploratory: a new sequence model that might cut loss by 15%. He sets a decision gate for June 1. The sequence model fails on the gate because of sparse data in the new market. The retuned model ships on time and cuts loss by 6%. The PM tells her VP that science “always does what it says”. Next quarter she funds two engineers for the sequence model once data matures.
Common failure modes
Selling the moonshot as a commitment. It fails and you lose trust for a year.
Talking in model metrics. Partners cannot weigh AUC against their other priorities.
Disappearing between launches. Partners hear from you only when you need something.
Taking all the credit. Engineering stops wanting to ship your models.
Say this in a manager interview: “I give partners a tiered roadmap: committed, likely and exploratory. They plan only on committed. Each exploratory bet has a decision date and a fallback. So when a bet fails, the product still ships, and partners keep trusting us.”
Technical leadership as a manager
Your team's science is only as good as the bar you set. You may not write the code, but you own the quality of what ships. A science manager who cannot catch a flawed experiment is a risk to the business.
Core principles
Set standards, then enforce them kindly. Rigor is a team habit, not a heroic act.
Review the question before the method. Most failures come from the wrong problem or the wrong metric.
Make rigor cheap. Templates, shared tools and checklists beat long lectures.
Grow reviewers. You cannot review everything. Teach seniors to review well.
Reviewing designs
What business problem does this solve? What is the metric?
What is the simplest baseline? Has it been tried?
Where do labels come from? How noisy or delayed are they?
What are the serving limits? Latency, memory, cost per call.
How will we know it works? Offline metric, then online test plan.
What could go wrong after launch? Drift, feedback loops, fairness gaps.
Reviewing experiments
Was the test planned before it started? Hypothesis, primary metric, sample size.
Is the sample big enough to detect the expected effect?
Any peeking or early stopping? Any metric shopping?
Do the guardrail metrics hold? Latency, revenue, user complaints.
Does the effect hold across key segments? New users, regions, devices.
Is there a novelty effect? Check the trend over the full run.
Raising rigor across the team
Write a one-page experiment standard. Share it in onboarding.
Use a design doc template with the questions above.
Hold a weekly model review. Rotate the presenter. Invite one engineer.
Run a monthly “failed experiments” session. Make it safe to show nulls.
Track rollbacks and false wins. Treat them as signals, not blame.
Scenario: the false win caught in review
Kim's team reports a 3.2% lift in purchases from a new ranking model. Launch is scheduled for Friday. In the weekly review, Kim asks to see the daily trend. The lift was 6% in week one and 0.5% in week two. She also notes the team checked the result four times and stopped on a good day. She asks for one more clean week with a fixed stop date. The final lift is 0.7%. The team still ships, since the cost is low. But they report the honest number. Two months later finance checks the impact. It matches 0.7%. Kim's team gains a reputation for numbers finance can trust.
Common failure modes
Rubber-stamp reviews. You approve everything to avoid conflict. Bad results ship.
Bottleneck reviews. Everything waits for you. Seniors never learn to review.
Method obsession. You debate model choice while the metric is wrong.
Blaming people for nulls. The team starts hiding failed tests.
Say this in a manager interview: “I own the rigor bar even when I do not write the code. We use a design template, a written experiment standard, and a weekly model review. I review the question and metric before the method. I train seniors to review, so I am not the bottleneck.”
Your first 90 days
The first 90 days set how the team sees you for a long time. Move too fast and you break things you do not understand. Move too slowly and people wonder why you are there. Use a simple order: listen, map, quick win, set direction.
Days 1 to 30: listen
Hold a long 1:1 with every report. Ask what works, what does not, and what they would change.
Meet each key partner. Ask what they need from science and where it fell short.
Meet your manager. Agree on what success looks like at 90 days and one year.
Read the last six months of design docs, experiment write-ups and launch posts.
Change nothing big yet. Take notes and look for patterns.
Days 31 to 60: map
Map the portfolio. List every bet, its owner, its odds, and its impact.
Map the people. Strengths, growth gaps, flight risks, and who is ready for more.
Map the partners. Who trusts the team? Who has given up on it?
Map the tools. How long does an experiment take from idea to result?
Write a short “what I learned” doc and share it with the team. Ask them to correct it.
Days 61 to 90: quick win and set direction
Pick one quick win the team already wants. Fix a slow pipeline, unblock a launch, or kill a zombie project.
Write a one-page direction doc. Mission, metrics owned, top three bets, what the team stops doing.
Review it with your manager and partners before you finalize it.
Start the rhythms. Weekly model review, monthly career 1:1s, quarterly portfolio review.
Give each person a growth goal tied to the new direction.
Scenario: a new manager on an inherited team
Fatima joins as manager of an eight-person NLP team. In month one, her 1:1s show three themes. Experiments take three weeks to set up. Two bets have run for nine months with no gate. And the PM partner thinks the team is “a black box”. In month two she maps the portfolio and finds 30% of capacity on the two stalled bets. In month three she kills one bet and sets a gate on the other. She gives two scientists four weeks to cut experiment setup time. It drops to four days. She shares a one-page direction doc with a tiered roadmap. The PM calls it the first science plan she could plan around.
Common failure modes
Reorg in week two. You change the team before you know why it works the way it does.
Proving yourself technically. You rewrite a model to show you are smart. The team feels judged.
Endless listening. Day 120 and still no direction. People lose patience.
Ignoring the previous manager's bets. Some were good. Killing all of them looks political.
Say this in a manager interview: “In my first 90 days I listen for a month and map for a month. Then I land one quick win and write a one-page direction doc. I share my learning doc early, so the team can correct me before I set direction.”
The thread through every topic. Science fails often by nature. A great science manager makes failure cheap, visible and useful. That frees the team to take the big bets that matter.
Recap
Science management means running uncertain bets on long, non-linear feedback loops.
Wear four hats: people, portfolio, partnerships and technical judgment. Audit your time against them.
Let go of the critical path. Stay technical by reading, reviewing and reproducing.
Hire for research taste and shipping, with a rubric that resists pedigree bias.
Grow people on scope, not just complexity. Sponsor them into visible work.
Judge the quality of bets and work, then weigh impact over the full year.
Own a metric, cap request work, and watch for failure streaks to keep people.
Give partners tiered roadmaps with decision dates and fallbacks.
Own the rigor bar through templates, standards and trained reviewers.
First 90 days: listen, map, quick win, set direction.