Part V Manager 1 Leadership

Being a Successful Applied Science Manager

Science managers bet on ideas that may not work. Your job is to place good bets, grow the people who make them, and turn the winners into product impact.

Most new science managers were strong scientists first. That helps, but it is not the job. The job is to build a team that does better science than you could alone. It also has to ship that science into products that matter. This page covers the full role. It is written from the seat of a manager who has run science teams at a large tech company. Each topic gives principles, practices, a short scenario, common failures, and a line you can use in a manager interview.

The loop for an Applied Science manager tests two things at once. It checks that you can still judge science. It also checks that you can lead people, plans and partners. Expect behavioral rounds on hiring, performance and conflict. Expect a technical round where you review a design or an experiment. Expect a round on strategy and roadmap.

Contents

  1. Why science management is different
  2. Your job in one sentence and the four hats
  3. The scientist-to-manager transition
  4. Hiring scientists
  5. Growing scientists
  6. Performance management for science
  7. 1:1s and coaching scientists
  8. Motivation and retention
  9. Team structure
  10. Building trust with product and engineering
  11. Technical leadership as a manager
  12. Your first 90 days

Why science management is different

An engineering manager mostly runs work that is known to be possible. The open questions are cost and time. A science manager runs work where the answer itself is unknown. Many projects will fail, and that is normal. This one fact changes how you plan, judge and report.

Core principles

Basic research new methods, papers horizon: 2+ years Applied research prove an idea on our data horizon: 6-18 months Applied science ship models, move metrics horizon: 1-6 months Product eng build known things well horizon: weeks certainty of shipping rises, uncertainty and horizon fall most AS teams live here A healthy AS portfolio keeps one foot in applied research.
The research-to-product spectrum. Applied Science sits near product, but a team with no research bets runs dry within a year.

Concrete practices

Scenario: the ranking team with a quiet quarter

Priya runs a team of seven scientists on feed ranking. In Q2 three of four bets fail offline. Her VP asks why the team “did nothing”. Priya does not argue effort. She shows the decision log. Each failed bet had a cost cap of three scientist-weeks and stopped on time. The team spent nine weeks total on failures, not thirty. One failure showed that a popular feature leaked label data. Fixing that leak gave a 0.4% engagement gain on its own. The fourth bet is now in an A/B test at 1.1% gain. The VP leaves with a clear story: cheap failures, one real win, one bug fixed.

Common failure modes

Say this in a manager interview: “In science, many good bets fail. So I manage the portfolio, not single projects. Every bet has a hypothesis, a cost cap and a kill date. I report confidence levels to partners, so a failed bet is a planned outcome and not a surprise.”

Your job in one sentence and the four hats

The job in one sentence. Turn the right scientific bets into durable business impact, through a team that keeps getting stronger.

Each part of that sentence matters. “Right bets” means you choose the problems. “Durable impact” means it ships and stays shipped. “Keeps getting stronger” means people grow and stay. You wear four hats to do this. On any day one hat matters most. Over a quarter you need all four.

The four hats

Concrete practices

Scenario: the manager stuck in one hat

Marco manages nine scientists on ads quality. His calendar audit shows 60% technical reviews, 25% 1:1s, 10% partner meetings, and 5% planning. Two partner teams have stopped inviting him to roadmap talks. He hands the weekly model review to his strongest senior scientist, Lena. That frees six hours a week. He uses four for partner syncs and two for a monthly portfolio review. In one quarter, his team gets two new asks from product. Lena gets a clear growth story for her promotion case.

Common failure modes

Say this in a manager interview: “I think of the job as four hats: people, portfolio, partnerships and technical judgment. I track my time against them each month. When one gets thin, I delegate from another rather than drop it.”

The scientist-to-manager transition

The hardest part of the switch is not new skills. It is letting go of the old ones. You were rewarded for your own models. Now you are rewarded for other people's models. Your output is the team's output.

Core principles

How much to code

Ways to stay technical without blocking

Scenario: the new manager who kept the model

Wei was promoted to manage the five-person search relevance team. He kept ownership of the main ranking model. After two months, the model launch slipped three weeks. Wei had spent his time in hiring loops and planning. Two juniors had waited on his code reviews for days. He handed the model to Aisha, a senior scientist, with a clear launch goal. He took a side analysis on query intent drift instead. The model shipped four weeks later with a 2.3% gain in click-through. Aisha used the launch as the core of her promotion packet.

Common failure modes

For more on handing off work, see Delegating.

Say this in a manager interview: “When I moved into management, I gave away the main model. I kept small analyses that nobody waited on. I stay technical by reading every experiment write-up and reproducing one key result each quarter.”

Hiring scientists

Hiring is the highest-leverage thing you do. One great scientist can define a team for years. One poor hire can cost a year of your time. Science hiring is also hard to get right. Papers and degrees are easy to see. Taste and shipping ability are not.

Core principles

Role definition checklist

Loop design

Assessing research taste

Assessing shipping ability

Avoiding pedigree bias

Closing candidates

Scenario: the hire that the brand almost hid

Dana has one senior scientist slot on a fraud team. Two finalists remain. Candidate A has a top PhD and six papers at major venues, but no shipped models. Candidate B has a master's degree from a smaller school and four years at a fintech firm. B shipped a fraud model that cut chargebacks by 18%. In the applied round, A proposes a graph neural net and cannot say how to label the data. B asks about label delay first, then proposes a simple model with a plan to improve it. Dana's rubric weights framing and shipping for this role. She hires B. A year later B leads the team's main model and mentors two juniors.

Common failure modes

Say this in a manager interview: “I define the role by what the person must own in month six. My loop tests research taste and shipping, not just theory. I use a written rubric and blind scoring so the debrief is about evidence, not pedigree.”

Growing scientists

Your team gets stronger only if each person grows. Science growth is less obvious than engineering growth. It is not just bigger systems. It is better problem choice, sharper rigor, and wider influence.

Core principles

Scope vs complexity vs impact

A common stuck case is high complexity, low scope. The scientist does hard work on a narrow problem. They need a broader problem, not a harder one.

Concrete practices

Publication policy

Scenario: the strong scientist stuck at one level

Omar has been at the same level for two years. His models are some of the best on the team. His manager, Grace, reads the ladder with him. The gap is scope, not skill. Omar only works on the one model his tech lead hands him. Grace gives him a new problem: unify three teams' separate churn models into one shared model. It needs buy-in from two other managers. Omar struggles at first with the meetings. Grace coaches him on writing a one-page proposal. Nine months later the shared model saves 40% of training cost and lifts recall by 3 points. Omar is promoted in the next cycle.

Common failure modes

For more, see Growing leaders and Coaching senior people.

Say this in a manager interview: “I separate scope, complexity and impact when I coach. Most stuck seniors do hard work in narrow scope. So I find them a cross-team problem, sponsor them into it, and coach them on influence.”

Performance management for science

Judging science work is hard because outcomes are noisy. A good scientist can have a bad year of results. A weak one can ride a lucky launch. Your job is to judge the quality of the bets and the work, and then weigh the impact fairly.

Core principles

Judging work when experiments fail

Rewarding good bets, not lucky outcomes

Writing calibration packets

Handling underperformance

Scenario: two scientists, two very different years

Sam and Jo both work on recommendations. Sam shipped one model with a 2% gain. But the review showed his baseline had a bug. The fair gain was closer to 0.6%. He also skipped the holdout test on a second launch. Jo ran three bets. Two failed offline within their caps. Her write-ups stopped two other teams from trying the same ideas. Her third bet is in an A/B test showing 1.4%. Their manager, Ken, rates Jo above Sam. In calibration he explains the baseline bug and Jo's saved effort in dollar terms. He also gives Sam clear feedback on rigor with a plan for next half.

Common failure modes

For more, see Underperformance and Hard feedback.

Say this in a manager interview: “I judge science on the quality of the bet and the work, then weigh impact over the full year. I ask for a short pre-registration note before big experiments. That lets me reward a sound bet that failed, and question a lucky win.”

1:1s and coaching scientists

The 1:1 is the main tool you have for each person. It belongs to them, not to you. Done well, it catches problems early and grows people. Done badly, it is a status meeting that wastes 30 minutes.

Core principles

A simple 1:1 agenda

  1. Their topics. What is on your mind?
  2. Blockers. What is slowing you down?
  3. Feedback both ways. One thing that went well, one thing to change.
  4. Your topics. Context they need, decisions coming.
  5. Once a month: career and growth only.

Technical vs career 1:1s

Giving feedback on research quality

Scenario: catching a flaw without crushing a junior

Nina is six months into her first job. She shows her manager, Tom, a model with a 9% offline lift. Tom suspects leakage, since gains on this task are usually 1 to 2%. He does not say that. He asks, “What is the biggest single feature by importance?” It is a timestamp field. He asks how that field is set. Nina realizes it is filled after the label event. She fixes it, and the lift falls to 1.3%. Tom praises her for tracing it fast. In the next team review, Nina presents the leak as a lesson. Two other scientists find the same field in their models.

Common failure modes

Say this in a manager interview: “My 1:1s start with the scientist's agenda. I keep career talks separate and monthly, so technical topics do not crowd them out. When I see a flaw in an analysis, I ask questions that lead them to it. They learn more that way.”

Motivation and retention

Scientists are hard to hire and easy to lose. Most leave for one of three reasons. They lost interest in the work. They stopped growing. Or they felt like a service desk. Pay matters, but it is rarely the only reason.

Core principles

Concrete practices

Avoiding the ticket factory

A ticket factory is a team that only takes requests. Product sends tickets. Scientists tune models. No one owns a problem end to end. Top people leave first.

Burnout

Scenario: turning around a team that felt like a service desk

Elena takes over a six-person team that supports four product groups. Attrition was 40% last year. Exit interviews say “no ownership”. She agrees with her director that the team will own one metric: recommendation diversity. She caps support requests at 40% of capacity, enforced through a simple weekly intake. Each scientist gets one problem they own from start to launch. She funds two conference trips and starts a reading group. In the next year attrition drops to one person. The team ships a diversity model that lifts long-term retention by 0.8%.

Common failure modes

For more, see On-call and burnout.

Say this in a manager interview: “Scientists stay for ownership, growth and recognition. So I make my team own a metric, not a ticket queue. I cap request work at a set share of capacity. I watch for failure streaks and make sure the next project has a good chance to land.”

Team structure

Where scientists sit shapes what they work on. There is no single right model. Each one trades depth for closeness to product. Good managers know the trade-offs and change the model as the org grows.

The three common models

Embedded

Centralized

Hub-and-spoke

Team size

Ratio of scientists to engineers

Scenario: moving from embedded to hub-and-spoke

A growing marketplace has 22 scientists spread across nine product teams. Each team uses its own A/B test method. Two launches were rolled back after false wins. Science promotion rates lag engineering by half. The new science director, Ravi, moves to hub-and-spoke. Scientists now report into three science managers but keep their desks and standups with product. He sets one shared experiment standard. Within a year, false-win rollbacks drop to zero. Promotion rates match engineering. Product leads say they still get the same daily support.

Common failure modes

For more on growing headcount, see Scaling a team.

Say this in a manager interview: “At scale I prefer hub-and-spoke. Scientists sit with product for context and report into science for rigor and growth. I watch how long proven models wait to ship. That tells me if my engineer ratio is right.”

Building trust with product and engineering

Science only creates value when it ships. That needs partners who trust you. Product and engineering leads plan in quarters with fixed dates. Science gives uncertain answers on uncertain timelines. Your job is to bridge that gap.

Core principles

Roadmaps with uncertainty

Concrete practices

Moonshots big if it works, likely fails 20-30% of capacity Core bets big and probable about 50% of capacity Drop small and unlikely near 0% Quick wins small but sure, builds trust about 20% of capacity chance of success size of impact
The portfolio quadrant. Quick wins buy partner trust. Moonshots keep the team ahead. Core bets pay the bills. See M2 for how to run the portfolio week to week.

Scenario: the roadmap that survived a failed bet

Lucas runs science for a payments product. His PM partner wants a new risk model in Q3 for a market launch. Lucas splits his plan into tiers. Committed: retune the current model on new market data for a 5% cut in fraud loss. Exploratory: a new sequence model that might cut loss by 15%. He sets a decision gate for June 1. The sequence model fails on the gate because of sparse data in the new market. The retuned model ships on time and cuts loss by 6%. The PM tells her VP that science “always does what it says”. Next quarter she funds two engineers for the sequence model once data matures.

Common failure modes

For more, see Engineering and product friction and Influence without authority.

Say this in a manager interview: “I give partners a tiered roadmap: committed, likely and exploratory. They plan only on committed. Each exploratory bet has a decision date and a fallback. So when a bet fails, the product still ships, and partners keep trusting us.”

Technical leadership as a manager

Your team's science is only as good as the bar you set. You may not write the code, but you own the quality of what ships. A science manager who cannot catch a flawed experiment is a risk to the business.

Core principles

Reviewing designs

Reviewing experiments

Raising rigor across the team

Scenario: the false win caught in review

Kim's team reports a 3.2% lift in purchases from a new ranking model. Launch is scheduled for Friday. In the weekly review, Kim asks to see the daily trend. The lift was 6% in week one and 0.5% in week two. She also notes the team checked the result four times and stopped on a good day. She asks for one more clean week with a fixed stop date. The final lift is 0.7%. The team still ships, since the cost is low. But they report the honest number. Two months later finance checks the impact. It matches 0.7%. Kim's team gains a reputation for numbers finance can trust.

Common failure modes

For related practice, see the Practitioner's Toolkit and Review philosophy.

Say this in a manager interview: “I own the rigor bar even when I do not write the code. We use a design template, a written experiment standard, and a weekly model review. I review the question and metric before the method. I train seniors to review, so I am not the bottleneck.”

Your first 90 days

The first 90 days set how the team sees you for a long time. Move too fast and you break things you do not understand. Move too slowly and people wonder why you are there. Use a simple order: listen, map, quick win, set direction.

Days 1 to 30: listen

Days 31 to 60: map

Days 61 to 90: quick win and set direction

Scenario: a new manager on an inherited team

Fatima joins as manager of an eight-person NLP team. In month one, her 1:1s show three themes. Experiments take three weeks to set up. Two bets have run for nine months with no gate. And the PM partner thinks the team is “a black box”. In month two she maps the portfolio and finds 30% of capacity on the two stalled bets. In month three she kills one bet and sets a gate on the other. She gives two scientists four weeks to cut experiment setup time. It drops to four days. She shares a one-page direction doc with a tiered roadmap. The PM calls it the first science plan she could plan around.

Common failure modes

For a broader view, see the notes on The Manager's Path and Leading through ambiguity.

Say this in a manager interview: “In my first 90 days I listen for a month and map for a month. Then I land one quick win and write a one-page direction doc. I share my learning doc early, so the team can correct me before I set direction.”
The thread through every topic. Science fails often by nature. A great science manager makes failure cheap, visible and useful. That frees the team to take the big bets that matter.

Recap

← T6 — The Practitioner's Toolkit M2 — Running Science: Portfolio, Process and Impact →