Part V Manager 2 Execution

Running Science: Portfolio, Process and Impact

A science manager runs a portfolio of bets. The job is to pick good bets, kill bad ones early, and prove the winners paid off.

The first manager page covered people. This page covers the machine. It explains how a strong manager plans, reviews, measures and scales science work. Each section gives principles, a template you can reuse, a short story with numbers, and the ways it goes wrong. The last section has ten interview questions with strong answers.

The viewpoint is a manager of 8 to 40 scientists at a large tech company. The team ships models into real products. Leaders judge it on business results, not papers.

Contents

  1. Managing a research portfolio
  2. Planning under uncertainty
  3. The experiment review culture
  4. Measuring and communicating impact
  5. Roadmaps and OKRs for science
  6. Operating rhythm
  7. Working with engineering and product
  8. Build vs buy, and foundation models
  9. Managing compute and data budgets
  10. Publishing, patents and external presence
  11. Managing risk and responsible AI reviews
  12. Scaling the org
  13. Manager interview questions

Managing a research portfolio

A science team is not a feature team. Most of its ideas will fail. That is normal. The manager's job is to make the whole set of bets pay off, even when single bets lose.

Research portfolio. The full set of projects a team funds at once. Each one has a size, a time horizon, a chance of success and a payoff. You manage the mix, not just the parts.

The 70/20/10 split

Most mature teams split effort into three buckets. The numbers are a starting point, not a law.

The right split depends on the team's stage. A new team with no shipped model may run 90/10/0. It must earn trust first. A team that owns a mature ranker may run 60/25/15. Its core gains are getting smaller each half.

70% Core 20% Adj. 10% Share of scientist time Odds of success 60 to 80% 30 to 50% 5 to 20% Time to payoff 1 quarter 2 to 4 quarters 4+ quarters Small slices carry the biggest upside. Protect them from being eaten by the core.
A 70/20/10 portfolio. Odds fall and horizons grow as you move right.

Horizons

Tag every project with a horizon. It sets how you review it and when you expect a result.

Expected value of a project

You cannot rank projects without a number. A rough one is fine. Use this simple form.

EV = P(success) × payoff − cost. Then divide by scientist-months to get value per month. Rank on that. A good guess beats no guess. The act of writing the number exposes weak plans.

Bet sizing

Scenario: the ads ranking team

Priya runs 14 scientists on ads ranking. Her list has nine candidate projects for the half. She scores each one.

She funds the feature refresh fully. She funds the multi-task model with a gate at week 8. She funds the LLM bet with one scientist for six weeks. The LLM bet fails its first gate. She shifts that scientist to the multi-task work, which passes its gate. The half ends with $14M in shipped gains and one clean failure.

Failure modes

Say this out loud: “I run the team as a portfolio. About 70% goes to core models, 20% to adjacent bets, and 10% to frontier work. I rank projects by expected value per scientist-month. I fund risky bets small and grow them only after a gate.”

Planning under uncertainty

Science plans fail when they look like engineering plans. You cannot promise “model done by June” when nobody knows if the model will work. You can promise to answer key questions by set dates.

Milestones as questions answered

Write each milestone as a question with a yes or no answer. The answer decides the next step.

A question milestone can fail and still be a success. A clear “no” saves months. Celebrate it.

Time-boxed spikes

Spike. A short, fixed-length test of one risky idea. Usually one or two scientists for two to six weeks. The goal is learning, not shipping.

Kill criteria up front

Decide how you will stop before you start. Once people have spent months, they always find a reason to keep going.

Kill criteria template. “We will stop this project if any of these are true by DATE. (1) Offline gain is below X. (2) Latency at p99 exceeds Y ms. (3) Online test shows no lift above Z with 80% power. (4) Data we need is not approved by legal.” Sign it with the product partner.

Stage gates

Large bets move through gates. Each gate has a clear bar and a small review. Staffing grows only after a gate passes.

Idea 1-page memo Spike 1 sci, 4 wks Prototype 2-3 sci, 8 wks Online test + eng, 6 wks Launch full team gate 1~50% killed gate 2~50% killed gate 3~30% killed gate 4~20% killed Cost per stage grows about 3x. Kill early, when it is cheap. Of 20 ideas, about 3 reach launch. That is a healthy funnel, not a broken one.
Stage gates. Staffing grows only after a bet passes a gate. Most kills should happen at the first two gates.

Scenario: the fraud model rewrite

Marcus leads 6 scientists on payment fraud. His director wants a new sequence model by Q3. Marcus does not promise a date for the model. He promises three answers instead.

Week 4 says yes, with a 7% gain. Week 10 says no, at 45 ms. The team spends three weeks on distillation and gets to 18 ms with a 6% gain. Week 16 confirms it. They launch in week 20. The director gets a later date than hoped but no surprises.

Failure modes

Say this out loud: “I plan science as questions with dates. Each milestone has a yes or no answer. We write kill criteria before we start and sign them with product. Staffing grows only when a bet passes a gate.”

The experiment review culture

A team's results are only as good as its review habits. One bad launch based on a false win can cost more than a year of real gains. Rigor is a team habit, not a personal trait.

Pre-registration of metrics

Before an A/B test starts, the owner writes down what will count as success. This stops people from picking the best-looking metric after the fact.

Pre-registration template.

The review checklist

Every launch review runs the same list. A senior scientist who did not work on the project runs it.

  1. Sample ratio. Is the split what you planned? A mismatch means a bug.
  2. Pre-period check. Did the groups look the same before the test? An A/A test helps.
  3. Novelty. Does the effect fade over the weeks? Plot it by day.
  4. Interference. Can treated users affect control users? Common in markets and social apps.
  5. Multiple tests. How many metrics and segments were checked? Adjust for it.
  6. Offline to online. Does the online gain match the offline gain in sign and rough size?
  7. Guardrails. Any guardrail outside its limit blocks the launch.
  8. Long-term risk. Could this help now and hurt in six months? Plan a holdout.

Launch criteria

Avoiding p-hacking in teams

P-hacking is rarely fraud. It is usually tired people under pressure trying one more cut. Build a system that makes it hard.

The peeking trap. A scientist checks the test daily and stops it the first day p < 0.05. This can push the false win rate from 5% to over 25%. Use a fixed run length or a sequential test built for peeking.

Scenario: the search launch that was not

A search team ran a new ranker for two weeks. The owner reported a 1.2% lift in clicks. The reviewer checked the pre-registration. The primary metric was purchases, not clicks. Purchases were flat at +0.1%, with an interval of −0.4% to +0.6%. The reviewer also found 14 segments in the deck. Only the three best were shown on the first slide.

The manager did not blame the owner. She changed the template. Every deck now opens with the pre-registered primary metric. The team reran the test with a fix for a known position bias. Purchases rose 0.5% and the model shipped.

Failure modes

Say this out loud: “Every test is pre-registered with one primary metric and guardrails. A scientist outside the project runs a fixed review checklist. We reward clean tests, including nulls. If our win rate gets too high, I treat that as a warning sign.”

Measuring and communicating impact

Science that leaders cannot see does not get funded. You must turn model gains into business terms. You must do it in a way that holds up when finance checks it.

From offline gains to business dollars

Build a chain from the model metric to money. Each link should rest on a past test, not a guess.

  1. Offline metric. AUC up 0.8%, or NDCG up 2%.
  2. Online metric. A/B test shows conversions up 0.6%.
  3. Business metric. 0.6% of conversions is 120,000 more orders a month.
  4. Dollars. At $8 margin per order, that is about $11.5M a year.

Agree on the conversion rates with finance once. Then use the same rates in every report. This stops each team from inflating its own math.

Attribution

Many teams touch the same metric. If every team claims the full gain, the total is three times the real number. Leaders notice.

The impact doc

Impact doc template (one page).

Executive updates

Scenario: the recommendations team

Lena's team shipped four model changes in a half. Each test showed lift. Added up, they claimed 6.2% more watch time. The long-term holdout showed 3.9%. Some gains overlapped and one faded.

Lena reported 3.9% to her VP. She explained the gap in one line. Finance trusted her numbers from then on. The next planning cycle, her team got two of the three extra heads it asked for. A peer team that reported raw sums got none.

Failure modes

Say this out loud: “I tie every launch to dollars using rates we agreed with finance. I report holdout numbers, not summed test lifts. Every impact doc includes cost and names everyone who helped.”

Roadmaps and OKRs for science

A science roadmap must be honest about risk. It must also show leaders what they get. OKRs (goals plus measurable key results) help when written well.

Outcome vs output

Write key results as outcomes when you can. That leaves room to find the best method. Use output key results for platform work or learning goals. Even then, say what the output enables.

How to write science OKRs

OKR template.

The science roadmap

Show confidence on each item. A simple high, medium or low works. Update the roadmap every month. Leaders forgive change. They do not forgive surprise.

Scenario: the trust and safety team

A safety team had an OKR that read “Ship v3 classifier.” They shipped it on time. Harmful content views did not drop. The new manager rewrote the OKR as “Cut harmful views per million by 15%.” The team found that half the gap came from new content types the model never saw. Better labeling did more than a new model. Views fell 18% that half.

Failure modes

Say this out loud: “I write science OKRs as outcomes, plus one learning goal for risky bets. Committed goals should land 90% of the time. Stretch goals 50%. The roadmap shows now, next and later, with confidence on each.”

Operating rhythm

A good rhythm makes work visible without slowing it down. It gives people regular places to get feedback, share ideas and make decisions.

The core meetings

Weekly science review format

Science review template.

Protect deep work

Scenario: a team drowning in syncs

A new manager joined a team of 11. Scientists sat in 9 hours of recurring meetings a week. Three meetings covered the same project status. She replaced them with one written weekly tracker and one science review. Meeting time fell to 4 hours. Experiments launched per scientist rose from 1.1 to 1.7 a month over the next quarter.

Failure modes

Say this out loud: “I run a weekly science review and a weekly launch review. I add a reading group, a demo day and quarterly planning. Status goes in writing. Meetings are for decisions.”

Working with engineering and product

Science alone ships nothing. A model reaches users through engineering and product. Most failed science projects fail at this seam, not in the math.

RACI for a model launch

RACI names who is Responsible, Accountable, Consulted and Informed for each task. Write it at the start of any project that crosses teams.

Handoff vs embedded

Own the launch, not just the model. A scientist who stays with a model through launch ships twice as often. Make launch ownership part of your team's job, and part of how you rate people.

Infra asks

Science often needs engineering help. A feature store, a new logging field, faster training. These asks compete with product features.

Scenario: the stalled pricing model

A pricing science team built a model with a 4% margin gain offline. It sat for two quarters. Engineering had not planned for it. The new manager wrote a RACI with the engineering lead. She moved one scientist to sit with the serving team for eight weeks. The scientist wrote the feature pipeline. Engineering owned serving. The model launched in nine weeks and delivered a 2.7% margin gain online.

Failure modes

Say this out loud: “I write a RACI at the start of every cross-team project. Product owns the metric, engineering owns serving, and science owns model quality. My scientists own the launch, not just the model. Infra asks come with a dollar value.”

Build vs buy, and foundation models

Many problems no longer need a custom model. A large language model or a vendor API may get you 80% of the way in a week. The manager must decide when that is enough.

Questions to ask

The usual path

  1. Prompt a foundation model. Fast baseline in days. Good for labeling, prototypes and low-volume tasks.
  2. Add retrieval or few-shot examples. Often closes much of the gap.
  3. Fine-tune a mid-size model. When volume is high or quality must rise.
  4. Distill to a small custom model. When latency and cost per call matter at scale.
  5. Train from scratch. Rare. Only for core models with huge data and a clear edge.

A cost check

Do the math early. Say a task gets 50 million calls a day. An LLM API at $0.002 per call costs $100,000 a day. That is $36M a year. A distilled model on owned hardware might cost $2M a year. If the quality gap is small, the custom model wins. At 50,000 calls a day, the API costs $36,000 a year. Buy.

Scenario: support ticket routing

A support team wanted a new ticket classifier. The old custom model hit 82% accuracy. A scientist tried a prompted LLM in three days. It hit 88%. Volume was 200,000 tickets a day. The API cost was about $150,000 a year. The team shipped the LLM version in four weeks. They used its labels to train a small model six months later. That model hit 89% at a tenth of the cost.

Failure modes

Say this out loud: “I start with the cheapest thing that might work, often a prompted foundation model. I move to fine-tuning or a custom model when volume, latency, cost or edge demand it. I do the cost math in week one.”

Managing compute and data budgets

Compute is now one of the biggest costs on a science team. GPUs can cost more than salaries. A manager who cannot explain the compute bill will lose it.

GPU allocation

Cost per experiment

Track what each experiment costs. Make it visible. People spend less when they can see the bill.

Data access and privacy

Scenario: the runaway training bill

A recommendations team spent $1.8M on GPUs in one quarter. The budget was $1.1M. The manager tagged jobs and found 40% of spend came from hyperparameter sweeps at full scale. Two sweeps had run for a week after the owner left on vacation. He set a rule. Sweeps run at 10% data first. Any run over $20,000 needs a lead's approval. Next quarter's spend was $1.05M. The number of launches did not change.

Failure modes

Say this out loud: “I treat compute like headcount. Production gets reserved capacity, research gets quotas, and big bets draw from a gated pool. Every job is tagged, so I know the cost per project. Data approvals go into the plan from day one.”

Publishing, patents and external presence

External work can help hiring, retention and the company's name. It can also leak secrets and distract from the product. A manager needs a clear policy.

When publishing helps the business

When it does not

A simple policy

  1. Product work comes first. Papers come from work you were doing anyway.
  2. Every paper goes through legal and comms review before submission.
  3. Remove business numbers. Use public datasets for headline results when you can.
  4. File patents before you publish, if the method is worth protecting.
  5. Count papers in reviews only when they tie to product or hiring value.

Patents

Scenario: the conference paper debate

A scientist wanted to publish a new ranking loss. It had added 0.7% to revenue. The manager asked legal and the product lead. They agreed the loss was general but the features were not. The team filed a patent on the loss. They wrote the paper using two public datasets. Revenue numbers were left out. The paper was accepted. Two strong candidates later named it as why they applied.

Failure modes

Say this out loud: “We publish when it helps hiring or quality and does not give away our edge. Papers come from product work, not side projects. Legal reviews every paper, and we file patents first when the method is worth protecting.”

Managing risk and responsible AI reviews

Models can harm users, break laws, or embarrass the company. A science manager owns these risks for the team's models. Good process catches them early, when fixes are cheap.

The main risk types

The responsible AI review

Review checklist.

Make it part of the flow

Scenario: the hiring screen model

A team built a model to rank job applicants for recruiters. Offline accuracy was strong. The fairness check showed women were ranked 12% lower on average at equal outcomes. The cause was a feature tied to past job titles. The team removed it and added a group-wise check to the launch criteria. Accuracy fell 1%. The gap fell to under 2%. The launch went ahead six weeks late but cleared legal review.

Failure modes

Say this out loud: “We run a light risk check at the first gate and a full responsible AI review before any online test. We report results by group, not just averages. Every risk has an owner in a register, and monitoring is part of launch.”

Scaling the org

What works for 5 scientists breaks at 15. What works at 15 breaks at 50. The manager must change the structure before it breaks, not after.

Stages of growth

5 to 10 scientists

10 to 25 scientists

25 to 50 scientists

Key roles

Two ladders, equal respect. Strong scientists must be able to grow without managing. Make the staff and principal path as visible and rewarded as the manager path. Otherwise your best ICs become unhappy managers.

Picking your first managers

Scenario: from 8 to 35 in two years

Ravi started with 8 scientists on a growing marketplace. At 14 he split into ranking and pricing, each with a tech lead. At 18 he hired an outside manager for pricing and promoted a senior scientist to manage ranking. At 28 he added a staff scientist for shared modeling and a program manager. At 35 he had four managers, two staff scientists, and a shared eval platform. Attrition stayed under 8% a year. Launches per quarter rose from 3 to 11.

Failure modes

Say this out loud: “I change structure before it breaks. Past about 12 people I split into sub-teams with tech leads. Past 25 I hire managers and add staff scientists. I keep the IC ladder as strong as the manager ladder.”

Manager interview questions

These are common questions for science manager roles. Each strong answer uses a first-person story with numbers. Adapt it to your own experience. Never memorize someone else's story.

1. How do you prioritize a science roadmap? Medium

What they are testing

A strong answer

“I use a portfolio approach. Each half, I gather every candidate project from the team, product and leadership. Last cycle that was 22 ideas for 16 scientists.

I score each one on four things. Chance of success, annual payoff in dollars, cost in scientist-months, and time to payoff. I use dollar rates we agreed with finance, so the numbers compare fairly. Then I rank by expected value per scientist-month.

I do not just take the top of the list. I aim for about 70% core work, 20% adjacent and 10% frontier. Last half, the top eight items were all core. I pulled in one frontier bet on LLM features, funded at one person for six weeks.

I review the draft with my product partners and two senior scientists. They catch bad estimates. Then I publish the list with what we are not doing and why. That list matters as much as the yes list.

The result was 11 funded projects. Seven shipped, three were killed at gates, and one carried over. Shipped gains were about $19M a year.”

Follow-ups

Say this out loud: “I rank by expected value per scientist-month, then shape the mix to 70/20/10. I publish what we will not do. Last half, 7 of 11 projects shipped for about $19M a year.”

2. Tell me about a project you killed Medium

What they are testing

A strong answer

“We had a project to replace our search ranker with a graph neural network. It had three scientists and strong support from a senior leader. I set kill criteria at the start. We needed a 1.5% offline NDCG gain by week 10 and p99 latency under 25 ms.

At week 10 we had a 0.9% gain. Latency was 60 ms. The team asked for six more weeks. I asked what would change. Their plan was more tuning, with no new idea for the latency gap.

I killed it. First I told the three scientists in person. I explained the criteria we had all signed. I made clear this was a project decision, not a judgment of their work. Then I told the senior leader with a one-page memo. It showed the numbers, the cost of six more weeks, and what we learned.

We moved two of the scientists to a feature project. One graph feature from the dead project became an input there. It added 0.4% to purchases.

The lead scientist got credit in her review for the clean test and the memo. She told me later that the quick decision helped her. She did not lose a year to a dead end.”

Follow-ups

Say this out loud: “I set kill criteria up front and signed them with the team. When we missed them, I stopped the project, told the team first, and gave leaders a memo. We reused one piece and credited the scientists for a clean result.”

3. A scientist whose work is brilliant but never ships Hard

What they are testing

A strong answer

“I had a scientist, call him Dev, with two strong papers from his first year. Neither idea had reached production. His peers saw him as the smartest person on the team. His rating was at risk because the role expects shipped impact.

I first looked for the cause. I read his last three project docs and talked to his engineering partners. The pattern was clear. He stopped at the offline result. Then he moved to the next interesting problem. Nobody owned the path to launch.

I had a direct talk with him. I said his ideas were the best on the team. I also said our job is to change the product, and his work was not doing that yet. I showed him the rating guide.

We agreed on a plan. He would take his best idea all the way to an A/B test in one quarter. I paired him with a strong ML engineer. I set check-ins on launch steps, not on model gains.

It took 14 weeks. The model lifted engagement 1.1% and shipped. Dev found he liked seeing users react. Over the next year he shipped three more models. He still wrote one paper a year, now based on shipped work.”

Follow-ups

Say this out loud: “I find the real cause first. Here it was no ownership past the offline result. I set a goal to take one idea to an A/B test. I paired him with an engineer and tracked launch steps. He shipped in 14 weeks and kept shipping.”

4. How do you measure your team's impact? Medium

What they are testing

A strong answer

“I measure impact at three levels.

The first is shipped business value. Every launch has an A/B test with a pre-registered primary metric. We convert lifts to dollars using rates set with finance. For example, one point of conversion is worth $7M a year on our surface.

The second is honest attribution. Summed test lifts overstate the truth. So we keep a 2% long-term holdout from all our launches. Last year our tests summed to 5.8% engagement. The holdout showed 4.1%. I report 4.1%.

The third is enabling work. Platform work does not move a metric by itself. I measure it by what it unlocked. Our new eval pipeline cut the time from idea to test from 19 days to 6. That doubled the number of tests we could run.

I also track health signals. Test velocity, launch rate, cost per test, and the share of tests that win. A win rate over 60% tells me our bar is too low.

Each quarter I write a one-page impact summary. It names every person who contributed, including partners in engineering.”

Follow-ups

Say this out loud: “I measure shipped value in dollars and report holdout numbers, not summed lifts. I credit platform work by what it unlocks. Last year our tests summed to 5.8%, but I reported the 4.1% holdout.”

5. Conflict between research and product deadlines Hard

What they are testing

A strong answer

“Our product team had promised a new personalized home feed for a holiday launch. The date was fixed by marketing. My team's new model was eight weeks from ready. The launch was in five.

I met with the product lead to understand what was truly fixed. The date was. The full model was not. What mattered was that the feed felt personal for returning users.

I proposed two tracks. Track one was a simple version for the launch. We would use our existing model with three new features. That gave about 60% of the expected gain. I was confident it could pass review in four weeks. Track two kept the full model on its path, with a test planned for January.

I wrote this up with risks and gains for each option. I shared it with product, engineering and my director in one meeting. We agreed on the plan in 30 minutes.

The simple version launched on time and lifted engagement 2.1%. The full model shipped in January and added another 1.6%. I did not cut the experiment review on either one. That was my line, and I said so early.”

Follow-ups

Say this out loud: “I find what is truly fixed. Then I offer a smaller version that meets the date and keep the full work on its own track. I wrote both options up and got a decision in one meeting. I never cut the review step.”

6. Hiring a senior scientist: what do you look for? Medium

What they are testing

A strong answer

“For a senior scientist, I look for four things.

First, depth. They must know one area far better than anyone on my team. I test this with a deep dive on their own past work. I keep asking why until they reach the edge.

Second, problem framing. Given a vague business problem, can they turn it into a clear ML problem? I give a real but simplified problem from our team. Strong people ask about the metric and the data before the model.

Third, shipped impact. They should have taken models to production and measured the result online. I ask for numbers and what broke after launch.

Fourth, multiplier behavior. Seniors make others better. I ask for times they mentored someone, set a technical direction, or changed a team's methods.

I also hire for gaps. Last year my team was strong in ranking and weak in causal inference. I wrote the role around that. The person I hired built our uplift modeling. Her first project improved coupon spend efficiency by 22%.

I use a written rubric for every round and hold a debrief with evidence, not feelings.”

Follow-ups

Say this out loud: “I look for depth, problem framing, shipped impact, and multiplier behavior. I hire for the team's gaps. I use a written rubric and evidence-based debriefs. My last senior hire filled our causal inference gap and cut coupon waste by 22%.”

7. A failed launch you owned Hard

What they are testing

A strong answer

“We launched a new delivery time model at a food app. The A/B test showed a 0.8% lift in orders. We shipped to everyone. Within a week, support tickets about late orders rose 30%.

The model was more optimistic in busy hours. That pulled orders, but many arrived late. Our test ran for two weeks during a quiet season. Our guardrails tracked cancellations but not lateness.

I owned it. I rolled back within 24 hours of finding the cause. I sent a short note to product and support leaders the same day. It said what happened, what we did, and what came next.

Then I ran a blameless review with the team. We found three gaps. First, no lateness guardrail. Second, a test window that missed peak load. Third, no check on calibration by time of day.

We fixed all three. Lateness became a standard guardrail for every logistics model. Tests now must cover at least one peak period. Calibration by segment joined our review checklist. We relaunched six weeks later. Orders rose 0.6% and late deliveries fell 4%.

The lesson for me was that the review checklist is a living thing. I now update it after every incident.”

Follow-ups

Say this out loud: “I owned it, rolled back in a day, and told leaders the same day. A blameless review found three gaps in our process. We fixed all three and relaunched with a better result. I now update our checklist after every incident.”

8. How do you raise scientific rigor? Medium

What they are testing

A strong answer

“When I joined my current team, launch reviews were informal. Each owner chose their own metrics. In my first month I re-analyzed ten past launches. Three of the claimed wins did not hold up. One had a sample ratio mismatch. Two had picked the best metric after the test.

I made four changes. First, a one-page pre-registration for every test, filed before launch. Second, a fixed review checklist run by a scientist outside the project. Third, a shared results log, so we could see all past tests. Fourth, I changed how we rate people. Clean tests count, even nulls.

I knew process could slow us down. So I added a light path for small changes. It uses a checklist of five items and needs no meeting.

After two quarters, our win rate fell from 70% to 45%. That sounds bad but it meant the bar was real. Our holdout gain rose from 2.1% to 3.4% a year. Fewer fake wins meant more real ones. Review time per launch stayed under three days.”

Follow-ups

Say this out loud: “I audited past launches to show the gap. Then I added pre-registration, an outside reviewer, a results log and credit for nulls. Our win rate fell, but holdout gains rose from 2.1% to 3.4%. Real rigor means more real wins.”

9. How do you handle a low performer? Hard

What they are testing

A strong answer

“I had a mid-level scientist, call her Ana, who missed three milestones in a row. Her peers were starting to carry her work.

First I looked for causes. I asked her directly how things were going. It turned out she was struggling with a new area. She came from NLP and was now on tabular ranking. She was also hesitant to ask for help.

I was clear with her. I said her work was below the bar for her level, and gave three concrete examples. I also said I wanted her to succeed and believed she could.

We wrote a plan together with three goals over eight weeks. Each was concrete. Ship one feature test. Write a design doc reviewed by a senior. Present at science review. I paired her with a senior scientist for two hours a week. I met her weekly to check progress.

She met two of three goals in eight weeks, and the third two weeks later. Her next rating was at the bar. In other cases, the plan has not worked. Then I move to a formal process with HR. I am honest with the person about where it is heading. Waiting too long is unfair to them and to the team.”

Follow-ups

Say this out loud: “I act early, find the cause, and give clear feedback with examples. We set concrete goals with support and weekly check-ins. If it works, great. If not, I move to a formal process and stay honest with the person.”

10. Why should a scientist join your team? Medium

What they are testing

A strong answer

“I give candidates four reasons, and one honest warning.

First, the problems are real and big. Our models touch 200 million users a day. A good idea can show up in a test within three weeks.

Second, the work ships. Last year, 70% of our funded projects reached an online test. Scientists own their models through launch, so they see the results.

Third, the bar is high and fair. Every test is pre-registered and reviewed. People get credit for clean nulls, not just wins. That makes it a safe place to try hard ideas.

Fourth, people grow. I hold a career talk with each person every quarter. In the last two years, six of my scientists were promoted. Two moved to staff. We also keep a 10% frontier slice, and we publish when it does not give away our edge.

The warning is that we are not a pure research lab. If you want to write papers with no product link, this is not the right fit. I would rather say that now.

The best way to judge is to meet the team. I always offer a call with two scientists who are not in the loop.”

Follow-ups

Say this out loud: “Big real problems, work that ships, a high and fair bar, and real growth. Six promotions in two years. I am honest that we are not a pure research lab. Candidates can talk to my team directly.”

Recap

← M1 — Being a Successful Applied Science Manager All topics →