Lakshya

Data Platform & Analytics

Data Scientist

Data · statistics in 44% of the reader sample

The archetype most often confused with two others. It is not ML engineering and it is not analytics engineering, and the interview tests the difference.

Data scientist is the least well-defined title in this atlas, which is precisely why it is worth pinning down. In practice the postings cluster into three quite different jobs and you should establish which one you are interviewing for in the first call, because preparing for the wrong one is a wasted month.

FlavourWhat you actually doAdjacent to
Product / experimentationA/B tests, causal inference, metric definition, telling product what happened and whyAnalytics Engineer
ModellingBuild models that predict something, hand to engineering or ship yourselfML Engineer
Research / decision scienceAnswer open business questions with whatever method fits; often forecasting or optimisationStrategy and finance

A day — product flavour, which is the most common

TimeWhat you are actually doing
09:00An experiment read-out. The primary metric is flat and three secondary ones moved. The job is refusing to report the secondary ones as the finding.
10:30Someone wants a test that would take fourteen weeks to reach power. You say so, and propose what could actually settle the question instead.
13:00Segment analysis. The effect reverses when split by platform, which is either Simpson's paradox or a real heterogeneous effect, and the difference matters.
14:30Writing. The output of this role is a document that changes a decision, not a notebook.
16:00Definition work with analytics engineering, because half of your disagreements are really about what 'activated' means.

The decisive round

“The experiment is flat. What do you tell the team?”

Interviewer: “You ran a two-week test on a new onboarding flow. The primary metric is flat. Product wants to ship it anyway because two secondary metrics improved. What do you say?”

A weak answer. “I'd explain that the primary metric didn't move, but note the positive secondary signals, and suggest we could ship and monitor.” Concedes the argument while appearing to hold it. It also treats twenty comparisons as if one of them being significant were evidence.

A strong answer. “First I would establish whether flat means ‘no effect’ or ‘this test could never have detected the effect we care about’, because those lead to opposite decisions and people conflate them constantly. If the minimum detectable effect at this sample size was five percent and we would have been happy with two, the test has told us almost nothing and the honest report is that it was underpowered — which is my mistake at design time, not a finding.

On the secondary metrics: with enough of them, something is always significant. If we did not name a primary metric beforehand, we do not have a result, we have a search. I would report them as exploratory and explicitly not as evidence — and if one of them is genuinely interesting, the correct next step is a new test with that as the primary, not shipping on it.

Then I would separate the statistics from the decision, because they get muddled and it damages trust. The data says we have no evidence this improves the primary metric. Whether to ship anyway is a product judgement, and there are legitimate reasons — strategic bet, reduces support load, unblocks something else. I would rather they ship it with clear eyes than have me manufacture a justification.

What I would not do is let ‘flat but directionally positive’ enter the record, because that phrase is how organisations accumulate a portfolio of changes none of which did anything.”

Distinguishing underpowered from null, refusing multiple-comparisons evidence, and separating the statistical claim from the product decision. That last separation is what makes a data scientist trusted rather than routed around.

What to build

  • Run one real experiment end to end — even on a small side project. Power calculation first, primary metric named in advance, and a written read-out.
  • Do one causal analysis without randomisation: difference-in-differences on public data is a weekend, and it is the method most often needed in practice.
  • Have a result you refused to report. This is the strongest single story for this archetype and almost nobody brings one.
  • Write for a non-technical reader. The output of this job is a decision changed, and a notebook has never changed one.

Red flags

Data scientist as analyst. If the work is dashboards and ad-hoc SQL with no experimentation or modelling, the title is inflated and the next role will be harder to get.

No experimentation infrastructure. Ask what fraction of launches are tested. If it is near zero, you will spend your time on retrospective analyses that nobody acts on.

No say in metric definitions. If someone else defines the metrics and you only measure them, you are downstream of the interesting decisions.

Compensation

MarketBandNotes
United States$130k – $230k baseProduct data science below ML engineering; research and decision science vary widely
India₹15L – ₹50L totalLarge market; product companies pay far above services analytics
BothEstablish which of the three flavours the role is before preparing. It is the single most useful question in the recruiter screen

The book for this field

Data Platform & Analytics

What the job actually is now, modelling and grain, data contracts and who gets paged, quality that is not a dashboard, streaming, governance, cost, serving the AI workload, Databricks versus Snowflake, and statistics without the degree.

CHAPTER 10Statistics and experimentation, without the degreeNamed in 44% of data postings in the reader sample. The part of the field that separates a data scientist from an analyst.READ THE CHAPTER →CHAPTER 2Modelling, and why it still decides everythingThe least fashionable skill in the field and the one that determines whether the platform is usable in three years.READ THE CHAPTER →CHAPTER 8Serving the AI workloadAI appears in 92% of these postings. What it actually asks of a data platform is specific.READ THE CHAPTER →

All 10 chapters in Data Platform & Analytics →

Cross-cutting

Skills every archetype tests

These are shared across every field on this site — the same question asked in different vocabulary — so preparation here compounds rather than being spent once. The book above is specific to your field.

The controlOne habit separates strong candidates from plausible ones more reliably than any technical depth.READ THE CHAPTER →Measuring what resists measurementTested in every field on this site under a different name — evals, SLOs, DORA, ablations. Named in 59% of AI and 42% of platform postings, and almost nobody studies it deliberately.READ THE CHAPTER →Discovery before solutionThe decisive round in three separate archetypes, with the lowest pass rate of any stage.READ THE CHAPTER →

Practice questions across all themes →  ·  Back to Data Platform & Analytics →