Data Scientist

Where the title means research, where it means analytics, and how to tell before you apply.

The title moved while nobody was looking

Ten years ago Data Scientist described someone who did everything: pulled the data, cleaned it, modelled it, and presented the result. The title carried a premium because the combination was rare.

It has since split. What was one role is now analytics, analytics engineering, machine learning engineering, and research, and companies allocate the old title across all four inconsistently. Some now call the analytics job Data Scientist and the modelling job Machine Learning Engineer. Others do the reverse.

The result is that two people with the same title and the same years can have almost no overlapping skills, and a candidate can be simultaneously overqualified and underqualified for the same posting. Reading which version is being advertised is the first task.

Product, algorithms, or research

Three broad flavours, and postings usually reveal which.

Product data science. Experimentation, causal questions, metric design, and helping product teams decide. Heavy on statistics and communication, light on production code. The tell is language about A/B testing, experimentation platforms, and working with product managers.

Machine learning in production. Building models that run as part of the product, and often owning them once they do. Closer to software engineering than most candidates expect. The tell is deployment, latency, monitoring, and named serving infrastructure.

Research. Novel methods, often with a publication expectation. Rare, concentrated in a small number of employers, and usually requiring a doctorate.

Applying to all three with one resume is why strong candidates get no response. The product data scientist who leads with model architectures reads as someone who will be unhappy running experiments. The machine learning engineer who leads with stakeholder communication reads as someone who cannot ship.

Notebooks are not evidence any more

A portfolio of notebooks applying standard models to public datasets was a differentiator once. It is now the default submission, and reviewers stopped reading them some time ago.

What still carries signal is work where the messy parts were not removed. You defined the question rather than receiving it. You gathered or joined data that was not designed to be joined. You made a decision under uncertainty and said what would change your mind.

Better still is anything that ran and was used by someone. A small model serving predictions through an endpoint, with monitoring, teaches you more and demonstrates more than a superior model in a notebook, because the difficulty in this field has moved from the modelling to everything around it.

If you have industry data you cannot share, describe the shape of the problem and the decision it informed without the data itself. Reviewers understand confidentiality.

The academic to industry translation

A large share of people in this field arrive from research, and the same three problems recur.

The resume describes methods rather than outcomes. Publications and techniques take the space, and the reader cannot tell what changed because of the work. Industry hiring managers are asking what decision your analysis affected.

The work is presented at a level of rigour the job does not require. Industry decisions are made on partial data under time pressure, and interviewers screen hard for whether a researcher can say "this is good enough to act on" rather than pursuing certainty that does not exist.

Engineering practice is absent. Version control, testing, writing code someone else can run. This is the most common gap and the easiest to close, and it is worth closing before applying rather than after being rejected.

The advantage is real though. Researchers are usually far stronger than average on experimental design and on knowing when a result is not what it appears to be, and those are the skills that prevent expensive mistakes.

What to ask about the data before you accept

The single largest determinant of whether a data science job is satisfying is the state of the data, and you can find out most of it in the interview.

Ask whether there is a data warehouse and who maintains it. If the answer is that the data scientists maintain it, a meaningful share of your week is data engineering, whatever the title says. That may be fine, but it should not be a surprise in month two.

Ask how experiments are run today. A company with an experimentation platform is one where your results will get used. A company where experiments are run by hand in spreadsheets is one where you will spend a year building the capability before anyone acts on a finding.

Ask what happened to the last piece of analysis the team produced. This is the most revealing question available. If they can name it and say what changed, the function has influence. If the answer is vague, you would be joining a team that produces reports nobody reads, which is the most common way data scientists become unhappy.

Ask who the work goes to. Reporting into a product or business unit usually means your findings meet decisions. Reporting into a central team that serves requests usually means less influence and more ticket queue.

The case study round

Most processes include a business case: here is a situation, how would you approach it. Reduce churn, decide whether a feature worked, forecast demand.

The failure is reaching for a model immediately. A candidate who answers "I would build a churn prediction model" has skipped the questions that determine whether the model is useful, and interviewers are listening for exactly that jump.

The structure that works: clarify what decision this is meant to support, ask what data exists, say what you would look at descriptively first, propose the simplest thing that could work, and say how you would know it was helping. Then, if a model is warranted, explain what you would do with the predictions, because a churn model nobody acts on has no value.

Being willing to say a problem does not need machine learning is a strong signal rather than a weak one. Plenty of senior people in this field spend their time talking teams out of models.

Communication is the constraint, not the mathematics

Ask hiring managers what separates the data scientists who progress and the answer is rarely technical. It is whether they can make a room of non-specialists act on a result.

That means leading with the conclusion rather than the method. Being clear about uncertainty without hiding behind it. Knowing when a chart is the answer and when a sentence is. Saying "we do not know yet, and here is what it would take to find out" without embarrassment.

On a resume this shows up as results framed by their consequence. "Found that the retention drop was concentrated in one acquisition channel, which changed how the marketing budget was allocated the following quarter" beats any description of the technique used to find it.

In an interview it shows up in how you answer the case study. Candidates who explain their method and stop are common. Candidates who say what they would do about it are not.

Why two data scientists earn differently

We do not print salary figures, because a number from San Francisco means nothing in Toronto, Bangalore or Nairobi.

The pattern worth understanding is that this title has one of the widest bands anywhere, and the spread is not mostly about years or credentials. It tracks how close the work sits to something the company sells.

A data scientist whose models run inside a product that generates revenue is compensated differently from one producing internal analysis, at the same title, the same qualifications and the same city. The same applies to whether you own the system in production or hand it to someone else.

The moves that reprice you: analysis into production machine learning, internal reporting into product decisions, and moving from a local company to a multinational. A doctorate matters in research roles and matters much less elsewhere, where three years of shipped work usually outweighs it.

For what the role pays where you live, our salary calculator takes your city and your years of experience.

Decide which one you are

Before the next application, write down which of the three flavours you want and which you can evidence. If those are different, that gap is your actual project, and it is more useful to work on than another model on a public dataset.

Our job search builder searches every board you trust at once, and the ATS scanner shows what a filter reads first.

Ready to apply? Tailor your resume to the role in a few minutes.

Open the resume builder
Keep exploring

Related career guides

Roles close to Data Scientist, and the same treatment for each: what the job involves, what employers screen for, and how to write for it.