What Does a Data Scientist Actually Do Day-to-Day?

A data scientist’s job title is one of the most misrepresented in the tech industry. The popular image — making world-changing discoveries from massive datasets — is real for a small fraction of practitioners at a small number of organizations. For most data scientists, most of the time, the work looks quite different. This guide describes what data scientists actually spend their time on, based on O*NET task data and practitioner accounts, without embellishment.

What O*NET Says Data Scientists Do

O*NET (occupation code 15-2051.00, Data Scientists) lists the following among core tasks: analyzing large, complex datasets to identify trends and patterns; developing and implementing data models, algorithms, and predictive models; presenting findings and recommendations to stakeholders; collecting and cleaning data; and collaborating with cross-functional teams on data-driven projects.

The Bureau of Labor Statistics (Occupational Outlook Handbook, 2022–2032) reports that data scientists typically need a bachelor’s degree — with many positions preferring a master’s — in computer science, statistics, applied mathematics, or a related field. The median annual wage was $108,020 as of May 2023 (BLS, United States). Employment is projected to grow 35% from 2022 to 2032.

How the Day Actually Breaks Down

Data Cleaning and Preparation (Often 40–60% of Time)

Real data is messy. Before any analysis can happen, the data needs to be understood, cleaned, reshaped, and validated. This means identifying missing values and deciding how to handle them, checking for duplicate records, normalizing inconsistent formats, merging data from multiple sources that don’t join cleanly, and investigating anomalies that might be errors or might be real signals. This work is not glamorous, but a data scientist who does it carelessly builds models on broken foundations.

Exploratory Data Analysis (EDA)

Before building a model, a data scientist explores the data: looking at distributions, checking correlations, visualizing relationships, and forming hypotheses about what patterns might be predictive. This phase generates questions more than it answers them. A skilled data scientist knows when EDA is done versus when it’s avoidance of the harder modeling work.

Building and Evaluating Models

This is the part most people picture. A data scientist selects an approach — regression, classification, clustering, time series forecasting, or more advanced methods — implements it, evaluates its performance on held-out test data, and iterates. Model building is fast with modern tools; the harder work is knowing which models are appropriate for the problem, understanding what the evaluation metrics actually mean in context, and interpreting why a model behaves the way it does.

Communicating Findings

A data scientist who can build good models but can’t communicate what they mean to non-technical stakeholders is less useful than their resume suggests. This work involves preparing reports, building visualizations, presenting to business leaders, and — critically — framing the uncertainty in model predictions honestly. A model that’s 75% accurate is not the same as a model that’s always 75% accurate; the distribution of errors matters and it’s the data scientist’s job to communicate it.

Meetings and Coordination

Most data scientists in non-research roles attend regular meetings with product managers, business stakeholders, engineering teams, and other analysts. The cadence and volume depend heavily on the organization. At companies that have integrated data science into product development, this can be significant — several hours per day. Understanding what problems the business actually needs solved, and which of those are tractable with the data available, is a substantial part of the actual work.

What Varies by Organization

The role description above assumes a mid-size or large organization with reasonably mature data infrastructure. In practice, the daily experience varies significantly:

  • Early-stage startup: Data scientists often also handle data engineering — building and maintaining the pipelines that collect and store data. The work is broader and scrappier.
  • Large tech company: More specialization. Separate data engineering, ML engineering, and research teams. A data scientist focuses more narrowly on analysis and modeling.
  • Research organization or university: Longer-horizon projects, more academic rigor, often publishing findings. Less emphasis on production deployment and more on novel methods.
  • Non-tech industry (healthcare, finance, retail): Domain knowledge matters more; the statistical methods may be more traditional; regulatory constraints on model use are often significant.

A Realistic Monday Morning

A data scientist at a mid-size e-commerce company might start the week reviewing the performance metrics of a purchase prediction model that was deployed the previous quarter. They notice the model’s accuracy has drifted down over the past 30 days — a common phenomenon called model drift — and add it to the week’s priority list. They spend a morning meeting with the product team discussing an upcoming A/B test of a new recommendation feature and what success looks like from a measurement standpoint. The afternoon involves writing SQL queries to pull the dataset they’ll need to investigate the drift, cleaning the results, and starting the exploratory analysis. They won’t finish this in one day.

Questions to Ask a Data Scientist Mentor

  • What percentage of your time is actually spent on modeling versus data preparation and meetings?
  • What does “good” data infrastructure look like, and how rare is it?
  • What’s the most important thing you do that isn’t in any job description?
  • Which technical skills have depreciated fastest in your time in this field, and which have stayed relevant?

See our full Data Scientists career guide for O*NET task data, salary information, and career path details.

Frequently Asked Questions

Is the data scientist role being replaced by AI tools?

Automated machine learning (AutoML) tools handle some of the routine model-building work more efficiently than manual approaches. But the problem formulation, stakeholder communication, data quality judgment, and results interpretation that make up a large part of the job are not easily automated. The BLS projects 35% employment growth through 2032 despite the existence of these tools.

Do data scientists mostly work remotely?

Remote work availability in data science is broadly similar to other tech roles. Many positions offer hybrid or fully remote options. The trend toward in-office requirements has varied significantly by company and sector since 2022.

Is Python or R more important to learn?

Python is the dominant language in most industry data science roles. R is more common in academic research and certain statistical fields like biostatistics. SQL is essential regardless of whether you work primarily in Python or R.