AI Made Data Scientists Faster. Now It’s Expanding the Job.

AI Made Data Scientists Faster. Now It’s Expanding the Job.

···ContextLast year, I wrote an article How AI Is Rewriting the Day-to-Day of Data Scientists. In that article, I summarized the AI impact on data scientists as eliminating low-value tasks and accelerating high-value work.A year has passed. AI has advanced from GPT 4.5/Claude 4 to GPT 6/Claude 5.5, and I’ve changed jobs (twice). It is time to revisit what has changed — not just how data scientists use AI today, but how AI has transformed the job itself.···How Data Scientists Use AI TodayI. AI eliminated much of the mechanical work — but not the need for data scientistsLast year, I predicted AI would help data scientists move from “working as data interpreters and gatekeepers” to “empowering stakeholders to self-serve so we can minimize low-value work”. Over the past year, I have seen this happen on multiple fronts.1. I have barely written any SQL or Python manually in the past six months.To be honest, I was a bit skeptical of “Text to SQL” last year. But now, the only time I write SQL from scratch is for a quick SELECT * LIMIT 100 table inspection or a simple aggregation. AI writes most analysis code. I cannot remember the last time I opened Stack Overflow or Kaggle to copy-paste code and debug. My role has changed from the coder to the reviewer. 2. Repeated workflows are increasingly packaged into Agent Skills.In my past article, I walked through this process by showing how I turned my weekly data visualization process into a skill. Here are some more examples: You can build a skill to conduct metric movement root cause analysis, based on the business context, past work, and common cuts. Or teach AI how to write an experiment readout from reading the design doc, connecting to the experimentation tool MCP, interpreting the numbers, running analysis, and compiling a stakeholder-facing report. These skills significantly reduce the time spent digging up old processes or relearning how something is done, while encouraging more knowledge-sharing within the data org.3. As a result, stakeholders pull more data and run more analysis themselves.I have seen some stakeholders (especially the technical ones) quickly embracing this new era and starting to run analysis themselves with Claude Code and similar AI tools. This is common when they need some numbers to plug into a doc or a quick exploratory analysis, and when data scientists have bandwidth constraints. 4. Semantic layers become more important for AI reliability.But then a question comes up constantly: “Is this number correct”? Data scientists generally have better judgment when reviewing generated code, so they can still correct the analysis — if they are patient and careful enough, which is not always possible. But many stakeholders are not trained to review analytical code or sense-check the output in the same way.As a result, I’ve seen two opposite reactions, both pretty common — some stakeholders lose trust after seeing AI make mistakes and still go to data scientists for everything, while others trust everything vibe-coded and end up quoting the wrong numbers in meetings and docs.To solve this, better context is the key. Setting up a data semantic layer is always a good starting point, as it documents tables and defines business metrics and dimensions. With it, agents no longer need to inspect the data and piece together the data definition from different tools, but know exactly how to pull “Retention Rate”. This not only improves output accuracy, but saves tokens.The bottleneck moved from “Can someone get the data?” to “Can we trust what they got?”···II. AI expanded the definition of data scienceLast year, when I used AI, I usually had an analysis plan and scripts in mind and asked it to fix a syntax error or write a specific piece of code — I used it more as a coding assistant. However, things have changed at an extremely fast pace, with far more depth and breadth than I expected. 1. AI dramatically reduced the analysis turnaround time. Let’s take geo-testing projects as an example. One year ago, I was using a Databricks notebook template to run market selection. I needed to plug in the metrics calculation, replace the eligible markets list, and set various thresholds and rules manually. Yet it threw errors from time to time. Then interpreting the test results with DiD analysis and regression adjustment was another heavy lift. Now, I simply hand the template code to AI and describe my constraints in plain English. It runs the simulations and finds the best sets of treatment and control markets accordingly. AI is also extremely helpful for brainstorming different causal inference approaches to make the measurement more robust.2. AI now sits across the end-to-end analytics workflow. As I discussed in my article earlier this year, if you are still only using AI for code generation, you are missing lots of productivity gains. Before starting an analysis, I use various tools’ MCPs to collect and summarize discussion and past research. Then I use the plan mode to debate approaches and settle on an analysis plan. AI then executes the plan, surfaces findings along the way, and eventually generates a stakeholder-facing write-up that is ready to share.3. AI makes data scientists more capable of doing engineering work. Six years ago, as an early Data Scientist at a startup, I learned dbt for the first time to build new tables — DS at startups usually means Data Engineer + Data Analyst + Data Scientist. I had a checklist for every step: which files to edit, which fields belonged in the YAML file, how to start Docker, which terminal commands to run for dev testing, how to trigger the DAG in Airflow, and how to monitor it in Datadog. Not to mention all the time that went into setting up the developer environment correctly. Now, as the first Data hire at another startup, these data engineering tasks are easier than ever. I really just need to figure out the right data model and table logic, then Claude Code can take it from there and deploy it into production with proper unit tests. The same applies to machine learning model development and other engineering work. With proper agentic development tools and code reviews, data scientists can implement simple features or model changes, and A/B test them quickly.AI doesn’t just make the same DS job faster. It changes what one data scientist can reasonably own.···III. AI created a new layer of complexityAll of this seems to be moving in the right direction. There is less repetitive work, tools are easier to use, stakeholders are more hands-on, and data scientists can take on a much broader scope. But it comes with new challenges and risks.1. Governance becomes more importantSemantic layers require governance. Setting up a semantic layer is only the first step — data definitions, table structures, and business context all change over time, so it needs active maintenance. This can be done by an agent that runs periodically to verify the documented table structure and collect new data knowledge from GitHub, Google Docs, Slack, and other internal tools. Human review and approval are still required.In practice, this becomes an intensive, continuous data governance effort. I was highly involved in building the semantic layer at my last company. And as I just joined a startup as the first data hire, I have reflected a lot on the best practices to build it from 0 to 1. This topic itself is worth its own article to discuss in detail, so I will save it for later :)Agentic workflows also require governance. Agent skills are handy, and sometimes way too easy to build. I’ve seen a team GitHub repo grow to 100+ skills within three months. The same applies to the shared context files. Managing the agentic workflow effectively becomes a challenge.Because these files are meant to be reused, they require higher quality and reproducibility — I recommend people read and test them extensively before sharing. Similar to the semantic layer management process, you can also use another agent to flag overlapping skills and clean up context.2. Rethinking data productsBI tools are undergoing a dramatic revamp. Historically, people used traditional BI for source-of-truth metrics reporting and drag-and-drop style self-serve exploration. However, both needs are being challenged today — AI has given people more flexible ways to build and share dashboards, either using AI-native tools like Claude Artifacts or Databricks Apps, or hosting HTML-based dashboards internally; And if you can play with data by simply chatting with an agent, why use a restrictive drag-and-drop explorer? As a result, the value of traditional BI tools is increasingly being questioned. In response, BI tools are either adding AI agents to help users build and interpret dashboards more easily (which I think still does not solve the “existential crisis” of BI tools), or stepping into the AI agent game to enable more flexible and reliable data exploration.Meanwhile, managing source-of-truth dashboards becomes more difficult. Now that everyone can build dashboards easily with AI, even with the semantic layer in place, you can easily run into multiple versions of similar dashboards. What makes it worse is that 1. AI-built dashboards published by Claude Artifacts or simply in HTML sometimes do not embed the underlying code, making it hard to validate and maintain. 2. They might have different numbers, but none of them may actually be wrong — they are just filtered to different users, products, or a specific time window.This is not a new problem, but AI just made it worse. In the past, data teams would publish verified SoT dashboards for each domain, and leave others for teams to use at their own risk. This approach still works. Meanwhile, I think for any shared dashboard, owners should either commit the code to GitHub or at least embed it in the dashboard, so humans or agents can go review the code, add detailed descriptions, and flag potential issues — Using agents to clean up problems created by other agents is becoming a surprisingly common pattern.3. Larger scope = more context switching and higher workloadThese days, I constantly have at least three agents running in parallel for different projects — one writing a revenue data pipeline, one building a retention dashboard, and another running a product usage analysis. Imagine that you used to be an IC with blocks of heads-down time on a single project. Now you are more like a manager “micro-managing” three ICs at once.This means continuous context switching. One minute I might be checking SQL logic; the next, I am debating an analysis insight. It is not that I don’t want to sequence these asks, but I don’t feel comfortable sitting around waiting for an agent to respond, so I end up multitasking. At the same time, people naturally expect us to be more productive with AI (which is very fair) and set tighter timeline expectations. However, this requires a lot of brainpower and creates a different kind of mental load. It sounds contradictory, but I actually think burnout may be easier in the AI age. That makes prioritization and workload management even more important.···What This Means for Data Science CareersLast year I thought AI would remove low-value work so data scientists could spend more time on high-value analysis. That happened — but only partially. What I underestimated was that AI wouldn’t simply compress the existing job. It expanded it.Now, what does this mean for the data science career?1. Execution becomes less differentiatingKnowing SQL/Python is still necessary to understand how things work, but simply being able to write the query or code is less differentiating when AI can do it too. In my own job searches and conversations, I’ve seen more companies replace traditional SQL coding interviews with AI-assisted coding or code-review rounds —something I predicted last year in this article.2. Judgment becomes more differentiatingInstead, judgment is the key skill for a data scientist to stand out. That includes technical judgment in areas like data modeling and causal inference, but also understanding business context, identifying the right question, and knowing how to sense-check AI outputs.Many people now have the impression of “I can build everything with AI”. This is not wrong, but I would like to expand it a bit — “I can build everything in a hacky way”. Taking my data engineer work as an example, of course, I can now build whatever table I need to answer an ad hoc question super fast. But doing that repeatedly just creates tech debt — it is not sustainable and will eventually lead to duplicate and conflicting data. Instead, I still need to design the right data model and decide how it fits into the broader semantic layer. This is what technical judgment means — knowing the right steps to take and the right questions to ask.3. The path from junior to senior becomes less obviousA lower execution barrier and greater emphasis on technical judgment may also help explain why companies seem to be hiring fewer juniors, but still looking for experienced data science candidates. Though sometimes this makes me wonder — if we don’t train juniors today, how can we have seniors tomorrow? And for juniors who just started their career, how will they develop their technical judgment with AI doing so much of the execution for them? Honestly, I don’t have the answer :(4. Role boundaries blurI’ve seen data teams evolve from combined DE+DA+DS roles into specialized teams as companies scale. Now I increasingly see movement in the opposite direction. DS, engineering, and PM scopes are also starting to overlap more — they are not replacing one another (yet), but being able to cross those boundaries and help the business move faster is increasingly valuable.AI is not shrinking the DS job. It is stretching it.

Original Source

Read the full article at Towardsdatascience →

KhanList aggregates and links to publicly available news content. We do not host full articles from third-party sources. Always verify important information with original sources.