A data article a week
Over the past few months I have been sharing links to articles and posts with my colleagues working in data science and data engineering. Some of these links might be of interest to a wider audience, so I share them here as well together with a brief note. The first link was shared on July 6, 2026.
Week 1: Encoding Your Domain Expert: The Context Layer Behind Spotify's Data Assistant. On how Spotify is working with data assistants, SQL queries, and domain experts (one interesting stat is that ~87% of the executed queries are not useful for subsequent work).
Week 2: Metric Semantic Layer: How Lyft Governs and Scales Key Data Definitions. On how to structure and save business logic and metric definitions as code (at Lyft). Not sure why they decided to build a Python package from scratch for this purpose (no explanation is provided in the post), but there are some interesting ideas on metric configurations and ownership.
Week 3: Guide to data tools landscape for developers. A longer post on some of the concepts and tools data teams rely on and how they fit into the broader data landscape (Redshift, dbt, Airflow, etc.). Some of the terms are not relevant for everybody, so do not expect all signal and no noise (but still sensible to be familiar with many of the concepts being described here).
Week 4: Ruff v0.16.0. A new version of Ruff. As I use Ruff as the linter for all my Python projects, it is good to be familiar with the basics of Ruff and in particular what it does. Ruff 0.16.0 enables 413 rules by default (up from 59). As agents now write more Python code, Ruff is a great tool to confirm that code being written with AI complies with specific standards across projects.
Week 5: Fluent, not native: agents translating pandas to Polars. As I have been refactoring a lot of old pandas code to Polars using AI, this is an interesting post on how good LLMs are at translating pandas code to idiomatic Polars code. The takeaway message is that AI is doing a great job at translating pandas to Polars, but it is still important to validate and check for mistakes. Do also check out the associated skill created to improve the translations.
Week 6: Many AI analysts, one dataset: Navigating the agentic data science multiverse. An academic paper by researchers from Amazon. Definitely no need to read it all but good to be familiar with this kind of work. It is behind a paywall, but an ungated version is available on arXiv. The methodological choices we make when analysing data matter for the conclusions we get, and as AI agents can make different decisions, they can also end up with different conclusions. There is nothing unique about this for AI agents that is not the case for humans as well (several studies have demonstrated how different researchers, when analysing the same data, can end up with different conclusions). However, the challenge is that when people can easily and cheaply use AI agents to analyse data, it makes it a lot easier to create many analyses and thereby, deliberately or by mistake, cherry-pick the preferred conclusions. The issue to consider is how we can enable stakeholders to use AI agents to analyse data while also taking into account the risks in how AI agents can potentially lead to people ending up with their preferred data-driven but unreliable conclusions.
Week 7: MCP tool design: Practical approaches and tradeoffs. An article on some of the issues that can arise with setting up and scaling MCPs. Specifically, the post deals with MCP tool design issues and the reasons why they fail (namely bloat and confusion). The interesting aspect is that, as always, we face trade-offs in how we set up these tools. For example, by addressing confusion we might increase the bloat. There are several good considerations on how to improve the quality of MCP tool design (schema constraints, multiple tools, server-side LLM inference). There is also a walkthrough of the different approaches and the associated trade-offs to consider.
Week 8: Python Polars: The Definitive Cheatsheet. A new reference guide for working with data in Polars. In general, cheatsheets and similar documents are less relevant today than a few years ago as AI tools can easily provide the answers and write the necessary code. However, it is still important to be familiar with the most relevant methods, functions, data types, etc. available in Polars to get work done. So this is a good supplement to the available books and video courses.
Week 9: How well does AI peer review work?. Peer review is an integral part of good data work. The author of this post introduces a series of errors in empirical psychology papers and examines the extent to which AI models can find those errors (though keep in mind that they are not the most recent state-of-the-art models). A key point is that ensembling is a good approach as different models can help identify different errors. There are certain errors that are not caught by any of the models, and those are the errors that remove information. This is a good reminder that AI models are more likely to work with material that is present rather than consider all the information that might be absent. So when working with data, AI can help catch errors, but do not expect it to catch all possible errors.
Week 10: Making Your Data Ready for Agentic AI. A longer, detailed post on things to consider when setting up data pipelines to be used and maintained by AI agents (including reflections on how agents can actively curate data). The abstract provides a good summary of what is covered by the post: "For thirty years we built data systems for human analysts, who supply the context, judgment, and skepticism to work around data that's incomplete or wrong. Autonomous agents supply none of that. They act on whatever they're handed, confidently. For data to be AI-ready we need to build a series of layers: a data foundation that makes data trusted, a context layer to apply proper meaning, and an access layer that supports and controls how agents operate on that data. While doing this we need continuous attention to observability that ensures the data is properly governed and we have an auditable trace of its use in decision-making."
Week 11: Pandas Should Go Extinct. Not that we are in dire need of another post demonstrating that Polars is better than pandas, but here we go. A sensible comparison of how pandas, Polars, and DuckDB work with (big) data, but keep in mind that the performance benchmarks are from last year (the post is based on a conference talk from last year), and pandas has incorporated some of the ideas from Polars and DuckDB since then.
Week 12: Lessons learned from scaling data scientists with AI. A post from the beginning of the year, but I believe the conclusion is even more relevant today than back then: "The biggest lesson is that LLMs don’t replace data scientists. Instead, they expose how critical they are and give them more bandwidth to do the highest-leverage work. To get there, though, data teams will need to shift their views on “good documentation” from a cherry-on-top feature to a production requirement with measurable impact. The work looks different from building reports and dashboards, but it’s an evolution rather than a wholesale change."
Week 13: The Shape and Feel of the Post-AI Data Stack. A post on the evolution of the data stack with a focus on how to establish a single source of truth on business-critical metrics and enable stakeholders to work with data and AI. If you only read one section, let it be the "Post-AI Data Stack" section. Primarily relevant for some of the ideas and general principles rather than the technical implementation.
Week 14: What Jev will do to data engineering. There has been a lot of attention to Jev over the past few weeks. In simple terms, it is a cost-efficient way to perform classification with LLMs (so not useful for generating text, but great at categorising based on textual data). It is difficult to say whether people will even talk about Jev in a year from now (or even a few weeks from now), but there are some strong points on data pipelines, semantic operators, and decision-making, including relevant examples on the LLMBranchOperator in Airflow.