Data jobs and tools (Q2 2026)
Disappointment is when reality doesn't meet expectations
That's been my guiding principle in life for a very long time. I usually follow it up with ".. and it's much easier to change expectations". Especially helpful when doing sports or raising kids.
But let's talk about jobs and roles in the data world. The following is a look at what those expectations are for getting paid to work with data - based on real data.
A snapshot of the European data job market in Q2 2026 (April to mid-June):
The roles are distinct in specific ways, so let's start by looking at one in depth, then how the others compare. The second most prevalent role/family and one I have most experience with is data engineering -
So what makes data engineering distinct in job descriptions?
Data engineering
Here's what data engineering postings name most, each bar the share that mentions it (plus aggregated "A cloud" and "AI tools" bars).
SQL and Python, tied at the top, very prevalent - they've stopped being skills and turned into assumptions, the way reading and writing are for a journalist. About 15% of postings don't bother naming them; those ask for Kafka, Airflow, Snowflake or Spark instead and just assume the basics. (Nobody writes "Linux" into a job ad anymore either. They did, 10-15 years ago.)
Then a cloud - at least one is mentioned in 61% of ads. Breakdown is Azure, AWS or GCP, at 36%, 33% and 26%, close enough that the choice barely matters. Know one well enough to run infrastructure on it and the concepts carry over. There are even these official pages for familiarizing yourself fast - Azure literally has one for AWS. The expectation is that you know a bit about one. It feels like it's also becoming more of an assumption that everyone has some experience with these. Not that on-prem is going away, but exclusively on-prem is becoming more rare.
Below that, the role starts to look like itself: Spark and Airflow turn up again and again, dbt in about a fifth of postings, and a warehouse somewhere in the mix - Databricks (27%), Snowflake (23%), Fabric (20%), BigQuery (13%).
That's the floor. And here's the thing about a floor: almost everyone in data is standing on the same one.
What Comes After SQL and Python
SQL and Python don't define data engineering. Everyone in data has them - the analyst, the scientist, and the ML engineer all start on the same SQL-and-Python base. What defines the role is the next tool. And that same choice is what routes everyone else to their own corner of the field.
The diagram below is a Sankey - a flow diagram where the width of every ribbon is proportional to the quantity flowing through it, so a thicker ribbon just means more postings. It follows that fork. Everyone starts on the SQL+Python trunk on the left - 82% of data-role postings sit on it. From there, each tool a posting asks for pulls toward a role; follow a tool rightward to see where it leads, and how often. A single posting that mentions 5 tools contributes 5 units to the chart, so the total is definitely greater than the number of job ads.
Data engineering is the pipeline path. Spark, Airflow, Kafka, Snowflake, and Databricks - the tools for moving, storing, and orchestrating data at scale - all run toward the DE node. That cluster is the job: pick up any of them and you're walking toward data engineering, not away from it.
The other tools are doors to other rooms. Power BI and Tableau lead to the analyst - a data engineer who picks up Power BI is drifting toward analytics, not deeper into engineering. scikit-learn and pandas lead to the data scientist; PyTorch, TensorFlow, and LangChain to the ML/AI engineer; PostgreSQL and Oracle to the database specialist. A few tools are shared corridors rather than doors - Docker and Kubernetes split almost evenly between engineering and ML, the one place those two genuinely overlap.
And a handful fork several ways at once. dbt is the clearest: it lands mostly in data engineering, but a thick strand peels off to analytics and another to analytics engineering - the same tool, three jobs. The diagram is busy on purpose. There's no single path through it, and that's the point: trace the tools you already know, and the door to the next room is wherever they bunch up.
Now a bit of data-nerdery, because the diagram hides something. Remember, a Sankey's grammar is width = quantity, so everything above is absolute volume - raw counts of mentioned tools. On volume, DE wins almost every tool, and not from any bias: DE ads simply have more mentioned tools - 10,917 across 4,838 postings, nearly double the 5,853 across the larger pool of 5,511 analyst postings. So DE swamps the chart.
But "how many" is only one way to count a tool. Take dbt and read it three ways - the answer changes each time.
By volume, dbt is a data-engineering tool. Of every posting that asks for dbt, most are DE roles: 941 of them, against 232 for analytics engineering. Learn dbt, and the jobs using it are mostly DE.
By rate, dbt is an analytics-engineering tool. 59% of analytics-engineering ads ask for dbt; only 19% of DE ads do. Per ad, it's three times more an AE expectation than a DE one.
By over-representation, dbt is an analytics-engineering signal. Analytics engineering is just 3% of the field, but 14% of all dbt ads - a 5x over-representation. If an ad names dbt, that's real evidence it's an AE role.
Same tool, same data, three honest answers. Here's the volume-versus-rate flip in two pictures - the same dbt numbers, ranked each way:
And the 2x2 of ads vs dbt mentions for DE and AE visualized:
There Are Jobs Out There
It's a wide field, and a busy one. Over the quarter these roles came from
The chart tracks new postings each week, by role. One caveat: these dates are when I first saw a posting, not when it went live, so read each line as a discovery rate rather than a precise count.
Analytics leads the flow - around 500 new postings a week, steadily. Data engineering runs close behind, then ML/AI engineering and data science. The specialist roles - database, governance, architecture, analytics engineering - are smaller, but none of them ever go quiet. Whatever corner of data you're aiming for, fresh postings show up every single week.
What This Means in Practice
The core expectations have been stable for years, under all the churn of new tools and shifting terminology: SQL, Python, a cloud platform, Power BI, Tableau, dbt, etc. Yes, dbt is newer, but still.
My main takeaway is to do whatever you find most interesting - there are larger and smaller niches for every combination. Anecdotally, I'd say data management/governance might become more popular as that context is what leads to AI being able to use company-internal data. But it's the one corner of the data field that's not really tool-dependent.
Methodology
This analysis covers
Geographic scope: European job markets - postings located in a European country, plus fully-remote roles (identified by title or a remote-specialist board) open to applicants anywhere.
Role definition: A posting is classified as "data engineering" if its title contains any of "data engineer," "data platform engineer," "ETL developer/engineer," "data pipeline engineer," "dataops engineer," "data warehouse engineer," "data infrastructure engineer," or "big data engineer/developer" (case-insensitive). This captures the infrastructure family across naming conventions - "ETL Developer" is largely how Central/Eastern-European and enterprise markets name the same role. Similar logic is applied for other role families.
Time window: All analyses use postings from Q2 2026, dated April 1 to June 15 (the quarter to date).