Data Platform & Analytics
Head of Data Engineering
In a reader-supplied sample, 22 of 38 postings were Lead, Principal, VP or Head-of. This is where data leadership actually sits.
The senior end of data is unusually well represented in postings, and the remit is broader than engineering leadership elsewhere: it typically includes owning an expensive platform portfolio, carrying regulatory accountability, and now defining how AI enters the data estate.
One posting in the sample — Head of Enterprise Data Engineering at a bank — states the shape plainly: enterprise data strategy and modernisation roadmap, architectural authority, executive accountability for the Databricks and Snowflake portfolio, defining how AI and generative AI are incorporated, and regulatory and risk leadership. That combination is the job.
| Time | What you are actually doing |
|---|---|
| 08:30 | A modernisation programme that is 70% done. The remaining 30% has no funding and both platforms are now permanent unless you change that. |
| 10:00 | Cost review. The warehouse and lakehouse bills together are the second-largest line in the technology budget and you own the number. |
| 13:00 | A regulator question about lineage for a reported figure. The answer needs to be reconstructable, not asserted. |
| 14:30 | AI governance. Where models may read from, what may leave the estate, and who approves it. |
| 16:00 | Hiring. Two senior people who will define the next two years, in a market where good data engineers have options. |
The decisive round
“We run both Databricks and Snowflake. What is your plan?”
Interviewer: “We have both platforms, significant spend on each, and teams attached to both. What would you do?”
A weak answer. “I'd assess both against our requirements and consolidate onto whichever fits better, to reduce cost and complexity.” The answer everyone gives, and it describes a two-year programme that will stall at 70% and leave both running anyway — which is the situation they are already in.
A strong answer. “My default assumption is that consolidation is the wrong first move, because migrations of this size stall and you end up carrying both plus a half-finished migration, which is worse than either.
So the first question is why we have both. If it is an acquisition, or two org units that each chose correctly for their workload, then coexistence is a legitimate architecture and the work is to make it cheap rather than to end it. If it is drift with no owner, that is different.
What I would do regardless, and quickly: one storage layer in an open table format and one place where conformed entities are defined, with both engines reading it rather than each holding a copy. That removes the failure that actually costs us — two customer tables with two definitions of active, so every number depends on which engine you asked. Iceberg and Delta interoperability make that genuinely possible now.
Then attribution before optimisation, because I would expect that today nobody can say what each platform costs per team, which means nobody owns it.
Then, only if consolidation still looks right, I would fund the finish before starting — including the long tail, and with a decommission date I intend to hold. And I would be honest with the board that the saving is real but slower than the business case will claim, because these business cases always assume the last 30% happens.”
Refusing the obvious consolidation answer, naming the duplicate-definition failure as the thing that actually costs money, and being honest about migration economics.
What to have ready
- A platform cost number you owned and moved, with the attribution work behind it.
- A modernisation programme you finished — or an honest account of one that stalled and what you learned about funding the finish.
- A regulatory conversation you handled: lineage for a reported figure, a residency question, an audit finding.
- An AI position — where models may read from in the data estate, what may leave it, and who approves. This appears in 55% of the enterprise postings in the sample.
- A hiring story at scale, because senior data engineers have options and this is a constrained market.
Red flags
Accountability for the portfolio without authority over spend. Ask who signs the platform contracts. If it is not you and you own the number, you own a report rather than a lever.
A modernisation programme already declared successful. Ask what percentage of workloads have actually moved, and what the decommission date is. If neither has an answer, you are inheriting a permanent dual estate described as a migration.
Regulatory accountability with no lineage capability. You will be asked to attest to things you cannot evidence.
Compensation
| Market | Band | Notes |
|---|---|---|
| United States | $220k – $360k base | VP and Head-of at banks and large enterprises run above |
| India | ₹80L – ₹1.8Cr total | Common in BFSI GCCs, where the data estate is often global from India |
The book for this field
Databricks & Snowflake
From first principles to production: the problem both solve, columnar storage, how Delta and Iceberg turn files into tables, how each engine executes a query, each platform in practice, performance, cost control, governance, and choosing between them.
Cross-cutting
Skills every archetype tests
These are shared across every field on this site — the same question asked in different vocabulary — so preparation here compounds rather than being spent once. The book above is specific to your field.
Practice questions across all themes → · Back to Data Platform & Analytics →