Book
Databricks & Snowflake
From first principles to production: the problem both solve, columnar storage, how Delta and Iceberg turn files into tables, how each engine executes a query, each platform in practice, performance, cost control, governance, and choosing between them.
Built because 66% of enterprise data postings name one of these by name and 30% name both. Assumes no prior knowledge of either.
CHAPTER 1
The problem both are solving
Warehouse, lake, lakehouse — the three architectures, why each appeared, and what was actually wrong with the previous one.
CHAPTER 2
Files, columns and compression
Why analytics runs on Parquet. The single idea that makes everything above it possible.
CHAPTER 3
Delta and Iceberg — how files become a table
The mechanism that makes a lakehouse possible. Understand this and both products stop being magic.
CHAPTER 4
How a query actually executes
Spark's model and Snowflake's, side by side. The mental model that makes performance work intelligible.
CHAPTER 5
Databricks in practice
Workspaces, clusters, jobs, Delta Live Tables and Unity Catalog — what each is and when it bites.
CHAPTER 6
Snowflake in practice
Virtual warehouses, micro-partitions, clustering and caching — and the three ways people waste money.
CHAPTER 7
Making queries fast on each
The specific moves that matter, and the diagnostic that tells you which one you need.
CHAPTER 8
Cost control, which is the skill people get hired for
Half of data-platform postings name cost. This chapter is the practical part.
CHAPTER 9
Governance and security on both
Named in 68% of postings. What each platform gives you and where the gaps are.
CHAPTER 10
Choosing, migrating, and the interview
How to answer the question you will be asked, and what to actually learn.