Iceberg, Hudi, Delta, and Paimon at company scale — compaction ops, format wars, catalogs, and the S3-native future.
The freshest slice of the corpus: Adobe's Iceberg arc, Ancestry's 100-billion-row table, Kakao's operational takeaways, Whoop's schema-migration tooling — plus the deepest expert vein in the archive (Vanlightly's consistency-model series). This is also the natural partner-track cohort.
Write-Audit-Publish on branches, streaming ingestion's freshness-files-semantics triangle, and object storage growing database powers.
At 3 a.m. an upstream service ships a unit bug — amounts arrive multiplied by 100. Your pipeline runs green: schema matches, volumes look normal, the load succeeds. By 9 a.m. three dashboards have consumed the numbers and a VP has already forwarded one. You spend the day rebuilding tables and credibility. Here's the uncomfortable question: your *application* code couldn't have shipped that bug — CI would have caught it in staging before production ever saw it. Why does your data platform deploy straight to prod?
The final week ties the format machinery you now own to its two frontiers: quality gates built out of snapshots, and the ingestion and file-format edges where the lakehouse is still being invented.
The full week 4 brief is part of LeetData Pro.