Lakehouse · medium · ~14 min
The events_iceberg table has two writers: a streaming job upserting micro-batches every 60 seconds, and a nightly compaction job that rewrites small files.
events_iceberg
For the past week, compaction has failed almost every night — burning 25 minutes of compute per attempt, retrying four times, then giving up. The table keeps degrading. The on-call suggestion: "just add a lock so only one job can write."
Explain what's actually happening in the commit protocol, and design a fix that keeps BOTH jobs running.