Spark · hard · ~12 min
You own logging for a Spark platform that runs 250K+ jobs a day across hundreds of thousands of executors, emitting up to 200TB of unstructured logs daily (~5.38PB per 30 days).
Last quarter's cost fix: retention cut to 3 days and verbosity dropped from INFO to WARN. Since then, three incident retros have stalled on "the logs we needed were already gone." The HDFS line still runs ~$180K/yr at 3-day retention — a month would be ~$1.8M/yr — and the small-write pattern is chewing through SSDs.
Two proposals are on the table:
Finance wants a decision; on-call wants INFO back.