Shuffle services, fleet-scale upgrades, cost engineering, and performance forensics — from Uber's 2M jobs to Wix's −60% bill.
Spark tuning folklore is endless; this cohort sticks to what companies actually did. Uber upgraded 2M+ jobs fleet-wide with shadow execution, Wix cut costs 60% moving 5,000 workflows to EMR-on-EKS, LinkedIn rebuilt shuffle around sequential reads, PayPal cut a flagship job's cost 70%, and Amazon rebuilt its exabyte compactor on Ray — each taught as a mechanism first, with the company as proof.
Symptom → subsystem: the UDF codegen cliff, join algorithms diagnosed from where they fail, memory forensics at both ends, and why shuffle physics eats 10–20% of clusters.
The job ran in forty minutes yesterday and four hours today. Same code, same cluster, roughly the same data volume. The Spark UI shows one stage holding everything: 197 of its 200 tasks finished in seconds, and three have been running for hours, spilling to disk the whole time. A colleague suggests doubling executor memory; another suggests doubling the cluster. Both suggestions cost real money, and neither names a mechanism. Before you spend anything: what, specifically, made three tasks different from the other 197?
The full week 2 brief is part of LeetData Pro.