Shuffle services, fleet-scale upgrades, cost engineering, and performance forensics — from Uber's 2M jobs to Wix's −60% bill.
Spark tuning folklore is endless; this cohort sticks to what companies actually did. Uber upgraded 2M+ jobs fleet-wide with shadow execution, Wix cut costs 60% moving 5,000 workflows to EMR-on-EKS, LinkedIn rebuilt shuffle around sequential reads, PayPal cut a flagship job's cost 70%, and Amazon rebuilt its exabyte compactor on Ray — each taught as a mechanism first, with the company as proof.
Three attacks on the bill — fewer fatter tasks, cheaper cores that survive reclaims, vectorized cores behind a fallback boundary — each with the arithmetic that predicts its payoff.
The job is healthy. It finishes inside its window every night, nothing spills, nothing retries. Then finance tags you: this one pipeline is a third of the platform's compute bill. You open the Spark UI looking for a villain and find none — just an input stage that ran two hundred thousand tasks, each completing in under two seconds, on a cluster of a hundred and forty machines that all went home on time. Nothing is broken. It just *costs*. Where, exactly, is the money going?
The full week 3 brief is part of LeetData Pro.