Shuffle services, fleet-scale upgrades, cost engineering, and performance forensics — from Uber's 2M jobs to Wix's −60% bill.
Spark tuning folklore is endless; this cohort sticks to what companies actually did. Uber upgraded 2M+ jobs fleet-wide with shadow execution, Wix cut costs 60% moving 5,000 workflows to EMR-on-EKS, LinkedIn rebuilt shuffle around sequential reads, PayPal cut a flagship job's cost 70%, and Amazon rebuilt its exabyte compactor on Ray — each taught as a mechanism first, with the company as proof.
Working-set arithmetic decides single-node vs cluster; copy-by-reference compaction and driver-serial fan-out mark where specialists beat generality; table formats hand you a maintenance loop.
Your nightly job aggregates eight gigabytes of events. It runs on a six-worker cluster that takes four minutes to start, executes for eleven, and shuts down. The pipeline has run this way for two years, nobody complains, and the code is clean. Then a new teammate asks the question you realize you've never actually answered: *why is there a cluster here at all?* You start to say "because it's Spark," and stop. What would it take to answer with arithmetic instead of habit?
The full week 4 brief is part of LeetData Pro.