Spark · medium · ~10 min
After migrating to beefy 96GB nodes, someone configured one giant executor per node ('use all the memory!'). Jobs now run 2.4x slower than on the old modest nodes, with executors spending nearly half their time in garbage collection and occasional multi-second pauses that trigger shuffle fetch timeouts.
Bigger machines, worse performance. Fix the executor topology.