Spark · easy · ~10 min
A well-meaning refactor added .cache() after nearly every transformation in the nightly ETL 'for performance'. Since then the job is slower, executors GC constantly, and two runs died with executor OOMs.
.cache()
Caching made things worse. Explain why, and set the caching policy.