Spark · medium · ~10 min
The hourly enrichment job joins a 4TB fact table to three dimension tables. The plan shows sort-merge joins everywhere, and 80% of runtime is shuffling the fact table three times. Two dims are tiny; the third is 90MB raw — but the job only uses four of its columns.
Re-plan the joins. The threshold setting is part of the story.