Pipeline · medium · ~12 min
Between 00:05 and 00:47 last night, nothing was scheduled. Not failed — *not started*. Forty-two minutes of silence across 1,900 DAGs, discovered at 07:30 when an analyst asked why every morning table was empty.
The infuriating part: monitoring shows a green night. The scheduler's health check passed every 30 seconds throughout — the process was up, the port answered, the standby never took over. The scheduler was alive by every probe you had, and dead by the only measure that matters.
Leadership wants two things by Friday: an explanation, and a guarantee this class of failure pages within five minutes.