Pipeline · medium · ~12 min
Kubernetes evicted three worker pods at 02:14 (node pressure). Since then, five tasks show running for 6+ hours — their processes died with the pods. Two hold slots in the critical ETL pool. Worse: when the scheduler finally reaps one and retries it, the load_ledger task DOUBLE-WRITES, because its first incarnation had already inserted rows before dying.
running
load_ledger
Untangle the zombie mechanics and both layers of fix.