The nightly job had run for three years. One Sunday in March it didn't, and nothing failed.
01Symptom
On Monday, March 9, at 09:40, a partner's finance team emails you. Their reconciliation shows settlement files for March 6 and March 8, and nothing for March 7. Your dashboards are green. The scheduler has had 100% uptime, job history shows no failed runs, and nobody was paged over the weekend. You look at Sunday's run history and find the settlement job isn't listed as failed. It isn't listed at all.
02Constraints
- Scheduler servers run in UTC; schedules are written in America/New_York local time
- The schedules were moved from UTC to local time in January, so reports "line up with the business day"
- Jobs, in local time: report-rollup at 01:30, stale-session-cleanup at 02:00, settlement-export at 02:30, partner-sync at 02:45
- Each job takes no arguments, and settlement-export exports "yesterday" relative to when it runs; Sunday's 02:30 run is supposed to export Saturday
- Monitoring alerts on a run that fails or exceeds 30 minutes; failed runs retry 3 times
- Nothing records which runs were expected
03Evidence
- Sunday's run history shows report-rollup at 01:30, then nothing until the 03:00 hourly health ping
- All three jobs scheduled between 02:00 and 02:59 local are missing from Sunday; jobs at 01:30 and at 03:00 or later all ran
- Every other night, all three jobs ran on time, including Monday morning
- Scheduler uptime is 41 days, with no restarts and no errors in its log that night; server clocks are in sync
- The scheduler's "next runs" list the evening before shows 01:30 and then 03:00, with no 02:xx entries
- The same jobs ran on the equivalent Sunday last year, when they were scheduled in UTC
- The partner has no Saturday file, but it does have Sunday's, which Monday's run produced
→The question
Why did exactly those three jobs vanish without a trace, why did nothing alert, and what would you change so that neither this Sunday nor the November one hurts you?
04Your prediction
0 / 600 chars