Runtimeclocksid-generationntpupsertsilent-corruptionmedium

Duplicate primary keys on one node, for about two seconds, once a month

01Symptom

Your order service generates 64-bit IDs in the application, with no database sequence. At 03:14:07 UTC, 231 inserts into orders fail with duplicate-key errors, all within a two-second window. Then everything is fine. The same thing happened five weeks ago, and nobody found a cause. This time a reconciliation job also reports that some rows in order_events don't match their parent orders. Nothing was deployed, and the error window is gone before anyone can attach a debugger.

02Constraints

  • 12 app nodes, each generating IDs locally with no coordination
  • ID layout: 41 bits of millisecond timestamp, 10 bits of node ID, 12 bits of per-millisecond sequence, reset to zero whenever the millisecond changes
  • Each node issues about 600 IDs per second
  • orders uses plain inserts; order_events uses an upsert (ON CONFLICT (id) DO UPDATE) for idempotent writes
  • Nodes are VMs on a cloud provider that does live migration for host maintenance
  • Time sync is a standard NTP client with default settings

03Evidence

  • All 231 duplicate-key errors come from requests handled by node 9, between 03:14:07.0 and 03:14:08.8; the other 11 nodes show nothing unusual
  • At 03:14:05, the cloud provider's event log shows a live migration of node 9's VM
  • Node 9's logs show the system clock stepped backward by about 1.8 seconds shortly after the VM resumed
  • IDs issued by node 9 in the failing window decode to timestamps about 1.8 seconds earlier than real time
  • For 41 events, the stored row now belongs to a different order than the one that originally wrote it, and no error was logged for those writes
  • The ID generator reads the wall clock for every ID and doesn't compare it to the last timestamp it used

→The question

Why did only one node produce duplicates, and why were the upserted rows worse than the failed inserts?

04Your prediction

01What caused the duplicate IDs?
02Why did only node 9 produce duplicates?
03Why were the 41 overwritten events more dangerous than the 231 errors?
0 / 600 chars