Runtime · 2026-10-08 · clocks / id-generation / ntp / upsert / silent-corruption · medium

Duplicate primary keys on one node, for about two seconds, once a month

01Symptom

Your order service generates 64-bit IDs in the application, with no database sequence. At 03:14:07 UTC, 231 inserts into orders fail with duplicate-key errors, all within a two-second window. Then everything is fine. The same thing happened five weeks ago, and nobody found a cause. This time a reconciliation job also reports that some rows in order_events don't match their parent orders. Nothing was deployed, and the error window is gone before anyone can attach a debugger.

02Constraints

  • 12 app nodes, each generating IDs locally with no coordination
  • ID layout: 41 bits of millisecond timestamp, 10 bits of node ID, 12 bits of per-millisecond sequence, reset to zero whenever the millisecond changes
  • Each node issues about 600 IDs per second
  • orders uses plain inserts; order_events uses an upsert (ON CONFLICT (id) DO UPDATE) for idempotent writes
  • Nodes are VMs on a cloud provider that does live migration for host maintenance
  • Time sync is a standard NTP client with default settings

03Evidence

  • All 231 duplicate-key errors come from requests handled by node 9, between 03:14:07.0 and 03:14:08.8; the other 11 nodes show nothing unusual
  • At 03:14:05, the cloud provider's event log shows a live migration of node 9's VM
  • Node 9's logs show the system clock stepped backward by about 1.8 seconds shortly after the VM resumed
  • IDs issued by node 9 in the failing window decode to timestamps about 1.8 seconds earlier than real time
  • For 41 events, the stored row now belongs to a different order than the one that originally wrote it, and no error was logged for those writes
  • The ID generator reads the wall clock for every ID and doesn't compare it to the last timestamp it used
→

Recent

Previously diagnosed.