Eng design review · Sep 9, 2026

Job sync —
retry queue

docs/rfc-retry-queue.md · coldharbour/sync-gateway
coldharbour rfc-retry-queue · rev 031 / 6
fig. 02 · current
Current architecture

Retries live inside the request

mobile-sync offline cache on device sync-gateway inline retry ×3, in memory job-service writes the job record ! after attempt 3 the update is dropped — techs in dead zones lose edits
coldharbour rfc-retry-queue · rev 032 / 6
fig. 03 · proposal
Proposal

A durable queue between gateway and service

mobile-sync unchanged sync-gateway enqueue only sync-queue (new) durable · backoff + jitter idempotency keys job-service unchanged dead-letter after 24h → support alert
coldharbour rfc-retry-queue · rev 033 / 6
tbl. 01 · tradeoffs
Tradeoffs · from the RFC

Against the alternatives

Inline retries (today)Cron sweepManaged brokerDurable queue
Lost updatesYes, after 3 triesWindow between sweepsNoNo
Ordering per jobBest effortNoneConfig-dependentGuaranteed
New infra to runNoneNoneA broker clusterOne table + worker
coldharbour rfc-retry-queue · rev 034 / 6
plan · rollout
Rollout

Shadow first, one region at a time

rollout.plan3 steps
  1. 01Two weeks in shadow — enqueue everything, keep serving from the inline path, diff the outcomes.
  2. 02Pacific Northwest techs first, the same ramp shape as the auth cutover.
  3. 03Kill switch is one flag back to the inline path; nothing is deleted until the queue has been boring for two weeks.
coldharbour rfc-retry-queue · rev 035 / 6
q · open
Open questions

Decide in the review

  • q01Max offline window before dead-letter — 24 hours or 72? Rural crews can be out for a weekend.owner · eng
  • q02Ordering scope — per job, or per tech? Per tech is simpler; per job is what dispatch actually assumes.owner · dispatch
  • q03Who owns the dead-letter rota once support gets the alerts?owner · tbd
coldharbour rfc-retry-queue · rev 036 / 6