Skip to main content

Persistence & Readiness

n8n-sync is event-driven, but its ordering guarantees depend on small file-backed state stores. Treat these files as sync state, not cache.

State Files​

SidePathPurpose
PublisherSYNC_PUBLISHER_STATE_PATHStores the configured SYNC_SOURCE_ID, the source-scoped event sequence, and per-entity revision counters.
Publisher<SYNC_PUBLISHER_STATE_PATH>.lockBest-effort local lock that rejects another live publisher process using the same path.
SubscriberSYNC_SUBSCRIBER_STATE_PATHStores the last applied revision for each [sourceId, entityKind, entityId], including delete tombstones.
Subscriber<subscriber-state-basename>.executions.jsonStores source-execution to target-execution mappings when execution sync is enabled.

Put these files on persistent storage. If the publisher state is lost, the publisher may reuse old revisions. If subscriber state is lost, the target can forget tombstones and stale-event protection until a newer event arrives.

Readiness​

The subscriber mounts three routes under SYNC_ROUTE_BASE (default /rest/sync/v1):

RouteAuthResponse
GET /healthnone200 { "ok": true } once routes are mounted.
GET /readynone200 { "ok": true, "ready": true } while required file-backed state is loaded and writable; otherwise 503.
POST /eventsHMAC or tokenApplies one sync event. Returns 503 { "ok": false, "ready": false } while readiness is false.

/ready checks current sync state load/write capability. It does not prove that n8n database writes and JSON checkpoints can commit atomically.

Crash Windows​

Current persistence is file-backed. JSON state writes are atomic per file, but they are not atomic with n8n database mutations or other JSON files.

  • A crash after a database write but before recording the subscriber checkpoint can cause a redelivery to hit row timestamp guards or revision conflicts.
  • A crash after an execution row write but before updating the execution identity file can require operator repair.
  • A publisher restart loses events that were still queued in memory and not yet delivered.
  • Exceeding SYNC_MAX_QUEUE_SIZE drops the oldest queued event for that target and logs a warning before enqueueing the new event.

Supported production topology is one publisher process for a given SYNC_SOURCE_ID / SYNC_PUBLISHER_STATE_PATH and one subscriber process for a given route/state path. The publisher lock is best-effort and local to the state path/PID namespace; it is not a distributed lock.

Recovery Notes​

  • 409 SYNC_REVISION_CONFLICT means the target already committed that source/entity revision under a different eventId. This usually indicates duplicate publishers for the same SYNC_SOURCE_ID or a publisher restored with stale state.
  • Stop duplicate publishers before removing a stale .lock file. Remove only the lock file after verifying the original process is gone.
  • To intentionally change SYNC_SOURCE_ID, stop the publisher, back up and move aside the old publisher state file, configure the new source ID, and run a full subscriber resync or source-retirement procedure.
  • Subscriber state format 1 is refused instead of migrated because its colon-separated keys are ambiguous. Back it up, then restore from unambiguous metadata if available or reset subscriber sync state and perform a full source resync.
  • Do not delete subscriber tombstone state unless you accept that stale source events can be applied again.
  • Invalid publisher order state (parsed JSON matches no known shape; supported publisher state versions 1, 2, 3) is quarantined via atomic rename to <statePath>.corrupt.<UTC-timestamp>.bak — the original path is renamed, never overwritten or deleted in place. With the default SYNC_PUBLISHER_INVALID_STATE=fail, the publisher stays degraded (invalid_state) without reiniting counters. With quarantine-reset, counters reinit from zero only as an epoch rotation: the configured SYNC_SOURCE_ID must differ from the quarantined file's stored sourceId, and a full subscriber resync is mandatory before trusting convergence. A same-identity reset — or a reset when the quarantined file carries no usable stored identity — is refused at startup, because reused eventId/entityRevision values under one source identity are rejected by the subscriber as stale/conflict (409 SYNC_REVISION_CONFLICT), causing silent divergence.
  • Publisher invalid-state runbook: back up the live file and any .corrupt.*.bak backup first; inspect the error's parsed version, stored sourceId preview, and entity-key count (a version outside 1, 2, 3 against an older bundle usually means bundle skew, not corruption); on skew, rebuild/redeploy current bundles so the valid file loads untouched; on genuine corruption, restore the quarantined backup after upgrading or start a new epoch and resync.
  • Successful publisher boot logs publisherStateVersion, publisherStateSourceId, publisherNextEventSequence, publisherEntityKeyCount, and invalidStateMode at info; the epoch-reset warn carries previousSourceId, newSourceId, quarantinedBackupPath, and invalidStateMode.

Topology (Kubernetes)​

  • Prefer a dedicated ReadWriteOnce (RWO) PVC for sync state over a shared ReadWriteMany (RWX) volume. A shared RWX volume preserves the state file across pod generations, so a redeployed older bundle can boot against a newer on-disk format and halt sync on the first hook.
  • Run a single publisher replica with the Recreate strategy so two publisher processes never share one SYNC_SOURCE_ID / SYNC_PUBLISHER_STATE_PATH.
  • Deploy digest-pinned images so every pod generation runs the bundle that matches the on-disk state format.
  • Keep SYNC_SOURCE_ID stable for the lifetime of the state file; rotate it only deliberately alongside a quarantined backup and a full subscriber resync.
  • Environment Variables — SYNC_PUBLISHER_STATE_PATH, SYNC_PUBLISHER_INVALID_STATE, SYNC_SUBSCRIBER_STATE_PATH, and SYNC_MAX_QUEUE_SIZE.
  • Architecture — how event revisions and row timestamp guards fit together.
  • Limitations — residual constraints and unsupported topologies.