Offline-first and sync

What is delta sync?

Delta sync is a way of keeping copies of data in step by sending only what changed since they last agreed, instead of the whole data set. Each side remembers a marker for the last state it synced, such as a version, offset, or token, and the other side answers with the changes after that point.

Learning objectives

After reading this article you will be able to:

  • Explain why sending only changes beats resending the whole data set
  • Describe the markers sync systems use to know where they left off
  • Identify when a delta is not possible and a full sync is needed

Why send only the changes

The simplest way to sync is to send everything. The device downloads the full data set, compares it with what it has, and keeps the newer copy. That works for small data and becomes expensive as data grows, because most of what is sent has not changed.

The IETF made this argument for the web in RFC 3229, Delta encoding in HTTP, published in 2002. Many requests fetch a slightly modified version of something the client already has cached, and the modifications are typically much smaller than the whole resource, so it is more efficient to send a description of the changes. On a slow, metered, or intermittent link, which is the normal case for an offline-first app, the difference decides whether sync finishes before the connection drops.

Three levels of delta

Delta sync appears at different levels depending on what the data looks like.

Bytes in a file or document. RFC 3229 lets an HTTP client name the version it holds, using the entity tag in an If-None-Match header, and list the delta formats it can apply in an A-IM header, such as vcdiff. A server that can compute the difference replies with status 226 (IM Used) and the delta; one that cannot simply sends the whole resource. The rsync algorithm, described by Andrew Tridgell and Paul Mackerras, solves a harder version of the problem: it finds which parts of a file the other machine already has and sends only the parts that cannot be matched, without ever having both files on the same machine.

Records and fields. Most app sync works on rows or documents. A sync service tracks which records changed and sends those. It can go finer: PowerSync’s client upload queue records an update as the row identifier and the value of each changed column, not the whole row.

Operations. Some systems record each change as an operation, such as “add this item” or “insert these characters here”, and sync the operations a peer has not yet seen. Replicated data types such as CRDTs are built so that these operations can be applied in any order and still produce the same result.

Knowing where you left off

A delta only makes sense relative to a starting point, so both sides need to agree on what the receiver already has. RFC 3229 puts it directly: a client cannot request a delta without identifying which version it holds. Sync systems use different markers:

  • A version token from the server. In Replicache, every pull request carries a cookie, a value the client treats as opaque that identifies the server state it last received. The server uses it to compute a patch that brings the client up to date.
  • A position in a log. Electric streams changes from a shape log. The first request asks for the whole log, and later requests pass the offset the client has reached, so each response contains only newer entries.
  • A change notification. In Android’s description of push-based sync, the app downloads a baseline on first start, then the server tells it which data is stale and the app fetches only that.

For two devices syncing directly, with no server keeping the order, each side summarises what it has seen from every other device, and comparing the summaries shows what is missing. What is a vector clock? explains the structure behind this.

When a delta is not possible

A delta needs the receiver’s starting point and enough history to build the difference from it. Sometimes neither side has that:

  • The history was trimmed. Servers cannot keep every old version forever. When a client’s marker is too old, the server has to send a full copy. Electric’s API has a must-refetch control message telling the client to throw away its local data and sync again from scratch.
  • The delta would be larger than the data. RFC 3229 does not require a server to send a delta, and gives this as one example of when not to.
  • Deletions. A record that no longer exists cannot be sent as a changed row. Sync protocols send an explicit delete, as PowerSync’s queue does, or keep a marker that the record was removed.

So every delta sync design also needs a full sync path, and should be tested with a device that comes back after a long time away.

Delta sync on the device

On the receiving side, a delta has to be applied atomically and saved before the interface treats it as real, or a crash halfway through leaves a copy that matches neither version. Offline Protocol’s replicated documents follow this order: the SDK emits its data_changed event only after a change’s delta is durable on the device, so a screen that refreshes on that event shows state that survives a crash. How does offline sync work? covers how the copies converge once the changes arrive.

Frequently asked questions

Is delta sync the same as incremental backup?

They share the idea of copying only what changed since a known point. Backup copies one way to an archive; sync usually runs in both directions and has to merge changes made on each side.

What happens if a device has been offline too long?

The other side may no longer hold the history needed to build a delta. A well-designed sync protocol then tells the client to discard its copy and download a full snapshot, as Electric's must-refetch message does.

Sources

Build it with Offline Protocol

Offline Protocol's replicated documents sync between devices and merge deterministically when copies meet. The shared state guide shows how to scope which documents replicate and how to test convergence after a peer partition.

Read the shared state guide