Local-first data

What is a replicated document?

A replicated document is a structured piece of shared data, such as a checklist, a note or a form, with a full copy on every device that shares it. Each device reads and edits its own copy, even offline, and the copies exchange changes and merge them, so devices that have seen the same changes hold the same content.

Learning objectives

After reading this article you will be able to:

  • Describe how copies of a replicated document track, send and merge changes
  • Choose document boundaries using access, size and merge behaviour
  • Explain what replication leaves unsolved, such as business rules and racing deletes

A copy on every device

In a client-server app, a record lives on the server and each device fetches it, edits it through the server and waits for the reply. A replicated document turns this around. Every device that shares the document keeps a complete copy, reads from it and writes to it locally, and the copies are reconciled later.

Kleppmann and Beresford describe this setting in their JSON CRDT paper: a full copy of the document is replicated on several devices, each device can change its own copy optimistically, and the changes are sent to the others asynchronously. Their focus is on phones and laptops with intermittent connectivity, and they make no distinction between devices belonging to one person and devices belonging to different people.

Because every write is local, the app keeps working with no network at all. What changes is the question of when everyone sees the same thing. The answer depends on how the copies exchange and merge their changes.

What a document holds

A replicated document usually looks like a familiar data structure built from CRDTs:

  • Maps of named fields, such as the title, status and assignee of a job.
  • Lists of ordered entries, such as checklist items or comments.
  • Text that several people can type into at once.
  • Counters that add up increments from every device.

Automerge, for example, describes a document as “a combination of a JSON object and a git repository”. Like a JSON object, it is a map whose values can be nested maps, lists or simple values. Like a git repository, it has a history of changes, and that history plus a set of merge rules means any two copies can always be merged. Yjs takes a similar approach with a Y.Doc that holds named shared types.

The document is also the unit of sharing. Automerge calls it the “unit of change” and gives each document a URL that peers use to request it. Sync protocols in both libraries work per document.

How the copies stay in step

Keeping copies aligned comes down to sending each device the changes it has not seen and merging them in.

  • Tracking what each side has. Yjs keeps a state vector that records how many changes from each client a copy already holds. Two clients can swap state vectors and send only the missing differences, at the cost of one extra round trip.
  • Sending changes. Changes travel as compact updates. Yjs encodes them in a binary format; Automerge has its own compact storage format and sync protocol.
  • Merging. The receiving device applies the updates. Because the merge rules produce the same result in any order, it does not matter which device’s changes arrive first.
  • Persisting. The device saves the merged copy locally so it survives a restart.

None of this depends on a particular network. A document can sync through a server, directly between two phones, or across a mesh with store-and-forward relays. Saving an edit locally and having it reach another device are separate events, and an app should not treat the first as the second.

Deciding what goes in one document

Document boundaries matter more than they first appear.

  • Access. Everyone who holds a document can read all of it. Put data with different audiences in different documents, or in different groups of documents. In Offline Protocol, documents live in a space, which is an existing MLS session or group, and the shared state guide advises separate spaces for different access boundaries.
  • Size. A document carries its history. The Ink & Switch team found in their prototypes that performance and memory use became a problem because CRDTs store all history, down to individual keystrokes, and that history is hard to truncate because someone may reconnect months later.
  • Merge behaviour. Fields that should merge independently belong in separate keys, not packed into one value that is replaced as a whole.

What a replicated document does not do

Replication gives every device the same content. It does not decide whether that content is acceptable. Two devices can each assign the last spare part, and the copies will converge on a state with the part assigned twice. Those rules need an authority, such as a backend, that checks them; how sync conflicts are resolved covers when to hand a decision to one.

Deletion needs care too. Kleppmann and Beresford show that if one device deletes a to-do item while another marks it done, the merged result keeps an item with no title, because the concurrent edit survived. Treat removal as a change that can race with other changes, not as a guarantee that data is gone everywhere.

Frequently asked questions

Is a replicated document the same as a database row?

Not quite. A row usually lives on a server and devices request it. A replicated document lives in full on each device that shares it, and the devices merge their copies. An app can keep both, with replicated documents for shared working state and a database for records outside that scope.

Sources

Build it with Offline Protocol

The shared state guide walks through creating a replicated document in the Offline Protocol Mesh SDK, writing and reading it, choosing which documents a device replicates, and testing edits made on both sides of a partition.

Read the shared state guide