Telemetry and edge data

How do you keep event order when devices are offline?

You keep event order offline by not relying on wall-clock time alone. Each device numbers its own events with a counter, logical clocks such as Lamport timestamps or vector clocks record which events could have influenced which, and a hybrid logical clock adds a value that stays close to real time. Together they let a backend put late, out-of-sequence uploads back in order.

Learning objectives

After reading this article you will be able to:

  • Explain why arrival order and device timestamps do not give true event order
  • Describe how per-device sequence numbers give order, gap detection and a unique key
  • Compare Lamport timestamps, vector clocks and hybrid logical clocks for ordering across devices

Why arrival order is not event order

When a connected device sends each event as it happens, the order events reach a server roughly follows the order they occurred. An offline device breaks that. It keeps working, stores what happens, and uploads everything when it next has a link. Events from several devices then arrive in bursts, interleaved, hours or days after they happened, and a retry can deliver a batch twice.

The obvious fix is to sort by the timestamp each device attached. That works only as well as the device’s clock. Android’s reference documentation warns that the standard wall clock “can be set by the user or the phone network”, so the time “may jump backwards or forwards unpredictably”, and a device without a connection cannot fetch network time to correct it. Two devices that never talk to each other can drift apart without either noticing, which is the subject of clock drift.

So the design question is what to attach to each event so that order can be reconstructed without trusting wall clocks.

Order within one device: sequence numbers

On a single device the answer is simple. Lamport’s 1978 paper starts from the assumption that the events of a single process form a sequence. A device can make that explicit by keeping a counter, stored durably so it survives restarts, and stamping each event with the next number.

A per-device sequence number gives three things a timestamp cannot:

  • Exact order for that device’s events, even if its clock jumped in between.
  • Gap detection. If a number is missing between two the server already holds, it knows exactly which event to ask for again.
  • A natural unique key. The pair of device identifier and sequence number names each event once, so a re-sent batch can be recognised and deduplicated.

For durations, such as how long a pump ran, use a monotonic clock rather than the wall clock. Android’s elapsedRealtime counts time since boot, including deep sleep, and is guaranteed to be monotonic, so it never runs backwards. It restarts from zero at boot, so it measures intervals within one boot, not dates.

Order across devices: logical clocks

Sequence numbers say nothing about how one device’s events relate to another’s. For that, distributed systems use logical clocks.

A Lamport timestamp is a counter that each device increments for every event and moves past any value it sees in an incoming message. If one event could have caused another, the cause has the smaller number. Breaking ties with the device identifier gives every event a single agreed order. Automerge, for example, orders operations by a counter and uses the actor ID only to break ties, rather than using wall-clock time.

A Lamport timestamp cannot show whether two events were independent. A vector clock keeps one counter per device and can, which matters when two technicians edit the same record without seeing each other’s change.

Logical clocks only capture influence that passes through the system. Lamport’s paper gives the example of a person who issues a request on one computer and then telephones a friend to issue a second request on another: the system has no way of knowing the first came before the second. If the order of such events matters, the application has to carry it explicitly, for instance by having the second request reference the first.

Close to real time: hybrid logical clocks

Logical clocks order events but cannot answer “what happened during this hour”. Hybrid logical clocks (HLC), proposed by Kulkarni, Demirbas and colleagues in 2014, combine the two. An HLC timestamp captures causality like a logical clock while staying close to the physical clock, fits in the 64-bit NTP timestamp format, and tolerates NTP adjustments and uncertainty. When a device receives a message from a peer whose clock is ahead, its HLC moves forward rather than issuing a timestamp that appears to come before the message it just read.

A record that can be reordered later

Putting these together, an event from a device that may be offline can carry:

FieldWhat it gives you
Device identifier and sequence numberExact per-device order, gap detection, a unique key
Logical or hybrid timestampOrder across devices that exchanged messages
Device wall-clock timeAn estimate of when it happened, for people to read
Server receive timeWhen the upload got through, for operations

Sort each device’s events by sequence number first. Use the logical or hybrid timestamp to interleave devices that communicated, and treat wall-clock time as an estimate that can be wrong. Where two changes are truly concurrent, a merge rule such as last writer wins or a CRDT decides the result, and that rule should be the same on every replica.

Frequently asked questions

Can I just sort by the server's receive time?

Only if you want arrival order. For a device that was offline, receive time records when its upload got through, not when anything happened, so a day of readings can share almost the same receive time. Keep it as a separate field alongside the device's own ordering information.

Sources

Build it with Offline Protocol

Offline Protocol's replicated documents do not depend on the order changes arrive in. Each collection type has a fixed merge rule, and concurrent list insertions survive in a deterministic order on every device. The shared state guide lists the rules and how to test them across a partition.

Read the shared state guide