Delivery and store-and-forward

What is message deduplication?

Message deduplication is the discarding of extra copies of a message that a device, relay or queue has already handled, so each message is processed once even when it arrives several times. It works by giving every message a unique identifier and remembering recently seen identifiers for a limited window.

Learning objectives

After reading this article you will be able to:

  • Identify where duplicate messages come from in at-least-once delivery and meshes
  • Describe how message IDs and a seen-list drop copies while still acknowledging them
  • Distinguish message deduplication from idempotency and explain why apps need both

Where duplicates come from

Duplicates are a side effect of trying hard to deliver. A sender that must not lose messages uses at-least-once delivery: keep sending until someone confirms receipt. Several things then produce extra copies.

  • Lost acknowledgments. The message arrived, but the acknowledgment did not make it back, so the sender sends it again.
  • Reconnections. MQTT 5.0 requires a client and server that resume a session to resend every unacknowledged message with its original packet identifier, and to set the DUP flag to show it might be a re-delivery.
  • Several paths. In a mesh, a message can reach the same device along two routes, and in flooding every relay rebroadcasts it, so a device can hear it from several neighbours.
  • Relays holding copies. With store-and-forward, a relay may keep a message and pass it on later, after another copy has already arrived by a faster route.

None of these is a fault. They are the price of not losing messages, and deduplication is how the price is paid.

How deduplication works

The mechanism is simple. Every message carries an identifier that is unique to it and stays the same on every copy and every retry. Each device, relay or queue keeps a list of the identifiers it has handled recently. When a message arrives, it checks the list: a new identifier is processed and added; a known one is dropped.

The important detail is what happens to the duplicate. The receiver should usually still acknowledge it, because the duplicate may exist precisely because the first acknowledgment was lost. MQTT’s exactly-once level (QoS 2) does this: until the exchange completes, the receiver must answer any repeated message with the same packet identifier with another acknowledgment, and must not deliver it onward again. Amazon SQS FIFO queues behave the same way: a later message with a deduplication ID the queue has already accepted is acknowledged but not delivered to consumers.

In a mesh, every relay can deduplicate as well as the final recipient. That stops a flooded message from circling, alongside a hop limit, and saves relays from forwarding the same message twice.

Choosing the identifier

There are three common ways to give a message its identity.

  • A random identifier created by the sender, such as a UUID. Simple and collision resistant, but it must be created once and stored with the message, not regenerated on retry.
  • Sender plus sequence. The delay-tolerant Bundle Protocol (RFC 9171) identifies a bundle by its source node and creation timestamp, which includes a sequence number for bundles created at the same moment.
  • A fingerprint of the content. The IETF idempotency key draft allows a fingerprint generated from the request payload alongside the key. A fingerprint used alone treats two genuinely separate but identical messages as one, so it suits checking a key’s misuse better than replacing the key.

How long to remember

A device cannot remember every identifier forever, so every deduplication scheme has a window. Amazon SQS keeps a deduplication ID for 5 minutes, and keeps tracking it even after the message has been received and deleted. MQTT holds a packet identifier only until that one exchange completes, then reuses it.

The window has to be longer than the time a duplicate can take to arrive. On a fast server link that is short. In an offline mesh, where a copy may sit on a relay until it meets the next device, it can be much longer, and the list has to be bounded in both count and age so it fits on a phone. As one example, the Offline Protocol mesh SDK documents defaults of 2,000 remembered message identifiers kept for 24 hours (v0.27.0).

Deduplication is not idempotency

Deduplication works on network messages. It has three blind spots: a copy that turns up after the window has expired, a list lost when an app restarts if it was not saved, and the same piece of work sent again as a brand new message with a new identifier, for example after a user taps send twice or an app resubmits on restart.

That is why applications also make their operations idempotent, using an operation ID that belongs to the work rather than to the message. Deduplication keeps the network efficient; idempotency keeps the result correct.

Frequently asked questions

Does deduplication give exactly-once delivery?

Not on its own. It removes copies within the window it remembers, and only on the device that keeps the record. A copy that arrives after the window, or work that is resent as a new message, gets through, so operations with side effects also need to be idempotent.

Should a receiver acknowledge a duplicate?

Usually yes. If the first acknowledgment was lost, the sender will keep retrying until it hears one. MQTT and Amazon SQS both acknowledge a repeated message while declining to deliver it again.

Sources

Build it with Offline Protocol

The configuration reference lists the mesh SDK's deduplication cache size and retention next to the acknowledgment, retry and outbox settings, so you can size them together for how long your devices may be out of contact.

Read the configuration reference