Delivery and store-and-forward

What is exponential backoff?

Exponential backoff is a retry strategy in which a client waits longer after each failed attempt, multiplying the delay by a fixed factor up to a maximum. It gives a congested network or an overloaded server room to recover, and adding random jitter to each delay stops many clients from retrying at the same moment.

Learning objectives

After reading this article you will be able to:

  • Explain why a client should wait longer after each failed attempt, up to a cap
  • Describe how full, equal and decorrelated jitter stop clients retrying in step
  • Identify how backoff settings change for offline devices with no path to the recipient

Why not retry straight away

When a request fails or an acknowledgment does not arrive, the obvious move is to send it again at once. That is often the worst choice. Many failures happen because a network is congested or a server is overloaded, and an instant retry adds load to exactly the thing that is struggling. If many clients fail together, they all retry together, and the extra traffic can keep the system down.

Network standards ask senders to slow down instead. RFC 8085, the IETF’s guidance for applications that use UDP, says such applications should use a retransmission interval that is exponentially backed off when packets are deemed lost.

How the delay grows

Exponential backoff has a few parts:

  • An initial delay, the wait before the first retry.
  • A multiplier, applied after each failure. A multiplier of 2 doubles the wait every time.
  • A cap, the longest the client will ever wait between attempts, so the delay does not grow without limit.
  • A stopping rule, either a maximum number of attempts or a total deadline, after which the client gives up and reports the failure.

TCP is the classic example. RFC 6298 says a TCP sender should start with a retransmission timeout of 1 second until it has measured the round-trip time, and must double the timeout each time it expires. It may put a maximum on the timeout, provided that maximum is at least 60 seconds. Client libraries use the same shape: Google Cloud Storage’s Go library, for example, starts with a 1 second delay, multiplies it by 2.0 after each attempt, and caps it at 30 seconds.

The cap matters as much as the growth. Without one, a client that has failed many times can end up waiting so long that it misses the moment the service comes back.

Why jitter matters

Backoff on its own spaces out each client’s retries, but it does not stop clients from moving in step. If many clients fail at the same moment and all use the same schedule, they all retry at the same moments too, in bursts with quiet gaps between.

Jitter fixes this by adding randomness to each delay. Marc Brooker’s post on the AWS Architecture Blog simulated clients competing for the same resource and compared three variants:

  • Full jitter picks a random delay between zero and the current backoff value.
  • Equal jitter keeps half of the backoff fixed and randomises the other half, which avoids very short waits.
  • Decorrelated jitter picks each delay at random between the base delay and three times the previous delay, up to the cap.

In that simulation, exponential backoff without jitter did more work and took longer than any of the jittered versions. Full jitter used the least client work, and decorrelated jitter finished slightly faster. The post concludes that jittered backoff should be a standard approach for remote clients. RFC 8085 makes a related recommendation for periodic keep-alive messages: add slight random variation to their timing so traffic from different hosts does not stay synchronised.

Retry only what is safe, and listen to the server

Backoff decides when to retry, not whether. Two other checks come first.

Is the error worth retrying? Google Cloud Storage’s retry guide treats HTTP 408, 429 and 5xx responses, socket timeouts and dropped TCP connections as temporary problems worth retrying. Errors that mean the request is wrong, such as a failed authorisation, will not improve with time.

Is the request safe to repeat? A retry after a lost reply runs the request a second time, so it should be idempotent or carry an idempotency key. Google’s guide makes the same point: retrying requests that are not idempotent can lead to race conditions and conflicts.

When a server says how long to wait, believe it. HTTP’s Retry-After header, defined in RFC 9110, tells a client how long to wait before its next request, for example while a service is unavailable.

Backoff when devices are offline

On a phone in a mesh or a field device with no signal, many failures mean there is simply no path to the recipient, not that anything is overloaded. Backoff still helps: it stops a device from spending battery and radio time on attempts that cannot succeed. But the right values look different.

  • Longer caps. A path may not return for minutes or hours, so the cap can be long without hurting anyone.
  • Retrying on a new path. When a new neighbour or a connection appears, there is a fresh reason to try, whatever the timer says.
  • Durable state. If the app may be closed between attempts, pending messages and their retry counts belong in an outbox on disk, not in memory.
  • A defined end. After the last attempt, the message should be reported as failed or held for later delivery, not dropped silently. A message TTL sets how long that can go on.

As a concrete example, the Offline Protocol mesh SDK documents these defaults for v0.27.0: an acknowledgment timeout of 10 seconds, and up to 10 retries with backoff from 1 to 300 seconds.

Frequently asked questions

What is the difference between exponential and linear backoff?

Linear backoff adds the same amount of time after each failure. Exponential backoff multiplies the delay, so it backs away from a struggling service much faster. Both should have a cap and a limit on attempts.

Should every failed request be retried with backoff?

No. Retry errors that are likely to be temporary, such as timeouts, dropped connections and overload responses. Errors that mean the request itself is wrong, such as a failed authorisation, will fail the same way every time, and a request that is not idempotent may do its work twice.

Sources

Build it with Offline Protocol

The mesh SDK emits a retrying event each time it schedules another attempt after a send error or acknowledgment timeout, with the retry count and the time of the next attempt, so an app can show that a message is still pending rather than failed.

Read the event reference