Delivery and store-and-forward

What is a dead-letter queue?

A dead-letter queue (DLQ) is a separate queue where a messaging system puts messages it could not deliver or process, instead of retrying them forever or dropping them without a trace. Messages usually land there after too many failed attempts, after they expire, or after a receiver rejects them, so that someone can inspect them, fix the cause and send them again.

Learning objectives

After reading this article you will be able to:

  • Explain what a dead-letter queue offers over endless retries or silent discards
  • Compare what moves a message to the DLQ in SQS, RabbitMQ and Azure Service Bus
  • Describe how an offline outbox can report expired messages as visible failed work

The problem it solves

A queue hands each message to a consumer and expects the consumer to finish with it. Sometimes it never can. The message may be malformed, refer to a record that no longer exists, or trigger a bug in the consumer. If the queue keeps redelivering it, the message fails over and over, wasting work and sometimes holding up the messages behind it. If the queue simply discards it, the data is gone and nobody knows why.

A dead-letter queue is the third option. After a set rule says the message has failed, the system moves it to a separate queue where it stays until someone looks at it. The main queue keeps moving, and the failed message is kept as evidence.

How messages end up there

Each broker has its own rules:

BrokerWhat moves a message to the dead-letter queue
Amazon SQSBeing received more times than the maxReceiveCount in the queue’s redrive policy
RabbitMQRejection by a consumer without requeue, per-message TTL expiry, the queue’s length limit, or more returns to a quorum queue than its delivery-limit
Azure Service BusA delivery count above the maximum (default 10), expiry when dead-lettering on expiry is enabled, a few system limits, or an explicit dead-letter call from the application

The attempt limit is the rule all three share. In SQS, maxReceiveCount is the number of times a consumer can receive a message before it is moved, and AWS advises setting it high enough to allow sufficient retries. Set it too low, and AWS notes that a single failed receive can move a message aside. Azure applies the same idea with a maximum delivery count, and moves a message when its count exceeds that limit.

What happens after

A dead-lettered message carries a record of why it is there. RabbitMQ adds an x-death header naming the queue it came from, the reason and how many times it has been dead-lettered. Azure Service Bus sets a dead-letter reason, such as MaxDeliveryCountExceeded or TTLExpiredException, and a description.

From there an operator can examine logs, fix the consumer or the data, and send the messages back. SQS calls this redrive. Azure notes there is no automatic cleanup: messages stay in its DLQ until they are explicitly retrieved and completed.

Retention needs care. For SQS standard queues, a message’s expiry is based on when it was first enqueued, not when it was moved, so AWS recommends giving the dead-letter queue a longer retention period than the source queue.

Pitfalls

  • Order. Moving one message aside lets later messages overtake it. AWS advises against using a DLQ with a FIFO queue when the exact order of messages or operations matters.
  • Dead-lettering can fail too. RabbitMQ points out that dead-lettering is a form of publishing, and messages can be lost if the target queue is not available. Its quorum queues offer at-least-once dead-lettering to close that gap.
  • Loops. A misconfigured setup can send a message round in a circle. RabbitMQ detects such cycles and drops the message if there was no rejection anywhere in the cycle.
  • Nobody looking. A DLQ that nobody watches is a slower way of losing data. Alert on any message arriving there.

Dead letters in an offline outbox

An app that works offline keeps unsent messages in an outbox and retries them with exponential backoff while no path exists. That outbox has to be bounded, so entries eventually expire, much like a message with a TTL in a broker.

What matters is what happens at expiry. A message that silently disappears from the outbox is the offline version of a dropped message. The better pattern borrows from the dead-letter queue: the sending layer reports a terminal failure, and the app moves the item to a local list of failed work that the user or a support process can see, retry or discard.

Offline Protocol’s mesh SDK is one documented example. A message waiting for an unreachable recipient stays pending, and the SDK may report it as undeliverable many times. It settles either with message_delivered or, when its outbox lifetime runs out, with message_failed, which is terminal: the SDK has stopped retrying and no delivery event follows. The app decides where that failure goes next.

Whatever the stack, make resending safe. A message can be reported as failed after the receiver already processed it, for example when only the acknowledgment was lost, so a resent dead letter should be idempotent.

Frequently asked questions

Should every queue have a dead-letter queue?

Any queue whose messages matter should have somewhere for failures to go and someone watching it. The exception is work where order is everything, since moving one message aside lets later ones overtake it.

Is a dead-letter queue the same as a retry queue?

No. A retry queue holds messages that will be tried again automatically, often after a delay. A dead-letter queue holds messages the system has stopped trying, until a person or a separate process decides what to do.

Sources

Build it with Offline Protocol

The Offline Protocol events reference explains which delivery events settle a message and which keep it pending, including the terminal failure an app receives when a message runs out of time in the outbox.

Read the events reference