The problem it solves
A queue hands each message to a consumer and expects the consumer to finish with it. Sometimes it never can. The message may be malformed, refer to a record that no longer exists, or trigger a bug in the consumer. If the queue keeps redelivering it, the message fails over and over, wasting work and sometimes holding up the messages behind it. If the queue simply discards it, the data is gone and nobody knows why.
A dead-letter queue is the third option. After a set rule says the message has failed, the system moves it to a separate queue where it stays until someone looks at it. The main queue keeps moving, and the failed message is kept as evidence.
How messages end up there
Each broker has its own rules:
| Broker | What moves a message to the dead-letter queue |
|---|---|
| Amazon SQS | Being received more times than the maxReceiveCount in the queue’s redrive policy |
| RabbitMQ | Rejection by a consumer without requeue, per-message TTL expiry, the queue’s length limit, or more returns to a quorum queue than its delivery-limit |
| Azure Service Bus | A delivery count above the maximum (default 10), expiry when dead-lettering on expiry is enabled, a few system limits, or an explicit dead-letter call from the application |
The attempt limit is the rule all three share. In SQS, maxReceiveCount is the number of times a consumer can receive a message before it is moved, and AWS advises setting it high enough to allow sufficient retries. Set it too low, and AWS notes that a single failed receive can move a message aside. Azure applies the same idea with a maximum delivery count, and moves a message when its count exceeds that limit.
What happens after
A dead-lettered message carries a record of why it is there. RabbitMQ adds an x-death header naming the queue it came from, the reason and how many times it has been dead-lettered. Azure Service Bus sets a dead-letter reason, such as MaxDeliveryCountExceeded or TTLExpiredException, and a description.
From there an operator can examine logs, fix the consumer or the data, and send the messages back. SQS calls this redrive. Azure notes there is no automatic cleanup: messages stay in its DLQ until they are explicitly retrieved and completed.
Retention needs care. For SQS standard queues, a message’s expiry is based on when it was first enqueued, not when it was moved, so AWS recommends giving the dead-letter queue a longer retention period than the source queue.
Pitfalls
- Order. Moving one message aside lets later messages overtake it. AWS advises against using a DLQ with a FIFO queue when the exact order of messages or operations matters.
- Dead-lettering can fail too. RabbitMQ points out that dead-lettering is a form of publishing, and messages can be lost if the target queue is not available. Its quorum queues offer at-least-once dead-lettering to close that gap.
- Loops. A misconfigured setup can send a message round in a circle. RabbitMQ detects such cycles and drops the message if there was no rejection anywhere in the cycle.
- Nobody looking. A DLQ that nobody watches is a slower way of losing data. Alert on any message arriving there.
Dead letters in an offline outbox
An app that works offline keeps unsent messages in an outbox and retries them with exponential backoff while no path exists. That outbox has to be bounded, so entries eventually expire, much like a message with a TTL in a broker.
What matters is what happens at expiry. A message that silently disappears from the outbox is the offline version of a dropped message. The better pattern borrows from the dead-letter queue: the sending layer reports a terminal failure, and the app moves the item to a local list of failed work that the user or a support process can see, retry or discard.
Offline Protocol’s mesh SDK is one documented example. A message waiting for an unreachable recipient stays pending, and the SDK may report it as undeliverable many times. It settles either with message_delivered or, when its outbox lifetime runs out, with message_failed, which is terminal: the SDK has stopped retrying and no delivery event follows. The app decides where that failure goes next.
Whatever the stack, make resending safe. A message can be reported as failed after the receiver already processed it, for example when only the acknowledgment was lost, so a resent dead letter should be idempotent.