One hop at a time
A live call or a web request needs a complete path from sender to receiver at the moment it runs. Store-and-forward drops that requirement. The message moves one hop at a time, and between hops it waits in storage on whichever device or server currently holds it. If the next hop is reachable, the wait is short. If it is not, the message stays put until it is. The sender does not need the recipient to be online, and no single link has to be up for the whole journey.
What happens at each hop
Every node in a store-and-forward system runs the same cycle.
- Accept. The node receives the message and checks it: is it well formed, is it addressed somewhere this node can help with, is there room to keep it?
- Store. The node writes the message to storage. For long waits this needs to be persistent: the Delay-Tolerant Networking architecture (RFC 4838) expects most nodes to keep bundles on disk or flash so that they survive a restart.
- Wait for a path. The node holds the message until a suitable next hop is available. RFC 4838 calls each window in which a link can carry data a “contact”.
- Forward and hand over. The node sends the message on. It keeps its own copy until the next node confirms it has taken the message, usually with an acknowledgment, and then it can delete its copy.
On the very first hop, the sending app usually does the same thing locally: it writes the message to an outbox before trying to send it, so that nothing is lost if the app closes.
How long a message waits
The defining question for any store-and-forward system is how long a node is prepared to hold a message.
- Internet routers store and forward packets too, but RFC 4838 points out that the storage is only expected to last about as long as queuing and transmission delays.
- Email waits much longer. SMTP (RFC 5321) requires mail that cannot be sent immediately to be queued and retried. It says the retry interval should generally be at least 30 minutes and that the give-up time generally needs to be at least 4-5 days.
- Message brokers hold messages for clients that are offline. In MQTT 5.0 a broker keeps session state for a disconnected client, including messages pending transmission to it, until the Session Expiry Interval runs out.
- Delay-tolerant networks may hold data for much longer, across links that only exist at scheduled or unpredictable moments. The delay-tolerant networking article covers this case in depth.
Whatever the system, the message carries or is given a limit. Past that point it is deleted and, ideally, the sender is told. Without a limit, messages for recipients that never return would fill storage forever.
Handing over responsibility
The important moment in store-and-forward is not when a message is sent but when another node accepts responsibility for it. SMTP makes this explicit: once a receiving server replies with a success code at the end of the message data, it must either deliver the message or report the failure. From that point the sending server is no longer responsible for it.
RFC 4838 describes an equivalent for delay-tolerant networks called custody transfer, in which nodes that agree to take on reliable delivery become “custodians” and retransmit if needed. The pattern is the same in both: the old holder keeps the message until the new holder confirms, so there is never a moment when nobody has it.
The flip side is that a confirmation can be lost. If the receiving node took the message but its confirmation never arrived, the sender sends it again and the receiver ends up with two copies. Many systems therefore give each message an identifier and drop copies they have already seen, a technique covered in message deduplication.
What makes it hard
Holding messages for a long time turns storage into a resource that has to be managed. RFC 4838 notes that long-term storage brings problems of congestion management and denial-of-service mitigation, because a node that keeps everything it is given can be filled up.
Order is another casualty. RFC 4838 says the relative order of messages might not be preserved: two messages can wait at different nodes or take different paths, and the later one can arrive first. Applications that care about order need their own sequence numbers or timestamps.
Finally, “the message left my device” says little about what happened next. A store-and-forward system can report several distinct outcomes: stored locally, handed to the next hop, delivered to the recipient, and accepted by the recipient’s application. Good interfaces show which of these the user is looking at.
Where you meet it
Email is the classic store-and-forward system, relaying through a chain of servers. Message brokers do it for devices with intermittent connections. NASA describes delay/disruption tolerant networking as a store-and-forward approach in which each node stores data until the next node becomes available, and it now runs as an operational service on NASA’s networks. Phone meshes use it too: a phone holds a message until a neighbour comes into range, and the message moves closer through multi-hop relay as devices meet.