The idea
In a network built around a single router or tower, every device depends on that one point. When it fails, everyone behind it loses service until someone repairs it. A self-healing network is designed so that no single link is essential: when one fails, traffic moves to another path on its own.
In wireless meshes the term describes something continuous: links between devices appear and vanish all the time as people move, walls get in the way, and batteries run down, and the network keeps adjusting its routes to match.
Why a mesh can heal
A mesh network connects each device to several neighbours instead of one central point. That redundancy is what makes healing possible. If a message from A to D usually goes A to B to D, and B leaves, a path such as A to C to D may still exist. Healing is the network discovering that path and using it.
It only works where the redundancy is real. A device with a single neighbour has no alternative when that neighbour goes away, and a group of devices connected to the rest only through one link is cut off when that link breaks. Healing cannot create a path that does not physically exist; it can only find the ones that do.
How a mesh notices a failure
Before it can route around a failure, a node has to know about it. Mesh protocols use a few mechanisms:
- Neighbour discovery. Nodes send short periodic messages announcing themselves. The IETF’s MANET Neighborhood Discovery Protocol (RFC 6130) uses HELLO messages so each router learns its 1-hop and symmetric 2-hop neighbours. When HELLOs from a neighbour stop arriving, the link is treated as lost.
- Link-layer feedback. A radio that fails to deliver a frame after its retries is a strong hint that the neighbour has gone.
- Missing acknowledgments. At a higher layer, a message that is never acknowledged suggests the path is broken.
How it reroutes
What happens next depends on the routing style.
Reactive (on-demand) routing. The Ad hoc On-Demand Distance Vector protocol, AODV (RFC 3561), finds routes only when they are needed. Nodes monitor the next hops on their active routes. When a link in an active route breaks, a Route Error (RERR) message notifies the nodes that were using it, so they invalidate the route, and a new route discovery finds another path. The RFC describes this as letting mobile nodes “respond to link breakages and changes in network topology in a timely manner”.
Proactive (table-driven) routing. OLSRv2 (RFC 7181) exchanges topology information regularly, so every router keeps routes to all destinations ready. When a link disappears from the topology, the routes are recomputed from the updated tables, often before any traffic needs them. The cost is the background control traffic.
Flooding. Some meshes do not keep routes at all: each node rebroadcasts what it has not seen before, within a hop limit. Flooding heals implicitly, because a message takes every surviving path at once. The cost is airtime.
Store-and-forward. When no path exists at all, a node can hold the message and send it once a path reappears. That turns a temporary partition into a delay rather than a loss. See store-and-forward messaging.
What healing does not fix
Self-healing is about reachability, and it has limits:
- Messages in flight can be lost when a link breaks mid-delivery. Retries and acknowledgments at a higher layer recover them; routing alone does not.
- Healing takes time. Detection depends on how often neighbours announce themselves, and rerouting depends on the protocol. Short outages may pass before the network reacts; long ones are handled by rerouting or holding messages.
- A new path may be worse. It may be longer, slower, or use a device with a low battery.
- Partitions stay partitioned. If no physical path exists, the two halves cannot talk until a device moves between them or a link returns.
Designing for it
Systems that rely on a mesh healing itself should give it room to do so. Place devices so most have more than one neighbour. Keep retry and outbox limits long enough to outlast the disconnections you expect. Make operations idempotent, so a message retried over a new path does not take effect twice. And test by removing devices while traffic is flowing, not only with every device in place.