Mesh networking

What is a self-healing network?

A self-healing network is one that detects a failed link or device and moves traffic onto another working path without a person reconfiguring it. Mesh networks heal this way because each device has several neighbours: when one link breaks, the routing protocol notices, tells the nodes that depended on it, and finds or already knows another route.

Learning objectives

After reading this article you will be able to:

  • Explain why redundant paths let a mesh route around a broken link
  • Describe how on-demand routing, proactive routing and flooding each recover from a failure
  • List what self-healing does not fix, such as messages lost in flight

The idea

In a network built around a single router or tower, every device depends on that one point. When it fails, everyone behind it loses service until someone repairs it. A self-healing network is designed so that no single link is essential: when one fails, traffic moves to another path on its own.

In wireless meshes the term describes something continuous: links between devices appear and vanish all the time as people move, walls get in the way, and batteries run down, and the network keeps adjusting its routes to match.

Why a mesh can heal

A mesh network connects each device to several neighbours instead of one central point. That redundancy is what makes healing possible. If a message from A to D usually goes A to B to D, and B leaves, a path such as A to C to D may still exist. Healing is the network discovering that path and using it.

It only works where the redundancy is real. A device with a single neighbour has no alternative when that neighbour goes away, and a group of devices connected to the rest only through one link is cut off when that link breaks. Healing cannot create a path that does not physically exist; it can only find the ones that do.

How a mesh notices a failure

Before it can route around a failure, a node has to know about it. Mesh protocols use a few mechanisms:

  • Neighbour discovery. Nodes send short periodic messages announcing themselves. The IETF’s MANET Neighborhood Discovery Protocol (RFC 6130) uses HELLO messages so each router learns its 1-hop and symmetric 2-hop neighbours. When HELLOs from a neighbour stop arriving, the link is treated as lost.
  • Link-layer feedback. A radio that fails to deliver a frame after its retries is a strong hint that the neighbour has gone.
  • Missing acknowledgments. At a higher layer, a message that is never acknowledged suggests the path is broken.

How it reroutes

What happens next depends on the routing style.

Reactive (on-demand) routing. The Ad hoc On-Demand Distance Vector protocol, AODV (RFC 3561), finds routes only when they are needed. Nodes monitor the next hops on their active routes. When a link in an active route breaks, a Route Error (RERR) message notifies the nodes that were using it, so they invalidate the route, and a new route discovery finds another path. The RFC describes this as letting mobile nodes “respond to link breakages and changes in network topology in a timely manner”.

Proactive (table-driven) routing. OLSRv2 (RFC 7181) exchanges topology information regularly, so every router keeps routes to all destinations ready. When a link disappears from the topology, the routes are recomputed from the updated tables, often before any traffic needs them. The cost is the background control traffic.

Flooding. Some meshes do not keep routes at all: each node rebroadcasts what it has not seen before, within a hop limit. Flooding heals implicitly, because a message takes every surviving path at once. The cost is airtime.

Store-and-forward. When no path exists at all, a node can hold the message and send it once a path reappears. That turns a temporary partition into a delay rather than a loss. See store-and-forward messaging.

What healing does not fix

Self-healing is about reachability, and it has limits:

  • Messages in flight can be lost when a link breaks mid-delivery. Retries and acknowledgments at a higher layer recover them; routing alone does not.
  • Healing takes time. Detection depends on how often neighbours announce themselves, and rerouting depends on the protocol. Short outages may pass before the network reacts; long ones are handled by rerouting or holding messages.
  • A new path may be worse. It may be longer, slower, or use a device with a low battery.
  • Partitions stay partitioned. If no physical path exists, the two halves cannot talk until a device moves between them or a link returns.

Designing for it

Systems that rely on a mesh healing itself should give it room to do so. Place devices so most have more than one neighbour. Keep retry and outbox limits long enough to outlast the disconnections you expect. Make operations idempotent, so a message retried over a new path does not take effect twice. And test by removing devices while traffic is flowing, not only with every device in place.

Sources

Build it with Offline Protocol

The Offline Protocol mesh SDK scores the eligible paths for each device and switches between them as conditions change, with retries and an outbox holding bounded work while connectivity changes. The transport and routing page describes how it routes and what it needs from the deployment.

Read transport and routing