Mesh networking

What is a network partition?

A network partition is a split in which a network divides into groups of nodes that can still communicate within each group but not with the other groups. Each side keeps running with only the nodes it can reach. When the links return, the groups have to merge again, including any data that changed on both sides in the meantime.

Learning objectives

After reading this article you will be able to:

  • Explain why a node inside a partition cannot tell a crash from a split
  • Describe the choice between consistency and availability that the CAP theorem sets out
  • List ways to keep an app working through a partition and merge afterwards

What a partition looks like

Picture a team spread across a building, their phones linked in a mesh. One person in the stairwell is the only device in range of both floors. When they walk out, the mesh splits in two. Phones upstairs can still reach each other, and so can phones downstairs, but nothing crosses between the floors. That is a network partition.

Partitions happen in every kind of network. In data centres, a failed switch can separate racks. On the internet, a cut cable can separate regions. In a mesh network, they happen whenever devices move apart, a relay node leaves, or a gateway goes offline.

The Thread mesh standard describes partitions plainly. OpenThread’s primer says a Thread network might be composed of partitions when a group of devices can no longer communicate with another group. Each partition operates as a distinct network with its own leader, while every device keeps the same security credentials. When the partitions regain connectivity, they merge into one.

Why partitions are hard

From inside a partition, you cannot see the other side. A node that stops hearing from another cannot tell whether that node crashed, switched off, moved away, or is still running happily on the far side of a broken link. Gilbert and Lynch use exactly this point in their proof sketch of the CAP theorem: a server in one part of a partition cannot distinguish between the different writes that might have happened on the other side.

That uncertainty matters most for shared data. If both sides keep accepting changes, their copies drift apart. If one side refuses changes to stay safe, its users are blocked until the network heals.

The CAP theorem

In a keynote at the PODC conference in 2000, Eric Brewer described a trade-off between three properties of a shared-data system: consistency, availability and tolerance to network partitions. His slide states that you can have at most two of them.

Gilbert and Lynch later restated it more carefully. Partition tolerance, they write, can be seen as a statement about the underlying system: communication between servers is unreliable, and the servers may be split into groups that cannot communicate. Since partitions will happen, the real choice is what to give up while one lasts.

  • Favour consistency. Refuse or delay operations that cannot be confirmed with the other side. Brewer’s examples include majority protocols, which make minority partitions unavailable.
  • Favour availability. Keep answering and accepting changes on each side, and reconcile afterwards. Brewer lists optimistic updates and conflict resolution as traits of this choice, with DNS and web caching as examples.

Gilbert and Lynch also point out that a system does not have to make one choice for everything. Read-only operations can stay available during a partition while updates wait, and a purchase can demand consistency while a query returns out-of-date data.

Designing for partitions

Treat a split as a normal state, not an error. For each operation, decide in advance what happens when the other side is unreachable.

Keep working locally where you can. Notes, checklists, readings and messages to nearby devices can usually carry on during a partition. Changes are stored and synced later.

Hold back what needs an authority. Booking the last seat or assigning the last spare part should not be decided by two groups that cannot see each other. Queue the request, or let the person know it will be confirmed later.

Carry data across the gap. Delay-tolerant networking was designed for this. RFC 4838 describes an architecture for occasionally connected networks that may suffer frequent partitions, using storage inside the network to move data store-and-forward, without assuming an end-to-end path exists.

Plan the merge. When the groups reconnect, each side brings changes the other has not seen. A merge rule decided ahead of time, such as last writer wins per field or a CRDT, lets every device reach the same result. How are sync conflicts resolved? covers the options.

Do not trust every split. RFC 4593 lists partition among the consequences of attacks on routing: part of the network can be made to believe it is cut off when it is not. Authenticated routing makes this harder.

A partition is also different from simply being offline. A phone with no signal may still be in a busy local network with the phones around it. Offline vs partitioned explains why apps should handle the two separately.

Frequently asked questions

Is a network partition the same as a network outage?

Not quite. In an outage, a device or link stops working. In a partition, the nodes on each side keep working and keep talking to each other; they just cannot reach the other side. Both sides may carry on making changes.

Can you avoid partitions entirely?

No. Links fail, radios go out of range and devices move. Gilbert and Lynch describe partition tolerance as a statement about the underlying system, in which communication is unreliable. The design question is what each side should do while split.

Sources

Build it with Offline Protocol

The Offline Protocol shared state guide shows how replicated documents merge changes when replicas communicate again, and which merge rule applies to each collection type.

Read the shared state guide