Delivery and store-and-forward

What is data muling?

Data muling is a way of moving data across a gap that no network link covers. A mobile device, the data mule, collects data when it passes close to a source, stores it while it moves, and hands it over when it comes within range of a destination or an access point. The term comes from the 2003 data MULEs work by Shah, Roy, Jain and Brunette on sparse sensor networks.

Learning objectives

After reading this article you will be able to:

  • Describe how a data mule picks up, carries and drops off data
  • Compare tier-to-tier and end-to-end acknowledgments for data carried by mules
  • Explain the latency, power and infrastructure trade-offs of data muling

Where the idea comes from

The term comes from a paper by Rahul Shah, Sumit Roy, Sushant Jain and Waylon Brunette, published in the journal Ad Hoc Networks in 2003. They looked at sensor networks spread so thinly over a large area that the sensors could not form a connected network, and where installing base stations to cover them all would be expensive.

Their answer was to use things that already move through the area. They called these mobile entities MULEs, for Mobile Ubiquitous LAN Extensions. In a traffic monitoring system, a MULE could be a car or bus fitted with a radio. In habitat monitoring, the paper suggests that animals could play the role. When a MULE passes close to a sensor, it picks up the sensor’s data, keeps it in a buffer, and drops it off when it later passes an access point with a wired connection.

The paper describes this as a three-tier architecture:

  • A top tier of access points connected to the wider network, set up where power and connectivity are available.
  • A middle tier of MULEs, which move in ways that cannot be predicted in advance and carry data between the other two tiers.
  • A bottom tier of fixed sensors with short-range radios.

How a data mule works

Data muling is store-and-forward where the “forward” step includes physical travel. Each handover happens over a short link that exists only while the two devices are near each other.

  1. Pick up. The mule comes within radio range of a source and receives whatever data the source has waiting.
  2. Carry. The mule stores the data while it moves. The paper assumes MULEs have much more storage than sensors and renewable power.
  3. Drop off. The mule meets an access point, or another device closer to the destination, and hands the data over.

The paper also notes that MULEs can exchange data with each other, forming a multi-hop MULE network that can cut waiting time and adds copies of the data for reliability. Copies bring their own work: in the paper, access points can synchronise through a central data warehouse so they can detect duplicates, which is the same job message deduplication does in any system that may receive one message more than once.

Knowing the data arrived

Because the source and the destination are never connected at the same time, confirming delivery is harder than on a live network. The paper discusses two options for acknowledgments:

  • Tier to tier. The mule acknowledges the sensor, and the access point acknowledges the mule. This is simple, but a mule can fail after acknowledging data and before delivering it.
  • End to end. The final destination acknowledges the source. This confirms what actually arrived, but the paper notes that the high variability in end-to-end delay makes it hard to decide when to retransmit.

The trade-offs

The paper compares data muling with two alternatives for sparse sensor networks: base stations that cover the whole area, and enough sensors to form a connected ad hoc network that routes data over many hops. Its qualitative summary rates the MULE approach as:

  • Higher latency, because data waits until a mule happens to pass.
  • Lower sensor power, because sensors only transmit over short range and do not route other sensors’ traffic.
  • Lower infrastructure cost, because far fewer fixed access points are needed.

It also rates the MULE approach lower than base stations on the share of data that reaches an access point. The authors argue that for data needed only for later analysis, on the order of hours or even a day, the extra delay is acceptable. They also flag a cost their analysis leaves out: a sensor may have to keep listening to notice when a mule passes, and listening uses energy too.

Failures are gentler than in a fixed network. No sensor depends on one particular mule, so losing a mule increases delay and lowers the share of data delivered rather than cutting a sensor off.

Data mules and delay-tolerant networking

Data muling is one case of delay-tolerant networking, which is built for paths that are never complete end to end at one moment. RFC 4838, the DTN architecture, expects nodes to keep messages in persistent storage, such as disk or flash, so they survive restarts while waiting for a link. It also describes opportunistic contacts, links that appear without being scheduled, with the example of a handheld device brought near a kiosk. A mule’s meetings with sensors and access points are contacts of that kind, and opportunistic networking studies how to route data over them.

NASA describes DTN as a store-and-forward approach in which each node can hold data until the next node becomes available, comparing it to emails waiting in an outbox until a connection is established.

Designing for it

A system that relies on mules has to plan for long waits:

  • Storage limits. A mule may carry data for many sources at once, so its buffer needs a bound and a rule for what to drop when it is full.
  • Expiry. Data that arrives too late can be useless. RFC 9171 gives every bundle a lifetime after which nodes need no longer keep or forward it, and a mule should apply the same idea through a message TTL.
  • Confidentiality. A mule carries data it was not meant to read. Encrypting the content end to end means a lost or curious mule exposes only ciphertext.

Frequently asked questions

Is a data mule the same as a relay?

Not quite. A relay usually forwards a message straight on while it is connected to both sides. A data mule holds the message while it moves and delivers it later, often after a long gap, so it needs storage and a message that is still worth delivering when it arrives.

Can a phone act as a data mule?

Yes, if the app keeps messages for other devices in durable storage and passes them on when it meets the next device or a network. The app also needs expiry, duplicate checks and end-to-end encryption, because the phone carries data it was not meant to read.

Sources

Build it with Offline Protocol

The mesh SDK configuration page lists how long a device keeps outgoing messages and how many it holds, so you can check that retention covers the gaps between contacts you expect.

Read the configuration docs