Industry problems

How do robots coordinate without a network?

Robots coordinate without a central server by discovering each other directly on a shared local network. ROS 2 does this through DDS, where nodes with the same domain ID announce themselves and connect, and Zenoh offers peer-to-peer scouting or routers. When robots share no working link at all, each one has to act on local rules and reconcile its records when contact returns.

Learning objectives

After reading this article you will be able to:

  • Distinguish losing the internet from having no shared link between robots
  • Explain how ROS 2 nodes discover each other through DDS and domain IDs
  • Compare Zenoh peer scouting with the router-based discovery in rmw_zenoh

Two different meanings of “no network”

Coordinating robots without a network can mean one of two situations. In the first, there is no internet or cloud connection, but the robots still share a local network such as site Wi-Fi or a wired backbone. In the second, there is no shared link at all between some of the robots, because one has driven out of range or the access point has failed.

The first case is the one robot middleware is designed for. The second is a network partition, and no middleware can deliver a message across it. It can only help each side notice, keep working, and catch up later.

How ROS 2 robots find each other

ROS 2 is built on DDS, a standard from the Object Management Group that defines a data-centric publish-subscribe model: producers publish data, and the middleware delivers it to matching consumers. Several vendors implement it, and the ROS 2 documentation lists implementations from eProsima, Eclipse and RTI among the supported options.

Discovery happens through that middleware, without a central registry. The ROS 2 documentation describes the process:

  • When a node starts, it advertises its presence to other nodes on the network with the same ROS domain.
  • Other nodes reply with information about themselves so connections can be made.
  • Nodes keep advertising periodically, so they can find nodes that appear later, and they announce when they go offline.
  • Nodes only connect if their Quality of Service settings are compatible.

The domain is set with the ROS_DOMAIN_ID environment variable. Nodes on the same domain can discover and message each other; nodes on different domains cannot. Every node uses domain 0 by default, so two robot teams on one network should use different IDs. The documentation suggests choosing a value between 0 and 101 inclusive to stay clear of port conflicts.

By default ROS 2 tries to find all nodes on all hosts on the same subnet, which for DDS means anything reachable by multicast. The ROS_AUTOMATIC_DISCOVERY_RANGE variable can narrow this to the local machine or turn it off, and ROS_STATIC_PEERS lists specific addresses to try, which helps where multicast is not available.

The ROS 2 documentation names lossy wireless networks as one reason it exposes DDS Quality of Service policies, which let each stream choose how to behave:

  • Reliability. Best effort may lose samples; reliable retries until they arrive. The documentation’s sensor data profile uses best effort, because timely readings matter more there than receiving every one.
  • Durability. Transient local makes the publisher keep samples for subscriptions that join late.
  • Liveliness. A publisher must show it is alive within a set period, or the system treats it as lost, which is one way to detect that a teammate has gone quiet.

These settings shape behaviour on a degraded link. They do not create a path where none exists.

Zenoh: peers, clients and routers

Eclipse Zenoh is another option, and ROS 2 can use it through the rmw_zenoh middleware. Plain Zenoh applications run in peer mode by default and talk directly to each other on the local network. They find each other by multicast scouting, or by gossip scouting from a configured entry point when multicast is unavailable. Client mode keeps a single session to another node, and router mode can route data for others in a mesh of routers that are configured with each other’s addresses.

The ROS 2 integration makes a different default choice. In rmw_zenoh, multicast scouting is disabled in the node’s session configuration, so nodes discover each other through a Zenoh router. Without the router they cannot find each other unless multicast scouting is enabled. A design that relies on that router has a coordination point to keep running on each host, and teams should plan for it.

Discovery protocols assume the robots share some network. When a robot leaves coverage, coordination has to become local:

  • Each robot needs rules it can follow alone, such as stopping at a boundary or holding a task until it can confirm it.
  • Records made while apart need the robot’s own timestamps and sequence numbers, so they can be ordered when devices reconnect.
  • Exclusive decisions, such as which robot takes a narrow aisle or a charger, need an explicit owner or a reservation that expires. Merging data after the fact cannot undo two robots acting on the same claim.

Device-to-device software can narrow the gap by carrying short coordination messages over any link two machines still share, including Bluetooth LE between a robot and a nearby phone, and relaying through other devices. Offline Protocol’s Python binding, which is built from source rather than installed from a package index, runs a peer on a Linux host over a local IP network and keeps its identity and state across restarts. It sits beside the robot middleware, not in place of it.

Frequently asked questions

Does ROS 2 need a central server to work?

Not with its DDS middleware. Nodes discover each other through the middleware on the local subnet. The Zenoh middleware for ROS 2 is different by default, because nodes learn about each other through a Zenoh router unless multicast scouting is turned on.

What is ROS_DOMAIN_ID for?

It separates logical ROS 2 networks that share one physical network. Nodes on the same domain can discover and message each other, and nodes on different domains cannot. The default is 0.

Sources

Build it with Offline Protocol

The Python and Linux gateway page shows how to build the Offline Protocol Python binding from source and run a peer on a Linux host over a local IP network, with persistent identity and state.

Read the Python and Linux gateway docs