Telemetry, at the edge
Telemetry is data a system reports about itself. OpenTelemetry, the open observability project, calls these reports signals: “system outputs that describe the underlying activity” of software running on a platform. It currently supports traces, metrics, logs and baggage, which is context passed between signals. The first three carry the data itself:
- Traces, the path of a request through an application.
- Metrics, measurements captured at runtime, such as memory use or temperature.
- Logs, recordings of individual events.
Edge telemetry is the same data produced at the edge of the network: on phones, sensors, vehicles, gateways and other devices that sit with their users rather than in a data centre. NIST’s fog computing model describes this edge as the layer that includes end devices and their users, such as a sensor or a meter with its own computing (see what edge computing is). A single record, such as “message delivered through a relay” or “battery at a given level”, is a telemetry event.
Why the edge changes the problem
A server in a data centre can send its telemetry to a collector a short, reliable network hop away. An edge device cannot assume that. Its link to the wider network, its backhaul, may be missing for minutes or days. It may be on battery, on a metered mobile plan, or behind a gateway that is only reachable some of the time. And the device itself may restart, be killed by the operating system, or run out of storage while it waits.
So edge telemetry is mostly a delivery problem. The interesting design questions are where records wait, how they are grouped, when they are sent, and what happens to them when the device cannot keep everything.
How an edge telemetry pipeline works
An edge telemetry pipeline has four steps.
- Record locally. The device writes each event to a buffer before trying to send it. For data that must survive a restart, that buffer lives on disk. OpenTelemetry’s Android project lists “offline buffering of telemetry via disk persistence” among its features.
- Batch. Events are grouped before sending. The OpenTelemetry Collector’s batch processor documentation says batching “helps better compress the data and reduce the number of outgoing connections”, and it closes a batch either when it reaches a set size or when a set time has passed. Fewer, larger uploads also mean fewer radio wake-ups (why batching saves battery and bandwidth).
- Send and retry. The OTLP specification separates failures into two kinds. A retryable failure, such as a server that is temporarily unavailable, may be retried, and the client should space its retries with exponential backoff. A non-retryable failure, such as data the server cannot parse, must not be retried: the client drops that data and should keep a count of what it dropped.
- Collect and route. On the receiving side, a collector accepts the data, processes it and exports it to one or more backends. The OpenTelemetry Collector is one vendor-agnostic implementation of this role.
What makes it hard
- Finite storage. A device that stays offline long enough will fill its buffer. The pipeline needs a rule for what to drop, such as the oldest records or the newest, and a counter so the backend knows data is missing rather than assuming nothing happened.
- Clocks. Timestamps come from the device’s own clock, which can be wrong or drift while offline. See clock drift.
- Volume. Recording everything can cost more battery, storage and data than the insight is worth. Sampling keeps a representative share.
- Privacy. Telemetry can reveal who a person is, where they were and who they talked to. Deciding what should never leave the device belongs at design time, not after launch.
Opt-in by design
Because telemetry runs on someone else’s device, a careful design keeps it off until the app developer, and where needed the user, turns it on, and publishes exactly what it collects. Offline Protocol’s mesh SDK is one documented example. It collects nothing until the app calls enableTelemetry with a key and App ID. It then uploads a fixed inventory of fourteen event types, such as deliveries, relaying, routing decisions, transport changes and MLS session health, with no message content and no identifier fields, to one endpoint over TLS. Separately, onEvent gives the app every protocol event locally, with no key.
Whatever tools you use, the same checks apply: a durable local buffer with a clear limit, batched uploads, retries that back off, a count of dropped records, and a written list of what leaves the device.