Telemetry and edge data

How do you sample telemetry?

You sample telemetry by keeping a representative fraction of it and dropping the rest, deciding either when an operation starts (head sampling) or after it has finished (tail sampling). Head sampling is cheap and can run on the device itself. Tail sampling can keep every error or slow operation, but it needs a component that holds the whole trace before it decides.

Learning objectives

After reading this article you will be able to:

  • Distinguish head sampling from tail sampling and where each one can run
  • Explain how adjusted counts let a backend estimate totals from a sample
  • Choose between sampling and on-device aggregation for different kinds of device events

Why sample at all

Every event a device records costs something: battery and radio time to send it, and storage and processing wherever it lands. OpenTelemetry’s documentation calls sampling one of the most effective ways to reduce the costs of observability without losing visibility, because a well-chosen sample still represents the whole. Filtering and aggregating also cut volume, but they do not keep that property.

The terms are easy to mix up. In OpenTelemetry’s words, a span or trace that is sampled is processed and exported; one that is not sampled is dropped. “Sampling out” data is not a correct use of the word.

Sampling is not always worth it. OpenTelemetry suggests avoiding it when you generate very little data, when you only ever use the data in aggregate and can pre-aggregate instead, or when regulation forbids dropping data. It also lists the costs: compute for the sampler, the engineering time to maintain sampling rules, and the risk of missing something important with a poor strategy.

Head sampling

Head sampling makes the keep-or-drop decision as early as possible, before the operation is complete. The common form is consistent probability sampling: the decision is computed from the trace ID and a target percentage, so every service that sees the same trace makes the same choice and whole traces are kept, with no missing spans.

OpenTelemetry’s tracing SDK specification defines built-in samplers for this. AlwaysOn keeps everything, AlwaysOff keeps nothing, a ratio-based sampler keeps a fixed probability, and ParentBased follows whatever decision the parent span already made. The default is ParentBased with AlwaysOn at the root. The ratio-based algorithm must be deterministic, so a given trace ID always gets the same answer.

Head sampling is easy to understand, easy to configure and cheap, and it can run at any point in the pipeline, including on the device. That last property matters for phones and sensors: a decision made on the device saves the battery, radio time and upload bytes for every dropped event, not only the backend storage. Its weakness is that it decides before it knows how things turned out, so on its own it cannot guarantee that every trace containing an error is kept.

Tail sampling

Tail sampling decides after seeing all or most of the spans in a trace. That allows rules such as keeping every trace with an error, keeping traces above a latency, or sampling more heavily from a newly deployed service.

The cost is state. The OpenTelemetry Collector’s tail sampling processor holds spans in memory until a decision wait has passed, 30 seconds by default, and then applies its policies, which include status code, latency, attribute and rate-limiting rules. OpenTelemetry notes that tail samplers can be difficult to implement and to operate, because they must accept and store large amounts of data before deciding.

For device telemetry, this means tail sampling happens after upload. The device still sends everything to a collector, so tail sampling reduces what you store and pay to keep, but not what the device spends to send. The two can be combined: head sampling first to protect the pipeline, then tail sampling for the finer choices.

Count what you dropped

A sample is only useful if you can scale it back up. OpenTelemetry defines the adjusted count as the reciprocal of the sampling probability: the number of original items each kept item represents. If you keep one event in N, each kept event stands for N. Record the rate on each event, or in the batch that carries it, so the backend can estimate totals instead of reporting the sample as if it were everything.

Sampling also changes what can be joined. If one device keeps some of a user’s events and drops others at random, a session can no longer be reconstructed. The Collector’s probabilistic sampling processor can take its randomness from a chosen log record attribute rather than the trace ID, which lets you sample per session or per install: use a session value as that attribute, and the whole session is kept or dropped together.

Aggregate instead of sampling

Some questions do not need individual events at all. If you only want to know how many messages failed or how long uploads took, count and summarise on the device and send the summary. Prometheus defines the building blocks: a counter that only goes up or resets to zero, a gauge that can move either way, and a histogram that counts observations into configurable buckets and keeps their sum. A histogram of upload durations sent once per period replaces every individual measurement while still supporting quantile estimates.

A practical policy for telemetry events from devices combines these ideas:

  • Keep failures, crashes and rare state changes at full rate.
  • Sample routine, high-volume success events by session or install, not one event at a time.
  • Send counters and histograms for things you only need in aggregate.
  • Put a rate limit on each device so a bug that emits in a loop cannot flood the pipeline or the telemetry bill.

Frequently asked questions

Should crash reports be sampled?

Treat them separately from routine events. Failures are the rare signal you collect telemetry to find, and OpenTelemetry's first example of a tail sampling rule is always keeping traces that contain an error. Sample the healthy, high-volume events instead.

Is dropping debug logs the same as sampling?

No. Filtering removes a whole category of data, so what remains no longer represents the original population. OpenTelemetry describes filtering and aggregation as cost controls that, unlike sampling, do not keep representativeness.

Sources

Build it with Offline Protocol

The telemetry controls page documents how the SDK batches and uploads its optional telemetry, the counters for buffered, accepted and dropped events, and the switch that stops new collection for a user setting.

Read the telemetry controls