Telemetry and edge data

What is observability for mobile apps?

Observability for mobile apps is the ability to understand how an app behaves on real users' devices from the telemetry it sends back, such as crashes, hangs, performance traces, logs and metrics. Unlike a server, the app runs on hardware you do not control and on networks that come and go, so its telemetry has to be buffered, batched, kept small and limited to what users have agreed to share.

Learning objectives

After reading this article you will be able to:

  • Identify the signals a mobile app sends, such as crashes, ANRs and performance timing
  • Distinguish platform reports like Android vitals and MetricKit from in-app SDKs
  • Describe how an offline app buffers and sequences telemetry it cannot send yet

What observability means

OpenTelemetry describes observability as the ability to understand a system from the outside, by asking questions about it without knowing how it works inside. That only works if the code is instrumented: it has to emit signals, which OpenTelemetry groups as traces, metrics and logs. Telemetry is the general name for that emitted data.

On a server, the team owns the machine, the network and the logs. A mobile app is different. OpenTelemetry’s guidance for client-side apps points out that they run on devices you do not control, with varying network conditions, hardware and user behaviour. The questions are also different: not only “is the API up?” but “how long did the app take to start on this phone?”, “why did this screen freeze?” and “what was the user doing before it crashed?”

Why mobile is harder

OpenTelemetry lists the factors that change when instrumentation moves onto a phone:

  • Resource constraints. CPU, memory and battery are limited, so collecting telemetry must not slow the app down.
  • Network variability. Users may be on slow, intermittent or no connectivity, so telemetry needs offline buffering and batched exports.
  • Sessions. Grouping telemetry by user session makes it possible to follow a journey across several app launches.
  • Privacy and consent. Client data can fall under privacy law, so plan for data minimisation, consent and redaction of attributes. See what data should never leave the device.
  • Volume. A large user base produces more data than is useful to keep, which makes sampling telemetry a cost decision as much as a technical one.

The signals a mobile app sends

  • Crashes. The app stops unexpectedly. Firebase Crashlytics describes itself as a crash reporter that groups crashes and highlights the circumstances that led to them, so a team can see whether one crash affects a lot of users.
  • Hangs and ANRs. The app is still running but stops responding; on Android this is reported as an ANR, for “application not responding”. OpenTelemetry’s Android agent bundles ANR detection alongside crash reporting, startup timing and slow or frozen frame detection.
  • Performance. Start-up time, screen load and rendering, and network request timing, as real users experience them rather than as measured in a test lab.
  • Logs and telemetry events. Timestamped records of things that happened, such as a sign-in, a failed upload or a change of network.
  • Resource use. Battery drain, wake locks and memory, which matter to users and to the stores.

Platform reports and in-app SDKs

Some telemetry comes from the operating system or the store, not from your code.

Android vitals is collected by devices whose users allow it and reported by Google Play. Its core vitals are user-perceived crash rate, user-perceived ANR rate, excessive partial wake locks, memory usage and bitmap memory usage, and they affect an app’s visibility on Google Play. The overall bad behaviour thresholds are 1.09% for user-perceived crash rate and 0.47% for user-perceived ANR rate.

MetricKit on Apple platforms provides on-device diagnostics and power and performance metrics that the system captures. The system delivers metric reports covering the previous 24 hours to the app at most once per day, and diagnostic reports arrive immediately on iOS 15 and later.

Everything else comes from code you add: a crash reporter, a performance SDK, or OpenTelemetry’s Android and Swift libraries. These run inside the app and send data to a backend you choose. OpenTelemetry’s Android agent, for example, buffers telemetry to disk while offline and can redact span attributes before export.

Observability for apps that work offline

For an app built to work without internet, some of the failures worth investigating happen exactly when there is no connection to report them. That shapes the design:

  • Buffer to disk, not memory. Events recorded in a dead zone must survive the app being closed or the phone restarting before they can be uploaded.
  • Record the device time and a session. Events will arrive late and out of order, so each one needs its own timestamp and enough context to put it back in sequence.
  • Keep a local view. Developers and support staff can read events on the device itself, without any upload, when diagnosing a problem in the field.

Offline Protocol’s mesh SDK is one example of that split. Its onEvent handler gives the app every protocol event locally with no key, while hosted telemetry stays off until the app calls enableTelemetry with a key and App ID, and then sends a fixed inventory of fourteen event types with no message content and no identifier fields.

Connecting the phone to the backend

Client telemetry is most useful when it can be joined to what happened on the server. OpenTelemetry recommends propagating trace context through HTTP requests, with traceparent and tracestate headers, so that a slow screen on the phone and the slow database query behind it appear in the same trace.

Frequently asked questions

Is crash reporting the same as observability?

It is one part. Crash reports tell you that something broke and where. Observability also covers hangs, slow screens, failed requests and the sequence of events that led there, so you can investigate problems you did not anticipate.

Sources

Build it with Offline Protocol

The portal analytics page explains what each Mesh SDK number counts once an app uploads telemetry, from first-try sends and deliveries to latency percentiles and average hops, and when a day's figures become final.

Read portal analytics and usage