An occurrence and its record
The CloudEvents specification separates two ideas. An occurrence is the capture of a statement of fact during the operation of a software system: a signal raised or observed, a state change, a timer elapsing. Its own example is a device going into an alert state because its battery is low. An event is the data record that expresses that occurrence and its context, so it can be sent to whoever needs to know.
OpenTelemetry’s conventions describe events as named occurrences at a meaningful point in time. They suggest using events for user interactions, state transitions, feature flag evaluations, lifecycle moments such as startup or shutdown, and exceptions raised during an operation. Product analytics uses the same idea: Google Analytics for Firebase describes events as insight into what is happening in an app, such as user actions, system events or errors, and lets developers log their own events with parameters.
The parts of an event
OpenTelemetry represents an event as a log record whose EventName field is set. The fields that matter for an event are these:
| Field | What it holds |
|---|---|
EventName | The class or type of event, which should uniquely identify its structure |
Timestamp | When the event occurred, by the clock at the source |
ObservedTimestamp | When the collection system observed it |
Attributes | Structured details about this occurrence |
Resource | The entity that produced it, such as the app or the device model |
SeverityNumber | How serious it is, from trace and debug up to error and fatal |
TraceId, SpanId | Links to the request or operation it happened inside, if any |
CloudEvents covers the same ground for events moving between systems. Every CloudEvent must carry an id, a source, a specversion and a type. Producers must make source plus id unique for each distinct event, and a consumer may treat two events with the same source and id as duplicates, which is what lets a resent event be recognised.
Events, logs, metrics and spans
The boundaries are practical rather than strict. OpenTelemetry’s guidance is:
- Use a span for an operation that has a duration and a clear start and end, such as an upload.
- Use a span attribute for a property of the whole operation, known when it starts.
- Use an event for a distinct occurrence that needs its own timestamp, may happen several times, or happens outside any operation.
- Use a plain log record for unstructured diagnostic messages nobody will query by name.
Metrics are different again: OpenTelemetry defines them as aggregations of numeric data over a period, such as an error rate. A device can turn a stream of events into a metric by counting them locally, which is one way to send less, alongside sampling telemetry.
Designing events for devices that go offline
Events from phones and field devices can arrive late, in batches, and sometimes twice. A few habits keep them usable:
- Record when it happened, not only when it arrived. Set the occurrence time on the device. OpenTelemetry’s event conventions require
Timestampto be the time of the occurrence and leaveObservedTimestampto whatever component received it. The gap between the two shows how long the event waited, but device clocks drift, so treat the device time with care. - Give each event a unique identifier. A device that retries an upload after a dropped connection will send some events again. A stable identifier lets the backend discard the copy, the same principle as idempotency for any retried request.
- Keep names fixed and values in attributes. OpenTelemetry says event names must not include dynamic values.
upload.failedwith an attribute for the error type can be counted and grouped; a name that embeds a file path or a user cannot. - Write down the schema. Each event name should map to one documented set of attributes, so that a dashboard built today still reads events sent by an old app version next year.
- Leave out what you do not need. OpenTelemetry asks convention authors to document attributes that may contain sensitive information. Counting failures or measuring latency does not need message bodies, free text or stable personal identifiers. See what data should never leave the device.
A fixed inventory of events
A small, published list of event types is easier to audit, document and disclose than open-ended logging. Offline Protocol’s mesh SDK takes this approach for its optional telemetry: once the app calls enableTelemetry with a key and App ID, it uploads a fixed inventory of fourteen event types, such as deliveries, relaying, routing decisions, transport changes and MLS session health, with no message content and no identifier fields.