Two places data can live
NIST defines cloud computing as a model for on-demand network access to a shared pool of configurable computing resources, such as networks, servers, storage, applications and services, that can be provisioned and released quickly with little management effort. Cloud data is data held in that pool: in a provider’s data centres, reached over the network, and managed by someone other than the device that produced it.
Edge data sits at the other end. NIST’s fog computing model describes the edge as the network layer that encompasses end devices and their users, providing local computing capability on a sensor, a meter or another device. Data at the edge is created there and, for at least part of its life, stored and processed there: a reading in a sensor’s buffer, a form saved on a phone, an event log on a vehicle gateway. Between the two, NIST places fog computing, a hierarchical layer of nodes that brings computing, storage and control closer to the devices than a distant data centre.
How they differ
| Edge data | Cloud data | |
|---|---|---|
| Where it lives | On the device or a nearby gateway | In a provider’s data centres |
| Without a connection | Still readable and usable locally | Out of reach until the link returns |
| Time to act on it | No network round trip | At least one round trip |
| Scope | One device or one site | Every device that has reported |
| Capacity and retention | Limited by device storage | Elastic, priced by volume and time |
| What is lost with the device | Anything not yet sent | Nothing held centrally |
Why keep data at the edge
NIST’s model starts from a problem: traditional cloud-based IoT systems are challenged by large scale, heterogeneity and the high latency seen in some cloud ecosystems. Its answer is to move applications, management and analytics into the network, closer to the devices.
Vendor edge runtimes are built on the same reasoning. Microsoft describes Azure IoT Edge as bringing analytics closer to devices for faster insights and offline decision making, with the example of running anomaly detection at the edge to respond to an emergency on a production line as quickly as possible. AWS describes IoT Greengrass as software that lets devices act locally on the data they generate.
The reasons fall into three groups:
- It keeps working offline. A device that decides from its own data does not stop when the uplink does. That is the starting point of an offline-first architecture.
- It responds faster. A local decision does not wait for a round trip.
- It can stay private. Data that never leaves the device cannot leak from a server. The GDPR’s data minimisation principle asks for personal data to be limited to what a purpose needs, which favours processing locally and sending results. See what data should never leave the device.
Why send data to the cloud
The edge sees one device or one site. Questions that span a fleet, such as which firmware version crashes more or how a route performed over a year, need data from every device in one place. Cloud storage also scales with demand and survives the loss, theft or reset of any single device, while edge storage is bounded by whatever the device has spare.
The cloud also suits work that is too heavy for the device: training models, joining with other business data, and keeping records for audit.
Moving data from edge to cloud
The hard engineering sits in the transfer between the two. Data recorded at the edge has to wait for a link, so it is buffered on durable storage, batched, and sent over whatever uplink is available when it appears, a pattern called store-and-forward. The backend then has to accept late, out-of-order and sometimes repeated records, which is why events carry stable identifiers and their own timestamps.
Two questions decide the design:
- Which copy is authoritative? If the device is the source of truth, the cloud holds a replica or a summary. If the backend is, the device holds a working copy that the server may correct.
- What is sent: raw data or results? Sending summaries, counts and alerts keeps the uplink small and exposes less. Sending raw data keeps every option open for later analysis but costs more to move, store and protect.
Deciding per kind of data
A system does not have to put everything on one side. A practical split looks like this:
- Keep at the edge only: secret keys, raw sensor streams that are only needed for local control, and personal data with no purpose off the device.
- Keep at the edge and send a summary: routine readings, health metrics and usage counts.
- Send in full when a link exists: completed work records, transactions, alerts, and anything that must be audited or shared with other sites.
The GDPR’s storage limitation principle applies on both sides: data kept in identifiable form should be held no longer than its purpose needs, on the device or in the cloud.