What a time series is
A time series is a sequence of values for one thing, each stamped with the time it was measured: a pump’s pressure every second, a phone’s battery level every minute, the count of requests a server handled. A time-series database is built around that shape. Every record has a timestamp, and queries ask questions over time ranges, such as the average for each hour last week.
Three well-documented systems describe the same idea in their own terms:
- Prometheus stores every series as a stream of timestamped values, identified by a metric name and a set of key-value labels. Each sample is a 64-bit floating-point value (or a native histogram) with a millisecond-precision timestamp. Changing any label value, or adding or removing a label, creates a new series.
- InfluxDB (version 2 documentation) organises a point into a measurement, a set of tags, a set of fields and a timestamp, stored on disk in epoch nanoseconds. Tags are indexed metadata and fields hold the values. The combination of measurement and tag set is the series key.
- TimescaleDB keeps the data in PostgreSQL. A hypertable is an ordinary-looking table that the extension automatically partitions by time into chunks, each holding a set time range, such as one day or one week.
How the data is stored
The storage designs differ, but they share a habit of splitting data by time.
Prometheus groups incoming samples into two-hour blocks on disk, each with its own index. The current block is kept in memory and protected against crashes by a write-ahead log, and older blocks are later compacted into larger ones in the background. Its documentation says local storage averages only 1 to 2 bytes per sample, and notes that samples within a series are compressed together.
TimescaleDB uses its time chunks the same way: a query for a time range only has to visit the chunks that cover it. In InfluxDB, a bucket combines a database with a retention period, the length of time each point is kept.
Splitting by time makes retention cheap. Prometheus removes whole expired blocks rather than deleting individual rows, and both time-based and size-based retention can be set.
Tags, labels and the cost of new series
Because each unique combination of labels or tags is a separate series, the choice of labels matters. InfluxDB indexes tags but not fields, so queries that filter on tags avoid scanning every value, while a filter on a field must scan them all. In Prometheus, a label whose value differs for almost every event, such as a request ID or a device’s free-form message, creates a new series each time. Values like that belong in a field or a log, not a label.
For device data, that suggests putting stable descriptors in tags or labels (device model, firmware version, site) and the readings themselves in fields or sample values.
Late and out-of-order data from devices
A device that was offline uploads its readings when it reconnects, so data arrives late, in bursts, and with timestamps earlier than points the database already holds. Time-series databases handle this in different ways, and it is worth checking before choosing one:
- Prometheus accepts an out-of-order sample only within a configurable out-of-order time window measured back from the newest data it holds. Older samples are rejected as too old. Its separate backfilling tool can create blocks from historical data, with care around the time range it is still writing.
- InfluxDB identifies a point by its measurement, tag set and timestamp. If the same point is written again, it merges the field sets and keeps the value from the latest write for any field that appears in both. A device that re-sends a batch with the original timestamps therefore overwrites its earlier points instead of creating duplicates, which makes retries idempotent as long as the timestamps do not change.
Both behaviours depend on the timestamp the device attached, which raises a second question: whose clock is it? A device’s clock can be wrong, especially after time without a network, so it helps to store both the device’s own timestamp and the time the server received the reading. The clock drift article explains why the two disagree, and keeping event order offline covers ordering events without trusting either clock completely.