Ch. 19 · System Design

System Design: Time-Series Metrics

Ingest metrics with bounded cardinality, roll up raw samples, and set retention tiers so storage stays predictable.

~2 min readadvancedupdated Oct 5, 2026

A metrics system ingests samples such as request counts and latencies, then answers queries over time. Its cost is dominated by cardinality — the number of unique label combinations — and by how long raw samples are kept, so the design is mostly about bounding both.

Before you start

You should understand counters, gauges and basic aggregation. This article is conceptual.

Step-by-step walkthrough

Step 1: Bound label cardinality

A time series is identified by its metric name plus labels. Labels with unbounded values, such as a raw user id or a full URL, multiply series and memory until ingestion falls over. Keep labels to bounded sets such as route, status and region, and move high-cardinality detail to logs.

Step 2: Roll up raw samples

Keep high-resolution raw samples for a short window and pre-aggregate into coarser buckets for longer windows. Dashboards querying a week of data read the rollup, not millions of raw points, which keeps query time and storage bounded.

Step 3: Set retention tiers

Define retention by resolution: raw for hours, minute rollups for weeks, hourly rollups for a year. Tiering matches value to cost and prevents unbounded growth. Align to a consistent timestamps grid so rollups are cheap to compute.

Worked scenario

The rollup turns raw per-second samples into per-minute averages.

raw:    req_latency{route=/api,status=200} sample every 10s   -> keep 6h
rollup: req_latency{route=/api} avg/min                       -> keep 30d
rollup: req_latency{route=/api} avg/hour                      -> keep 1y
Text

Walk through the example

The raw tier answers recent debugging questions at full resolution. Older queries read the minute or hour rollup, which has far fewer points. By dropping the high-cardinality status label in the rollup, storage shrinks while the common queries still work.

Common mistake

Using a request id or user id as a label, which creates unbounded series, and keeping raw high-resolution data forever because it is convenient. Another is querying raw data for long ranges, which is slow and expensive.

Verify the behavior

Measure series count with and without a high-cardinality label and observe the difference. Query a long range against raw versus rollup and compare latency. Confirm retention drops data by resolution as configured.

Interview exercise

Why do high-cardinality labels break a metrics system?

Answer and reasoning

Each unique label combination is a separate time series with its own storage and index entry, so cardinality multiplies memory and ingestion work. An unbounded label such as a user id creates a new series per user, which explodes the count and can exhaust the system. Metrics need bounded label sets; per-entity detail belongs in logs or traces.

Continue learning

Compare observability design in Observability design and Metric cardinality. Read the Prometheus best practices on labels and try the System design interview questions.

More in System Design

read ✓System Design · hard

System Design: Bloom Filters

Use a Bloom filter to skip lookups with a tiny memory footprint, and understand its false-positive-only guarantee.

~2 min readread →
esc