Metrics and percentileseasy0-2 years

A teammate adds `.tag("userId", userId)` to an existing HTTP request timer "so we can see per-user latency." What happens to the metrics platform, and why does this one line matter more than it looks like it should?

Every distinct combination of tag values on a metric is a separate time series, each with its own storage and index entry — so adding a tag whose values come from a large, unbounded set (every user id) multiplies the number of series by the number of distinct users, not by a fixed, small factor the way a status code or an HTTP method does. A metric that was a manageable few thousand series across endpoints, methods and status codes can become billions once it's crossed with 100,000 user ids, and no time-series database stores that. The practical effect isn't a clean error for that one metric — the platform's memory climbs until it degrades for every team using it, caused by one line of code in one service. Tags need to come from a small, fixed set of values; a user id belongs in a log or a trace, where one value per event costs almost nothing extra.

The lesson behind it →
More on Metrics and percentiles