Reading Latency Percentiles Without Fooling Yourself
By Marcus Hale · April 11, 2026 · Operations
Averages conceal what percentiles reveal: the request mix's tail. A service averaging fifty milliseconds can still be sending one in twenty users through a two-second odyssey, and those users - the ones hitting the cold paths, the expired sessions, the overloaded shard - are the ones filing the tickets.
Percentiles have traps of their own. Comparing a p95 across deploys makes sense only within the same traffic mix; a marketing burst skews the distribution and your 'regression' can be pure arithmetic. And the beloved p99 across a one-minute window is sightreading a single bar of music, not hearing the song.
The durable practice is histograms kept cheaply at the edge, aggregated later. You can always recompute the percentile of the week from a histogram; from an average you can recompute nothing at all.