Engineering
P99 Latency
P99 latency is the response time value below which 99% of requests complete — equivalently, the slowest 1% of requests take longer than this figure. It is a percentile measurement, distinct from and generally far higher than average (mean) latency.
Average latency is a poor proxy for user experience because it's dominated by the bulk of fast, uneventful requests and can look fine even while a meaningful number of real users experience serious slowness. Percentile metrics (p50, p95, p99) are used specifically because they surface that tail behaviour instead of hiding it.
At scale, a "rare" p99 event stops being rare in absolute terms: a service handling a million requests a day still serves 10,000 requests a day at or beyond its p99 latency, which is why teams operating high-traffic systems track and optimise p99 (and sometimes p999) directly rather than relying on averages.
Example
A service with 100ms average latency but 1,200ms p99 latency is fast for the typical request but has a meaningful share of users experiencing a multi-second wait.