Metrics Design: The RED and USE Methods
Structure production telemetry: RED Method (Rate, Errors, Duration) for application services, USE Method (Utilization, Saturation, Errors) for infrastructure, Prometheus TSDB mechanics, and preventing high-cardinality crashes.
RED Method vs USE Method Telemetry Framework
Hierarchical monitoring architecture: Request-driven microservice metrics (RED) paired with resource-driven infrastructure metrics (USE) ingested into Prometheus and Grafana.
01.1. The RED Method for Request-Driven Services
Introduced by Tom Wilkie, the RED Method defines the three essential metric dimensions that every request-driven microservice and API endpoint must expose:
- Rate (Traffic): The number of requests handled per second.
- Metric Type: Monotonically increasing counter (e.g.,
http_requests_total). - PromQL Calculation:
sum(rate(http_requests_total[5m])) by (service, route)
- Metric Type: Monotonically increasing counter (e.g.,
- Errors (Failures): The number of requests that fail per second (typically HTTP 5xx responses or unhandled application exceptions).
- Metric Type: Monotonically increasing counter (e.g.,
http_requests_total{status=~"5.."}). - Error Ratio Calculation:
sum(rate(http_requests_total{status=~"5.."}[5m])) / sum(rate(http_requests_total[5m]))
- Metric Type: Monotonically increasing counter (e.g.,
- Duration (Latency): The distribution of how long requests take from receipt to completion.
- Metric Type: Histogram (e.g.,
http_request_duration_seconds_bucket). - Quantile Calculation: Never use averages! Calculate the 90th, 99th, and 99.9th percentiles using:
promql
histogram_quantile(0.99, sum(rate(http_request_duration_seconds_bucket[5m])) by (le, service))
- Metric Type: Histogram (e.g.,
The RED method focuses entirely on the customer experience: if Rate is steady, Errors are zero, and Duration p99 is within SLA, the service is functioning correctly from the user's perspective, regardless of internal resource state.
02.2. The USE Method for Infrastructure & Resource Constraints
Formulated by Brendan Gregg (Netflix performance architect), the USE Method applies to any resource with finite physical capacity (CPUs, Memory, Disks, Network Interfaces, Database Connection Pools, Thread Pools):
- Utilization: The percentage of time that the resource was busy servicing work over a measured window, or the percentage of capacity currently occupied.
- Examples: CPU % busy, Memory allocated (GB / Total GB), Disk bandwidth % utilized.
- Saturation: The amount of extra work that has queued up because the resource is at 100% utilization and cannot service it immediately.
- Examples: Linux OS run queue length (Load Average > vCPU count), TCP backlog socket queue drops, JVM thread pool queue depth, HikariCP database connection wait queue.
- Critical Rule: Saturation is the earliest warning indicator of severe performance degradation before error rates spike.
- Errors: The count of error events at the resource level.
- Examples: Network interface dropped packets (
netstat -sdrops), disk write retry faults, ECC memory correction events, OS out-of-memory (OOM) killer invocations.
- Examples: Network interface dropped packets (
03.3. The High-Cardinality Explosion Disaster in TSDBs
Time-Series Databases (such as Prometheus, VictoriaMetrics, or InfluxDB) store metrics as time-stamped floating-point values indexed by a set of key-value labels:
texthttp_requests_total{service="checkout", method="POST", status="200", region="us-east-1"} 41920
The Mathematics of Metric Cardinality
Every unique combination of metric name and label values creates a distinct time series in RAM. If a developer adds dynamic, unbounded labels:
service: 10 valuesstatus: 5 valuesuser_id:1,000,000unique active usersorder_id:5,000,000unique orders
Total Active Time Series = 10 Γ 5 Γ 1,000,000 Γ 5,000,000 = 2.5 Γ 10^{14} series
The Outcome: TSDB OOM Crash Loop
Prometheus allocates ~ 2 - 4 KB of RAM per active series for index chunking. Ingesting millions of high-cardinality series exhausts host memory, triggering the Linux OOM Killer. On reboot, Prometheus attempts to replay the Write-Ahead Log (WAL), runs out of memory again, and enters an unrecoverable crash loop.
Cardinality Rules:
- Allowed in Metrics: Low-cardinality enums with
< 50distinct values (e.g., HTTP status code, HTTP method, cloud region, service name). - Forbidden in Metrics: High-cardinality identifiers (e.g.,
user_id,email,ip_address,order_id, un-sanitized URL paths with UUIDs like/users/88219/orders). These belong exclusively in Distributed Traces and Structured Logs.
04.4. Prometheus Pull vs Push Architecture
Monitoring systems differ fundamentally in their telemetry collection model:
1. The Prometheus Pull Model (Scraping via HTTP)
- The Prometheus server periodically queries a standard
/metricsHTTP endpoint exposed by each microservice (every 15β30 seconds). - Service Discovery: Prometheus queries the Kubernetes API server to discover active pod IP addresses dynamically.
- Benefits: Built-in health checking (if a target cannot be scraped, Prometheus instantly detects it as
UP = 0); central control over scraping frequency; decoupled from microservice uptime.
2. The Push Model (OpenTelemetry / StatsD)
- The application actively pushes metric UDP/gRPC packets to a collector daemon or external gateway.
- When Push is Required (Ephemeral Batch Jobs): Short-lived cron jobs or AWS Lambda functions run for only 200 millisecondsβtoo fast for a Prometheus scrape interval. They push metrics to a Prometheus Pushgateway or OpenTelemetry Collector before exiting.
05.5. Long-Term TSDB Storage & Downsampling (Thanos & VictoriaMetrics)
A single Prometheus server is designed for local, short-term storage (typically 15 to 30 days) on high-speed SSDs. Keeping 1β2 years of metric history for capacity planning and year-over-year SLO analysis requires a distributed long-term engine:
- Thanos / Cortex Architecture:
- Thanos Sidecar: Uploads immutable 2-hour Prometheus data blocks directly to cheap Object Storage (AWS S3 / GCS).
- Thanos Compact: Periodically merges historical blocks in S3 and applies mathematical downsampling:
- Raw data (15-second resolution): Retained for 30 days.
- 5-minute resolution averages: Retained for 6 months.
- 1-hour resolution averages: Retained for 2+ years.
- Thanos Querier: Executes PromQL queries across both local real-time Prometheus instances and long-term S3 data with global deduplication across replicated clusters.
- VictoriaMetrics: A highly optimized single-binary or clustered TSDB offering
10Γhigher compression efficiency and lower RAM usage than standard Prometheus.
βοΈArchitectural Trade-offs & Production Realities
Architectural Advantages
- The RED method provides standard, intuitive Golden Signal dashboards for all microservices
- The USE method pinpoints physical hardware and queue saturation bottlenecks before user outages occur
- Prometheus pull-based architecture provides automatic target liveness detection and rate control
Trade-offs & Constraints
- Accidental inclusion of high-cardinality labels (like user IDs) can instantly crash TSDB memory
- Histograms require careful bucket threshold selection (e.g. 10ms, 50ms, 250ms, 1s) to calculate accurate p99 quantiles
Grafana Labs and Netflix implement standard dashboard templates across their global fleets: every microservice automatically renders a RED overview row (QPS, 5xx rate, p99 latency), followed by a USE drill-down row displaying CPU saturation, thread pool queuing, and memory headroom.
π― Staff+ Engineering Takeaways
- Use the RED method (Rate, Errors, Duration) to monitor customer-facing microservice health.
- Use the USE method (Utilization, Saturation, Errors) to monitor hardware and queue constraints.
- Always track latency using percentile histograms (p95, p99, p99.9) rather than misleading arithmetic means.
- Protect TSDBs by strictly prohibiting unbounded high-cardinality strings in Prometheus metric labels.
Topic Knowledge Assessment π§
Step through 3 scenario questions to test your staff-level grasp.
Why is tracking the arithmetic mean (average) of request latency considered an SRE anti-pattern compared to monitoring p99 percentiles?
How clear and staff-actionable was this system breakdown?