Limited Offer

30% OFF Lifetime Access ($139) with code SYSTEM30

TOPIC #176Intermediate 9 min read

Metrics Design: The RED and USE Methods

πŸ’‘
Core Architecture Summary

Structure production telemetry: RED Method (Rate, Errors, Duration) for application services, USE Method (Utilization, Saturation, Errors) for infrastructure, Prometheus TSDB mechanics, and preventing high-cardinality crashes.

Key Glossary Concepts in this TopicAll Glossary Terms

RED Method vs USE Method Telemetry Framework

Hierarchical monitoring architecture: Request-driven microservice metrics (RED) paired with resource-driven infrastructure metrics (USE) ingested into Prometheus and Grafana.

RED Method vs USE Method Telemetry Framework
100%
Rendering visual architecture flowchart...

01.1. The RED Method for Request-Driven Services

Introduced by Tom Wilkie, the RED Method defines the three essential metric dimensions that every request-driven microservice and API endpoint must expose:

  1. Rate (Traffic): The number of requests handled per second.
    • Metric Type: Monotonically increasing counter (e.g., http_requests_total).
    • PromQL Calculation: sum(rate(http_requests_total[5m])) by (service, route)
  2. Errors (Failures): The number of requests that fail per second (typically HTTP 5xx responses or unhandled application exceptions).
    • Metric Type: Monotonically increasing counter (e.g., http_requests_total{status=~"5.."}).
    • Error Ratio Calculation: sum(rate(http_requests_total{status=~"5.."}[5m])) / sum(rate(http_requests_total[5m]))
  3. Duration (Latency): The distribution of how long requests take from receipt to completion.
    • Metric Type: Histogram (e.g., http_request_duration_seconds_bucket).
    • Quantile Calculation: Never use averages! Calculate the 90th, 99th, and 99.9th percentiles using:
      promql
      histogram_quantile(0.99, sum(rate(http_request_duration_seconds_bucket[5m])) by (le, service))

The RED method focuses entirely on the customer experience: if Rate is steady, Errors are zero, and Duration p99 is within SLA, the service is functioning correctly from the user's perspective, regardless of internal resource state.

02.2. The USE Method for Infrastructure & Resource Constraints

Formulated by Brendan Gregg (Netflix performance architect), the USE Method applies to any resource with finite physical capacity (CPUs, Memory, Disks, Network Interfaces, Database Connection Pools, Thread Pools):

  1. Utilization: The percentage of time that the resource was busy servicing work over a measured window, or the percentage of capacity currently occupied.
    • Examples: CPU % busy, Memory allocated (GB / Total GB), Disk bandwidth % utilized.
  2. Saturation: The amount of extra work that has queued up because the resource is at 100% utilization and cannot service it immediately.
    • Examples: Linux OS run queue length (Load Average > vCPU count), TCP backlog socket queue drops, JVM thread pool queue depth, HikariCP database connection wait queue.
    • Critical Rule: Saturation is the earliest warning indicator of severe performance degradation before error rates spike.
  3. Errors: The count of error events at the resource level.
    • Examples: Network interface dropped packets (netstat -s drops), disk write retry faults, ECC memory correction events, OS out-of-memory (OOM) killer invocations.

03.3. The High-Cardinality Explosion Disaster in TSDBs

Time-Series Databases (such as Prometheus, VictoriaMetrics, or InfluxDB) store metrics as time-stamped floating-point values indexed by a set of key-value labels:

text
http_requests_total{service="checkout", method="POST", status="200", region="us-east-1"} 41920

The Mathematics of Metric Cardinality

Every unique combination of metric name and label values creates a distinct time series in RAM. If a developer adds dynamic, unbounded labels:

  • service: 10 values
  • status: 5 values
  • user_id: 1,000,000 unique active users
  • order_id: 5,000,000 unique orders

Total Active Time Series = 10 Γ— 5 Γ— 1,000,000 Γ— 5,000,000 = 2.5 Γ— 10^{14} series

The Outcome: TSDB OOM Crash Loop

Prometheus allocates ~ 2 - 4 KB of RAM per active series for index chunking. Ingesting millions of high-cardinality series exhausts host memory, triggering the Linux OOM Killer. On reboot, Prometheus attempts to replay the Write-Ahead Log (WAL), runs out of memory again, and enters an unrecoverable crash loop.

Cardinality Rules:

  • Allowed in Metrics: Low-cardinality enums with < 50 distinct values (e.g., HTTP status code, HTTP method, cloud region, service name).
  • Forbidden in Metrics: High-cardinality identifiers (e.g., user_id, email, ip_address, order_id, un-sanitized URL paths with UUIDs like /users/88219/orders). These belong exclusively in Distributed Traces and Structured Logs.

04.4. Prometheus Pull vs Push Architecture

Monitoring systems differ fundamentally in their telemetry collection model:

1. The Prometheus Pull Model (Scraping via HTTP)

  • The Prometheus server periodically queries a standard /metrics HTTP endpoint exposed by each microservice (every 15–30 seconds).
  • Service Discovery: Prometheus queries the Kubernetes API server to discover active pod IP addresses dynamically.
  • Benefits: Built-in health checking (if a target cannot be scraped, Prometheus instantly detects it as UP = 0); central control over scraping frequency; decoupled from microservice uptime.

2. The Push Model (OpenTelemetry / StatsD)

  • The application actively pushes metric UDP/gRPC packets to a collector daemon or external gateway.
  • When Push is Required (Ephemeral Batch Jobs): Short-lived cron jobs or AWS Lambda functions run for only 200 millisecondsβ€”too fast for a Prometheus scrape interval. They push metrics to a Prometheus Pushgateway or OpenTelemetry Collector before exiting.

05.5. Long-Term TSDB Storage & Downsampling (Thanos & VictoriaMetrics)

A single Prometheus server is designed for local, short-term storage (typically 15 to 30 days) on high-speed SSDs. Keeping 1–2 years of metric history for capacity planning and year-over-year SLO analysis requires a distributed long-term engine:

  1. Thanos / Cortex Architecture:
    • Thanos Sidecar: Uploads immutable 2-hour Prometheus data blocks directly to cheap Object Storage (AWS S3 / GCS).
    • Thanos Compact: Periodically merges historical blocks in S3 and applies mathematical downsampling:
      • Raw data (15-second resolution): Retained for 30 days.
      • 5-minute resolution averages: Retained for 6 months.
      • 1-hour resolution averages: Retained for 2+ years.
    • Thanos Querier: Executes PromQL queries across both local real-time Prometheus instances and long-term S3 data with global deduplication across replicated clusters.
  2. VictoriaMetrics: A highly optimized single-binary or clustered TSDB offering 10Γ— higher compression efficiency and lower RAM usage than standard Prometheus.

βš–οΈArchitectural Trade-offs & Production Realities

Architectural Advantages

  • The RED method provides standard, intuitive Golden Signal dashboards for all microservices
  • The USE method pinpoints physical hardware and queue saturation bottlenecks before user outages occur
  • Prometheus pull-based architecture provides automatic target liveness detection and rate control

Trade-offs & Constraints

  • Accidental inclusion of high-cardinality labels (like user IDs) can instantly crash TSDB memory
  • Histograms require careful bucket threshold selection (e.g. 10ms, 50ms, 250ms, 1s) to calculate accurate p99 quantiles
Production Implementation in Big Tech
Grafana Labs & Netflixβ€’ Standardized Golden Signals & Resource Telemetry

Grafana Labs and Netflix implement standard dashboard templates across their global fleets: every microservice automatically renders a RED overview row (QPS, 5xx rate, p99 latency), followed by a USE drill-down row displaying CPU saturation, thread pool queuing, and memory headroom.

🎯 Staff+ Engineering Takeaways

  • Use the RED method (Rate, Errors, Duration) to monitor customer-facing microservice health.
  • Use the USE method (Utilization, Saturation, Errors) to monitor hardware and queue constraints.
  • Always track latency using percentile histograms (p95, p99, p99.9) rather than misleading arithmetic means.
  • Protect TSDBs by strictly prohibiting unbounded high-cardinality strings in Prometheus metric labels.

Topic Knowledge Assessment 🧠

Step through 3 scenario questions to test your staff-level grasp.

Question 1 of 30 answered
#1

Why is tracking the arithmetic mean (average) of request latency considered an SRE anti-pattern compared to monitoring p99 percentiles?

Rate This Architecture Chapter4.9 / 5.0 (38 ratings)

How clear and staff-actionable was this system breakdown?