Autoscaling Triggers: CPU vs Custom Metrics vs Queue Lag
Scale on leading indicators: CPU/Memory lagging metric limits, KEDA event-driven autoscaling, Kafka consumer group lag, and SQS queue backlog depth.
01.1. The Flaw in Lagging Metrics: CPU & Memory Limits
Traditional cloud autoscaling triggers evaluate resource consumption metrics such as CPU Utilization and Memory Usage. While effective for CPU-bound synchronous web workloads, they fail fundamentally for asynchronous worker pools:
Why CPU Scaling Fails for Worker Queues:
- I/O-Bound Bottlenecks: Worker processes (e.g., invoice generators, payment processors, video transcoders) often spend the majority of execution time waiting on external network I/O (relational database queries, third-party Stripe APIs, cloud storage writes).
- The CPU Paradox: A worker fleet of 5 pods processing a sudden spike of
100,000messages might register only20\%CPU utilization. Because CPU remains far below the standard70\%scaling threshold, the cluster fails to auto-scale, causing message backlogs to accumulate and violating customer Service Level Agreements (SLAs). - Memory Scaling Pitfalls: Modern runtimes with garbage collection (Java JVM, Go, Node.js) allocate memory dynamically and do not immediately release heap back to the OS. Memory metrics stay elevated even when traffic drops to zero, causing expensive scaling thrashing.
Lagging Metric (CPU) vs Leading Metric (Queue Backlog) Autoscaling π
Lagging Metric (CPU) vs Leading Metric (Queue Backlog) Autoscaling π
Scaling proactively on queue depth backlog with KEDA vs reactively after worker pools are overwhelmed.
Unlock Topic #207: Autoscaling Triggers: CPU vs Custom Metrics vs Queue Lag
You are viewing a preview. The full in-depth engineering deep dive, interactive simulators, architecture flowcharts, and self-assessment quizzes for this topic are available with Pro or Lifetime Access.
Failure modes, high-throughput bottlenecks, and real FAANG implementation decisions.
Interactive system topology diagrams, live parameter simulators, and downloadable SVG charts.
Staff-level multiple-choice quiz questions with instant feedback and answer explanations.
Firebase Google authentication automatically syncs your completed topics and quiz scores.
How clear and staff-actionable was this system breakdown?