Load Balancing Algorithms (Round Robin, Least Connections, Hash, Weighted)
Deep dive into traffic distribution algorithms: Static vs Dynamic scheduling, Consistent Hashing, Maglev, Weighted Least Connections, and Peak EWMA.
Load Balancing Traffic Simulator ⚖️
Simulate real-time traffic routing across backend servers with dynamic health checks.
App Pod Alpha
App Pod Beta
App Pod Gamma
💡 Try toggling one server to DOWN and send a 20x burst to observe automated health check failover in action.
Core Load Balancing Scheduling Algorithms ⚖️
Comparative analysis of Round Robin, Least Connections, Consistent Hashing, and Response-Time (EWMA) traffic distribution strategies.
01.1. Static vs Dynamic Scheduling Algorithms
Load balancing algorithms determine how incoming requests or network flows are distributed across a pool of available backend servers. They are categorized into two fundamental classes:
- Static Algorithms: Make routing decisions based purely on mathematical sequences or deterministic hashing, without considering the real-time CPU load, memory utilization, or connection count of backend servers.
- Dynamic Algorithms: Continuously collect runtime telemetry (active TCP connection counts, response latency histograms, CPU metrics) and adjust routing decisions dynamically in real time.
02.2. Deep Dive: The 5 Essential Load Balancing Algorithms
1. Round Robin & Weighted Round Robin (WRR)
- Round Robin: Sequentially assigns request 1 to Server A, request 2 to Server B, request 3 to Server C, and wraps around.
- Weighted Round Robin (WRR): Assigns integer weights based on server hardware capacity (e.g., a 64-core server gets weight 4; a 16-core server gets weight 1). The algorithm sends 4 requests to the 64-core server for every 1 request sent to the 16-core server.
- Best For: Homogeneous stateless web APIs with uniform, short request processing times.
- Failure Mode: If some requests take 10 seconds (complex reports) and others take 5ms, Round Robin can pile heavy requests onto the same server, triggering CPU overload.
2. Least Connections & Weighted Least Connections (WLC)
- The load balancer tracks the exact number of active, open TCP connections on each backend.
- New requests are dispatched to the server currently handling the fewest active connections.
- Best For: Long-lived connections (WebSockets, database connection pools, file uploads, large video streaming).
3. IP Hash & Consistent Hashing
- Computes a cryptographic hash of the client's IP address or session token:
Server Index = hash(Client IP) \pmod N. - Guarantees that requests from the same user are consistently routed to the exact same backend server (Session Affinity / Sticky Sessions).
- Best For: In-memory user sessions, stateful local caching, chunked file upload pipelines.
4. Least Response Time & Peak EWMA (Exponentially Weighted Moving Average)
- Combines the number of active connections with the server's rolling average latency:
Score = Active Connections × EWMA Latency
- The load balancer routes traffic to the server that produces the lowest score.
- Best For: Microservices with heterogeneous workloads (used heavily by Twitter/X Finagle and Envoy Proxy).
5. Google Maglev Hashing
- A consistent hashing algorithm developed by Google that ensures deterministic lookup in
O(1)time while minimizing key reshuffling when backends crash or scale out, maintaining consistent packet flows across Line-Rate L4 switches.
# Least Connections Algorithm
upstream dynamic_backend {
least_conn;
server 10.0.1.11:8080 weight=3;
server 10.0.1.12:8080 weight=1;
}
# IP Hash for Session Stickiness
upstream sticky_backend {
ip_hash;
server 10.0.1.11:8080;
server 10.0.1.12:8080;
}03.3. Architectural Tradeoff Matrix: Choosing the Right Algorithm
| Algorithm | State Required on LB? | Hardware Heterogeneity? | Workload Pattern | Key Vulnerability |
|---|---|---|---|---|
| Round Robin | Stateless (O(1) pointer) | Poor (Treats all equal) | Uniform fast requests | Heavy request clumping |
| Weighted RR | Stateless | Excellent | Uniform fast requests | Cannot sense real-time lag |
| Least Connections | Stateful (Tracks active sockets) | Good | Long-lived / variable requests | Slow start on new nodes |
| IP Hash | Stateless | Moderate | Sticky user sessions | Uneven distribution behind NAT |
| Peak EWMA | Stateful (Tracks latency) | Best | Microservice RPCs (gRPC) | Metric collection overhead |
⚖️Architectural Trade-offs & Production Realities
Architectural Advantages
- Weighted Least Connections prevents server overload during long-running database queries and WebSocket sessions.
- Consistent Hashing maintains local cache affinity without centralized Redis lookups.
- Peak EWMA adapts dynamically to tail latency spikes across heterogeneous microservice clusters.
Trade-offs & Constraints
- IP Hashing fails behind corporate proxy networks where thousands of distinct users share a single public NAT IP address.
- Dynamic algorithms (Least Connections, EWMA) require state synchronization across distributed load balancer fleets.
Twitter engineered the Finagle RPC framework using Peak EWMA load balancing, directing requests to backends based on rolling latency moving averages to smooth out garbage collection latency spikes. Google uses Maglev consistent hashing across all Google Cloud data centers to balance billions of queries per second across server clusters with zero connection drops during maintenance.
🎯 Staff+ Engineering Takeaways
- Round Robin is fast and stateless; ideal for uniform, short-lived stateless web requests.
- Least Connections routes to the server with the fewest active sockets; ideal for WebSockets and heavy DB queries.
- IP Hash routes by client IP; vulnerable to hotspots behind corporate NAT proxies.
- Consistent Hashing and Maglev minimize key redistribution when server clusters scale dynamically.
- Peak EWMA combines active connection count with real-time response latency.
Topic Knowledge Assessment 🧠
Step through 2 scenario questions to test your staff-level grasp.
Why can the IP Hash load balancing algorithm cause severe server hotspots in production?
How clear and staff-actionable was this system breakdown?