Limited Offer

30% OFF Lifetime Access ($139) with code SYSTEM30

TOPIC #28Intermediate 9 min read

Load Balancing Algorithms (Round Robin, Least Connections, Hash, Weighted)

💡
Core Architecture Summary

Deep dive into traffic distribution algorithms: Static vs Dynamic scheduling, Consistent Hashing, Maglev, Weighted Least Connections, and Peak EWMA.

Key Glossary Concepts in this TopicAll Glossary Terms

Load Balancing Traffic Simulator ⚖️

Simulate real-time traffic routing across backend servers with dynamic health checks.

Algorithm:

App Pod Alpha

Active Sockets:4 conns
Total Requests:0 (0%)
Hardware Weight:1x
CPU Core Utilization:25%

App Pod Beta

Active Sockets:12 conns
Total Requests:0 (0%)
Hardware Weight:2x
CPU Core Utilization:60%

App Pod Gamma

Active Sockets:2 conns
Total Requests:0 (0%)
Hardware Weight:3x
CPU Core Utilization:15%
Total Requests Processed: 0

💡 Try toggling one server to DOWN and send a 20x burst to observe automated health check failover in action.

Core Load Balancing Scheduling Algorithms ⚖️

Comparative analysis of Round Robin, Least Connections, Consistent Hashing, and Response-Time (EWMA) traffic distribution strategies.

Core Load Balancing Scheduling Algorithms ⚖️
100%
Rendering visual architecture flowchart...

01.1. Static vs Dynamic Scheduling Algorithms

Load balancing algorithms determine how incoming requests or network flows are distributed across a pool of available backend servers. They are categorized into two fundamental classes:

  1. Static Algorithms: Make routing decisions based purely on mathematical sequences or deterministic hashing, without considering the real-time CPU load, memory utilization, or connection count of backend servers.
  2. Dynamic Algorithms: Continuously collect runtime telemetry (active TCP connection counts, response latency histograms, CPU metrics) and adjust routing decisions dynamically in real time.

02.2. Deep Dive: The 5 Essential Load Balancing Algorithms

1. Round Robin & Weighted Round Robin (WRR)

  • Round Robin: Sequentially assigns request 1 to Server A, request 2 to Server B, request 3 to Server C, and wraps around.
  • Weighted Round Robin (WRR): Assigns integer weights based on server hardware capacity (e.g., a 64-core server gets weight 4; a 16-core server gets weight 1). The algorithm sends 4 requests to the 64-core server for every 1 request sent to the 16-core server.
  • Best For: Homogeneous stateless web APIs with uniform, short request processing times.
  • Failure Mode: If some requests take 10 seconds (complex reports) and others take 5ms, Round Robin can pile heavy requests onto the same server, triggering CPU overload.

2. Least Connections & Weighted Least Connections (WLC)

  • The load balancer tracks the exact number of active, open TCP connections on each backend.
  • New requests are dispatched to the server currently handling the fewest active connections.
  • Best For: Long-lived connections (WebSockets, database connection pools, file uploads, large video streaming).

3. IP Hash & Consistent Hashing

  • Computes a cryptographic hash of the client's IP address or session token: Server Index = hash(Client IP) \pmod N.
  • Guarantees that requests from the same user are consistently routed to the exact same backend server (Session Affinity / Sticky Sessions).
  • Best For: In-memory user sessions, stateful local caching, chunked file upload pipelines.

4. Least Response Time & Peak EWMA (Exponentially Weighted Moving Average)

  • Combines the number of active connections with the server's rolling average latency:

Score = Active Connections × EWMA Latency

  • The load balancer routes traffic to the server that produces the lowest score.
  • Best For: Microservices with heterogeneous workloads (used heavily by Twitter/X Finagle and Envoy Proxy).

5. Google Maglev Hashing

  • A consistent hashing algorithm developed by Google that ensures deterministic lookup in O(1) time while minimizing key reshuffling when backends crash or scale out, maintaining consistent packet flows across Line-Rate L4 switches.
nginx— Nginx configuration comparing Least Connections and IP Hash
# Least Connections Algorithm
upstream dynamic_backend {
    least_conn;
    server 10.0.1.11:8080 weight=3;
    server 10.0.1.12:8080 weight=1;
}

# IP Hash for Session Stickiness
upstream sticky_backend {
    ip_hash;
    server 10.0.1.11:8080;
    server 10.0.1.12:8080;
}

03.3. Architectural Tradeoff Matrix: Choosing the Right Algorithm

AlgorithmState Required on LB?Hardware Heterogeneity?Workload PatternKey Vulnerability
Round RobinStateless (O(1) pointer)Poor (Treats all equal)Uniform fast requestsHeavy request clumping
Weighted RRStatelessExcellentUniform fast requestsCannot sense real-time lag
Least ConnectionsStateful (Tracks active sockets)GoodLong-lived / variable requestsSlow start on new nodes
IP HashStatelessModerateSticky user sessionsUneven distribution behind NAT
Peak EWMAStateful (Tracks latency)BestMicroservice RPCs (gRPC)Metric collection overhead

⚖️Architectural Trade-offs & Production Realities

Architectural Advantages

  • Weighted Least Connections prevents server overload during long-running database queries and WebSocket sessions.
  • Consistent Hashing maintains local cache affinity without centralized Redis lookups.
  • Peak EWMA adapts dynamically to tail latency spikes across heterogeneous microservice clusters.

Trade-offs & Constraints

  • IP Hashing fails behind corporate proxy networks where thousands of distinct users share a single public NAT IP address.
  • Dynamic algorithms (Least Connections, EWMA) require state synchronization across distributed load balancer fleets.
Production Implementation in Big Tech
Twitter / X & Google• Peak EWMA in Finagle & Google Maglev

Twitter engineered the Finagle RPC framework using Peak EWMA load balancing, directing requests to backends based on rolling latency moving averages to smooth out garbage collection latency spikes. Google uses Maglev consistent hashing across all Google Cloud data centers to balance billions of queries per second across server clusters with zero connection drops during maintenance.

🎯 Staff+ Engineering Takeaways

  • Round Robin is fast and stateless; ideal for uniform, short-lived stateless web requests.
  • Least Connections routes to the server with the fewest active sockets; ideal for WebSockets and heavy DB queries.
  • IP Hash routes by client IP; vulnerable to hotspots behind corporate NAT proxies.
  • Consistent Hashing and Maglev minimize key redistribution when server clusters scale dynamically.
  • Peak EWMA combines active connection count with real-time response latency.

Topic Knowledge Assessment 🧠

Step through 2 scenario questions to test your staff-level grasp.

Question 1 of 20 answered
#1

Why can the IP Hash load balancing algorithm cause severe server hotspots in production?

Rate This Architecture Chapter4.9 / 5.0 (38 ratings)

How clear and staff-actionable was this system breakdown?