Limited Offer

30% OFF Lifetime Access ($139) with code SYSTEM30

TOPIC #26Beginner 9 min read

Load Balancers: What, Why, & Single Point of Failure

💡
Core Architecture Summary

Learn the foundational role of load balancers: Eliminating single points of failure, horizontal traffic distribution, health checking, and active-passive failover.

Key Glossary Concepts in this TopicAll Glossary Terms

High-Availability Load Balancer Pair (Active-Passive VRRP) ⚖️

Preventing the load balancer from becoming a Single Point of Failure (SPOF) using Virtual Router Redundancy Protocol (VRRP) and automated backend health checking.

High-Availability Load Balancer Pair (Active-Passive VRRP) ⚖️
100%
Rendering visual architecture flowchart...

01.1. What is a Load Balancer and Why is it Essential?

A Load Balancer (LB) sits between incoming client traffic (mobile apps, web browsers, external API consumers) and backend server fleets. It acts as an authoritative traffic dispatcher, providing four core architectural guarantees:

  1. Horizontal Scalability: Add 10, 50, or 500 application server instances dynamically during flash sales or traffic spikes without modifying client DNS records or API configurations.
  2. High Availability & Fault Tolerance: Continuously monitors the operational health of downstream servers. If a server crashes or encounters memory exhaustion, the load balancer automatically detaches it within milliseconds and redirects traffic to surviving nodes.
  3. Zero-Downtime Deployments: Enables seamless Blue-Green Deployments and Rolling Restarts by bleeding off active connections (connection draining) from old server versions before cutting over to new container images.
  4. Security & Cryptographic Offloading: Centralizes SSL/TLS certificate management and hardware cryptographic termination, shielding backend compute nodes from expensive TLS handshakes and basic volumetric DDoS attacks.

02.2. Solving the Single Point of Failure (SPOF)

If a single load balancer instance is placed in front of 50 backend servers, the load balancer itself becomes a catastrophic Single Point of Failure (SPOF). If that single node crashes or experiences network link failure, the entire system goes offline.

In production engineering, this is resolved using two primary architectures:

1. Active-Passive Failover with Virtual IP (VRRP / Keepalived)

  • Two identical load balancer instances share a single floating Virtual IP (VIP) (e.g., 198.51.100.1).
  • The Active Load Balancer owns the VIP and processes all incoming packets.
  • The Standby Load Balancer continuously listens for heartbeat multicast packets from the Active node via VRRP (Virtual Router Redundancy Protocol) or CARP.
  • If the Standby node misses 3 consecutive heartbeats (e.g., 300ms), it immediately claims the VIP using a Gratuitous ARP (GARP) broadcast, taking over traffic in under 1 second without dropping client TCP sessions.

2. BGP Anycast Equal-Cost Multi-Path (ECMP)

In hyperscale cloud environments (AWS ALB, Cloudflare, Google Cloud Load Balancing), multiple physical load balancer appliances in different racks advertise the exact same public IP address via Border Gateway Protocol (BGP). Upstream routers use Equal-Cost Multi-Path (ECMP) hashing to distribute incoming network flows across dozens of active load balancers simultaneously (Active-Active at the network layer).

03.3. Health Checking Mechanics & Connection Draining

Load balancers maintain an up-to-date registry of healthy backends through periodic health checks:

  • TCP Health Checks (L4): The load balancer attempts to establish a 3-way TCP handshake (SYN, SYN-ACK, ACK) on the target port. Fast and low CPU overhead, but cannot detect if an application process has hung in an internal deadlock.
  • HTTP/gRPC Health Checks (L7): The load balancer sends an HTTP request (GET /healthz) every 5 seconds. The backend must return an HTTP 200 OK within a timeout (e.g., 2 seconds). If 3 consecutive checks fail, the instance is marked unhealthy.

Connection Draining (Deregistration Delay): When removing a server for maintenance or autoscaling down:

  1. The load balancer stops sending new incoming requests to the target server.
  2. It allows in-flight active TCP connections and long-running requests to complete gracefully within a configured draining window (e.g., 30–60 seconds).
  3. Only after all active connections drop to zero is the server terminated, achieving true zero-downtime updates.
nginx— HAProxy active health checking and graceful connection draining configuration
backend app_cluster
    balance roundrobin
    # Enable HTTP health checks every 2s, fail after 3 missed checks
    option httpchk GET /healthz
    http-check expect status 200
    
    server app1 10.0.1.11:8080 check inter 2000ms rise 2 fall 3 weight 100
    server app2 10.0.1.12:8080 check inter 2000ms rise 2 fall 3 weight 100
    server app3 10.0.1.13:8080 check inter 2000ms rise 2 fall 3 weight 100

⚖️Architectural Trade-offs & Production Realities

Architectural Advantages

  • Enables linear horizontal scaling and seamless zero-downtime rolling deployments.
  • Automated health checks eliminate manual operational interventions during hardware or container crashes.
  • Centralizes SSL/TLS certificate termination, rate limiting, and DDoS scrubbing.

Trade-offs & Constraints

  • Adds a network hop (~0.2 - 1.5ms latency depending on L4 vs L7 inspection).
  • Requires careful Active-Passive or Anycast setup to prevent the load balancer itself from becoming an SPOF.
Production Implementation in Big Tech
GitHub• GitHub Load Balancer (GLB Director)

GitHub open-sourced GLB Director, an Anycast L4 load balancer utilizing Linux eBPF/XDP. GLB distributes tens of gigabits of traffic across backend clusters without state synchronization, enabling zero-downtime rolling server maintenance and instant failover without dropping active TCP connections.

🎯 Staff+ Engineering Takeaways

  • Load balancers distribute traffic and execute automated active health checks.
  • Eliminate LB SPOF using Active-Passive VRRP failover pairs or BGP Anycast routing.
  • Connection Draining allows in-flight requests to complete before server termination.
  • L7 health checks verify true application layer readiness (`/healthz`).

Topic Knowledge Assessment 🧠

Step through 2 scenario questions to test your staff-level grasp.

Question 1 of 20 answered
#1

How do production environments prevent a software load balancer (like HAProxy or Nginx) from becoming a Single Point of Failure (SPOF)?

Rate This Architecture Chapter4.9 / 5.0 (38 ratings)

How clear and staff-actionable was this system breakdown?