Limited Offer

30% OFF Lifetime Access ($139) with code SYSTEM30

TOPIC #58Intermediate 9 min read

Synchronous vs Asynchronous vs Semi-Sync Replication

💡
Core Architecture Summary

Navigate the replication durability spectrum: Zero data loss (RPO=0) with Sync vs Maximum throughput with Async vs Balanced Semi-Synchronous replication.

Key Glossary Concepts in this TopicAll Glossary Terms

Asynchronous vs Synchronous vs Semi-Synchronous Replication Protocols ⏱️

Detailed sequence flow comparing Asynchronous (fast write, risk of data loss), Synchronous (zero data loss, network latency penalty), and Semi-Synchronous (balanced RAM ack).

Asynchronous vs Synchronous vs Semi-Synchronous Replication Protocols ⏱️
100%
Rendering visual architecture flowchart...

01.1. The Core Replication Tradeoff: Durability vs Latency

When configuring database replication, systems engineers must make a fundamental architectural choice balancing Recovery Point Objective (RPO) against Write Latency:

  • Recovery Point Objective (RPO): The maximum acceptable amount of data loss measured in time (e.g., RPO = 0 means zero data loss; RPO = 5s means up to 5 seconds of data can be lost during a disaster).
  • Recovery Time Objective (RTO): The maximum acceptable duration of system downtime required to restore service.

02.2. Deep Dive: The 3 Replication Protocols

1. Asynchronous Replication

  • The Primary writes to its local WAL, commits the transaction, and immediately returns 200 OK to the client.
  • The WAL stream is transmitted to replicas asynchronously in the background.
  • Pros: Maximum write throughput and minimal write latency (~1-2ms). The Primary continues operating smoothly even if all replicas crash or become unreachable.
  • Cons: If the Primary host suffers an unrecoverable hardware failure, un-replicated WAL bytes sitting in flight are permanently lost (RPO > 0).

2. Fully Synchronous Replication

  • The Primary writes the WAL, transmits it to the replica, and blocks, waiting for the replica to flush the WAL to physical disk and reply with an ACK before acknowledging the client.
  • Pros: Zero Data Loss Guarantee (RPO = 0). If the primary explodes, the replica is mathematically guaranteed to have 100% of all committed transactions.
  • Cons: Write latency increases by the full network Round-Trip Time (RTT) to the replica. Availability Risk: If a single synchronous replica freezes or experiences a network hiccup, all database writes on the Primary are completely blocked!

3. Semi-Synchronous Replication (The Cloud Production Standard)

  • Used heavily by MySQL (rpl_semi_sync), PostgreSQL (synchronous_commit = remote_write), and AWS Aurora.
  • The Primary commits the write as soon as at least ONE replica acknowledges that it received the WAL record in memory (Relay Log), without waiting for the replica to execute the slower disk fsync().
  • Configurable Fallback: If the replica does not ACK within a timeout (e.g., 500ms), the Primary temporarily degrades to Asynchronous mode to protect system availability.

03.3. Multi-AZ Cloud Architectures: Quorum Replication

Modern cloud distributed databases (AWS Aurora, Google Spanner, CockroachDB) use Quorum-Based Replication:

  • Writes are replicated to N Availability Zones (e.g., N = 6 storage nodes across 3 AZs).
  • The write is acknowledged as soon as a Write Quorum (W = 4 of 6 nodes) confirms storage.
  • A single lagging or crashed node in one data center does not delay write latencies or block transactions, combining RPO=0 durability with high availability.
sql— Configuring synchronous standby replication in PostgreSQL
-- PostgreSQL postgresql.conf
# Require at least 1 replica out of 3 to ACK WAL receipt synchronously:
synchronous_commit = on
synchronous_standby_names = 'FIRST 1 (replica_node_1, replica_node_2, replica_node_3)'

⚖️Architectural Trade-offs & Production Realities

Architectural Advantages

  • Asynchronous replication delivers the highest write throughput and insulates primary from replica slowdowns.
  • Synchronous replication guarantees complete financial data durability (RPO=0).
  • Semi-synchronous replication provides a pragmatic balance between low write latency and crash protection.

Trade-offs & Constraints

  • Synchronous replication adds network round-trip latency to every transaction commit.
  • A network partition between primary and synchronous replicas can halt all write operations across the entire application.
Production Implementation in Big Tech
AWS Aurora & GitHub• Multi-AZ Quorum Storage & Semi-Sync Replication

AWS Aurora writes each database log change across 6 storage nodes distributed across 3 Availability Zones, requiring a 4-of-6 quorum to commit. GitHub uses MySQL semi-synchronous replication with automated failover orchestrators to ensure zero data loss during cloud zone outages.

🎯 Staff+ Engineering Takeaways

  • Async: Fast (~2ms), Primary never blocks, risk of data loss on crash (RPO > 0).
  • Sync: Zero data loss (RPO = 0), but adds network RTT latency and blocks if replica dies.
  • Semi-Sync: Primary waits for replica in-memory RAM acknowledgment (fast + safe).
  • Quorum Replication (Aurora/Spanner) commits when a majority of nodes acknowledge.

Topic Knowledge Assessment 🧠

Step through 2 scenario questions to test your staff-level grasp.

Question 1 of 20 answered
#1

What is the primary operational risk of using fully Synchronous Replication with a single Read Replica in a distributed database cluster?

Rate This Architecture Chapter4.9 / 5.0 (38 ratings)

How clear and staff-actionable was this system breakdown?