Limited Offer

30% OFF Lifetime Access ($139) with code SYSTEM30

TOPIC #15Beginner 9 min read

Latency Numbers Every Programmer Should Know

💡
Core Architecture Summary

Master the iconic back-of-the-envelope latency benchmarks compiled by Jeff Dean: Scale hardware nanoseconds into human intuitive time scales.

Key Glossary Concepts in this TopicAll Glossary Terms

Latency Numbers Visualizer ⏱️

Scale nanosecond hardware delays into intuitive human time (1 CPU cycle = 1 second).

1 CPU Clock Cycle (3.3 GHz)
0.3 ns1 Second
L1 CPU Cache Reference
0.9 ns3 Seconds
Branch Misprediction Penalty
3.0 ns10 Seconds
L2 CPU Cache Reference
4.5 ns15 Seconds
Mutex Lock / Unlock
17.0 ns1 Minute
Main Memory Reference (DRAM)
100 ns5.5 Minutes
Read 1 MB Sequentially from RAM
3,000 ns (3 μs)2.8 Hours
NVMe SSD Random Read
16,000 ns (16 μs)15 Hours
Read 1 MB Sequentially from NVMe SSD
200,000 ns (200 μs)8 Days
Intra-Datacenter Network RTT
500,000 ns (0.5 ms)19 Days
Mechanical HDD Disk Seek
8,000,000 ns (8 ms)10 Months
Cross-Atlantic Network RTT (US to Europe)
150,000,000 ns (150 ms)16 YEARS
cpu Tier0.3 ns

1 CPU Clock Cycle (3.3 GHz)

The fundamental heartbeat of modern silicon processors executing a single instruction.

Human Time Analogy (1 CPU Cycle = 1 Sec)

1 Second

If fetching from CPU register takes 1 second, this operation feels like waiting 1 Second to the processor!

💡 Architectural Insight:

Reading from RAM is like a 5-minute coffee break; reading from spinning disk is waiting 10 months; cross-ocean network calls are a 16-year career!

Jeff Dean Latency Numbers & Human Scale Analogy ⏱️

Scaling raw hardware time (1 ns = 1 human second) reveals the massive orders-of-magnitude chasms between RAM, SSD, and Network.

Jeff Dean Latency Numbers & Human Scale Analogy ⏱️
100%
Rendering visual architecture flowchart...

01.1. The Canonical Latency Table for Distributed Engineers

Originally popularized by Google Senior Fellow Jeff Dean and Peter Norvig, these foundational latency benchmarks are the mathematical backbone of all system design estimations. Memorizing these orders of magnitude allows engineers to immediately evaluate whether an architectural proposal is physically feasible before writing code:

Hardware OperationActual LatencyScaled to Human Time (1ns = 1s)Key Architectural Implication
L1 Cache Reference0.5 - 1.0 ns1 secondInstantaneous register/cache compute
Branch Mispredict3.0 - 5.0 ns3 - 5 secondsPipeline flush overhead
L2 Cache Reference~4.0 ns4 seconds4x slower than L1
Mutex Lock / Unlock~15 - 25 ns15 - 25 secondsUncontended synchronization cost
Main Memory (DRAM) Reference~100 ns~1.7 minutesIn-memory cache lookup baseline (Redis)
Compress 1KB with Snappy~2,000 ns (2 μs)~33 minutesLightweight wire compression
Read 1 MB Sequentially from Memory~3,000 ns (3 μs)~50 minutesIn-memory linear table scan
NVMe SSD Random Read (4KB)~20 - 50 μs~7 to 17 hoursFast database block read (B-Tree lookup)
Read 1 MB Sequentially from NVMe SSD~250 μs (0.25 ms)~2.9 daysSequential data scanning (LSM compaction)
Intra-Datacenter Network Round-Trip~500 μs (0.5 ms)~5.8 daysInter-service microservice RPC (gRPC)
Mechanical HDD Random Seek~5 - 10 ms~2 to 4 monthsRandom disk I/O (death of database performance)
Read 1 MB Sequentially from HDD~20 ms~8 monthsDisk streaming throughput
Send Packet Across USA (SF to NYC)~40 - 50 ms~1.5 yearsCross-region US replication latency
Send Packet Across Atlantic (NYC to London)~70 - 80 ms~2.5 yearsTransatlantic network physics
Cross-Continental Round-Trip (SF to Tokyo)~150 - 200 ms~5 to 6.5 yearsGlobal active-active synchronization floor

02.2. The Three Cardinal Rules Derived from Latency Physics

Every distributed architectural pattern exists to bypass one of the massive latency cliffs in the hierarchy:

  1. Rule 1: RAM is 1,000x Faster than Storage.

    • Fetching a session from Redis (in-memory DRAM: 100ns) is 500x faster than fetching it from an NVMe SSD (50μs) and 100,000x faster than mechanical disk (10ms).
    • This single physical fact explains why multi-tier caching (Redis, Memcached, Application In-Memory Caches) is universal in modern high-throughput architectures.
  2. Rule 2: Sequential I/O is 10x to 100x Faster than Random I/O.

    • Reading 1MB sequentially from an NVMe SSD takes ~250μs, whereas performing 250 random 4KB reads takes ~5,000μs (20x slower).
    • On mechanical hard drives, sequential reads run at ~150-200 MB/s, while random seeks achieve less than 1-2 MB/s.
    • Architectural consequence: Databases use Write-Ahead Logs (WAL) and Log-Structured Merge (LSM) Trees (e.g., RocksDB, Cassandra, Kafka) to turn all incoming writes into sequential appends.
  3. Rule 3: Network Latency Dominates Everything.

    • Even within the same availability zone, an inter-service RPC (0.5ms) takes 5,000 times longer than reading from local memory (100ns).
    • Across regions (SF to Frankfurt: ~120ms), network propagation speed is bound by the speed of light in fiber optic glass (~200,000 km/s). No amount of CPU optimization can reduce this physical floor.

03.3. Applying Latency Numbers to Back-of-the-Envelope Calculations

During architecture design and interview sessions, apply these numbers to establish system feasibility:

Example: Can we design a system handling 50,000 writes/sec to a single MySQL instance?

  • If each write performs a random synchronous disk I/O without batching:

Throughput Limit = \frac{1}{Disk Latency} = \frac{1}{0.001 s} = 1,000 IOPS

  • Single-instance synchronous random writes will fail catastrophically at 50,000 writes/sec.
  • Required Architectural Solution: Introduce a message queue (Kafka) to buffer writes, use in-memory write-back caching (Redis), or employ a distributed LSM-tree database (Cassandra) that writes sequentially to disk.
typescript— Latency-aware batching vs individual network calls
// ANTI-PATTERN: N+1 Network RPC calls (100 items * 0.5ms RTT = 50ms latency)
for (const id of userIds) {
  const user = await userClient.getUser({ id }); // 100 individual network roundtrips
}

// LATENCY-OPTIMIZED: Single Batch RPC (1 network roundtrip = 0.6ms latency)
const users = await userClient.getUsersBatch({ ids: userIds }); // 80x faster!

⚖️Architectural Trade-offs & Production Realities

Architectural Advantages

  • Allows instantaneous mathematical validation of architectural proposals before writing code.
  • Prevents disastrous over-engineering (e.g., optimizing CPU algorithms when network latency accounts for 99.5% of total request time).
  • Provides objective criteria for choosing caching layers, database engines, and communication protocols.

Trade-offs & Constraints

  • Numbers shift incrementally as hardware evolves (PCIe Gen 5, CXL memory), though the relative orders of magnitude remain constant.
  • Tail latency (p99 / p99.9) under heavy contention often diverges significantly from average latency numbers.
Production Implementation in Big Tech
Google Search• Sub-200ms Global Search Query Execution

Google achieves sub-200ms global search response times by holding index shards in DRAM across thousands of machines, parallelizing search RPCs within the same data center (<0.5ms), and streaming results sequentially to avoid random disk seeks entirely.

🎯 Staff+ Engineering Takeaways

  • L1 Cache (1ns) → RAM (100ns) → NVMe SSD (20μs) → Datacenter Network (0.5ms) → WAN Network (150ms).
  • RAM is ~1,000x faster than NVMe SSD and ~100,000x faster than mechanical HDD.
  • Network latency dominates single-node compute times by orders of magnitude.
  • Sequential I/O is dramatically faster than random I/O on both SSDs and HDDs.
  • Batching network calls and disk writes amortizes latency overhead across thousands of operations.

Topic Knowledge Assessment 🧠

Step through 2 scenario questions to test your staff-level grasp.

Question 1 of 20 answered
#1

Roughly how much slower is fetching data across a transatlantic network (e.g. California to London, ~150ms) compared to reading data from local RAM (~100ns)?

Rate This Architecture Chapter4.9 / 5.0 (38 ratings)

How clear and staff-actionable was this system breakdown?