Latency Numbers Every Programmer Should Know
Master the iconic back-of-the-envelope latency benchmarks compiled by Jeff Dean: Scale hardware nanoseconds into human intuitive time scales.
Latency Numbers Visualizer ⏱️
Scale nanosecond hardware delays into intuitive human time (1 CPU cycle = 1 second).
1 CPU Clock Cycle (3.3 GHz)
The fundamental heartbeat of modern silicon processors executing a single instruction.
1 Second
If fetching from CPU register takes 1 second, this operation feels like waiting 1 Second to the processor!
💡 Architectural Insight:
Reading from RAM is like a 5-minute coffee break; reading from spinning disk is waiting 10 months; cross-ocean network calls are a 16-year career!
Jeff Dean Latency Numbers & Human Scale Analogy ⏱️
Scaling raw hardware time (1 ns = 1 human second) reveals the massive orders-of-magnitude chasms between RAM, SSD, and Network.
01.1. The Canonical Latency Table for Distributed Engineers
Originally popularized by Google Senior Fellow Jeff Dean and Peter Norvig, these foundational latency benchmarks are the mathematical backbone of all system design estimations. Memorizing these orders of magnitude allows engineers to immediately evaluate whether an architectural proposal is physically feasible before writing code:
| Hardware Operation | Actual Latency | Scaled to Human Time (1ns = 1s) | Key Architectural Implication |
|---|---|---|---|
| L1 Cache Reference | 0.5 - 1.0 ns | 1 second | Instantaneous register/cache compute |
| Branch Mispredict | 3.0 - 5.0 ns | 3 - 5 seconds | Pipeline flush overhead |
| L2 Cache Reference | ~4.0 ns | 4 seconds | 4x slower than L1 |
| Mutex Lock / Unlock | ~15 - 25 ns | 15 - 25 seconds | Uncontended synchronization cost |
| Main Memory (DRAM) Reference | ~100 ns | ~1.7 minutes | In-memory cache lookup baseline (Redis) |
| Compress 1KB with Snappy | ~2,000 ns (2 μs) | ~33 minutes | Lightweight wire compression |
| Read 1 MB Sequentially from Memory | ~3,000 ns (3 μs) | ~50 minutes | In-memory linear table scan |
| NVMe SSD Random Read (4KB) | ~20 - 50 μs | ~7 to 17 hours | Fast database block read (B-Tree lookup) |
| Read 1 MB Sequentially from NVMe SSD | ~250 μs (0.25 ms) | ~2.9 days | Sequential data scanning (LSM compaction) |
| Intra-Datacenter Network Round-Trip | ~500 μs (0.5 ms) | ~5.8 days | Inter-service microservice RPC (gRPC) |
| Mechanical HDD Random Seek | ~5 - 10 ms | ~2 to 4 months | Random disk I/O (death of database performance) |
| Read 1 MB Sequentially from HDD | ~20 ms | ~8 months | Disk streaming throughput |
| Send Packet Across USA (SF to NYC) | ~40 - 50 ms | ~1.5 years | Cross-region US replication latency |
| Send Packet Across Atlantic (NYC to London) | ~70 - 80 ms | ~2.5 years | Transatlantic network physics |
| Cross-Continental Round-Trip (SF to Tokyo) | ~150 - 200 ms | ~5 to 6.5 years | Global active-active synchronization floor |
02.2. The Three Cardinal Rules Derived from Latency Physics
Every distributed architectural pattern exists to bypass one of the massive latency cliffs in the hierarchy:
-
Rule 1: RAM is 1,000x Faster than Storage.
- Fetching a session from Redis (in-memory DRAM: 100ns) is 500x faster than fetching it from an NVMe SSD (50μs) and 100,000x faster than mechanical disk (10ms).
- This single physical fact explains why multi-tier caching (Redis, Memcached, Application In-Memory Caches) is universal in modern high-throughput architectures.
-
Rule 2: Sequential I/O is 10x to 100x Faster than Random I/O.
- Reading 1MB sequentially from an NVMe SSD takes ~250μs, whereas performing 250 random 4KB reads takes ~5,000μs (20x slower).
- On mechanical hard drives, sequential reads run at ~150-200 MB/s, while random seeks achieve less than 1-2 MB/s.
- Architectural consequence: Databases use Write-Ahead Logs (WAL) and Log-Structured Merge (LSM) Trees (e.g., RocksDB, Cassandra, Kafka) to turn all incoming writes into sequential appends.
-
Rule 3: Network Latency Dominates Everything.
- Even within the same availability zone, an inter-service RPC (0.5ms) takes 5,000 times longer than reading from local memory (100ns).
- Across regions (SF to Frankfurt: ~120ms), network propagation speed is bound by the speed of light in fiber optic glass (~200,000 km/s). No amount of CPU optimization can reduce this physical floor.
03.3. Applying Latency Numbers to Back-of-the-Envelope Calculations
During architecture design and interview sessions, apply these numbers to establish system feasibility:
Example: Can we design a system handling 50,000 writes/sec to a single MySQL instance?
- If each write performs a random synchronous disk I/O without batching:
Throughput Limit = \frac{1}{Disk Latency} = \frac{1}{0.001 s} = 1,000 IOPS
- Single-instance synchronous random writes will fail catastrophically at 50,000 writes/sec.
- Required Architectural Solution: Introduce a message queue (Kafka) to buffer writes, use in-memory write-back caching (Redis), or employ a distributed LSM-tree database (Cassandra) that writes sequentially to disk.
// ANTI-PATTERN: N+1 Network RPC calls (100 items * 0.5ms RTT = 50ms latency)
for (const id of userIds) {
const user = await userClient.getUser({ id }); // 100 individual network roundtrips
}
// LATENCY-OPTIMIZED: Single Batch RPC (1 network roundtrip = 0.6ms latency)
const users = await userClient.getUsersBatch({ ids: userIds }); // 80x faster!⚖️Architectural Trade-offs & Production Realities
Architectural Advantages
- Allows instantaneous mathematical validation of architectural proposals before writing code.
- Prevents disastrous over-engineering (e.g., optimizing CPU algorithms when network latency accounts for 99.5% of total request time).
- Provides objective criteria for choosing caching layers, database engines, and communication protocols.
Trade-offs & Constraints
- Numbers shift incrementally as hardware evolves (PCIe Gen 5, CXL memory), though the relative orders of magnitude remain constant.
- Tail latency (p99 / p99.9) under heavy contention often diverges significantly from average latency numbers.
Google achieves sub-200ms global search response times by holding index shards in DRAM across thousands of machines, parallelizing search RPCs within the same data center (<0.5ms), and streaming results sequentially to avoid random disk seeks entirely.
🎯 Staff+ Engineering Takeaways
- L1 Cache (1ns) → RAM (100ns) → NVMe SSD (20μs) → Datacenter Network (0.5ms) → WAN Network (150ms).
- RAM is ~1,000x faster than NVMe SSD and ~100,000x faster than mechanical HDD.
- Network latency dominates single-node compute times by orders of magnitude.
- Sequential I/O is dramatically faster than random I/O on both SSDs and HDDs.
- Batching network calls and disk writes amortizes latency overhead across thousands of operations.
Topic Knowledge Assessment 🧠
Step through 2 scenario questions to test your staff-level grasp.
Roughly how much slower is fetching data across a transatlantic network (e.g. California to London, ~150ms) compared to reading data from local RAM (~100ns)?
How clear and staff-actionable was this system breakdown?