Limited Offer

30% OFF Lifetime Access ($139) with code SYSTEM30

TOPIC #190Intermediate 10 min read

Back-of-the-Envelope Estimation Techniques

πŸ’‘
Core Architecture Summary

Master rapid capacity sizing and mental arithmetic: Powers of two vs powers of ten, the 100,000 seconds/day rule, Jeff Dean's latency numbers, and converting business scale into CPU, RAM, disk, and bandwidth requirements.

Key Glossary Concepts in this TopicAll Glossary Terms

Back-of-the-Envelope Estimation Pipeline & Latency Hierarchy πŸ“

Systematic conversion from business DAU metrics to hardware capacity planning constraints alongside critical latency baselines.

Back-of-the-Envelope Estimation Pipeline & Latency Hierarchy πŸ“
100%
Rendering visual architecture flowchart...

01.1. The Philosophy of System Design Estimations

Back-of-the-envelope estimations are not intended to yield exact accounting precision down to decimal places. Rather, their purpose is to determine the order of magnitude (O(10^x)) of a system's physical constraints:

  • Are we designing for 100 QPS (single node) or 100,000 QPS (distributed cluster)?
  • Is 5-year storage 50 GB (single PostgreSQL database) or 50 PB (sharded S3 data lake)?
  • Is egress bandwidth 10 Mbps (commodity network card) or 100 Gbps (dedicated CDN edge routing)?

By rapidly bounding the problem space, you immediately discard non-viable architectural designs (e.g., attempting in-memory joins on 50 TB of data) and justify distributed components (caching, sharding, message queues) before drawing boxes.

02.2. Core Mental Math Constants & Approximations

The 100,000 Seconds per Day Rule

A day has 24 Γ— 60 Γ— 60 = 86,400 seconds. In mental calculations, round this up to 100,000 (10^5) seconds:

Average QPS = \frac{Total Daily Requests}{100,000}

Peak QPS = Average QPS Γ— 2 (Standard Enterprise) \quad or \quad Average QPS Γ— 3 to 5 (Flash Sales / Consumer)

Instant QPS Benchmarks:

  • 1 Million requests/day = \frac{10^6}{10^5} = 10 QPS (Peak β‰ˆ 20-30 QPS)
  • 100 Million requests/day = \frac{10^8}{10^5} = 1,000 QPS (Peak β‰ˆ 2,000-3,000 QPS)
  • 1 Billion requests/day = \frac{10^9}{10^5} = 10,000 QPS (Peak β‰ˆ 20,000-30,000 QPS)

Powers of 2 vs Powers of 10 Approximation

Power of 2Exact ValuePower of 10 ApproxMetric Prefix
2^{10}1,02410Β³ (Thousand)1 Kilobyte (KB)
2^{20}1,048,57610^6 (Million)1 Megabyte (MB)
2^{30}1,073,741,82410^9 (Billion)1 Gigabyte (GB)
2^{40}1,099,511,627,77610^{12} (Trillion)1 Terabyte (TB)
2^{50}1,125,899,906,842,62410^{15} (Quadrillion)1 Petabyte (PB)

03.3. Latency Hierarchy: Jeff Dean's Numbers Every Engineer Must Know

Understanding the physics of latency across the computer storage and network hierarchy dictates why caching, indexing, and locality exist:

OperationLatencyScaled to Human Perspective (1ns = 1 sec)
L1 CPU Cache Reference0.5 - 1 ns1 second
Branch Mispredict3 ns3 seconds
L2 CPU Cache Reference3 - 4 ns4 seconds
Mutex Lock / Unlock17 ns17 seconds
Main Memory Reference (DRAM)100 ns1.5 minutes
Compress 1KB with Snappy2,000 ns = 2 Β΅s33 minutes
Read 1 MB sequentially from Memory3 Β΅s50 minutes
Read 4 KB random from NVMe SSD20 - 50 Β΅s8 hours
Read 1 MB sequentially from SSD200 Β΅s2.3 days
Datacenter LAN Roundtrip500 Β΅s = 0.5 ms5.8 days
Send 1 MB over 10 Gbps network800 Β΅s = 0.8 ms9.3 days
Read 1 MB sequentially from HDD (Disk)2,000 Β΅s = 2 ms23 days
Disk Seek (Spinning Spindle)10 ms4 months
Cross-Continent Roundtrip (NYC to London)70 - 80 ms2.5 years
Global Trans-Pacific (NYC to Tokyo)160 - 200 ms6 years

Architectural Insights from Latency Physics:

  • Memory vs SSD: DRAM is 200-500Γ— faster than NVMe SSDs for random reads.
  • Disk Seeks: Random disk seeks on spinning magnetic platters are 100,000Γ— slower than RAM, explaining why LSM-Trees use sequential log writes.
  • Speed of Light: Network roundtrips across continents (100+ ms) dwarf internal server processing times (< 5 ms), mandating regional edge PoPs and CDNs.

04.4. Critical Safety Margins & Sizing Overheads

When estimating hardware infrastructure from raw data math, real-world systems incur structural overheads that must be factored in:

  1. Replication Factor Multiplier (3Γ—): Enterprise distributed stores (HDFS, Kafka, Cassandra, Ceph) maintain 3 replicas across different availability zones or failure domains. Always multiply persistent raw storage by 3Γ—.
  2. Database Index & Metadata Overhead (+30\% - 50\%): B-Tree indexes, primary key lookups, and table metadata (PostgreSQL MVCC vacuum headers) add 30\% - 50\% storage on top of raw payload bytes.
  3. Network Bit/Byte Conversion Trap: Network capacity is measured in Bits per second (bps, Gbps), while storage/memory is measured in Bytes per second (B/s, GB/s). Always multiply Bytes by 8 to obtain network line rates:

Bandwidth (Gbps) = Throughput (GB/s) Γ— 8

  1. Headroom Buffer: Production systems should run at ≀ 65-70\% CPU and disk capacity to absorb sudden traffic surges, garbage collection pauses, and failover re-balancing.

βš–οΈArchitectural Trade-offs & Production Realities

Architectural Advantages

  • Allows architects and candidates to ground complex architectural designs in hard physics within 3 minutes
  • Quickly eliminates unscalable approaches (e.g. attempting to scan 100TB tables synchronously on HTTP requests)

Trade-offs & Constraints

  • Over-focusing on exact arithmetic during interviews wastes precious time needed for high-level distributed systems trade-offs
  • Assuming average load instead of peak load leads to severe under-provisioning during live flash events
Production Implementation in Big Tech
Google / Metaβ€’ Production Cluster Allocation Sizing

Google and Meta engineering teams require formal Back-of-the-Envelope Capacity Estimations in every Design Document (Design Doc) before provisioning compute allocations, datacenter power quotas, and cross-region network backbone bandwidth for new global services.

🎯 Staff+ Engineering Takeaways

  • Use 100,000 seconds per day for instant, clean mental division ($1\text{M}/\text{day} = 10\text{ QPS}$).
  • RAM is 500x faster than SSD and 50,000x faster than network roundtrips.
  • Multiply persistent storage by 3x for cross-AZ replication and add 30-50% for indexing overhead.
  • Always distinguish between Bytes ($B$) for storage and Bits ($b$) for network line bandwidth ($1\text{ MB/s} = 8\text{ Mbps}$).

Topic Knowledge Assessment 🧠

Step through 3 scenario questions to test your staff-level grasp.

Question 1 of 30 answered
#1

A photo-sharing service processes 300 million API read requests per day. Using standard back-of-the-envelope approximations, what is the estimated Average QPS and recommended Peak QPS capacity?

Rate This Architecture Chapter4.9 / 5.0 (38 ratings)

How clear and staff-actionable was this system breakdown?