Back-of-the-Envelope Estimation Techniques
Master rapid capacity sizing and mental arithmetic: Powers of two vs powers of ten, the 100,000 seconds/day rule, Jeff Dean's latency numbers, and converting business scale into CPU, RAM, disk, and bandwidth requirements.
Back-of-the-Envelope Estimation Pipeline & Latency Hierarchy π
Systematic conversion from business DAU metrics to hardware capacity planning constraints alongside critical latency baselines.
01.1. The Philosophy of System Design Estimations
Back-of-the-envelope estimations are not intended to yield exact accounting precision down to decimal places. Rather, their purpose is to determine the order of magnitude (O(10^x)) of a system's physical constraints:
- Are we designing for
100 QPS(single node) or100,000 QPS(distributed cluster)? - Is 5-year storage
50 GB(single PostgreSQL database) or50 PB(sharded S3 data lake)? - Is egress bandwidth
10 Mbps(commodity network card) or100 Gbps(dedicated CDN edge routing)?
By rapidly bounding the problem space, you immediately discard non-viable architectural designs (e.g., attempting in-memory joins on 50 TB of data) and justify distributed components (caching, sharding, message queues) before drawing boxes.
02.2. Core Mental Math Constants & Approximations
The 100,000 Seconds per Day Rule
A day has 24 Γ 60 Γ 60 = 86,400 seconds. In mental calculations, round this up to 100,000 (10^5) seconds:
Average QPS = \frac{Total Daily Requests}{100,000}
Peak QPS = Average QPS Γ 2 (Standard Enterprise) \quad or \quad Average QPS Γ 3 to 5 (Flash Sales / Consumer)
Instant QPS Benchmarks:
1 Million requests/day = \frac{10^6}{10^5} = 10 QPS(Peakβ 20-30 QPS)100 Million requests/day = \frac{10^8}{10^5} = 1,000 QPS(Peakβ 2,000-3,000 QPS)1 Billion requests/day = \frac{10^9}{10^5} = 10,000 QPS(Peakβ 20,000-30,000 QPS)
Powers of 2 vs Powers of 10 Approximation
| Power of 2 | Exact Value | Power of 10 Approx | Metric Prefix |
|---|---|---|---|
2^{10} | 1,024 | 10Β³ (Thousand) | 1 Kilobyte (KB) |
2^{20} | 1,048,576 | 10^6 (Million) | 1 Megabyte (MB) |
2^{30} | 1,073,741,824 | 10^9 (Billion) | 1 Gigabyte (GB) |
2^{40} | 1,099,511,627,776 | 10^{12} (Trillion) | 1 Terabyte (TB) |
2^{50} | 1,125,899,906,842,624 | 10^{15} (Quadrillion) | 1 Petabyte (PB) |
03.3. Latency Hierarchy: Jeff Dean's Numbers Every Engineer Must Know
Understanding the physics of latency across the computer storage and network hierarchy dictates why caching, indexing, and locality exist:
| Operation | Latency | Scaled to Human Perspective (1ns = 1 sec) |
|---|---|---|
| L1 CPU Cache Reference | 0.5 - 1 ns | 1 second |
| Branch Mispredict | 3 ns | 3 seconds |
| L2 CPU Cache Reference | 3 - 4 ns | 4 seconds |
| Mutex Lock / Unlock | 17 ns | 17 seconds |
| Main Memory Reference (DRAM) | 100 ns | 1.5 minutes |
| Compress 1KB with Snappy | 2,000 ns = 2 Β΅s | 33 minutes |
| Read 1 MB sequentially from Memory | 3 Β΅s | 50 minutes |
| Read 4 KB random from NVMe SSD | 20 - 50 Β΅s | 8 hours |
| Read 1 MB sequentially from SSD | 200 Β΅s | 2.3 days |
| Datacenter LAN Roundtrip | 500 Β΅s = 0.5 ms | 5.8 days |
| Send 1 MB over 10 Gbps network | 800 Β΅s = 0.8 ms | 9.3 days |
| Read 1 MB sequentially from HDD (Disk) | 2,000 Β΅s = 2 ms | 23 days |
| Disk Seek (Spinning Spindle) | 10 ms | 4 months |
| Cross-Continent Roundtrip (NYC to London) | 70 - 80 ms | 2.5 years |
| Global Trans-Pacific (NYC to Tokyo) | 160 - 200 ms | 6 years |
Architectural Insights from Latency Physics:
- Memory vs SSD: DRAM is
200-500Γfaster than NVMe SSDs for random reads. - Disk Seeks: Random disk seeks on spinning magnetic platters are
100,000Γslower than RAM, explaining why LSM-Trees use sequential log writes. - Speed of Light: Network roundtrips across continents (
100+ ms) dwarf internal server processing times (< 5 ms), mandating regional edge PoPs and CDNs.
04.4. Critical Safety Margins & Sizing Overheads
When estimating hardware infrastructure from raw data math, real-world systems incur structural overheads that must be factored in:
- Replication Factor Multiplier (
3Γ): Enterprise distributed stores (HDFS, Kafka, Cassandra, Ceph) maintain 3 replicas across different availability zones or failure domains. Always multiply persistent raw storage by3Γ. - Database Index & Metadata Overhead (
+30\% - 50\%): B-Tree indexes, primary key lookups, and table metadata (PostgreSQL MVCC vacuum headers) add30\% - 50\%storage on top of raw payload bytes. - Network Bit/Byte Conversion Trap: Network capacity is measured in Bits per second (bps, Gbps), while storage/memory is measured in Bytes per second (B/s, GB/s). Always multiply Bytes by
8to obtain network line rates:
Bandwidth (Gbps) = Throughput (GB/s) Γ 8
- Headroom Buffer: Production systems should run at
β€ 65-70\%CPU and disk capacity to absorb sudden traffic surges, garbage collection pauses, and failover re-balancing.
βοΈArchitectural Trade-offs & Production Realities
Architectural Advantages
- Allows architects and candidates to ground complex architectural designs in hard physics within 3 minutes
- Quickly eliminates unscalable approaches (e.g. attempting to scan 100TB tables synchronously on HTTP requests)
Trade-offs & Constraints
- Over-focusing on exact arithmetic during interviews wastes precious time needed for high-level distributed systems trade-offs
- Assuming average load instead of peak load leads to severe under-provisioning during live flash events
Google and Meta engineering teams require formal Back-of-the-Envelope Capacity Estimations in every Design Document (Design Doc) before provisioning compute allocations, datacenter power quotas, and cross-region network backbone bandwidth for new global services.
π― Staff+ Engineering Takeaways
- Use 100,000 seconds per day for instant, clean mental division ($1\text{M}/\text{day} = 10\text{ QPS}$).
- RAM is 500x faster than SSD and 50,000x faster than network roundtrips.
- Multiply persistent storage by 3x for cross-AZ replication and add 30-50% for indexing overhead.
- Always distinguish between Bytes ($B$) for storage and Bits ($b$) for network line bandwidth ($1\text{ MB/s} = 8\text{ Mbps}$).
Topic Knowledge Assessment π§
Step through 3 scenario questions to test your staff-level grasp.
A photo-sharing service processes 300 million API read requests per day. Using standard back-of-the-envelope approximations, what is the estimated Average QPS and recommended Peak QPS capacity?
How clear and staff-actionable was this system breakdown?