Estimating QPS, Storage, & Bandwidth: End-to-End Walkthrough
Work through a comprehensive, rigorous end-to-end capacity planning model: Sizing a 500M DAU global social platform from QPS and 5-year multi-tier storage to network egress bandwidth and Redis cluster RAM sizing.
Complete 500M DAU End-to-End Capacity Estimation Pipeline π
A structured, rigorous derivation from user activity metrics to compute instances, persistent storage tiers, network line rates, and in-memory cache allocations.
01.1. Problem Statement & Baseline Assumptions
Let us design the capacity and infrastructure requirements for a global social platform (e.g., a Twitter/X or Instagram hybrid) with the following product scale specifications:
- Daily Active Users (DAU):
500 Million - User Activity (Writes):
100 Million new posts/day(0.2 posts/user/day).100\%contain text & metadata (500 bytes).10\%of posts include an image (2 MBaverage size).
- User Activity (Reads):
10 Billion feed requests/day(Each active user opens the feed20 times/day).- Each feed request returns a page of 20 post summaries
β 10 KBresponse payload.
- Each feed request returns a page of 20 post summaries
- Read-to-Write Ratio:
10 Billion reads : 100 Million writes = 100:1 (Heavily Read-Heavy).
02.2. Mathematical Step-by-Step Capacity Derivation
Step 1: Traffic Estimation (QPS Modeling)
Using the 100,000 seconds/day mental approximation:
Average Write QPS = \frac{100,000,000 writes}{100,000 sec} = 1,000 Write QPS
Peak Write QPS = 1,000 Γ 2 = 2,000 Peak Write QPS
Average Read QPS = \frac{10,000,000,000 reads}{100,000 sec} = 100,000 Read QPS
Peak Read QPS = 100,000 Γ 2 = 200,000 Peak Read QPS
Step 2: Storage Sizing (Daily & 5-Year Projections)
- Metadata & Text Database Storage:
- Daily volume:
100M posts Γ 500 bytes = 50 GB/day. - 5-Year persistent volume:
50 GB/day Γ 365 days Γ 5 years = 91,250 GB β 91.25 TB. - With
3Γcross-AZ replication+ 30\%index overhead:
- Daily volume:
91.25 TB Γ 3 Γ 1.30 = 355.8 TB Raw Disk Capacity
- Blob / Object Storage (Photos & Media):
10\%of100M posts = 10M images/day.- Daily media ingestion:
10M images Γ 2 MB = 20,000,000 MB = 20 TB/day. - 5-Year media volume:
20 TB/day Γ 365 Γ 5 β 36.5 Petabytes (PB). - Architecture decision: Store images in Amazon S3 / Google Cloud Storage with automated lifecycle policies transitioning objects older than 90 days to S3 Glacier Deep Archive (
95\%cost savings).
Step 3: Network Bandwidth Estimation
-
Ingress Bandwidth (Writes):
- Media Ingress:
10M images/day Γ 2 MB = 20 TB/day = \frac{20 Γ 10^{12} bytes}{10^5 sec} = 200 MB/s. - Text Ingress:
1,000 QPS Γ 500 B = 0.5 MB/s. - Total Ingress Line Rate:
200.5 MB/s Γ 8 bits/byte = 1.6 Gbps.
- Media Ingress:
-
Egress Bandwidth (Reads):
- API Feed Egress:
100,000 Read QPS Γ 10 KB = 1,000,000 KB/s = 1 GB/s. - In Line Rate:
1 GB/s Γ 8 bits/byte = 8 Gbps Egress. - With a Global CDN (Cloudflare/CloudFront) caching
90\%of feed assets and media at the edge, origin datacenter egress drops from8 Gbpsto800 Mbps, protecting backend API infrastructure.
- API Feed Egress:
Step 4: Memory Cache Sizing (80/20 Pareto Working Set)
- Daily API read volume:
10 Billion views Γ 10 KB = 100 TB read data/day. - Applying the 80/20 Pareto Rule,
20\%of the unique daily content generates80\%of read traffic. - Required in-memory cache size:
RAM Cache Size = 100 TB Γ 0.20 = 20 TB RAM
03.3. Hardware & Cluster Sizing: Translating Math to Server Fleets
A senior system design interview requires connecting mathematical storage and bandwidth numbers to concrete physical node fleet counts:
-
API Application Fleet Sizing:
- A single optimized stateless Go/Java/Node.js API container handles
~ 2,000 QPSunder normal CPU load. - Peak Read Traffic =
200,000 QPS. - Required API instances =
\frac{200,000 Peak QPS}{2,000 QPS/instance} = 100 API Pods(Provision 130 pods for30\%headroom and redundancy).
- A single optimized stateless Go/Java/Node.js API container handles
-
Redis In-Memory Caching Cluster:
- Total RAM required =
20 TB. - Standard AWS memory-optimized node (
r6g.4xlarge):128 GB RAM. - Usable RAM per node (leaving
25\%for Redis copy-on-write BGSAVE overhead)β 96 GB. - Cluster Node Count =
\frac{20,000 GB}{96 GB/node} β 208 Redis Primary Nodes(+208replicas across AZs for high availability).
- Total RAM required =
-
Database Sharding Fleet:
- 5-Year Metadata Storage =
91.25 TB. - To keep PostgreSQL/MySQL B-Tree index scans fast and backup restore times under 30 minutes, limit each database shard size to
β€ 2 TB SSD. - Shard Count =
\frac{91.25 TB}{2 TB/shard} β 46 Shards(Round to64 shardsfor clean power-of-2 consistent hashing partition rings).
- 5-Year Metadata Storage =
04.4. The Capacity Estimation Summary Card
| Metric Dimension | Average Load | Peak Load (2x) | 5-Year Cumulative | Infrastructure Footprint |
|---|---|---|---|---|
| Write QPS | 1,000 ops/sec | 2,000 ops/sec | 182.5 Billion records | Kafka Ingestion Buffer (6 brokers) |
| Read QPS | 100,000 ops/sec | 200,000 ops/sec | β | 130 Stateless API Pods |
| Relational Storage | 50 GB/day | β | 91.25 TB (355 TB with 3x repl + index) | 64 Database Shards (2TB each) |
| Media Blob Storage | 20 TB/day | β | 36.5 Petabytes | S3 Object Store + Glacier Lifecycle |
| Network Egress | 1 GB/s (8 Gbps) | 2 GB/s (16 Gbps) | β | Edge CDN with 90% cache offload |
| Memory Cache (RAM) | β | β | β | 208 Redis Nodes (20 TB total working set) |
βοΈArchitectural Trade-offs & Production Realities
Architectural Advantages
- Gives concrete, defensible numbers for every architectural component (shards, pods, cache nodes, object tiers)
- Reveals critical scaling bottlenecks early (e.g. realizing media storage requires Petabyte-scale object storage rather than relational DBs)
Trade-offs & Constraints
- Assumptions may shift drastically during actual product growth (e.g. video introduction vs pure text)
- Failure to account for network line rate conversion ($8\times$) leads to severe network interface card (NIC) saturation
Twitter engineers use strict 5-year storage projections and 80/20 cache sizing formulas to determine SSD partition allocations across Manhattan distributed key-value stores and to size memory footprints across tens of thousands of Twemcache (Memcached) instances.
π― Staff+ Engineering Takeaways
- Always follow the standard 5-step capacity sizing sequence: QPS -> Storage -> Bandwidth -> Memory -> Node Counts.
- Use the 80/20 rule to size the RAM caching tier (cache 20% of daily read volume).
- Divide cumulative 5-year storage by 2TB to calculate the number of database shards required.
- Apply CDN edge caching to offload 85-95% of egress bandwidth away from origin microservices.
Topic Knowledge Assessment π§
Step through 3 scenario questions to test your staff-level grasp.
A photo platform receives 50 million photo uploads per day, with an average photo size of 1.5 Megabytes. Approximately how much raw object storage will be consumed over 5 years (ignoring compression)?
How clear and staff-actionable was this system breakdown?