Limited Offer

30% OFF Lifetime Access ($139) with code SYSTEM30

TOPIC #14Beginner 9 min read

Memory Hierarchy: CPU, RAM, Disk, Cache

💡
Core Architecture Summary

Explore the trade-offs between speed, cost, and capacity across CPU registers, L1/L2/L3 caches, Main Memory (DRAM), NVMe SSDs, and Magnetic Hard Drives.

Key Glossary Concepts in this TopicAll Glossary Terms

Latency Numbers Visualizer ⏱️

Scale nanosecond hardware delays into intuitive human time (1 CPU cycle = 1 second).

1 CPU Clock Cycle (3.3 GHz)
0.3 ns1 Second
L1 CPU Cache Reference
0.9 ns3 Seconds
Branch Misprediction Penalty
3.0 ns10 Seconds
L2 CPU Cache Reference
4.5 ns15 Seconds
Mutex Lock / Unlock
17.0 ns1 Minute
Main Memory Reference (DRAM)
100 ns5.5 Minutes
Read 1 MB Sequentially from RAM
3,000 ns (3 μs)2.8 Hours
NVMe SSD Random Read
16,000 ns (16 μs)15 Hours
Read 1 MB Sequentially from NVMe SSD
200,000 ns (200 μs)8 Days
Intra-Datacenter Network RTT
500,000 ns (0.5 ms)19 Days
Mechanical HDD Disk Seek
8,000,000 ns (8 ms)10 Months
Cross-Atlantic Network RTT (US to Europe)
150,000,000 ns (150 ms)16 YEARS
cpu Tier0.3 ns

1 CPU Clock Cycle (3.3 GHz)

The fundamental heartbeat of modern silicon processors executing a single instruction.

Human Time Analogy (1 CPU Cycle = 1 Sec)

1 Second

If fetching from CPU register takes 1 second, this operation feels like waiting 1 Second to the processor!

💡 Architectural Insight:

Reading from RAM is like a 5-minute coffee break; reading from spinning disk is waiting 10 months; cross-ocean network calls are a 16-year career!

The Modern Computer Memory Hierarchy 🔺

The physical trade-off continuum: As you move up toward the CPU register file, latency drops by orders of magnitude while cost-per-byte skyrockets.

The Modern Computer Memory Hierarchy 🔺
100%
Rendering visual architecture flowchart...

01.1. The Physical Necessity of the Memory Hierarchy

Modern CPU cores run at clock frequencies between 3.5 GHz and 5.5 GHz, completing single-cycle arithmetic operations in approximately 0.2 to 0.3 nanoseconds. However, electrical signals traveling across motherboard copper traces to main memory (DRAM) face significant physical and protocol delays, requiring approximately 100 nanoseconds to fetch data.

If a modern processor had to communicate directly with DRAM on every instruction, the execution pipeline would stall for 300 to 500 clock cycles on every single memory reference. This massive performance gap is known in computer architecture as the "Memory Wall".

To prevent CPU execution units from idling, hardware engineers stack a hierarchy of progressively smaller, faster, and more expensive SRAM (Static RAM) caches directly on the processor die. Each tier acts as a high-speed staging buffer for the tier below it:

  • Registers: Zero-cycle to half-nanosecond direct operand access for CPU Arithmetic Logic Units (ALUs).
  • L1 Cache (Level 1): Split into L1 Instruction Cache (L1i) and L1 Data Cache (L1d). Operates at ~1ns.
  • L2 Cache (Level 2): Dedicated per-core SRAM buffer (~3-4ns).
  • L3 Cache (Level 3 / LLC): Last Level Cache shared across all processor cores on the socket (~15-40ns).
  • Main Memory (DRAM): High-density volatile capacitive memory connected via DDR channels (~100ns).
  • Non-Volatile Storage (NVMe SSD / HDD): Block devices communicating over PCIe or SATA buses (microseconds to milliseconds).

02.2. Cache Lines, Spatial Locality, and Temporal Locality

Data is never transferred between DRAM and CPU caches as individual bytes. Instead, memory controllers fetch fixed-size chunks called Cache Lines, almost universally 64 bytes in size on modern x86 and ARM64 architectures.

Software performance in high-scale distributed systems is fundamentally governed by the Principle of Locality:

  1. Temporal Locality: If a particular memory location is accessed, it is highly likely to be accessed again in the near future (e.g., loop counter variables, accumulator totals, frequently accessed hash map buckets).
  2. Spatial Locality: If a particular memory location is accessed, memory locations with adjacent addresses are highly likely to be accessed soon (e.g., sequentially traversing an array or struct buffer).

When a program iterates over a contiguous array of 64-bit integers (8 bytes each), accessing the first element loads the entire 64-byte cache line from RAM into L1/L2/L3. The subsequent 7 integer accesses are guaranteed L1 cache hits (0.5-1ns), executing at hardware wire speed. In contrast, traversing a linked list scattered across random heap allocations incurs an LLC cache miss and DRAM latency (~100ns) on virtually every single pointer dereference.

03.3. NUMA (Non-Uniform Memory Access) in Multi-Socket Servers

Enterprise cloud servers (e.g., AWS c6i.32xlarge, bare-metal database nodes) use multi-socket motherboards with 2, 4, or 8 physical CPU processors. In these systems, memory is partitioned physically across sockets in a NUMA architecture:

  • Local Memory Access: When Core 0 on Socket 0 accesses DRAM directly wired to Socket 0's memory controller, memory latency is ~80-90ns.
  • Remote Memory Access: When Core 0 on Socket 0 must access DRAM wired to Socket 1, the request must traverse an inter-socket bus (Intel UPI / AMD Infinity Fabric), doubling latency to ~160-220ns.

High-throughput distributed systems (PostgreSQL, Redis, Kafka, Cassandra) must be NUMA-aware. Misconfigured thread placement can lead to random 2x-3x p99 latency spikes because worker threads continuously fetch remote NUMA pages.

bash— Inspecting NUMA topology and binding processes on Linux
# Check NUMA topology on a database server
numactl --hardware

# Output shows memory split across nodes:
# node 0 cpus: 0 1 2 3 4 5 6 7
# node 0 size: 65536 MB
# node 1 cpus: 8 9 10 11 12 13 14 15
# node 1 size: 65536 MB

# Run high-performance Redis instance pinned to single NUMA node
numactl --cpunodebind=0 --membind=0 redis-server /etc/redis/redis.conf

04.4. Architectural Implications for Distributed System Design

Understanding the memory hierarchy dictates major architectural decisions when designing databases, caches, and stream processors:

  1. Columnar Storage Formats (Parquet, ClickHouse): Storing data column-by-column rather than row-by-row guarantees that scanning a single column (e.g., user_age) reads contiguous memory addresses, achieving 100% spatial locality and enabling SIMD (Single Instruction, Multiple Data) vectorized execution.
  2. In-Memory Caches (Redis, Memcached): Keeping active datasets in DRAM bypasses OS storage stacks, eliminating context switches and disk block transfers to achieve sub-millisecond latencies.
  3. Log-Structured Merge Trees (LSM Trees in RocksDB/Cassandra): Converting random disk writes into sequential append-only writes maximizes drive controller buffer utilization and sequential write bandwidth.

⚖️Architectural Trade-offs & Production Realities

Architectural Advantages

  • Locality-aware software operates 10x to 100x faster by keeping hot working sets in L1/L2/L3 caches.
  • Sequential memory access allows hardware prefetchers to completely hide DRAM latency.
  • In-memory architectures provide predictable sub-millisecond latency for real-time applications.

Trade-offs & Constraints

  • DRAM is volatile: sudden power failure causes complete state loss without Write-Ahead Logging (WAL) or replication.
  • SRAM and high-speed DRAM are orders of magnitude more expensive per gigabyte than NVMe SSD storage.
  • NUMA traversal and false sharing can introduce subtle, hard-to-diagnose p99 latency regressions.
Production Implementation in Big Tech
Redis & ClickHouse• Maximizing Hardware Memory Throughput

Redis keeps all primary data structures directly in DRAM to achieve 100k+ operations per second per CPU core at <1ms latency. ClickHouse organizes analytical datasets into columnar blocks, processing millions of rows per second by keeping contiguous columns in CPU L1/L2 caches and executing SIMD vector instructions.

🎯 Staff+ Engineering Takeaways

  • The Memory Wall: CPUs execute in 0.3ns, while DRAM access takes ~100ns (300+ cycles stall without caches).
  • CPUs transfer memory in 64-byte Cache Lines; contiguous array traversals maximize spatial prefetching.
  • Temporal locality means recently accessed data will be accessed again; spatial locality means adjacent data will be accessed soon.
  • NUMA architectures introduce 2x latency penalties when accessing remote socket memory.
  • In-memory systems trade off high cost per GB and volatility for 1000x faster access speeds compared to storage.

Topic Knowledge Assessment 🧠

Step through 2 scenario questions to test your staff-level grasp.

Question 1 of 20 answered
#1

Why is iterating through a contiguous array in memory significantly faster than traversing a linked list of the exact same size?

Rate This Architecture Chapter4.9 / 5.0 (38 ratings)

How clear and staff-actionable was this system breakdown?