Storage pyramid -- why so many layers

Have you ever wondered why computers have both "memory" and "hard drive"? Why not put all data in the fastest place?

The answer lies in the design of computer storage systems--Storage pyramidThis lecture gives you a thorough understanding of the architect's most core trade-off wisdom.


Everyday analogy: kitchen vs. supermarket warehouse

Use the placement of ingredients while cooking to build an intuitive impression first.

Imagine you are cooking.

Salt and soy sauce, you put them by the stove, within reach--these are the computer'sregisterandL1 cache。

Ingredients in the fridge, you need to walk a few steps to get them--this is equivalent toMemory (RAM)。

Rice, flour, and oil stored in the basement, you have to make a special trip--this isSolid State Drive (SSD)。

And the stock in the supermarket warehouse, you wouldn't keep it at home at all--that isHard Disk Drive (HDD)and cloud storage.

This reveals a simple rule:Things used more often are placed closer, but closer space is limited and more expensive; things used less often are placed farther away, with larger capacity and also cheaper.

Computer storage systems are designed according to this simple rule.


Storage pyramid: the trade-off between speed and capacity

Inside a computer, storage devices form a pyramid structure according to the progressive relationship of "speed-capacity-cost".

example Storage pyramidallocateGraph — diagram-design Specification SVG:纸色底 + emit丝LineDividelayer,锈red accent onlymarkNote顶layerregister
Background halftone dot texture Layer 1: Registers — topmost layer, accent emphasizes (only focus) register ~1 KB · 0.3 nsLayer 2: L1 cache L1 cache ~64 KB · 1 nsLayer 3: L2 cache L2 cache ~256 KB · 4 nsLayer 4: L3 cache L3 cache ~8 MB · 12 nsLayer 5: Memory (RAM) Memory RAM ~16 GB · 100 nsLayer 6: Solid State Drive (SSD) Solid State Drive (SSD) ~1 TB · 100 μsLayer 7: Mechanical Hard Disk (HDD) — bottommost, widest Mechanical Hard Disk (HDD) ~10 TB · 10 msLeft and right side annotations (Geist Mono small text)Faster ↑ Slower ↓ Smaller capacity Larger capacity Mouse hover detail overlay

Storage pyramid diagram: the higher up, the faster and smaller; the lower down, the slower and larger. Hover over a trapezoid to see details for each layer.

Core rules

This pyramid diagram reveals an iron rule in hardware design:

  • Faster, more expensive, smallerRegisters are inside the CPU, made of the fastest transistors, but only a few kilobytes in size. Because chip area is extremely expensive.
  • Slower, cheaper, largerHard drives use magnetic heads and spinning platters, with extremely low cost, and can hold tens of terabytes of data.
  • Purpose of the pyramid structureUsing a combination of "small and fast + large and slow" achieves an experience close to the fastest storage at a reasonable cost.

If a computer used registers for all storage, not only would the price be astronomical, the chip area would also be too large to manufacture.

Conversely, if it used only hard drives, the computer would be too slow to use--you would have to wait hundreds of milliseconds to open any program.


Detailed comparison of each layer

Put the key attributes of the seven storage layers into one table for easy horizontal comparison.

TierTypical capacityAccess latencyRelative speedPhysical locationManufacturing cost
register~1 KB~0.3 ns1x (baseline)Inside the CPU coreextremely high
L1 cache~64 KB~1 ns3x slower than registersInside the CPU coreextremely high
L2 cache~256 KB~4 ns13x slowerInside the CPU coreVery High
L3 cache~8 MB~12 ns40x slowerOn the CPU chip, shared by multiple coresHigh
Memory (RAM)~16 GB~100 ns333x slowerOn the motherboard (separate chip)Medium
Solid State Drive (SSD)~1 TB~100 μs333,333x slowerInside the computer case (separate device)Relatively Low
Hard Disk Drive (HDD)~10 TB~10 ms33,333,333x slowerInside the computer case (separate device)Low

Notice the order-of-magnitude jumps: memory is 100 times slower than L1 cache, SSD is 1000 times slower than memory, and HDD is 100 times slower than SSD. The speed gap between every layer is measured in orders of magnitude.

Why is SSD so much faster than HDD?

The key difference lies in whether there are mechanical parts.

SSD uses flash memory chips, reading purely electronically, with no mechanical parts.

HDD uses spinning platters and a moving head--to read a piece of data, the head must first move to the correct track (seek time, about 5-10 ms), then wait for the platter to rotate to the correct position (rotational latency, about 2-4 ms).

This mechanical movement process is exactly the root cause of HDD being hundreds of times slower than SSD.


Interactive demo: simulating access times of multi-layer storage

Use a Python snippet to simulate the latency differences when reading the same address from each of the seven storage layers.

Example

"""
Memory Pyramid Access Latency Simulator (example Demo)
Simulate reading data blocks of the same size from different levels to intuitively compare the speed differences of each level
"""


class StorageHierarchy:
    """Simulate the memory hierarchy of a computer"""

    def __init__(self):
        # Simulated latency (in nanoseconds) and typical capacity of each storage layer
        self.layers = {
            'Register':  {'latency_ns': 0.3,  'capacity': '~1 KB',   'color': 'red'},
            'L1 Cache':  {'latency_ns': 1,    'capacity': '~64 KB',  'color': 'orange'},
            'L2 Cache':  {'latency_ns': 4,    'capacity': '~256 KB', 'color': 'gold'},
            'L3 Cache':  {'latency_ns': 12,   'capacity': '~8 MB',   'color': 'yellow'},
            'RAM':       {'latency_ns': 100,  'capacity': '~16 GB',  'color': 'green'},
            'SSD':       {'latency_ns': 100000, 'capacity': '~1 TB', 'color': 'blue'},
            'HDD':       {'latency_ns': 10000000, 'capacity': '~10 TB', 'color': 'purple'},
        }

    def read(self, layer_name, address):
        """
Simulate reading data from the specified layer
Return simulated data values
        """

        value = f"DATA_FROM_{layer_name.upper()}_{address}"
        return value

    def compare_access(self, address, num_accesses=3):
        """Compare the speed of reading the same address from each level."""
        print("=" * 65)
        print(f"Memory hierarchy access latency comparison (address: {address})")
        print("=" * 65)
        print(f{'Level':<12} {'Latency (ns)':>12} {'Capacity':<12} {'Relative to register':>12})
        print("-" * 65)

        baseline = self.layers['Register']['latency_ns']

        for name, info in self.layers.items():
            lat_ns = info['latency_ns']
            ratio = lat_ns / baseline
            cap = info['capacity']
            print(f"{name:<12} {lat_ns:>10,.1f} ns {cap:<12} {ratio:>10,.0f}x")


# Run demo
print(EXAMPLE Storage System Tutorial: Storage Pyramid Access Latency Comparison)
print()

store = StorageHierarchy()
store.compare_access("0x7FFF1234")

print()
print("=" * 65)
print("Conclusion analysis:")
print("=" * 65)

# Calculate key ratios
register_lat = 0.3          # ns
ram_lat = 100               # ns
ssd_lat = 100000            # ns
hdd_lat = 10000000          # ns

print(f1. Memory (RAM) is {ram_lat / register_lat:,.0f} times slower than registers)
print(f2. SSD is {ssd_lat / ram_lat:,.0f} times slower than memory (RAM))
print(f3. HDD is {hdd_lat / ram_lat:,.0f} times slower than memory (RAM))
print(f4. HDD is {hdd_lat / register_lat:,.0f} times slower than registers)
print()
print("If register access takes 1 second, then:")
print(f- Fetching from L1 cache takes {1/register_lat:.0f} seconds)
print(f"  - fromMemoryGetrequires {ram_lat/register_lat:,.0f} second(约 {ram_lat/register_lat/60:.0f} Divide钟)")
print(f"  - from HDD Getrequires {hdd_lat/register_lat:,.0f} second(约 {hdd_lat/register_lat/3600:,.0f} hours)")

Run result:

EXAMPLE 存储系统教学: 存储金字塔访问延迟对比

=================================================================
存储层次访问延迟对比(地址: 0x7FFF1234)
=================================================================
层级                 延迟(纳秒) 容量                  相对寄存器
-----------------------------------------------------------------
Register            0.3 ns ~1 KB                 1x
L1 Cache            1.0 ns ~64 KB                3x
L2 Cache            4.0 ns ~256 KB              13x
L3 Cache           12.0 ns ~8 MB                40x
RAM               100.0 ns ~16 GB              333x
SSD           100,000.0 ns ~1 TB           333,333x
HDD          10,000,000.0 ns ~10 TB       33,333,333x

=================================================================
结论分析:
=================================================================
1. 内存(RAM) 比 寄存器 慢 333 倍
2. SSD 比 内存(RAM) 慢 1,000 倍
3. HDD 比 内存(RAM) 慢 100,000 倍
4. HDD 比 寄存器 慢 33,333,333 倍

如果寄存器访问数据需要 1 秒,那么:
  - 从 L1 缓存获取需要 3 秒
  - 从内存获取需要 333 秒(约 6 分钟)
  - 从 HDD 获取需要 33,333,333 秒(约 9,259 小时)

Interactive demo: logarithmic comparison chart of access latency across layers

The chart below draws a bar chart on a logarithmic scale; adjacent intervals on the horizontal axis differ by 10 times. Hover over a bar to see capacity and latency details.

example storage hierarchy latency comparison — Plotly log-coordinate horizontal bar chart (standard selection: use Plotly for logarithmic coordinate system)

How the pyramid works: layer-by-layer caching strategy

The pyramid is not static--it has an automatic data movement mechanism.

Data flow rules

  1. When the CPU Needs DataFirst look in the fastest L1 cache; if not found, go to L2; if still not found, go to L3, and so on down to main memory.
  2. After reading data from the slow layerIt not only gives the data to the CPU, but also stores a copy in a faster layer. That way the next use is fast.
  3. What to do when the fast layer is full?According to a certain policy (such as LRU--evicting the least recently used), kick infrequently used data back to slower layers to make room for new data.

This process is completely transparent to programmers--you don't need to manually manage which cache layer when writing programs; hardware and the operating system automatically do all of this.

Real-world examples

Suppose you are editing a video file:

  • Video files existHDD or SSDAbove (bottom layer).
  • When you open a file, the operating system loads part of it intoMemory (RAM)In.
  • When you start playing, the CPU copies the data of the frames currently being processed toL3/L2/L1 cacheIn.
  • Pixel values currently being computed by the ALU are stored inregisterIn.

When you feel "very smooth" in editing software, it's because the CPU finds the data it needs in the cache most of the time.


Historical background: why the pyramid has this shape

The number of pyramid layers is a historical product of the ever-widening speed gap between CPU and memory.

In the 1980s, the speed gap between CPU and memory was not large. But as semiconductor manufacturing advanced, CPU speed grew at about 60% per year, while memory speed only grew at about 10% per year.

This ever-widening gap is called「Memory Wall」(Memory Wall)。

The storage pyramid is engineers' response to the "memory wall"--using multiple levels of cache to buffer the speed gap between CPU and main memory.

eraCPU frequencyMemory latencySpeed gapCache hierarchy
1980s~10 MHz~200 nsapproximately 2xNone or Level 1
1990s~200 MHz~70 nsAbout 14 timesL1 + L2
2000s~3 GHz~50 nsAbout 150 timesL1 + L2 + L3
2020s~5 GHz~80 nsAbout 400 timesL1 + L2 + L3

Note that memory latency has barely changed in essence over 40 years! It's not that memory hasn't improved, but that CPU has improved too fast. The physical limit of memory (capacitor charge/discharge speed) makes it very difficult to significantly shorten its latency.

other extensions