Three Ways of I/O -- Polling, Interrupt, DMA
In the previous lecture, we learned how the CPU and memory communicate via the bus. But a computer is not just CPU plus memory—it also needs to exchange data with various external devices (I/O devices) such as keyboards, mice, hard disks, and network cards.
There are three basic ways for the CPU to exchange data with external devices:Polling、Interrupt (Interrupt)、DMA (Direct Memory Access)The efficiency of the three methods differs enormously—which one you choose often determines the performance of the entire system.
In this lecture, we use three everyday scenarios to help you build intuition, and then use Python code to intuitively compare the efficiency differences among the three methods.
Three real-life scenario analogies
Scenario 1: Polling—Check every 5 seconds whether the water is boiling
You are boiling water in the kitchen, but the kettle has no whistle. Every 5 seconds, you run from the living room to the kitchen to check: "Is the water boiling? Is the water boiling? Is the water boiling?"
During this period, youCompletely Unable to Watch TV with Peace of MindBecause you have to make a trip every few seconds, you spend almost all your time on "checking the status."
This ispollingPolling: the CPU continuously loops to check whether the device is ready, wasting time on meaningless queries.
Scenario 2: Interrupt—Turn off the stove when the kettle's whistle sounds
This time, the kettle has a whistle. After putting the water on the stove, you go back to the living room and watch TV in peace. The water boils, and the whistle sounds sharply. You pause watching TV, go to the kitchen to turn off the stove, and then continue watching TV.
youThe vast majority of the time is spent doing meaningful things(watching TV), interrupted only briefly at critical moments.
This isinterruptInterrupt: the CPU executes its program normally; when the device is ready, it actively sends a signal to notify the CPU. The CPU pauses its current work, processes the device data, and then resumes its previous work.
Scenario 3: DMA—Just call a moving company, you don't have to lift a finger
You are moving, with a truckload of furniture to be transported from the old house to the new house. If you carry it piece by piece yourself (the interrupt way), you have to run back and forth dozens of times.
You hire a moving company (DMA controller) and tell them: "The old house address is X, the new house address is Y, and there are Z pieces of furniture in total." Then you continue with your own things. After the moving company handles everything, they call you and say "Done moving."
This isDMADMA: a dedicated DMA controller directly manages data transfer between memory and peripherals, and the CPU does not need to participate at all. After the transfer completes, the DMA controller sends an interrupt notification to the CPU.
The essential differences among the three methods
| Features | Polling | Interrupt | DMA |
|---|---|---|---|
| CPU involvement | Busy all the way, constantly checking | Briefly participate during transmission | Hardly participates, only receives completion notification |
| Who moves the data | CPU personally moves every byte | CPU personally moves every byte | The DMA controller does the moving; the CPU doesn't get involved |
| CPU Idle Time | Almost none | Idle most of the time | Idle almost all the time |
| Implementation complexity | Simplest | Requires interrupt controller hardware | Requires DMA controller hardware |
| Response latency | Depends on polling interval | Microsecond-level | Depends on total DMA transfer duration |
| Applicable data size | Extremely small amount | Small to medium amounts (KB level) | Large amounts (MB/GB level) |
| Typical scenarios | Simple embedded systems, extremely short waits | Keyboard input, mouse movement, network packet arrival | Disk reads/writes, graphics card framebuffer, large network packets |
Each time you press a key, an interrupt is triggered; when you copy a large file, DMA is used. Modern operating systems automatically choose the best scheme among the three methods based on the data volume and the scenario.
Timeline visualization: comparison of CPU busy vs. idle states
Suppose we need to read 4 data blocks from the hard disk into memory. We use an ASCII timeline to visually compare the CPU state under the three methods:
任务: 从硬盘读取 4 个数据块到内存
【轮询 - CPU 几乎全程忙碌】
时间: 0----1----2----3----4----5----6----7----8----9----10--->
CPU: [查][查][搬][查][查][搬][查][查][搬][查][查][搬]
忙 忙 忙 忙 忙 忙 忙 忙 忙 忙 忙 忙
↑ 几乎没有任何空闲
【中断 - CPU 大部分时间空闲】
时间: 0----1----2----3----4----5----6----7----8----9----10--->
CPU: [工作][等待...][搬][工作] [等待...][搬][工作]
忙 闲 忙 忙 闲 忙 忙
↑ 中断触发 ↑ 中断触发
【DMA - CPU 几乎全程空闲】
时间: 0----1----2----3----4----5----6----7----8----9----10--->
CPU: [设置DMA][工作][工作][工作] [被通知完成]
忙 忙 忙 忙 忙
DMA: [搬][搬][搬][搬] ↑ DMA 完成后中断
忙 忙 忙 忙
↑ CPU 完全不管传输过程
It's clear at a glance: under polling, the CPU is tied up; under interrupt, the CPU is only briefly interrupted; under DMA, the CPU is almost completely freed.
Code demonstration: simulate the three I/O methods in Python and print the timeline
The code below simulates the transfer of the same batch of data using polling, interrupt, and DMA respectively, calculates the proportion of CPU busy/idle time, and uses an ASCII-art timeline to intuitively show the differences.
Example
Comparison of three I/O methods: polling vs interrupt vs DMA (example demo)
This program simulates the scenario where the CPU transfers 10 data blocks from an 'external device' to memory,
Implement using polling, interrupt, and DMA respectively, and print a timeline to visually compare CPU states.
Core metric: CPU busy time ratio. The higher the ratio, the more severely the CPU is dragged down by I/O.
"""
import time
import random
class IODevice:
"""Simulate external I/O devices (such as hard disks, keyboards, etc.)"""
def __init__(self, name="Hard disk", response_time=4):
"""
name: device name
response_time: the number of 'time units' required for device response
"""
self.name = name
self.response_time = response_time # Device preparation time per data block
self.data_buffer = [] # Device internal data buffer
self.ready = False # Is Device Ready
def prepare_data(self, block_id):
"""
Device prepares data—simulating the device's internal processing.
In actual hardware, this step involves operations such as disk seeking, network card packet reception, etc.
"""
self.ready = False
data = f"DATA_BLOCK_{block_id:02d}"
self.data_buffer.append(data)
# Simulate that devices are ready after different time units
return data
def is_ready(self):
"""CPU checks this status when polling"""
return self.ready
def mark_ready(self):
"""Device finished preparing, sending interrupt signal"""
self.ready = True
def get_data(self):
"""Take a data block from the device buffer"""
if self.data_buffer:
return self.data_buffer.pop(0)
return None
class Timeline:
"""
Timeline logger — records the CPU and DMA state for each time unit
Used for final rendering of the ASCII timeline
"""
def __init__(self):
self.cpu_states = [] # CPU state per time unit: 'BUSY' or 'IDLE'
self.dma_states = [] # DMA status per time unit
self.events = [] # Event log: (time, description)
def record_cpu(self, state):
self.cpu_states.append(state)
def record_dma(self, state):
self.dma_states.append(state)
def log_event(self, time_unit, desc):
self.events.append((time_unit, desc))
def draw(self, title="Timeline"):
"""Draw an ASCII timeline"""
total = len(self.cpu_states)
if total == 0:
print(" (no time data)")
return
print(f"\n ╔══ {title} ══╗")
print(f" ║ Total time units: {total}")
# CPU Status Line
cpu_line = "CPU: "
for s in self.cpu_states:
cpu_line += "[" + ("Busy" if s == "BUSY" else "Idle") + "]"
print(f" ║ {cpu_line}")
# DMA status line (if DMA activity exists)
if any(s == "BUSY" for s in self.dma_states):
dma_line = "DMA: "
for s in self.dma_states:
dma_line += "[" + ("Busy" if s == "BUSY" else "Idle") + "]"
print(f" ║ {dma_line}")
# Event log
if self.events:
print(f" ╠══ Event Log ══╣")
for t, desc in self.events:
print(f" ║ T={t}: {desc}")
# Statistical summary
busy_count = sum(1 for s in self.cpu_states if s == "BUSY")
idle_count = total - busy_count
busy_pct = busy_count / total * 100 if total > 0 else 0
idle_pct = idle_count / total * 100 if total > 0 else 0
print(f" ╠══ Statistics ══╣")
print(f║ CPU busy: {busy_count}/{total} ({busy_pct:.0f}%))
print(f║ CPU idle: {idle_count}/{total} ({idle_pct:.0f}%))
print(f" ╚{'═' * 20}╝")
return busy_pct
class Memory:
"""Simulate main memory"""
def __init__(self):
self.data = {}
def write(self, addr, data):
self.data[addr] = data
def __repr__(self):
return fMemory({len(self.data)} data items)
class Controller:
"""CPU controller — simulation of executing different I/O modes"""
def __init__(self):
self.memory = Memory()
self.total_transferred = 0
def simulate_polling(self, device, block_count):
"""
Polling simulation:
The CPU checks whether the device is ready at each time unit.
If ready, move a data block; otherwise continue checking.
"""
print(f"\n{'#'*55}")
print(f"# Method 1: Polling")
print(f# Device: {device.name} | Block count: {block_count})
print(f"# Strategy: CPU checks device status every time unit")
print(f"{'#'*55}")
timeline = Timeline()
time_unit = 0
blocks_done = 0
device_ready_at = 0 # At which time unit the device is ready for the next block
# Initialization: the device starts preparing the first block of data
device.prepare_data(0)
device_ready_at = time_unit + device.response_time
while blocks_done < block_count:
# CPU checks device status (polling)
timeline.record_cpu("BUSY")
timeline.record_dma("IDLE")
if time_unit >= device_ready_at and device.is_ready():
# Device ready -> CPU transfers data
data = device.get_data()
self.memory.write(blocks_done, data)
timeline.log_event(time_unit, f"CPU moved data block {blocks_done}: {data}")
blocks_done += 1
self.total_transferred += 1
# Prepare next block of data (if any)
if blocks_done < block_count:
device.prepare_data(blocks_done)
device_ready_at = time_unit + device.response_time
else:
# Device not ready -> CPU wasted a trip
timeline.log_event(time_unit,
fCPU checking device... {'Ready' if device.is_ready() else 'Not ready'})
# Device background work (if in preparation phase)
if time_unit == device_ready_at:
device.mark_ready()
timeline.log_event(time_unit, Device ready, data ready)
time_unit += 1
print(f"\nTransfer completed: {self.total_transferred} data blocks)
busy_pct = timeline.draw("Polling Mode Timeline")
return busy_pct
def simulate_interrupt(self, device, block_count):
"""
Interrupt simulation:
CPU spends most of its time executing other tasks (denoted as IDLE),
Only after receiving the interrupt signal does it pause the current work, handle the data transfer, and then continue.
"""
print(f"\n{'#'*55}")
print(f# Method 2: Interrupt)
print(f# Device: {device.name} | Block count: {block_count})
print(f# Strategy: The CPU executes tasks normally, and sends an interrupt notification when the device is ready)
print(f"{'#'*55}")
timeline = Timeline()
time_unit = 0
blocks_done = 0
transferred_in_interrupt = 0
device_ready_at = 0
# Initialize
device.prepare_data(0)
device_ready_at = time_unit + device.response_time
while blocks_done < block_count:
# Default: CPU is executing other tasks (idle)
if time_unit == device_ready_at:
# Device ready, send interrupt signal
device.mark_ready()
timeline.log_event(time_unit, Interrupt triggered! Device data ready)
# CPU responds to the interrupt and transfers data (during this time, CPU is busy)
data = device.get_data()
for substep in range(2): # Interrupt handling requires 2 time units
if substep == 0:
timeline.record_cpu("BUSY") # Save context, enter interrupt handling
timeline.record_dma("IDLE")
timeline.log_event(time_unit + substep, f"CPU enters interrupt handling")
else:
timeline.record_cpu("BUSY") # Move data
timeline.record_dma("IDLE")
self.memory.write(blocks_done, data)
transferred_in_interrupt += 1
timeline.log_event(time_unit + substep, f"CPU moved data block {blocks_done}: {data}")
blocks_done += 1
# Interrupt handling complete, CPU resumes its previous work.
timeline.record_cpu("BUSY")
timeline.record_dma("IDLE")
timeline.log_event(time_unit + 2, CPU resumes previous work)
time_unit += 3 # Skip the time taken by interrupt handling
# Prepare next block
if blocks_done < block_count:
device.prepare_data(blocks_done)
device_ready_at = time_unit + device.response_time
continue
# Normal execution phase — CPU handles its own tasks
timeline.record_cpu("IDLE")
timeline.record_dma("IDLE")
timeline.log_event(time_unit, CPU executes other tasks (idle/user program))
time_unit += 1
print(f"\nTransfer completed: {self.total_transferred} data blocks)
busy_pct = timeline.draw("Interrupt Mode Timeline")
return busy_pct
def simulate_dma(self, device, block_count):
"""
DMA Mode Simulation:
The CPU only configures the DMA controller at the beginning, and then completely leaves the data transfer alone.
DMA controller autonomously moves all data blocks.
After the transfer is completed, DMA sends an interrupt to notify the CPU.
"""
print(f"\n{'#'*55}")
print(f# Method 3: DMA (Direct Memory Access))
print(f# Device: {device.name} | Block count: {block_count})
print(f# Strategy: CPU configures DMA and continues working, DMA controller automatically transfers data)
print(f"{'#'*55}")
timeline = Timeline()
time_unit = 0
dma_busy_until = 0
setup_done = False
blocks_to_transfer = list(range(block_count))
dma_buffer = []
# Prepare All Data Blocks
for i in range(block_count):
device.prepare_data(i)
while blocks_to_transfer or dma_buffer:
# Phase 1: CPU configures the DMA controller (only once, taking 2 time units)
if not setup_done:
for substep in range(2):
timeline.record_cpu("BUSY")
timeline.record_dma("IDLE")
if substep == 0:
timeline.log_event(time_unit,
CPU configures DMA: set source address, destination address, and transfer byte count)
else:
timeline.log_event(time_unit,
CPU initiates DMA transfer, then continues its own work)
time_unit += 1
setup_done = True
dma_busy_until = time_unit + block_count * device.response_time
# Mark Device as Ready
device.mark_ready()
# Pre-fill DMA buffer
while device.data_buffer:
dma_buffer.append(device.data_buffer.pop(0))
continue
# Second phase: DMA controller transfers data autonomously, CPU is idle
if time_unit < dma_busy_until:
timeline.record_cpu("IDLE")
timeline.record_dma("BUSY")
if dma_buffer:
data = dma_buffer.pop(0)
blocks_to_transfer.pop(0)
mem_addr = block_count - len(blocks_to_transfer) - 1
self.memory.write(mem_addr, data)
self.total_transferred += 1
timeline.log_event(time_unit, fDMA transfers data block {mem_addr}: {data})
else:
timeline.log_event(time_unit, "DMA controller working...")
else:
# Stage 3: DMA complete, sends interrupt to notify the CPU
timeline.record_cpu("BUSY")
timeline.record_dma("BUSY")
timeline.log_event(time_unit, "DMA transfer complete, interrupt notifies CPU")
timeline.record_cpu("BUSY")
timeline.record_dma("IDLE")
timeline.log_event(time_unit + 1, "CPU confirms transfer complete")
time_unit += 2
break
time_unit += 1
print(f"\nTransfer completed: {self.total_transferred} data blocks)
busy_pct = timeline.draw("DMA mode timeline")
return busy_pct
# ===== Main program: comparison of three approaches =====
if __name__ == "__main__":
print("=" * 60)
print(Efficiency comparison of three I/O modes — EXAMPLE Computer Organization Principles demo)
print("=" * 60)
print()
print(Scenario: Transfer 5 data blocks from an external device to memory)
print(Device response time (time units required to prepare each block of data): 4)
print()
BLOCK_COUNT = 5
RESPONSE_TIME = 4
results = {}
# ---- Method 1: Polling ----
dev1 = IODevice("Hard disk", RESPONSE_TIME)
ctrl1 = Controller()
pct1 = ctrl1.simulate_polling(dev1, BLOCK_COUNT)
results["Polling"] = pct1
# ---- Method 2: Interrupt ----
dev2 = IODevice("Hard disk", RESPONSE_TIME)
ctrl2 = Controller()
pct2 = ctrl2.simulate_interrupt(dev2, BLOCK_COUNT)
results["Interrupt"] = pct2
# ---- Method 3: DMA ----
dev3 = IODevice("Hard disk", RESPONSE_TIME)
ctrl3 = Controller()
pct3 = ctrl3.simulate_dma(dev3, BLOCK_COUNT)
results["DMA"] = pct3
# ---- Final comparison ----
print("\n" + "=" * 60)
print(Comparison of CPU busy time ratios among three methods)
print("=" * 60)
print(fPolling: CPU is busy {results['polling']:.0f}% of the time)
print(fInterrupt: CPU busy {results['interrupt']:.0f}% of the time)
print(fDMA: CPU busy {results['DMA']:.0f}% of the time)
print()
best = min(results, key=results.get)
worst = max(results, key=results.get)
print(f" Best method: {best} (lowest CPU usage)")
print(f" Worst method: {worst} (highest CPU usage)")
print()
print(fConclusion: For large data transfers, DMA reduces CPU usage compared to polling.)
print(fapproximately {results['polling'] - results['DMA']:.0f} percentage points.)
print(f" This is why modern computers use DMA when transferring large files.")
print("=" * 60)
Real-world practical application scenarios
Keyboard input — a typical example of interrupts
Every time you press a key, the keyboard controller sends an interrupt signal to the CPU. The CPU pauses the current program, reads the key code, and then resumes. This process is so fast that you don't notice any delay.
If the keyboard used polling—the CPU checks 1,000 times per second "Has the user pressed a key?"—your computer would freeze and you couldn't do anything.
Large file copying — a typical example of DMA
You copy a 2GB movie from drive C to drive D. The disk controller uses DMA to move data directly to memory; the CPU is hardly involved throughout the entire process. You can also browse the web and write code at the same time, completely unaffected.
Without DMA—the CPU would have to carry 2GB of data by itself—and during this time, the CPU basically cannot handle any other tasks.
Why hasn't polling been eliminated?
Although polling has the lowest efficiency, it isExtremely Simple Embedded SystemStill useful in:
- Lowest hardware cost—no interrupt controller or DMA controller is needed.
- For scenarios where "the device is ready almost instantly," polling has overhead comparable to interrupt, but is much simpler to implement.
- In some real-time systems, the determinism of polling (knowing exactly when to check) has an advantage over the randomness of interrupts.
Summary and verification
One-sentence summary:pollingYes CPU No停groundask「OKalreadynot」,interruptYesdevicesmain动喊「IOKalready」,DMA Yes找items帮hand替you干活——Dataquantity越big,DMA ofAdvantages越明显。
self-test questions
- You are using your computer to write a document while a large file is being copied in the background. Which I/O method is used for file copying? Which one is used for your keyboard input?
- If a device's data preparation time is extremely short (for example, only 1 microsecond), would polling still waste CPU? Why?
- After the DMA controller completes the data transfer, how does it notify the CPU? What mechanism does it use?
Reference answer: 1. File copying uses DMA, keyboard input uses interrupts. 2. It still wastes — because during polling the CPU cannot do other things; even if each poll is fast, the occupied time slice cannot be used to execute user programs. But if the device is truly very fast, the simplicity of polling implementation may be more advantageous. 3. After DMA completes, it notifies the CPU via an interrupt — this is the only moment when DMA requires CPU involvement.
other extensions