The Complete Journey of an Instruction
From the binary basics in Lecture 1 to the I/O system in the previous lecture, we have covered all the core content of computer organization principles. This lecture will string these knowledge points together into a single line, tracing the complete journey of an instruction from birth to completion.
We will implement a completevon Neumann Computer Simulator (FullComputer), it includes: four-level storage structure (hard disk, memory, cache, registers), DMA data transfer, four-stage pipeline of fetch-decode-execute-writeback, bus interaction, cache hit rate statistics. Finally, print a complete execution report.
Review: Knowledge Map of Six Modules
Before starting the end-to-end simulation, let's quickly review the core concepts of each of the six modules in this course:
| Module | topic | Core Concepts | Embodiment in this lecture |
|---|---|---|---|
| Module 1 | Binary and Information Representation | 0/1 encoding, base conversion, two's complement, floating-point numbers | Both instructions and data are stored in binary form in memory. |
| Module 2 | Logic Gates and Digital Circuits | AND/OR/NOT gates, Adders, ALU | The ALU operations in the execution stage are implemented by gate circuits. |
| Module 3 | Von Neumann architecture | Five Major Components, Stored Program, Instruction Cycle | What this lecture simulates is a complete von Neumann machine. |
| Module 4 | CPU principles | Registers, PC, IR, pipeline, four stages | All four stages of fetch-decode-execute-writeback are simulated. |
| Module 5 | Storage system | Memory Hierarchy, Cache, Principle of Locality | Hard disk → memory → cache → registers, the full chain of four-level storage. |
| Module 6 | Bus and I/O | Address/data/control buses, polling/interrupt/DMA | DMA loads the program, the bus transfers data, cache line filling. |
The Complete Lifecycle of an Instruction
We use a simple pseudo-instructionADD A, B(add the values of register A and B, store the result back to A) as an example to trace its complete journey:
Journey Step 1: The program is stored on the hard disk
The program is initially stored on the hard disk as a file. The hard disk ispersistent storage, data is not lost after power-off, but the read/write speed is the slowest (millisecond level).
Step 2 of the journey: DMA loads the program into memory.
When you double-click to run the program, the operating system, throughDMAMove the program from the hard disk to memory. The DMA controller takes over the data movement work, and the CPU can do other things during this time.
Step 3 of the journey: the CPU fetches an instruction via the bus.
The CPU places the value of the program counter (PC) onto theAddress busOn, atControl busand issues a "read" signal. After receiving the signal, the memory places the instruction at the corresponding address onto theData bus, and the CPU reads it into the instruction register (IR).
Step 4 of the journey: decode — the controller parses the instruction.
The ADD A, B in the instruction register is decomposed into three parts: opcode ADD, destination operand A, source operand B. The controller decides the next operation based on the opcode.
Step 5 of the journey: execute — the ALU performs the operation.
The controller sends control signals, causing the ALU to take values from registers A and B, perform the addition operation, and temporarily store the result at the ALU output.
Journey Step 6: Write-back — the result is stored into registers via the data bus
The ALU calculation results pass throughData busWrite back to register A. At this point, one instruction has completed execution. The PC is automatically incremented by 1 to point to the next instruction.
Code demonstration: FullComputer complete simulation.
The FullComputer class below fully simulates the entire process described above. It contains over 300 lines of code, covering the knowledge points of all six modules in this course. Each step prints detailed logs.
Example
FullComputer - Complete von Neumann Computer Simulator (example Teaching Demo)
This program implements a complete (simplified) von Neumann computer, including:
- Four-tier storage: Disk → Memory → Cache → Register
- DMA controller: automatically moves programs from hard disk to memory
- Three types of buses: address bus, data bus, control bus.
- fourPhaseExecute: 取指 (Fetch) → translatecode (Decode) → Execute (Execute) → write回 (WriteBack)
- Statistics system: cache hit rate, instruction count, bus usage count, execution time
Supported instruction set:
LOAD reg value - Load immediate value into register
STORE reg addr - Store the register value into the memory address
ADD dst src - dst = dst + src
SUB dst src - dst = dst - src
MUL dst src - dst = dst * src
CMP a b - Compare two registers, set flags
JMP addr - unconditional jump
JZ addr - Jump if zero flag is true
How to run: python fullcomputer.py
"""
import time
# ============================
# 1. Storage System - Four-layer Structure
# ============================
class Disk:
"""
Hard disk - the lowest-level persistent storage.
Slowest speed (millisecond level), largest capacity, data is not lost on power outage.
"""
def __init__(self):
self.sectors = {} # Sector storage, simulating hard disk sectors
self.read_latency = 5 # Read latency (simulated time unit)
self.write_latency = 8 # Write latency
def load_program(self, program_name):
Load the program stored on the hard disk. Return None if the program does not exist.
if program_name in self.sectors:
print(f" [硬disk] 找toCoursesequence '{program_name}',Size: {len(self.sectors[program_name])} 条Directive")
return self.sectors[program_name]
print(f[Hard disk] Program '{program_name}' not found!)
return None
def store_program(self, name, program):
"""Store the program into the hard disk"""
self.sectors[name] = program
print(f" [硬disk] 已StorageCoursesequence '{name}' ({len(program)} 条Directive) to扇区 0x{hash(name) & 0xFFFF:04X}")
class Memory:
"""
Main memory - second-level storage.
Medium speed (nanosecond level), large capacity, data is lost when power is off.
CPU can only directly access data in memory (cannot directly access the hard disk).
"""
def __init__(self, size=65536):
self.size = size # Memory size (bytes)
self.cells = {} # Memory unit: key is address, value is data/instruction
self.read_count = 0 # Read operation count
self.write_count = 0 # Write operation count
def read(self, addr):
Read data from the specified address
self.read_count += 1
if addr not in self.cells:
print(f" Warning: address {addr} is uninitialized, returning 0")
return 0
return self.cells[addr]
def write(self, addr, data):
Write data to the specified address
self.write_count += 1
self.cells[addr] = data
def load_block(self, start_addr, data_block):
"""
Load an entire block of data into a contiguous address region of memory.
For DMA batch transfer scenarios.
"""
for i, item in enumerate(data_block):
self.cells[start_addr + i] = item
return len(data_block)
class Cache:
"""
Cache (L1 Cache Simulation) - Third-level storage.
Extremely fast (same frequency as CPU), smallest capacity.
Using a simple direct-mapped cache model:
- Each cache line stores: (memory address, data content)
- Cache hit: data already in cache, return directly
Cache miss: need to load from memory, replace a cache line.
"""
def __init__(self, num_lines=8):
self.num_lines = num_lines
self.lines = {} # Cache line: {cache_line_index: (mem_addr, data)}
self.hit_count = 0
self.miss_count = 0
def read(self, mem_addr, memory):
"""
Read from cache. If hit, return directly; if miss, load from memory.
Return: (data, was_hit)
"""
# Simple mapping: cache line index = memory address % cache line count
line_index = mem_addr % self.num_lines
if line_index in self.lines:
cached_addr, cached_data = self.lines[line_index]
if cached_addr == mem_addr:
# Cache hit
self.hit_count += 1
print(f" [Caching] hit!line {line_index}, Address {mem_addr}, Data: {cached_data}")
return cached_data, True
# Cache Miss
self.miss_count += 1
data = memory.read(mem_addr)
self.lines[line_index] = (mem_addr, data)
print(f" [Caching] miss, fromMemory address {mem_addr} addloadData {data} → Cachingline {line_index}")
return data, False
@property
def hit_rate(self):
total = self.hit_count + self.miss_count
return self.hit_count / total * 100 if total > 0 else 0.0
class RegisterFile:
"""
Register file - the topmost storage.
Fastest (one clock cycle), smallest capacity.
General-purpose registers + Special-purpose registers (PC, IR, MAR, MDR, FLAGS)
"""
def __init__(self):
# General-purpose registers
self.gpr = {'A': 0, 'B': 0, 'C': 0, 'D': 0}
# Special-purpose registers
self.PC = 0 # Program counter: stores the address of the next instruction
self.IR = None # Instruction register: stores the currently executing instruction
self.MAR = 0 # Memory address register: stores the memory address to be accessed
self.MDR = 0 # Memory Data Register: stores data read from / to be written to memory
# Flags register
self.FLAGS = {'Z': False} # Zero flag: whether the previous operation result is 0
def __repr__(self):
gpr_str = ', '.join(f'{k}={v}' for k, v in self.gpr.items())
return f"GPR[{gpr_str}] PC={self.PC} IR={self.IR} FLAGS={self.FLAGS}"
# ============================
# 2. Bus System - Three Buses
# ============================
class SystemBus:
"""
System bus - includes address bus, data bus, control bus.
Provides a unified interface for CPU, memory, and DMA.
"""
def __init__(self):
self.address_bus = 0 # Address bus: address currently being transmitted
self.data_bus = 0 # Data bus: data currently being transmitted
self.control_bus = 'IDLE' # Control Bus: IDLE, READ, WRITE, INTERRUPT, DMA_REQ, DMA_ACK
self.transfer_count = 0 # Bus Transfer Count
def cpu_read(self, address, memory, cache):
"""
The CPU reads data from memory via the bus (through the cache).
Complete bus interaction process:
1. The CPU places the address on the address bus
2. CPU issues READ signal on the control bus.
3. The memory (through the cache layer) returns data to the data bus
4. CPU reads data from the data bus
5. Clear the bus
"""
self.address_bus = address
self.control_bus = 'READ'
self.transfer_count += 1
print(f[Bus] address bus ← {address}, control bus ← READ)
# Memory responds through the cache layer
data, was_hit = cache.read(address, memory)
self.data_bus = data
print(f" [bus] Data bus ← {data} {'(Cache hit)' if was_hit else '(come自Memory)'}")
# CPU reads data
result = self.data_bus
# Bus clear
self.address_bus = 0
self.data_bus = 0
self.control_bus = 'IDLE'
print(f[Bus] Bus cleared, entering idle state)
return result
def cpu_write(self, address, data, memory):
"""
CPU writes data to memory through the bus.
"""
self.address_bus = address
self.data_bus = data
self.control_bus = 'WRITE'
self.transfer_count += 1
print(f[Bus] Address Bus ← {address}, Data Bus ← {data}, Control Bus ← WRITE)
memory.write(address, data)
print(f[Bus] Data {data} has been written to memory address {address})
# Bus clear
self.address_bus = 0
self.data_bus = 0
self.control_bus = 'IDLE'
print(f[Bus] Bus cleared, entering idle state)
# ============================
# 3. DMA Controller
# ============================
class DMAController:
"""
DMA controller - responsible for directly transferring data between the hard disk and memory.
The CPU only needs to tell the DMA: source address, destination address, transfer length.
After DMA completes the transfer, it notifies the CPU via an interrupt.
"""
def __init__(self, bus):
self.bus = bus
self.transfer_count = 0 # DMA Transfer Count
self.total_bytes = 0 # Total Bytes Transferred
def transfer_program(self, disk, program_name, memory, load_addr=0):
"""
Load the program from hard disk to memory via DMA.
This is a practical application of Lecture 20 'DMA Mode' in this course.
"""
print(f"\n {'='*50}")
print(fDMA transfer starts: hard disk → memory)
print(f" {'='*50}")
print(f[DMA] Source: hard disk sector (program '{program_name}'))
print(f[DMA] Target: memory address 0x{load_addr:04X} start)
print(f[DMA] Requesting bus control...)
# DMA occupies the bus
self.bus.control_bus = 'DMA_REQ'
print(f[DMA] obtains bus control (control bus ← DMA_REQ))
program = disk.load_program(program_name)
if program is None:
print(f[DMA] Error: program does not exist, transfer aborted)
self.bus.control_bus = 'IDLE'
return 0
# Instruction-by-Instruction Transfer
for i, instruction in enumerate(program):
addr = load_addr + i
memory.write(addr, instruction)
self.bus.address_bus = addr
self.bus.data_bus = instruction
self.total_bytes += 1
self.transfer_count += 1
print(f" [DMA] transmission #{i}: Directive '{instruction}' → Memory address {addr} (0x{addr:04X})")
# DMA releases the bus and sends a completion interrupt
self.bus.control_bus = 'DMA_ACK'
print(f" [DMA] transmission完Cheng!common {len(program)} 条Directive, {self.total_bytes} timesbustransmission")
print(f[DMA] Release bus control (control bus ← IDLE))
# Interrupt notifies the CPU
print(f" [DMA] Send interrupt signal to notify CPU: program loading complete")
self.bus.control_bus = 'INTERRUPT'
# Processing after CPU receives an interrupt
time.sleep(0.001)
print(f[CPU Interrupt Response] DMA completion interrupt received, program is ready)
self.bus.control_bus = 'IDLE'
return len(program)
def cache_prefetch(self, memory, cache, start_addr, count):
"""
DMA prefetch: preload data that will be used soon from memory into the cache.
Simulate the hardware prefetch mechanism of a modern CPU.
"""
print(f"\n[DMA Prefetch] Prefetch memory addresses {start_addr}~{start_addr+count-1} into cache...)
for i in range(count):
addr = start_addr + i
data = memory.read(addr)
line_index = addr % cache.num_lines
cache.lines[line_index] = (addr, data)
print(f[DMA Prefetch] Complete, {count} instructions written to cache)
# ============================
# 4. ALU - Arithmetic Logic Unit
# ============================
class ALU:
"""
Arithmetic logic unit - performs all arithmetic and logical operations.
Content corresponding to the second module of the course, such as adders and logic gates.
"""
@staticmethod
def add(a, b):
print(f[ALU] Performing addition: {a} + {b} = {a + b})
return a + b
@staticmethod
def sub(a, b):
print(f[ALU] Performing subtraction: {a} - {b} = {a - b})
return a - b
@staticmethod
def mul(a, b):
print(f[ALU] Performing multiplication: {a} * {b} = {a * b})
return a * b
@staticmethod
def cmp(a, b):
result = (a == b)
print(f" [ALU] ExecuteCompare: {a} == {b} ? {'is' if result else 'no'} (Z={1 if result else 0})")
return result
# ============================
# 5. FullComputer - Complete Computer
# ============================
class FullComputer:
"""
Complete von Neumann computer simulator.
Components:
- disk: hard disk (persistent storage)
- memory: main memory
- cache: L1 cache
- registers: register group (including PC, IR, MAR, MDR, FLAGS)
- alu: arithmetic logic unit
- bus: system bus (address + data + control)
- dma: DMA controller
Execution flow:
1. Loading: DMA transfers the program from the hard disk to memory.
2. Instruction cycle: fetch → decode → execute → write back.
3. Generate complete execution report
"""
def __init__(self):
print("=" * 60)
print(FullComputer - von Neumann computer initializing...)
print("=" * 60)
# Four-Tier Storage System
self.disk = Disk()
self.memory = Memory()
self.cache = Cache(num_lines=8)
self.registers = RegisterFile()
# Computing unit
self.alu = ALU()
# Bus system
self.bus = SystemBus()
# DMA Controller
self.dma = DMAController(self.bus)
# Statistics
self.instructions_executed = 0
self.clock_cycles = 0
print(Components ready: hard disk, memory (64KB), cache (8-line L1), 4 general-purpose registers)
print(Components ready: ALU, system bus, DMA controller)
print()
def store_program(self, name, program):
"""Store the program into the hard disk"""
print(f[Step 0] Store the program '{name}' to the hard disk)
self.disk.store_program(name, program)
print()
def load_program(self, program_name, load_addr=0):
"""
Load program: DMA transfers from hard disk to memory.
"""
print(f[Step 1] DMA loads program into memory)
size = self.dma.transfer_program(
self.disk, program_name, self.memory, load_addr
)
print()
# Prefetch instructions to cache
print(f[Step 2] Cache warm-up: prefetch first {min(4, size)} instructions)
self.dma.cache_prefetch(self.memory, self.cache, load_addr, min(4, size))
print()
# Set PC to program start address
self.registers.PC = load_addr
print(f[Step 3] Set PC ← {load_addr} (program start address))
print()
def fetch(self):
"""
Stage 1: Fetch
- Put the PC value into MAR
- Send the address via the address bus
- Send READ signal via the control bus
- Fetch instruction from the data bus (via the cache layer)
- Store instruction into IR
- PC increment
"""
print(f"--- Stage 1: Instruction Fetch (Fetch) ---")
self.clock_cycles += 1
# MAR ← PC
self.registers.MAR = self.registers.PC
print(f" MAR ← PC = {self.registers.PC}")
# Read memory through the bus (via cache)
instruction = self.bus.cpu_read(
self.registers.MAR, self.memory, self.cache
)
# MDR ← data bus, IR ← MDR
self.registers.MDR = instruction
self.registers.IR = instruction
print(f" MDR ← {instruction}, IR ← MDR")
# PC increment
old_pc = self.registers.PC
self.registers.PC += 1
print(f" PC ← {old_pc} + 1 = {self.registers.PC}")
return instruction
def decode(self):
"""
Stage 2: Decode
- Controller decodes instruction in IR
- Separate opcode and operand
- Determine the operations of the execution stage
"""
print(f--- Stage 2: Decode ---)
self.clock_cycles += 1
instruction = self.registers.IR
parts = instruction.split()
opcode = parts[0]
operands = parts[1:] if len(parts) > 1 else []
print(f" Instruction: '{instruction}'")
print(fOpcode: {opcode}, operands: {operands})
return opcode, operands
def execute(self, opcode, operands):
"""
Stage 3: Execute
- Execute the corresponding operation based on the opcode
- ALU participates in arithmetic/logic operations
- May modify registers or set flag bits
"""
print(f--- Phase 3: Execute ---)
self.clock_cycles += 1
if opcode == 'LOAD':
# LOAD reg value: Load the immediate value into the register
reg, value = operands[0], int(operands[1])
self.registers.gpr[reg] = value
print(fOperation: load immediate {value} → register {reg})
print(f" ALU does not participate (direct load)")
elif opcode == 'STORE':
# STORE reg addr: Store register value into memory
reg, addr = operands[0], int(operands[1])
value = self.registers.gpr[reg]
print(fOperation: preparing to store the value {value} of register {reg} into memory address {addr})
# Data is ready; the write-back stage performs the actual store operation.
return ('STORE', reg, addr, value)
elif opcode == 'ADD':
# ADD dst src: dst = dst + src
dst, src = operands[0], operands[1]
old_val = self.registers.gpr[dst]
result = self.alu.add(old_val, self.registers.gpr[src])
self.registers.gpr[dst] = result
# Set zero flag
self.registers.FLAGS['Z'] = (result == 0)
print(f" Operation: {dst} ← {old_val} + {self.registers.gpr.get(src, src)} = {result}")
print(fZero flag Z ← {self.registers.FLAGS['Z']})
elif opcode == 'SUB':
# SUB dst src: dst = dst - src
dst, src = operands[0], operands[1]
old_val = self.registers.gpr[dst]
result = self.alu.sub(old_val, self.registers.gpr[src])
self.registers.gpr[dst] = result
self.registers.FLAGS['Z'] = (result == 0)
print(f" Operation: {dst} ← {old_val} - {self.registers.gpr.get(src, src)} = {result}")
elif opcode == 'MUL':
# MUL dst src: dst = dst * src
dst, src = operands[0], operands[1]
old_val = self.registers.gpr[dst]
result = self.alu.mul(old_val, self.registers.gpr[src])
self.registers.gpr[dst] = result
self.registers.FLAGS['Z'] = (result == 0)
print(f" Operation: {dst} ← {old_val} * {self.registers.gpr.get(src, src)} = {result}")
elif opcode == 'CMP':
# CMP a b: compare two register values, set zero flag
a, b = operands[0], operands[1]
is_equal = self.alu.cmp(
self.registers.gpr[a],
self.registers.gpr[b]
)
self.registers.FLAGS['Z'] = is_equal
elif opcode == 'JMP':
# JMP addr: Unconditional jump
target = int(operands[0])
print(fOperation: unconditional jump to address {target})
old_pc = self.registers.PC
self.registers.PC = target
print(fPC ← {old_pc} → {target} (jump))
elif opcode == 'JZ':
# JZ addr: jump if the zero flag is true
target = int(operands[0])
if self.registers.FLAGS['Z']:
print(fOperation: if zero flag is true (Z=1), jump to address {target})
self.registers.PC = target
else:
print(fOperation: zero flag is false (Z=0), no jump, sequential execution)
elif opcode == 'HALT':
# HALT: Halt
print(fOperation: Halt instruction, program ends)
return 'HALT'
else:
print(f" Unknown opcode: {opcode}")
return None
def writeback(self):
"""
Stage 4: WriteBack
- If there are operations that need to write back to memory (such as STORE), they are completed in this stage
- Transfer data via the data bus
"""
print(f"--- Stage 4: WriteBack ---")
self.clock_cycles += 1
# The result of a general instruction has already been written to the register during the execution stage
# STORE instruction needs to write memory at this stage
print(fRegister state: {self.registers.gpr})
print(fFlags register: {self.registers.FLAGS})
def writeback_store(self, addr, value):
"""Write-back operation of the STORE instruction — write to memory through the bus"""
self.bus.cpu_write(addr, value, self.memory)
print(fSTORE complete: {value} → memory address {addr})
def run(self):
"""
Run the loaded program.
Loop through fetch → decode → execute → write back until the program ends.
"""
print("=" * 60)
print(Program starts execution)
print("=" * 60)
while self.registers.PC < max(self.memory.cells.keys(), default=0) + 1:
# Check whether PC is in valid range
if self.registers.PC not in self.memory.cells:
# PC exceeds program range, possibly due to a jump to the end, terminate
print(f"\nProgram counter PC={self.registers.PC} exceeds the program range, execution ends)
break
self.instructions_executed += 1
print(f"\n{'#' * 55}")
print(f"### Directive #{self.instructions_executed} | PC = {self.registers.PC}")
print(f"{'#' * 55}")
# Instruction fetch
instruction = self.fetch()
# Decode
opcode, operands = self.decode()
# Execute
special = self.execute(opcode, operands)
# Handle the HALT instruction
if special == 'HALT':
self.clock_cycles += 1
break
# Write back
if special and special[0] == 'STORE':
_, reg, addr, value = special
self.writeback_store(int(addr), value)
else:
self.writeback()
# Safety check: prevent infinite loops
if self.instructions_executed > 1000:
print("\nWarning: Over 1000 instructions executed, forced stop (possible infinite loop))
break
print(f"\n{'=' * 60}")
print(f" Program execution completed")
print(f"{'=' * 60}")
def report(self):
"""
Generate a complete execution report.
Summarize all statistics: cache hit rate, instruction count, clock cycles,
Bus transfer count, memory read/write count, etc.
"""
print(f"\n{'=' * 65}")
print(fFullComputer execution report —— EXAMPLE)
print(f"{'=' * 65}")
print()
# Basic information
print(f" {'─' * 55}")
print(f" Basic Information")
print(f" {'─' * 55}")
print(fTotal instructions executed: {self.instructions_executed:>6d})
print(fTotal clock cycles: {self.clock_cycles:>6d} cycles)
if self.instructions_executed > 0:
cpi = self.clock_cycles / self.instructions_executed
print(fAverage CPI (cycles/instruction): {cpi:>6.2f})
print()
# Cache performance
print(f" {'─' * 55}")
print(f"Cache performance (L1 Cache)")
print(f" {'─' * 55}")
total_access = self.cache.hit_count + self.cache.miss_count
print(f" Cache hits: {self.cache.hit_count:>6d} times")
print(f" Cache miss: {self.cache.miss_count:>6d} times")
print(fTotal accesses: {total_access:>6d} times)
print(fHit rate: {self.cache.hit_rate:>6.1f}%)
print(fTotal cache lines: {self.cache.num_lines:>6d} lines)
print()
# Memory access
print(f" {'─' * 55}")
print(f" Memory access statistics")
print(f" {'─' * 55}")
print(fMemory read operations: {self.memory.read_count:>6d} times)
print(fMemory write operation: {self.memory.write_count:>6d} times)
print(f" MemorytotalVisit: {self.memory.read_count + self.memory.write_count:>6d} times")
print(fNumber of used memory addresses: {len(self.memory.cells):>6d})
print()
# Bus statistics
print(f" {'─' * 55}")
print(f" Bus Statistics")
print(f" {'─' * 55}")
print(f" CPU emit起ofbustransmission: {self.bus.transfer_count:>6d} times")
print(f" DMA emit起ofbustransmission: {self.dma.transfer_count:>6d} times")
total_bus = self.bus.transfer_count + self.dma.transfer_count
print(fTotal bus transfer count: {total_bus:>6d} times)
print()
# Final register state
print(f" {'─' * 55}")
print(f" Register final state")
print(f" {'─' * 55}")
print(fGeneral-purpose registers: {self.registers.gpr})
print(fProgram counter PC: {self.registers.PC})
print(fInstruction register IR: {self.registers.IR})
print(fFlag register FLAGS: {self.registers.FLAGS})
print()
# Memory data dump (only shows data at valid addresses)
print(f" {'─' * 55}")
print(f" Memory data dump (only shows program area)")
print(f" {'─' * 55}")
if self.memory.cells:
for addr in sorted(self.memory.cells.keys()):
val = self.memory.cells[addr]
print(fAddress {addr:>4d} (0x{addr:04X}): {val})
else:
print(f" (Memory is empty)")
print()
print(f"{'=' * 65}")
print(fEnd of report. Thank you for using the EXAMPLE FullComputer simulator!)
print(f"{'=' * 65}")
# ============================
# 6. Main Program - End-to-End Demo
# ============================
if __name__ == "__main__":
print()
print("╔" + "═" * 58 + "╗")
print("║" + The complete journey of one instruction — EXAMPLE FullComputer Demo.center(52) + "║")
print("╚" + "═" * 58 + "╝")
print()
# Create computer
computer = FullComputer()
# Define demo program: calculate (10 + 7) * 2 - 3 = ?
# Use LOAD to load the initial value, and ADD/MUL/SUB to perform the calculation
demo_program = [
"LOAD A 10", # A = 10
"LOAD B 7", # B = 7
"ADD A B", # A = A + B = 17
"LOAD C 2", # C = 2
"MUL A C", # A = A * C = 34
"LOAD D 3", # D = 3
"SUB A D", # A = A - D = 31
"STORE A 100", # Store result in memory address 100
"LOAD B 31", # B = 31 (expected result)
"CMP A B", # Compare A and B -> should be equal
"JZ 12", # If equal (Z=1), jump to HALT
"LOAD D 0", # (not executed) D = 0
"HALT", # Halt
]
# Step 1: Store the program on the hard disk
computer.store_program("EXAMPLE_DEMO", demo_program)
# Step 2: DMA load into memory
computer.load_program("EXAMPLE_DEMO", load_addr=0)
# Step 3: Run the program
computer.run()
# Step 4: Print full execution report
computer.report()
print(f"\n")
print(fReview: This demo covers all six modules of Computer Organization Principles:)
print(fModule 1 (Binary): Instructions and data are stored in memory in binary encoding.)
print(fModule 2 (Logic Gates): The ADD/SUB/MUL in the execution stage is completed by ALU gate circuits)
print(fModule 3 (von Neumann): program and data share memory, five major components work together)
print(fModule 4 (CPU Principles): PC fetch → IR decode → ALU execute → write back, complete four-stage pipeline.)
print(fModule 5 (Storage System): hard disk → memory → cache → registers, four-level storage hierarchy)
print(fModule 6 (Bus and I/O): Three-bus interaction, DMA program loading, interrupt notification)
print()
Interpreting the execution report
After running the above code, FullComputer outputs a detailed execution report. Let's interpret the key metrics:
cache hit rate
The cache hit rate shown in the report reflects the program'stemporal localityandspatial locality. Since our demo program is very small and the DMA prefetches the first 4 instructions, the hit rate will be relatively high. In actual large programs, the cache hit rate is usually above 90% — this is exactly the purpose of cache.
CPI (Clock cycles Per Instruction)
In our simplified simulation, each instruction takes a fixed 4 clock cycles (fetch-decode-execute-writeback), so the CPI is about 4. Modern CPUs use pipelining to reduce the CPI to close to 1 — that is, completing one instruction per clock cycle.
Number of bus transfers
Bus transfer count = transfers initiated by the CPU + transfers initiated by the DMA. The DMA completes a large number of transfers during the program loading phase, and the CPU initiates at least one read transfer during the instruction fetch stage of each instruction.
What we have learned — review of the six-week course.
Six weeks ago, we started from the simplest question: "Why do computers only understand 0 and 1?" Now, we are able to build a complete computer simulator from scratch. Let's walk through this learning path again:
| week number | Module | Core takeaways |
|---|---|---|
| Week 1 | Binary and Information Representation | Understood why computers use binary, how to perform base conversion, and how various information (numbers, text, images) is encoded into 0s and 1s. |
| Week 2 | Logic Gates and Digital Circuits | Mastered the three basic gate circuits: AND, OR, NOT; understood how to build adders and ALUs using them. |
| Week 3 | Von Neumann architecture | Understood the division of labor among the five major components (arithmetic unit, controller, memory, input, output) and the stored-program concept. |
| Week 4 | CPU principles | Mastered the four stages of the instruction cycle: fetch-decode-execute-writeback, and how pipelining improves efficiency. |
| Week 5 | Storage system | Understood the four-level storage pyramid from registers to hard disk, and how cache uses the principle of locality to speed up access. |
| Week 6 | Bus and I/O | Mastered the cooperative work of the three buses (address/data/control) and the efficiency differences among the three I/O methods (polling/interrupt/DMA). |
From theory to practice: where to go next.
This course has laid a solid foundation for you in computer organization principles. If you are interested in deeper learning, here are some recommended directions:
- Operating system: Learn how operating systems manage the CPU, memory, file systems, and I/O devices. The "interrupts," "DMA," and "memory hierarchy" learned in this course are core concepts of operating systems.
- Assembly language: Interact directly with CPU instructions. After completing this course, assembly language is no longer a "sealed book" — you know how each assembly instruction is executed at the hardware level.
- nand2tetris: A famous online course that takes you from a single NAND gate to gradually building a complete programmable computer. The content of this course is highly complementary to it.
- Computer architecture: Delve into advanced CPU design techniques such as pipelining, superscalar execution, out-of-order execution, and branch prediction.
- embedded system: Apply the knowledge from this course to microcontrollers (such as Arduino, STM32), writing programs that directly manipulate registers and peripherals.
Summary and final verification
One-sentence summary:Computemachineis not魔法——It isonelayeronelayer精心DesignofAbstraction, from晶body管openclosestart,逐步搭build出逻辑门、ALU、CPU、MemorySystem、bus,最终Compositionone台abilityExecuteAnyCoursesequenceof通usemachinedevice.UnderstandingTheseLevel,youJusttruecorrectUnderstandingalreadyComputemachine。
Comprehensive course self-test questions
- In the FullComputer simulator, during the execution of a LOAD instruction, in which steps are the address bus, data bus, and control bus used respectively?
- If FullComputer's cache has only 4 lines (instead of 8), what impact does it have on the program's execution efficiency? What is this phenomenon called in computer architecture?
- In what scenarios do modern operating systems use polling, interrupts, and DMA? Please give one example for each that you can observe when using a computer daily.
Reference answer: 1. Fetch stage: The address bus transmits the PC value (instruction address), the control bus sends a READ signal, and the data bus returns the instruction content. 2. More cache conflicts, lower hit rate (called cache thrashing or conflict misses). 3. Polling: progress check during firmware updates on some simple embedded devices; Interrupt: every keyboard key press triggers an interrupt; DMA: large file copying, graphics card rendering framebuffer transfer.
other extensions