The Complete Journey of an Instruction

From the binary basics in Lecture 1 to the I/O system in the previous lecture, we have covered all the core content of computer organization principles. This lecture will string these knowledge points together into a single line, tracing the complete journey of an instruction from birth to completion.

We will implement a completevon Neumann Computer Simulator (FullComputer), it includes: four-level storage structure (hard disk, memory, cache, registers), DMA data transfer, four-stage pipeline of fetch-decode-execute-writeback, bus interaction, cache hit rate statistics. Finally, print a complete execution report.


Review: Knowledge Map of Six Modules

Before starting the end-to-end simulation, let's quickly review the core concepts of each of the six modules in this course:

ModuletopicCore ConceptsEmbodiment in this lecture
Module 1Binary and Information Representation0/1 encoding, base conversion, two's complement, floating-point numbersBoth instructions and data are stored in binary form in memory.
Module 2Logic Gates and Digital CircuitsAND/OR/NOT gates, Adders, ALUThe ALU operations in the execution stage are implemented by gate circuits.
Module 3Von Neumann architectureFive Major Components, Stored Program, Instruction CycleWhat this lecture simulates is a complete von Neumann machine.
Module 4CPU principlesRegisters, PC, IR, pipeline, four stagesAll four stages of fetch-decode-execute-writeback are simulated.
Module 5Storage systemMemory Hierarchy, Cache, Principle of LocalityHard disk → memory → cache → registers, the full chain of four-level storage.
Module 6Bus and I/OAddress/data/control buses, polling/interrupt/DMADMA loads the program, the bus transfers data, cache line filling.

The Complete Lifecycle of an Instruction

We use a simple pseudo-instructionADD A, B(add the values of register A and B, store the result back to A) as an example to trace its complete journey:

Journey Step 1: The program is stored on the hard disk

The program is initially stored on the hard disk as a file. The hard disk ispersistent storage, data is not lost after power-off, but the read/write speed is the slowest (millisecond level).

Step 2 of the journey: DMA loads the program into memory.

When you double-click to run the program, the operating system, throughDMAMove the program from the hard disk to memory. The DMA controller takes over the data movement work, and the CPU can do other things during this time.

Step 3 of the journey: the CPU fetches an instruction via the bus.

The CPU places the value of the program counter (PC) onto theAddress busOn, atControl busand issues a "read" signal. After receiving the signal, the memory places the instruction at the corresponding address onto theData bus, and the CPU reads it into the instruction register (IR).

Step 4 of the journey: decode — the controller parses the instruction.

The ADD A, B in the instruction register is decomposed into three parts: opcode ADD, destination operand A, source operand B. The controller decides the next operation based on the opcode.

Step 5 of the journey: execute — the ALU performs the operation.

The controller sends control signals, causing the ALU to take values from registers A and B, perform the addition operation, and temporarily store the result at the ALU output.

Journey Step 6: Write-back — the result is stored into registers via the data bus

The ALU calculation results pass throughData busWrite back to register A. At this point, one instruction has completed execution. The PC is automatically incremented by 1 to point to the next instruction.


Code demonstration: FullComputer complete simulation.

The FullComputer class below fully simulates the entire process described above. It contains over 300 lines of code, covering the knowledge points of all six modules in this course. Each step prints detailed logs.

Example

"""
FullComputer - Complete von Neumann Computer Simulator (example Teaching Demo)

This program implements a complete (simplified) von Neumann computer, including:
- Four-tier storage: Disk → Memory → Cache → Register
- DMA controller: automatically moves programs from hard disk to memory
- Three types of buses: address bus, data bus, control bus.
- fourPhaseExecute: 取指 (Fetch) → translatecode (Decode) → Execute (Execute) → write回 (WriteBack)
- Statistics system: cache hit rate, instruction count, bus usage count, execution time

Supported instruction set:
LOAD reg value - Load immediate value into register
STORE reg addr - Store the register value into the memory address
  ADD dst src      - dst = dst + src
  SUB dst src      - dst = dst - src
  MUL dst src      - dst = dst * src
CMP a b - Compare two registers, set flags
JMP addr - unconditional jump
JZ addr - Jump if zero flag is true

How to run: python fullcomputer.py
"""


import time


# ============================
# 1. Storage System - Four-layer Structure
# ============================

class Disk:
    """
Hard disk - the lowest-level persistent storage.
Slowest speed (millisecond level), largest capacity, data is not lost on power outage.
    """

    def __init__(self):
        self.sectors = {}          # Sector storage, simulating hard disk sectors
        self.read_latency = 5      # Read latency (simulated time unit)
        self.write_latency = 8     # Write latency

    def load_program(self, program_name):
        Load the program stored on the hard disk. Return None if the program does not exist.
        if program_name in self.sectors:
            print(f"  [硬disk] 找toCoursesequence '{program_name}',Size: {len(self.sectors[program_name])} 条Directive")
            return self.sectors[program_name]
        print(f[Hard disk] Program '{program_name}' not found!)
        return None

    def store_program(self, name, program):
        """Store the program into the hard disk"""
        self.sectors[name] = program
        print(f"  [硬disk] 已StorageCoursesequence '{name}' ({len(program)} 条Directive) to扇区 0x{hash(name) & 0xFFFF:04X}")


class Memory:
    """
Main memory - second-level storage.
Medium speed (nanosecond level), large capacity, data is lost when power is off.
CPU can only directly access data in memory (cannot directly access the hard disk).
    """

    def __init__(self, size=65536):
        self.size = size           # Memory size (bytes)
        self.cells = {}            # Memory unit: key is address, value is data/instruction
        self.read_count = 0        # Read operation count
        self.write_count = 0       # Write operation count

    def read(self, addr):
        Read data from the specified address
        self.read_count += 1
        if addr not in self.cells:
            print(f" Warning: address {addr} is uninitialized, returning 0")
            return 0
        return self.cells[addr]

    def write(self, addr, data):
        Write data to the specified address
        self.write_count += 1
        self.cells[addr] = data

    def load_block(self, start_addr, data_block):
        """
Load an entire block of data into a contiguous address region of memory.
For DMA batch transfer scenarios.
        """

        for i, item in enumerate(data_block):
            self.cells[start_addr + i] = item
        return len(data_block)


class Cache:
    """
Cache (L1 Cache Simulation) - Third-level storage.
Extremely fast (same frequency as CPU), smallest capacity.

Using a simple direct-mapped cache model:
- Each cache line stores: (memory address, data content)
- Cache hit: data already in cache, return directly
Cache miss: need to load from memory, replace a cache line.
    """

    def __init__(self, num_lines=8):
        self.num_lines = num_lines
        self.lines = {}            # Cache line: {cache_line_index: (mem_addr, data)}
        self.hit_count = 0
        self.miss_count = 0

    def read(self, mem_addr, memory):
        """
Read from cache. If hit, return directly; if miss, load from memory.
Return: (data, was_hit)
        """

        # Simple mapping: cache line index = memory address % cache line count
        line_index = mem_addr % self.num_lines

        if line_index in self.lines:
            cached_addr, cached_data = self.lines[line_index]
            if cached_addr == mem_addr:
                # Cache hit
                self.hit_count += 1
                print(f"    [Caching] hit!line {line_index}, Address {mem_addr}, Data: {cached_data}")
                return cached_data, True

        # Cache Miss
        self.miss_count += 1
        data = memory.read(mem_addr)
        self.lines[line_index] = (mem_addr, data)
        print(f"    [Caching] miss, fromMemory address {mem_addr} addloadData {data} → Cachingline {line_index}")
        return data, False

    @property
    def hit_rate(self):
        total = self.hit_count + self.miss_count
        return self.hit_count / total * 100 if total > 0 else 0.0


class RegisterFile:
    """
Register file - the topmost storage.
Fastest (one clock cycle), smallest capacity.
General-purpose registers + Special-purpose registers (PC, IR, MAR, MDR, FLAGS)
    """

    def __init__(self):
        # General-purpose registers
        self.gpr = {'A': 0, 'B': 0, 'C': 0, 'D': 0}

        # Special-purpose registers
        self.PC = 0                # Program counter: stores the address of the next instruction
        self.IR = None             # Instruction register: stores the currently executing instruction
        self.MAR = 0               # Memory address register: stores the memory address to be accessed
        self.MDR = 0               # Memory Data Register: stores data read from / to be written to memory

        # Flags register
        self.FLAGS = {'Z': False}  # Zero flag: whether the previous operation result is 0

    def __repr__(self):
        gpr_str = ', '.join(f'{k}={v}' for k, v in self.gpr.items())
        return f"GPR[{gpr_str}] PC={self.PC} IR={self.IR} FLAGS={self.FLAGS}"


# ============================
# 2. Bus System - Three Buses
# ============================

class SystemBus:
    """
System bus - includes address bus, data bus, control bus.
Provides a unified interface for CPU, memory, and DMA.
    """

    def __init__(self):
        self.address_bus = 0       # Address bus: address currently being transmitted
        self.data_bus = 0          # Data bus: data currently being transmitted
        self.control_bus = 'IDLE'  # Control Bus: IDLE, READ, WRITE, INTERRUPT, DMA_REQ, DMA_ACK
        self.transfer_count = 0    # Bus Transfer Count

    def cpu_read(self, address, memory, cache):
        """
The CPU reads data from memory via the bus (through the cache).
Complete bus interaction process:
1. The CPU places the address on the address bus
2. CPU issues READ signal on the control bus.
3. The memory (through the cache layer) returns data to the data bus
4. CPU reads data from the data bus
5. Clear the bus
        """

        self.address_bus = address
        self.control_bus = 'READ'
        self.transfer_count += 1
        print(f[Bus] address bus ← {address}, control bus ← READ)

        # Memory responds through the cache layer
        data, was_hit = cache.read(address, memory)
        self.data_bus = data
        print(f"  [bus] Data bus ← {data} {'(Cache hit)' if was_hit else '(come自Memory)'}")

        # CPU reads data
        result = self.data_bus

        # Bus clear
        self.address_bus = 0
        self.data_bus = 0
        self.control_bus = 'IDLE'
        print(f[Bus] Bus cleared, entering idle state)

        return result

    def cpu_write(self, address, data, memory):
        """
CPU writes data to memory through the bus.
        """

        self.address_bus = address
        self.data_bus = data
        self.control_bus = 'WRITE'
        self.transfer_count += 1
        print(f[Bus] Address Bus ← {address}, Data Bus ← {data}, Control Bus ← WRITE)

        memory.write(address, data)
        print(f[Bus] Data {data} has been written to memory address {address})

        # Bus clear
        self.address_bus = 0
        self.data_bus = 0
        self.control_bus = 'IDLE'
        print(f[Bus] Bus cleared, entering idle state)


# ============================
# 3. DMA Controller
# ============================

class DMAController:
    """
DMA controller - responsible for directly transferring data between the hard disk and memory.
The CPU only needs to tell the DMA: source address, destination address, transfer length.
After DMA completes the transfer, it notifies the CPU via an interrupt.
    """

    def __init__(self, bus):
        self.bus = bus
        self.transfer_count = 0    # DMA Transfer Count
        self.total_bytes = 0       # Total Bytes Transferred

    def transfer_program(self, disk, program_name, memory, load_addr=0):
        """
Load the program from hard disk to memory via DMA.
This is a practical application of Lecture 20 'DMA Mode' in this course.
        """

        print(f"\n  {'='*50}")
        print(fDMA transfer starts: hard disk → memory)
        print(f"  {'='*50}")
        print(f[DMA] Source: hard disk sector (program '{program_name}'))
        print(f[DMA] Target: memory address 0x{load_addr:04X} start)
        print(f[DMA] Requesting bus control...)

        # DMA occupies the bus
        self.bus.control_bus = 'DMA_REQ'
        print(f[DMA] obtains bus control (control bus ← DMA_REQ))

        program = disk.load_program(program_name)
        if program is None:
            print(f[DMA] Error: program does not exist, transfer aborted)
            self.bus.control_bus = 'IDLE'
            return 0

        # Instruction-by-Instruction Transfer
        for i, instruction in enumerate(program):
            addr = load_addr + i
            memory.write(addr, instruction)
            self.bus.address_bus = addr
            self.bus.data_bus = instruction
            self.total_bytes += 1
            self.transfer_count += 1
            print(f"  [DMA] transmission #{i}: Directive '{instruction}' → Memory address {addr} (0x{addr:04X})")

        # DMA releases the bus and sends a completion interrupt
        self.bus.control_bus = 'DMA_ACK'
        print(f"  [DMA] transmission完Cheng!common {len(program)} 条Directive, {self.total_bytes} timesbustransmission")
        print(f[DMA] Release bus control (control bus ← IDLE))

        # Interrupt notifies the CPU
        print(f" [DMA] Send interrupt signal to notify CPU: program loading complete")
        self.bus.control_bus = 'INTERRUPT'

        # Processing after CPU receives an interrupt
        time.sleep(0.001)
        print(f[CPU Interrupt Response] DMA completion interrupt received, program is ready)
        self.bus.control_bus = 'IDLE'

        return len(program)

    def cache_prefetch(self, memory, cache, start_addr, count):
        """
DMA prefetch: preload data that will be used soon from memory into the cache.
Simulate the hardware prefetch mechanism of a modern CPU.
        """

        print(f"\n[DMA Prefetch] Prefetch memory addresses {start_addr}~{start_addr+count-1} into cache...)
        for i in range(count):
            addr = start_addr + i
            data = memory.read(addr)
            line_index = addr % cache.num_lines
            cache.lines[line_index] = (addr, data)
        print(f[DMA Prefetch] Complete, {count} instructions written to cache)


# ============================
# 4. ALU - Arithmetic Logic Unit
# ============================

class ALU:
    """
Arithmetic logic unit - performs all arithmetic and logical operations.
Content corresponding to the second module of the course, such as adders and logic gates.
    """

    @staticmethod
    def add(a, b):
        print(f[ALU] Performing addition: {a} + {b} = {a + b})
        return a + b

    @staticmethod
    def sub(a, b):
        print(f[ALU] Performing subtraction: {a} - {b} = {a - b})
        return a - b

    @staticmethod
    def mul(a, b):
        print(f[ALU] Performing multiplication: {a} * {b} = {a * b})
        return a * b

    @staticmethod
    def cmp(a, b):
        result = (a == b)
        print(f"    [ALU] ExecuteCompare: {a} == {b} ? {'is' if result else 'no'} (Z={1 if result else 0})")
        return result


# ============================
# 5. FullComputer - Complete Computer
# ============================

class FullComputer:
    """
Complete von Neumann computer simulator.

Components:
- disk: hard disk (persistent storage)
- memory: main memory
- cache: L1 cache
- registers: register group (including PC, IR, MAR, MDR, FLAGS)
- alu: arithmetic logic unit
- bus: system bus (address + data + control)
- dma: DMA controller

Execution flow:
1. Loading: DMA transfers the program from the hard disk to memory.
2. Instruction cycle: fetch → decode → execute → write back.
3. Generate complete execution report
    """


    def __init__(self):
        print("=" * 60)
        print(FullComputer - von Neumann computer initializing...)
        print("=" * 60)

        # Four-Tier Storage System
        self.disk = Disk()
        self.memory = Memory()
        self.cache = Cache(num_lines=8)
        self.registers = RegisterFile()

        # Computing unit
        self.alu = ALU()

        # Bus system
        self.bus = SystemBus()

        # DMA Controller
        self.dma = DMAController(self.bus)

        # Statistics
        self.instructions_executed = 0
        self.clock_cycles = 0

        print(Components ready: hard disk, memory (64KB), cache (8-line L1), 4 general-purpose registers)
        print(Components ready: ALU, system bus, DMA controller)
        print()

    def store_program(self, name, program):
        """Store the program into the hard disk"""
        print(f[Step 0] Store the program '{name}' to the hard disk)
        self.disk.store_program(name, program)
        print()

    def load_program(self, program_name, load_addr=0):
        """
Load program: DMA transfers from hard disk to memory.
        """

        print(f[Step 1] DMA loads program into memory)
        size = self.dma.transfer_program(
            self.disk, program_name, self.memory, load_addr
        )
        print()

        # Prefetch instructions to cache
        print(f[Step 2] Cache warm-up: prefetch first {min(4, size)} instructions)
        self.dma.cache_prefetch(self.memory, self.cache, load_addr, min(4, size))
        print()

        # Set PC to program start address
        self.registers.PC = load_addr
        print(f[Step 3] Set PC ← {load_addr} (program start address))
        print()

    def fetch(self):
        """
Stage 1: Fetch
- Put the PC value into MAR
- Send the address via the address bus
- Send READ signal via the control bus
- Fetch instruction from the data bus (via the cache layer)
- Store instruction into IR
- PC increment
        """

        print(f"--- Stage 1: Instruction Fetch (Fetch) ---")
        self.clock_cycles += 1

        # MAR ← PC
        self.registers.MAR = self.registers.PC
        print(f"  MAR ← PC = {self.registers.PC}")

        # Read memory through the bus (via cache)
        instruction = self.bus.cpu_read(
            self.registers.MAR, self.memory, self.cache
        )

        # MDR ← data bus, IR ← MDR
        self.registers.MDR = instruction
        self.registers.IR = instruction
        print(f"  MDR ← {instruction}, IR ← MDR")

        # PC increment
        old_pc = self.registers.PC
        self.registers.PC += 1
        print(f"  PC ← {old_pc} + 1 = {self.registers.PC}")

        return instruction

    def decode(self):
        """
Stage 2: Decode
- Controller decodes instruction in IR
- Separate opcode and operand
- Determine the operations of the execution stage
        """

        print(f--- Stage 2: Decode ---)
        self.clock_cycles += 1

        instruction = self.registers.IR
        parts = instruction.split()
        opcode = parts[0]
        operands = parts[1:] if len(parts) > 1 else []

        print(f" Instruction: '{instruction}'")
        print(fOpcode: {opcode}, operands: {operands})

        return opcode, operands

    def execute(self, opcode, operands):
        """
Stage 3: Execute
- Execute the corresponding operation based on the opcode
- ALU participates in arithmetic/logic operations
- May modify registers or set flag bits
        """

        print(f--- Phase 3: Execute ---)
        self.clock_cycles += 1

        if opcode == 'LOAD':
            # LOAD reg value: Load the immediate value into the register
            reg, value = operands[0], int(operands[1])
            self.registers.gpr[reg] = value
            print(fOperation: load immediate {value} → register {reg})
            print(f" ALU does not participate (direct load)")

        elif opcode == 'STORE':
            # STORE reg addr: Store register value into memory
            reg, addr = operands[0], int(operands[1])
            value = self.registers.gpr[reg]
            print(fOperation: preparing to store the value {value} of register {reg} into memory address {addr})
            # Data is ready; the write-back stage performs the actual store operation.
            return ('STORE', reg, addr, value)

        elif opcode == 'ADD':
            # ADD dst src: dst = dst + src
            dst, src = operands[0], operands[1]
            old_val = self.registers.gpr[dst]
            result = self.alu.add(old_val, self.registers.gpr[src])
            self.registers.gpr[dst] = result
            # Set zero flag
            self.registers.FLAGS['Z'] = (result == 0)
            print(f"  Operation: {dst} ← {old_val} + {self.registers.gpr.get(src, src)} = {result}")
            print(fZero flag Z ← {self.registers.FLAGS['Z']})

        elif opcode == 'SUB':
            # SUB dst src: dst = dst - src
            dst, src = operands[0], operands[1]
            old_val = self.registers.gpr[dst]
            result = self.alu.sub(old_val, self.registers.gpr[src])
            self.registers.gpr[dst] = result
            self.registers.FLAGS['Z'] = (result == 0)
            print(f"  Operation: {dst} ← {old_val} - {self.registers.gpr.get(src, src)} = {result}")

        elif opcode == 'MUL':
            # MUL dst src: dst = dst * src
            dst, src = operands[0], operands[1]
            old_val = self.registers.gpr[dst]
            result = self.alu.mul(old_val, self.registers.gpr[src])
            self.registers.gpr[dst] = result
            self.registers.FLAGS['Z'] = (result == 0)
            print(f"  Operation: {dst} ← {old_val} * {self.registers.gpr.get(src, src)} = {result}")

        elif opcode == 'CMP':
            # CMP a b: compare two register values, set zero flag
            a, b = operands[0], operands[1]
            is_equal = self.alu.cmp(
                self.registers.gpr[a],
                self.registers.gpr[b]
            )
            self.registers.FLAGS['Z'] = is_equal

        elif opcode == 'JMP':
            # JMP addr: Unconditional jump
            target = int(operands[0])
            print(fOperation: unconditional jump to address {target})
            old_pc = self.registers.PC
            self.registers.PC = target
            print(fPC ← {old_pc} → {target} (jump))

        elif opcode == 'JZ':
            # JZ addr: jump if the zero flag is true
            target = int(operands[0])
            if self.registers.FLAGS['Z']:
                print(fOperation: if zero flag is true (Z=1), jump to address {target})
                self.registers.PC = target
            else:
                print(fOperation: zero flag is false (Z=0), no jump, sequential execution)

        elif opcode == 'HALT':
            # HALT: Halt
            print(fOperation: Halt instruction, program ends)
            return 'HALT'

        else:
            print(f" Unknown opcode: {opcode}")

        return None

    def writeback(self):
        """
Stage 4: WriteBack
- If there are operations that need to write back to memory (such as STORE), they are completed in this stage
- Transfer data via the data bus
        """

        print(f"--- Stage 4: WriteBack ---")
        self.clock_cycles += 1

        # The result of a general instruction has already been written to the register during the execution stage
        # STORE instruction needs to write memory at this stage
        print(fRegister state: {self.registers.gpr})
        print(fFlags register: {self.registers.FLAGS})

    def writeback_store(self, addr, value):
        """Write-back operation of the STORE instruction — write to memory through the bus"""
        self.bus.cpu_write(addr, value, self.memory)
        print(fSTORE complete: {value} → memory address {addr})

    def run(self):
        """
Run the loaded program.
Loop through fetch → decode → execute → write back until the program ends.
        """

        print("=" * 60)
        print(Program starts execution)
        print("=" * 60)

        while self.registers.PC < max(self.memory.cells.keys(), default=0) + 1:
            # Check whether PC is in valid range
            if self.registers.PC not in self.memory.cells:
                # PC exceeds program range, possibly due to a jump to the end, terminate
                print(f"\nProgram counter PC={self.registers.PC} exceeds the program range, execution ends)
                break

            self.instructions_executed += 1

            print(f"\n{'#' * 55}")
            print(f"### Directive #{self.instructions_executed} | PC = {self.registers.PC}")
            print(f"{'#' * 55}")

            # Instruction fetch
            instruction = self.fetch()

            # Decode
            opcode, operands = self.decode()

            # Execute
            special = self.execute(opcode, operands)

            # Handle the HALT instruction
            if special == 'HALT':
                self.clock_cycles += 1
                break

            # Write back
            if special and special[0] == 'STORE':
                _, reg, addr, value = special
                self.writeback_store(int(addr), value)
            else:
                self.writeback()

            # Safety check: prevent infinite loops
            if self.instructions_executed > 1000:
                print("\nWarning: Over 1000 instructions executed, forced stop (possible infinite loop))
                break

        print(f"\n{'=' * 60}")
        print(f" Program execution completed")
        print(f"{'=' * 60}")

    def report(self):
        """
Generate a complete execution report.
Summarize all statistics: cache hit rate, instruction count, clock cycles,
Bus transfer count, memory read/write count, etc.
        """

        print(f"\n{'=' * 65}")
        print(fFullComputer execution report —— EXAMPLE)
        print(f"{'=' * 65}")
        print()

        # Basic information
        print(f"  {'─' * 55}")
        print(f" Basic Information")
        print(f"  {'─' * 55}")
        print(fTotal instructions executed: {self.instructions_executed:>6d})
        print(fTotal clock cycles: {self.clock_cycles:>6d} cycles)
        if self.instructions_executed > 0:
            cpi = self.clock_cycles / self.instructions_executed
            print(fAverage CPI (cycles/instruction): {cpi:>6.2f})
        print()

        # Cache performance
        print(f"  {'─' * 55}")
        print(f"Cache performance (L1 Cache)")
        print(f"  {'─' * 55}")
        total_access = self.cache.hit_count + self.cache.miss_count
        print(f" Cache hits: {self.cache.hit_count:>6d} times")
        print(f"  Cache miss:          {self.cache.miss_count:>6d} times")
        print(fTotal accesses: {total_access:>6d} times)
        print(fHit rate: {self.cache.hit_rate:>6.1f}%)
        print(fTotal cache lines: {self.cache.num_lines:>6d} lines)
        print()

        # Memory access
        print(f"  {'─' * 55}")
        print(f" Memory access statistics")
        print(f"  {'─' * 55}")
        print(fMemory read operations: {self.memory.read_count:>6d} times)
        print(fMemory write operation: {self.memory.write_count:>6d} times)
        print(f"  MemorytotalVisit:          {self.memory.read_count + self.memory.write_count:>6d} times")
        print(fNumber of used memory addresses: {len(self.memory.cells):>6d})
        print()

        # Bus statistics
        print(f"  {'─' * 55}")
        print(f" Bus Statistics")
        print(f"  {'─' * 55}")
        print(f"  CPU emit起ofbustransmission:  {self.bus.transfer_count:>6d} times")
        print(f"  DMA emit起ofbustransmission:  {self.dma.transfer_count:>6d} times")
        total_bus = self.bus.transfer_count + self.dma.transfer_count
        print(fTotal bus transfer count: {total_bus:>6d} times)
        print()

        # Final register state
        print(f"  {'─' * 55}")
        print(f" Register final state")
        print(f"  {'─' * 55}")
        print(fGeneral-purpose registers: {self.registers.gpr})
        print(fProgram counter PC: {self.registers.PC})
        print(fInstruction register IR: {self.registers.IR})
        print(fFlag register FLAGS: {self.registers.FLAGS})
        print()

        # Memory data dump (only shows data at valid addresses)
        print(f"  {'─' * 55}")
        print(f" Memory data dump (only shows program area)")
        print(f"  {'─' * 55}")
        if self.memory.cells:
            for addr in sorted(self.memory.cells.keys()):
                val = self.memory.cells[addr]
                print(fAddress {addr:>4d} (0x{addr:04X}): {val})
        else:
            print(f" (Memory is empty)")
        print()

        print(f"{'=' * 65}")
        print(fEnd of report. Thank you for using the EXAMPLE FullComputer simulator!)
        print(f"{'=' * 65}")


# ============================
# 6. Main Program - End-to-End Demo
# ============================

if __name__ == "__main__":
    print()
    print("╔" + "═" * 58 + "╗")
    print("║" + The complete journey of one instruction — EXAMPLE FullComputer Demo.center(52) + "║")
    print("╚" + "═" * 58 + "╝")
    print()

    # Create computer
    computer = FullComputer()

    # Define demo program: calculate (10 + 7) * 2 - 3 = ?
    # Use LOAD to load the initial value, and ADD/MUL/SUB to perform the calculation
    demo_program = [
        "LOAD A 10",    # A = 10
        "LOAD B 7",     # B = 7
        "ADD A B",      # A = A + B = 17
        "LOAD C 2",     # C = 2
        "MUL A C",      # A = A * C = 34
        "LOAD D 3",     # D = 3
        "SUB A D",      # A = A - D = 31
        "STORE A 100",  # Store result in memory address 100
        "LOAD B 31",    # B = 31 (expected result)
        "CMP A B",      # Compare A and B -> should be equal
        "JZ 12",        # If equal (Z=1), jump to HALT
        "LOAD D 0",     # (not executed) D = 0
        "HALT",         # Halt
    ]

    # Step 1: Store the program on the hard disk
    computer.store_program("EXAMPLE_DEMO", demo_program)

    # Step 2: DMA load into memory
    computer.load_program("EXAMPLE_DEMO", load_addr=0)

    # Step 3: Run the program
    computer.run()

    # Step 4: Print full execution report
    computer.report()

    print(f"\n")
    print(fReview: This demo covers all six modules of Computer Organization Principles:)
    print(fModule 1 (Binary): Instructions and data are stored in memory in binary encoding.)
    print(fModule 2 (Logic Gates): The ADD/SUB/MUL in the execution stage is completed by ALU gate circuits)
    print(fModule 3 (von Neumann): program and data share memory, five major components work together)
    print(fModule 4 (CPU Principles): PC fetch → IR decode → ALU execute → write back, complete four-stage pipeline.)
    print(fModule 5 (Storage System): hard disk → memory → cache → registers, four-level storage hierarchy)
    print(fModule 6 (Bus and I/O): Three-bus interaction, DMA program loading, interrupt notification)
    print()

Interpreting the execution report

After running the above code, FullComputer outputs a detailed execution report. Let's interpret the key metrics:

cache hit rate

The cache hit rate shown in the report reflects the program'stemporal localityandspatial locality. Since our demo program is very small and the DMA prefetches the first 4 instructions, the hit rate will be relatively high. In actual large programs, the cache hit rate is usually above 90% — this is exactly the purpose of cache.

CPI (Clock cycles Per Instruction)

In our simplified simulation, each instruction takes a fixed 4 clock cycles (fetch-decode-execute-writeback), so the CPI is about 4. Modern CPUs use pipelining to reduce the CPI to close to 1 — that is, completing one instruction per clock cycle.

Number of bus transfers

Bus transfer count = transfers initiated by the CPU + transfers initiated by the DMA. The DMA completes a large number of transfers during the program loading phase, and the CPU initiates at least one read transfer during the instruction fetch stage of each instruction.


What we have learned — review of the six-week course.

Six weeks ago, we started from the simplest question: "Why do computers only understand 0 and 1?" Now, we are able to build a complete computer simulator from scratch. Let's walk through this learning path again:

week numberModuleCore takeaways
Week 1Binary and Information RepresentationUnderstood why computers use binary, how to perform base conversion, and how various information (numbers, text, images) is encoded into 0s and 1s.
Week 2Logic Gates and Digital CircuitsMastered the three basic gate circuits: AND, OR, NOT; understood how to build adders and ALUs using them.
Week 3Von Neumann architectureUnderstood the division of labor among the five major components (arithmetic unit, controller, memory, input, output) and the stored-program concept.
Week 4CPU principlesMastered the four stages of the instruction cycle: fetch-decode-execute-writeback, and how pipelining improves efficiency.
Week 5Storage systemUnderstood the four-level storage pyramid from registers to hard disk, and how cache uses the principle of locality to speed up access.
Week 6Bus and I/OMastered the cooperative work of the three buses (address/data/control) and the efficiency differences among the three I/O methods (polling/interrupt/DMA).

From theory to practice: where to go next.

This course has laid a solid foundation for you in computer organization principles. If you are interested in deeper learning, here are some recommended directions:

  • Operating system: Learn how operating systems manage the CPU, memory, file systems, and I/O devices. The "interrupts," "DMA," and "memory hierarchy" learned in this course are core concepts of operating systems.
  • Assembly language: Interact directly with CPU instructions. After completing this course, assembly language is no longer a "sealed book" — you know how each assembly instruction is executed at the hardware level.
  • nand2tetris: A famous online course that takes you from a single NAND gate to gradually building a complete programmable computer. The content of this course is highly complementary to it.
  • Computer architecture: Delve into advanced CPU design techniques such as pipelining, superscalar execution, out-of-order execution, and branch prediction.
  • embedded system: Apply the knowledge from this course to microcontrollers (such as Arduino, STM32), writing programs that directly manipulate registers and peripherals.

Summary and final verification

One-sentence summary:Computemachineis not魔法——It isonelayeronelayer精心DesignofAbstraction, from晶body管openclosestart,逐步搭build出逻辑门、ALU、CPU、MemorySystem、bus,最终Compositionone台abilityExecuteAnyCoursesequenceof通usemachinedevice.UnderstandingTheseLevel,youJusttruecorrectUnderstandingalreadyComputemachine。

Comprehensive course self-test questions

  1. In the FullComputer simulator, during the execution of a LOAD instruction, in which steps are the address bus, data bus, and control bus used respectively?
  2. If FullComputer's cache has only 4 lines (instead of 8), what impact does it have on the program's execution efficiency? What is this phenomenon called in computer architecture?
  3. In what scenarios do modern operating systems use polling, interrupts, and DMA? Please give one example for each that you can observe when using a computer daily.

Reference answer: 1. Fetch stage: The address bus transmits the PC value (instruction address), the control bus sends a READ signal, and the data bus returns the instruction content. 2. More cache conflicts, lower hit rate (called cache thrashing or conflict misses). 3. Polling: progress check during firmware updates on some simple embedded devices; Interrupt: every keyboard key press triggers an interrupt; DMA: large file copying, graphics card rendering framebuffer transfer.

other extensions