CPU Architecture Lab — Interactive 5-Stage Pipeline Processor Simulator

Interactive 32-bit RV32I-style processor simulator — write and assemble your own program, step through a five-stage pipeline cycle by cycle, inspect the data cache and RAM, compare pipelined vs. non-overlapped execution, and run a 16-check architectural verification bench.

← Computer Science Labs
About this tool — how it works & FAQOpen ▾Close ▴

About the CPU Architecture Lab

This simulator executes a genuine, deterministic subset of RV32I-style integer instructions on an abstract five-stage, in-order, single-issue processor core — with real forwarding, interlocks, branch flushing, a configurable data cache and blocking RAM — so every register value, cache hit, stall and cycle count you see is computed, not scripted.

What the simulator shows

• A live 3D processor diagram (drag to orbit, pinch to zoom; Overview / Core view / Auto orbit / Expand controls) with selectable components and a component-info callout. • A microarchitecture configurator: five-stage pipeline vs. non-overlapped five-stage execution, an operand-forwarding toggle, a data-cache mode selector (64 B direct-mapped, 64 B 2-way LRU, or bypass), and RAM transaction latency (3, 8 or 20 cycles) — applied on Assemble & reset. • Run/​+1 cycle/​Next retirement/​Run up to 2,000 cycles controls with four animation speeds, a retirement breakpoint on program counter, and a program workbench with an editable assembly source box and preset programs. • A live pipeline view showing stage occupancy after each clock edge, ALU and bus activity, a register file with selectable hex/signed/unsigned display and last-write highlighting, plus a callout explaining each cycle's outcome. • A cache & memory tab: a 4-line × 16-byte data cache with tag/set/offset breakdown and an address inspector, a 1 KiB little-endian data RAM viewer with page paging, and an initial-state editor for presetting RAM words or registers before cycle 1. • An execution trace tab (instruction disassembly, a 60-cycle timeline table, and a diagnostic event log with JSON export), a tests tab with a 16-check verification bench and a four-configuration performance comparison, and an architecture-guide tab with the full supported instruction set reference.

How the five-stage pipeline executes an instruction

Each instruction moves through fetch, decode, execute, memory and writeback stages, and with the pipeline enabled, up to five instructions can be in different stages simultaneously — this is what pipelining buys you over the non-overlapped mode, where each instruction must fully complete before the next one starts. Operand forwarding routes a just-computed result directly to a dependent instruction's execute stage instead of waiting for it to reach the register file, cutting or eliminating load-use and arithmetic-dependency stalls; disabling forwarding doesn't change the final program result, since interlocks still preserve correctness, but it does increase stall cycles and total cycle count.

The data cache uses a 4-line, 16-byte-per-line, direct-mapped or 2-way-LRU organization with write-through, no-write-allocate stores — a store miss still writes to RAM but does not pull a new line into cache. Instruction fetch reads from a separate ideal ROM, so instruction-side memory timing is never a bottleneck in this model; only data-side loads and stores interact with the configurable cache and RAM latency.

What this model is and isn't

This is an executable, deterministic teaching core with real RV32I instruction encodings (aside from a lab-only HALT sentinel) — not a full RISC-V emulator and not a model of any specific commercial processor. There is no operating system, no privilege levels, no virtual memory, no interrupt controller, no instruction cache, no out-of-order or speculative execution, and no multicore coherence. Physical component layout and signal animation in the 3D view are illustrative; the instruction execution, stalls, cache behavior and cycle counts are the parts that are actually computed and can be verified against the built-in test bench.

Frequently asked questions

What is the difference between the ISA and the microarchitecture in this lab?

The instruction set architecture (ISA) defines the visible behavior of each RV32I-style instruction — what it does to registers and memory. The microarchitecture is this lab's specific five-stage, in-order implementation of that ISA, including its pipeline, forwarding and cache choices; a different microarchitecture could execute the identical ISA with different timing.

Does disabling operand forwarding change the final program result?

No. Forwarding affects timing, not correctness — when it is disabled, interlocks still stall dependent instructions until their operands are correctly available, so the architectural result is unchanged. What changes is the stall count and therefore the total cycle count and CPI.

What happens to a store that misses in the data cache?

This lab uses a no-write-allocate policy: a store miss still writes the data to RAM, but it does not allocate a new line in the cache for that address. Only load misses allocate a cache line.

Why does a taken branch remove some already-fetched instructions?

In a pipelined processor, instructions after a branch are fetched speculatively down the not-yet-known path before the branch resolves. When the branch is taken, those younger wrong-path instructions must be flushed from the pipeline before they can modify any architectural state, which is exactly what this lab's cycle timeline and execution log show happening.

Related tools & guides