Transformer Model Lab — Interactive Attention & Architecture Simulator

Interactive simulator of a real trained transformer block — a 3D diagram of embedding, attention, residual and feed-forward stages, a calculation studio exposing every operation on one token, controlled mechanism ablations, and an 18-check verification bench.

← Neural Networks & Transformers Labs
About this tool — how it works & FAQOpen ▾Close ▴

About the Transformer Model Lab

This simulator opens up a real, trained miniature decoder-only transformer block (1 block, 3 attention heads, width 24) and walks through every internal operation that turns a token into a prediction — embedding, attention, residual connections, layer normalization, and a position-wise feed-forward network — using the model's actual learned weights rather than illustrative placeholders.

What the simulator shows

• A live 3D decoder-transformer diagram (drag to rotate, pinch to zoom; Reset camera, Auto orbit, Expand, and Previous/Next stage + Walk-through controls) with tiles colored mint for positive and violet for negative feature values, and an output tile spanning vocabulary IDs 0–23 of 94. • A prompt workbench identical in function to the LLM Inference Lab — five teaching presets, a free-text prompt, Encode & run prefill, +1 next token / Generate up to 8, temperature/top-k/top-p sampling, a KV-cache toggle, and session reset/export. • A per-head scaled dot-product attention matrix inspector and a real embedding/residual/layer-norm activation inspector, plus a controlled-interventions panel covering residual paths, feed-forward output, layer normalization, the √dₕ score scaling, the causal mask, position embeddings, and per-head ablation. • A dedicated Calculation studio tab: pick a query token, key token, attention head and head feature, then step through six labeled stages — embedding lookup, learned projection, the dot-product score-to-weight calculation, the weighted value mix, the residual connection, and a chosen feed-forward hidden neuron (0–47) — each showing the actual numbers involved. • A controlled-comparison tool that reruns the same sequence from the shipped weights while changing exactly one mechanism at a time, reporting total-variation distance across all 94 output probabilities, plus a reference table contrasting encoder-only, decoder-only and encoder-decoder attention access patterns. • A How it works tab with the full 30-sentence training corpus and model card, an 18-check automated verification bench, and a knowledge-check quiz.

How the transformer block processes a token

Each token starts as a learned embedding vector combined with a learned position embedding, so the model knows both what the token is and where it sits in the sequence. Scaled dot-product attention then lets each token's query vector compare against every earlier token's key vector (masked so it can never see future tokens), producing attention weights that determine how much of each key token's value vector gets mixed into the output — the calculation studio's Step 03 and Step 04 panels show exactly this q·k score, its softmax weight, and the resulting weighted mix.

A residual connection adds the attention output back onto the original embedding (Step 05) rather than replacing it, which is what lets deep transformer stacks train stably. Layer normalization rescales features between stages, and a position-wise feed-forward network — one hidden neuron of which you can inspect directly by number (0–47) in Step 06 — processes each token's representation independently before producing the final next-token logits.

What this model is and isn't

This lab implements exactly one decoder block, trained on 30 short sentences with a simple lowercase word tokenizer — a real working transformer, but a tiny single-layer one, not a production multi-layer language model. Turning off the causal mask demonstrates bidirectional attention access within this one block; it does not convert the trained decoder into a trained encoder, and the lab does not implement cross-attention or a multi-layer stack. The 3D tiles represent software operations and actual computed features, not physical hardware or biological neurons, and there is no BPE tokenization, retrieval, chat tuning, or GPU timing simulated anywhere in the lab.

Frequently asked questions

What is the difference between this lab and the LLM Inference Lab?

Both run the same real miniature transformer and share the same prompt workbench and attention/activation inspectors. This lab adds a dedicated Calculation studio that exposes every internal operation — embedding lookup, projection, dot-product scoring, value mixing, residual addition, and a feed-forward neuron — for one chosen token, head and feature, plus a controlled-comparison tool for isolating individual mechanisms.

What does turning off the causal mask actually demonstrate?

It shows that this decoder block can technically attend to future tokens once the mask is removed, illustrating bidirectional attention access. It does not retrain or convert the model into an encoder — the lab implements one decoder block and does not simulate cross-attention or a full encoder-decoder architecture.

What does the residual connection do, and why does the model need it?

The residual connection adds each stage's output back onto its input rather than replacing it, so information from earlier in the computation is preserved rather than lost. Step 05 of the Calculation studio shows this addition using the model's actual numbers, and toggling "Residual paths" off in the interventions panel lets you see how removing it changes the next-token prediction.

How does the controlled-comparison tool measure the effect of a mechanism?

It reruns the identical token sequence from the same shipped weights, changing only one named mechanism at a time (such as the causal mask or layer normalization), then computes the total variation distance — from 0 (identical) to 1 (completely disjoint) — between the resulting probability distributions across all 94 vocabulary outputs.

Related tools & guides