LLM Inference Lab — Interactive Transformer Token Generation Simulator

Interactive simulator of a real trained miniature transformer generating text — prefill and autoregressive decoding, temperature/top-k/top-p sampling, a KV cache, attention-matrix and activation inspectors, controlled mechanism ablations, and a 12-check verification bench.

← Neural Networks & Transformers Labs
About this tool — how it works & FAQOpen ▾Close ▴

About the LLM Inference Lab

This simulator runs a real, trained miniature decoder-only transformer (1 block, 3 attention heads, width 24, trained on just 30 short sentences) entirely in your browser, and lets you step through exactly how it turns a prompt into a sequence of predicted next tokens — prefill, sampling, and autoregressive generation with a KV cache — using its actual learned weights rather than a simulated stand-in.

What the simulator shows

• A live 3D decoder-transformer diagram (drag to rotate, pinch to zoom, Reset camera / Auto orbit / Expand controls) with a prefill/decode phase badge and selectable pipeline stages. • A prompt workbench with five teaching presets (capital cities, how models work, engineering, mixing colors, and a deliberately out-of-corpus question), a free-text prompt box, Encode & run prefill, +1 next token and Generate up to 8 controls, a KV-cache reuse toggle, and a Reset session button. • Sampling controls — greedy/argmax vs. seeded sampling, a random seed field, a temperature slider (0.1–2.0), top-k (0 = disabled) and top-p (0.05–1.0) — plus live readouts for the most likely next token, sampling entropy in bits, token states computed this pass, and KV-cache memory in float32-equivalent bytes. • A next-token probability bar chart (softmax at T=1, then temperature/top-k/top-p resampling) and a scrollable, per-head scaled dot-product attention matrix you can click for the exact q·k score, scaled score and attention weight behind any cell. • An activation inspector showing real embedding, position, residual and layer-norm values per token, and a controlled-interventions panel — toggle the causal mask, position embeddings, or disable an individual attention head — that recomputes the full sequence and shows how the next-token distribution changes. • An inference log with session export, a How it works tab with the full 30-sentence training corpus and model card, a 12-check automated verification bench, and a knowledge-check quiz.

How autoregressive generation actually works

During prefill, the model encodes every token of your prompt at once, computing query/key/value projections and running them through masked scaled dot-product attention and a feed-forward network to produce a probability distribution over the vocabulary for the next token. Generation then proceeds one token at a time: the model samples (or greedily picks) the next token, appends it to the sequence, and repeats — this is why it's called autoregressive.

The KV cache stores each earlier token's key and value vectors so they don't need to be recomputed at every new step; reusing it is what makes multi-token generation fast, and the memory readout shows exactly how many float32-equivalent values that cache is holding. Temperature reshapes the probability distribution before sampling — lower values sharpen it toward the most likely tokens, higher values flatten it — while top-k and top-p further restrict which tokens are eligible to be sampled at all.

What this model is and isn't

The 24-dimensional, 3-head, single-block model was trained on only 30 short sentences using a simple lowercase word tokenizer — it is a real, working transformer, but a tiny one, not a general-purpose assistant. It has no BPE/subword tokenization, no retrieval, no chat tuning, no tool use, and no multi-layer reasoning, and it can repeat factual errors confidently (try the deliberately out-of-corpus "capital of Mars" preset). The 3D blocks represent software operations, not physical hardware, and the memory estimate uses 4 bytes per scalar purely for comparison to float32 storage — it is not a measurement of real GPU or server performance.

Frequently asked questions

Is this simulator running a real transformer, or is it a pre-scripted animation?

It runs a real, trained miniature decoder-only transformer (1 block, 3 heads, width 24) entirely in your browser. Every attention weight, activation and probability shown is computed from that model's actual learned weights, not scripted or pre-recorded.

What is the KV cache, and why does reusing it matter?

The KV cache stores each previously processed token's key and value vectors so the model does not need to recompute them at every new generation step. Reusing it is what makes multi-token autoregressive generation efficient; the lab shows its size as a float32-equivalent memory readout.

How do temperature, top-k and top-p change what gets generated?

Temperature rescales the raw softmax probabilities before sampling — lower values make the output more deterministic, higher values make it more random. Top-k then restricts sampling to only the k highest-probability tokens, and top-p restricts it to the smallest set of tokens whose cumulative probability reaches p, applied in that order after temperature.

Why does this model sometimes answer confidently but incorrectly?

It was trained on only 30 short sentences, so prompts outside that narrow training corpus — like the built-in "capital of Mars" preset — have no grounded answer, yet the model will still produce a confident-looking probability distribution. This illustrates why fluent, confident output is not the same as correct output.

Related tools & guides