Neural Network Lab — Interactive Backpropagation & Training Simulator

Interactive 3D neural network simulator — train a real feedforward classifier in your browser on XOR, circle or linear tasks, inspect neuron math and gradients, trace a forward pass and backprop update, and watch the decision boundary and loss curve change live.

← Neural Networks & Transformers Labs
About this tool — how it works & FAQOpen ▾Close ▴

About the Neural Network Lab

This simulator trains a real, small feedforward neural network directly in your browser using seeded synthetic data and full-batch gradient descent. Nothing is pre-recorded: every weight, activation, loss value and decision-boundary pixel you see comes from an actual running network, viewed through a 3D architecture diagram and a set of calculation inspectors.

What the simulator shows

• A live 3D network diagram (default architecture 2 → 4 → 3 → 1) with mint/orange edges for positive/negative weights and thickness proportional to magnitude — drag to rotate, pinch/scroll to zoom, click a neuron to inspect it, plus Reset view, Auto orbit and Expand controls. • Three learning tasks (XOR, Circle, Linear), a hidden-layer selector (4+3, 4, 8+4, or none for a linear boundary), four activation functions (tanh, ReLU, Sigmoid, Linear), and a learning-rate slider from 0.01 to 1.00. • Train/​+1 step/​Train 100 steps controls with live step count, parameter count and training-loss readouts (128 training examples per step; 96 held-out test examples that never update weights). • A prediction sandbox with x₁/x₂ input sliders, a predicted-probability bar, and Test this input / Random input buttons; a canvas-based decision-landscape map you can click to choose an input, with an optional test-points overlay. • A learning-curve chart, a Run held-out test button, and an adjustable classification threshold (0.10–0.90). • A black-box inspector — Trace forward pass, Next layer, and Backprop + update buttons — that shows a selected neuron's incoming connections (a × w), lets you drag its weight/bias sliders directly, and reveals the actual single-example gradient used in a parameter update. • A How it works tab (six numbered lessons on inputs, neuron math, nonlinearity, loss, backprop and generalization), an Experiments tab of guided one-variable-at-a-time challenges, and a knowledge-check quiz.

How the network learns

Each neuron computes z = Σ(aᵢ × wᵢ) + b, then applies an activation function before passing the result to the next layer. Binary cross-entropy loss, L = −[y log p + (1−y) log(1−p)], compares the network's output probability p against the true label y, penalizing confident wrong predictions more heavily than uncertain ones. Backpropagation uses the chain rule to compute how each weight and bias affects that loss, and gradient descent then updates every parameter by w ← w − η · ∂L/∂w, where η is the learning rate you control.

Stacking purely linear hidden layers still produces only a linear decision boundary no matter how many layers you add — this is why the XOR task (which is not linearly separable) requires a nonlinear activation like tanh or ReLU to solve. The held-out test set of 96 examples never contributes to training; it exists purely to check whether the network generalizes to data it hasn't seen, which is a different question from how well it fits its training data.

What this model is and isn't

This is a genuine small classifier — not a scripted animation — trained with real full-batch gradient descent on synthetic, seeded data. The 3D layout and moving pulses illustrate computation for teaching purposes; they do not represent biological brain activity, physical signal propagation speed, or the internals of a large language model. Training results depend on initialization, task, architecture, activation function and learning rate, so re-running with different settings can produce different outcomes, and probability outputs are not guaranteed to be well-calibrated.

Frequently asked questions

Is this a real neural network or a scripted animation?

It is a real, small feedforward classifier trained live in your browser using full-batch gradient descent on seeded synthetic data. Every weight, loss value and prediction shown comes from that actual running network, not a pre-recorded sequence.

Why does XOR need a nonlinear activation function to be solved?

Stacking linear hidden layers still produces only a linear decision boundary, and XOR (opposite-sign inputs labeled class 1) is not linearly separable. Nonlinear activations like tanh or ReLU let the network learn curved or disconnected decision regions that a purely linear model cannot represent.

What is the difference between training loss and test accuracy?

Training loss (binary cross-entropy) measures how confidently wrong or right the network is on the 128 examples it learns from. Test accuracy, measured on 96 held-out examples that never update the weights, measures how well the trained network generalizes to data it has never seen — repeatedly tuning settings based on the test score reduces its independence as a true generalization check.

What does the Backprop + update button actually compute?

It shows a real single-example gradient and the resulting parameter change for whichever connection you have selected, computed via the chain rule from that one example. Automatic training, by contrast, averages gradients over all 128 training examples before each step.

Related tools & guides