This simulator runs an actual k-means algorithm on a seeded two-dimensional point cloud. Choose how many clusters to ask for, how far apart the hidden source groups are, and how spread out they are, then watch points change membership and centroids move until the assignments stop changing.
• A real-time 3D scene with 4 inspectable parts (Unlabeled data cloud, Cluster centroids, Nearest-center membership and Centroid movement trails), with home view, focus-selected-part, auto-rotate, expand, instrument-cover and hide-labels scene tools, plus a model response curve beneath the scene. • Experiment controls: Number of clusters k (2-5); Source-group separation (1-4); Source spread (0.2-1); Repeatable random seed (1-99); show animated explanatory markers; pause/resume, 0.1 s and 1 s single-step buttons, four playback speeds and a restart button. • A Curves & measurements tab with a parameter-comparison chart, a live-measurements chart, the model equations and snapshot readouts (Completed Lloyd iterations; Within-cluster squared distance; Maximum centroid shift; Converged (1=yes)). • An Experiments tab with 2 guided presets (too few centers and extra centers) and a Model verification bench that runs independent fresh models, plus a timestamped event log and a copyable trial report. • A Learn & assess tab with guided lessons, a knowledge-check quiz with reset and a written model-scope statement linking to a technical reference.
Each Lloyd iteration assigns every point to its nearest center by Euclidean distance, then moves each center to the mean of its assigned points. The lab counts completed iterations, reports the within-cluster squared distance (inertia), the maximum centroid shift, and whether the run has converged.
Source-group separation and spread control how distinct the underlying groups are, so you can compare a clean, well-separated cloud with heavily overlapping groups.
K-means never sees the original source labels; it uses point positions and the requested k only. Lower inertia does not prove the right k, because adding clusters can always reduce inertia without revealing more useful structure. Use the too-few-centers and extra-centers experiments to see the trade-off.
The model uses seeded deterministic initialization and at most 20 iterations, and an empty cluster keeps its previous center. Local minima, unequal-density groups and non-spherical shapes can mislead, and the algorithm does not discover a uniquely true number of categories.
No. It is unsupervised: it uses only point positions and the number of clusters you request. The colors show its own assignments, which may or may not match the hidden source groups.
No. Adding more clusters nearly always lowers within-cluster squared distance, even when the extra centers just split a natural group. Lower training inertia alone is not a test for the right k.
It means assignments have stopped changing so centroids no longer shift. The run is also capped at 20 iterations, and the converged flag reports whether it settled before the cap.
Yes. It can settle in local minima depending on initialization and handles unequal-density or non-spherical groups poorly because it relies on Euclidean distance to a mean.