This simulator fits polynomials to twelve noisy training points generated from a known sine signal and scores them on a separate held-out dataset. Raise the degree, change the noise, and add a ridge penalty while comparing training error, held-out error and error against the noiseless truth.
• A real-time 3D scene with 4 inspectable parts (Training sample rail, Held-out sample rail, Polynomial prediction ribbon and Training / test error comparison), with home view, focus-selected-part, auto-rotate, expand, instrument-cover and hide-labels scene tools, plus a model response curve beneath the scene. • Experiment controls: Polynomial degree (1-9); Observation noise SD (0.02-0.3); Ridge penalty (0-0.1); Repeatable random seed (1-99); show animated explanatory markers; pause/resume, 0.1 s and 1 s single-step buttons, four playback speeds and a restart button. • A Curves & measurements tab with a parameter-comparison chart, a live-measurements chart, the model equations and snapshot readouts (Training mean squared error; Held-out mean squared error; MSE against noiseless signal; Polynomial degree). • An Experiments tab with 2 guided presets (flexible unregularized fit and add regularization) and a Model verification bench that runs independent fresh models, plus a timestamped event log and a copyable trial report. • A Learn & assess tab with guided lessons, a knowledge-check quiz with reset and a written model-scope statement linking to a technical reference.
A low-degree polynomial is too stiff to follow the signal, while a high-degree polynomial can pass near every training point and swing wildly between them. The lab reports training mean squared error, held-out mean squared error, and error against the noiseless signal f(x) = sin(pi x), so you can see a flexible fit driving training error down while held-out error rises.
Ridge regularization adds a penalty on the non-intercept coefficients, which shrinks them and tames the wiggle; the fit is computed by QR least squares on an augmented system.
Held-out data must never be used to fit the coefficients, otherwise the evaluation is contaminated. High training accuracy can coexist with poor generalization, because a flexible model can fit the noise rather than the signal.
The lab uses a single seeded training and test split with known synthetic truth, so trends can vary by seed and held-out error is not guaranteed to be monotone in degree. Fits are restricted to the sampled x range, and degree 9 uses 12 training observations with ridge optional.
No. Fitting on the held-out set would contaminate the evaluation and hide overfitting. In this lab only the twelve training points determine the coefficients, and the held-out rail is used purely for scoring.
Yes. A flexible high-degree polynomial can fit the noise in the training points, giving very low training error and a much larger held-out error. The flexible unregularized experiment shows this.
It adds a penalty proportional to the squared size of the non-intercept coefficients, shrinking them toward zero. That reduces wild oscillation and often lowers held-out error for a flexible model, at the cost of some extra training error.
Not always. The lab uses one seeded split, so the trend can vary with the seed, and there is no guarantee that test error is monotone in polynomial degree.