Model error · interactive essay

Why does a new training sample change the prediction?

Compare fitted models at the same input. Their average can miss the target systematically; individual predictions can also scatter around that average. These are different sources of error.

Compare before interpreting

Follow the fitted lines and their error bars. In the LOESS sandbox, hold smoothness fixed and randomize the training data before changing smoothness: separate sample variation from a change in the fitting rule.

Play first: complexity story Jump to the sandboxes

Separating Bias and Variance

The assumptions behind the decomposition

Hold the input x fixed and write a fresh target as Y = f(x) + ε, where f(x) is the conditional mean E[Y | X = x]. Assume finite second moments and zero-mean noise ε independent of the training process at this input. Average squared prediction error over repeated training samples (and any fitting randomness) and independent target noise. It equals squared bias + prediction variance + noise variance. Bias compares the average fitted prediction with f(x); variance measures squared spread around that average. Noise variance is irreducible for this fixed input and information set. This identity does not require either bias or variance to change monotonically with complexity.

Using our wildest imagination, we can picture a dataset consisting of features X and labels Y, as on the left. Also imagine that we’d like to generalize this relationship to additional values of X - that we’d like to predict future values based on what we’ve already seen before.

With our imagination now undoubtedly spent, we can take a very simple approach to modeling the relationship between X and Y by just drawing a line to the general trend of the data.

A Simple Model

Our simple model isn’t the best at modeling the relationship - clearly there's information in the data that it's failing to capture.

We'll measure the performance of our model by looking at the mean-squared error of its output and the true values (displayed in the bottom barchart). Our model is close to some of the training points, but overall there's definitely room for improvement.

The error on the training data is important for model tuning, but what we really care about is how it performs on data we haven't seen before, called test data. So let's check that out as well.

Low Complexity & Underfitting

Uh-oh, it looks like our earlier suspicions were correct - our model is garbage. The test error is even higher than the train error!

In this case, we say that our model is underfitting the data: our model is so simple that it fails to adequately capture the relationships in the data. The high test error is a direct result of the lack of complexity of our model.

An underfit model is one that is too simple to accurately capture the relationships between its features X and label Y.

A Complex Model

Our previous model performed poorly because it was too simple. Let's try our luck with something more complex. In fact, let's get as complex as we can - let's train a model that predicts every point in our training data perfectly.

Great! Now our training error is zero.

High Complexity & Overfitting

Wait a second... Even though our training error from our model was effectively zero, the error on our test data is high. What gives?

Unsurprisingly, our model is too complicated. We say that it overfits the data. Instead of learning the true trends underlying our dataset, it memorized noise and, as a result, the model is not generalizable to datasets beyond its training data.

Overfitting refers to the case when a model is so specific to the data on which it was trained that it is no longer applicable to different datasets.

In situations where your training error is low but your test error is high, you've likely overfit your model.

Test Error Decomposition

Our test error can come as a result of both under- and over-fitting our data, but how do the two relate to each other?

Under the fixed-input squared-error assumptions above, expected prediction error is squared bias plus prediction variance plus noise variance. The expectation averages over training samples and fresh target noise, not just the residuals from one fitted model.



Or, mathematically:



We can’t do much about the irreducible term, but we can make use of the relationship between both bias and variance to obtain better predictions.

Bias

Bias is the average fitted prediction minus the conditional mean f(x), at the same fixed input :



The term is a tricky one. It refers to the average prediction after the model has been trained over several independent datasets. We can think of the bias as measuring a systematic error in prediction.

These different model realizations are shown in the top chart, while the error decomposition (for each point of data) is shown in the bottom chart.

In this demonstration's low-complexity fit, squared bias is the larger contribution.

Variance

As with bias, the notion of variance also relates to different realizations of our model. Specifically, prediction variance is the average squared deviation from the mean prediction at a fixed input, across training samples :



In the displayed high-complexity example, variance contributes more than squared bias. Compare the spread of fitted lines at one input with the corresponding decomposition; this example is not a guarantee about every high-complexity model.

Finding A Balance

Compare the three displayed fits: the middle one captures the trend without following every training point. Its advantage belongs to this example, not to a universal rule that the best model has intermediate complexity.

Choose a fitting rule using validation performance, then evaluate it on held-out test data. Complexity alone does not determine bias, variance, or generalization.

Across Complexities

We just showed, at different levels of complexity, a sample of model realizations alongside their corresponding prediction error decompositions.

Let’s direct our focus to the error decompositions across model complexities.

For each level of complexity, we’ll aggregate the error decomposition across all data-points, and plot the aggregate errors at their level of complexity.

This aggregation applied to our balanced model (i.e. the middle level of complexity) is shown to the left.

The Bias Variance Trade-off

Aggregating over the displayed inputs gives this example's U-shaped error curve. That is an additional average across inputs, distinct from the fixed-input expectation used in the decomposition.

Here, squared bias dominates the simpler fits and variance dominates the more flexible fits. Other data distributions and fitting procedures can produce different curves.

The decomposition explains contributions to expected squared error; it does not prove that increasing complexity always lowers bias, raises variance, or yields a U-shaped curve.

What to watch for

Keep one input in view as you randomize training samples: does the fitted value move? Then change the fitting rule. A single train/test split does not by itself measure the expectation over repeated training samples.



Many models have parameters that change the final learned models, called hyperparameters. Let’s look at how these hyperparameters may be used to control the bias-variance tradeoff with two examples: LOESS regression and K-nearest neighbors.



LOESS Regression

LOESS (LOcally Estimated Scatterplot Smoothing) regression is a nonparametric technique for fitting a smooth surface between an outcome and some predictor variables. The curve fitting at a given point is weighted by nearby data. This weighting is governed by a smoothing hyperparameter, which represents the proportion of neighboring data used to calculate each local fit.

Small neighborhoods can make the fit sensitive to which training points were sampled; large neighborhoods can smooth away local structure. These are tendencies to investigate in this dataset, not guaranteed changes in bias and variance for every smoothing value.

Below a LOESS curve is fit to two variables. Randomize the training data to observe the effect different model realizations have on variance, and control the smoothness to observe the tradeoff between under- and over-fitting (and thus, bias and variance).


K-Nearest Neighbors

K-nearest neighbors (KNN) classification is a simple technique for assigning class membership to a data point with some majority vote of its K-nearest neighbors. For example, when K = 1, the data point is simply assigned to the class of that single nearest neighbor. If K = 69, the data point is assigned to the majority class of its 69-nearest neighbors.

Change K to compare local votes with votes over larger neighborhoods. Randomize data at the same K to inspect how the decision regions change across samples. Jagged boundaries alone do not measure variance. This classification view illustrates sensitivity and smoothing; the squared-error identity above is not a decomposition of classification error under zero-one loss.

Explore the trade-off for yourself below. The plot on the left shows the training data. The plot on the right shows decision regions based on the current value of K. Deeper colors reflect more confidence in the classification. Hover over a point to see its classification to the right, and the K-nearest neighbors used for consideration to the left.


Related Concept: Double Descent

You may also encounter the phenomenon known as double descent, where the classic U-shaped test-error curve is followed by a second dip at higher complexity.

This does not contradict the squared-error identity: that identity asserts no universal monotonic relationship between complexity, bias, and variance. The fitting procedure and data distribution determine how those terms behave.



The visualization below is kept from the published page so the local route still illustrates how the classical tradeoff connects to that later regime.


Wrapping Up

Thanks for reading. The main lesson here is to balance underfitting and overfitting rather than optimizing only for training performance.

Some adjacent topics, such as regularization, are outside the scope of this page, but the charts above should give you a compact mental model for reasoning about them.




References

This page draws on the following references and libraries: