Mathematics
Bias and Variance Tradeoff in Estimators
Quick fact
The bias-variance tradeoff is elegantly captured by the equation: Expected Test Error = Bias² + Variance + Irreducible Noise.
Why this is interesting
You’ve probably heard that “you can’t have your cake and eat it too.” In machine learning, that’s literally true for bias and variance—why can’t we just make both as small as possible?
Read the full explanation
Understanding Bias and Variance Tradeoff in Estimators
Imagine you’re trying to hit a bullseye on a dartboard. The bullseye is the true relationship you want to learn. Your throws are your predictions from different models. Bias is how far your average throw lands from the bullseye—if you always overshoot to the right, you’re biased. Variance is how spread out your throws are—if sometimes you hit the left edge and sometimes the right, you have high variance. A good model is like a thrower who aims accurately (low bias) and is consistent (low variance). But here’s the catch: if you try too hard to hit the bullseye every time (by memorising every past throw), you’ll over-adjust to random noise and become inconsistent (high variance). If you only aim with a simple, rigid strategy, you’ll be consistent but often miss the true bullseye (high bias). The tradeoff forces you to choose a sweet spot where the combination of being close on average and being consistent is minimal.
A deeper explanation
In statistical estimation, we want to approximate a target function or parameter. Suppose we have training data and we choose a model. The error of our estimate, evaluated on new data, can be decomposed into three parts: the squared bias, the variance, and irreducible noise. Bias arises from simplifying assumptions—for example, fitting a straight line to a curved relationship. Variance arises from the model’s sensitivity to the particular training sample—a very flexible model (like a high-degree polynomial) will produce wildly different estimates if the sample changes slightly. The irreducible noise is the randomness in the data itself that no model can remove. The tradeoff exists because increasing model complexity typically reduces bias (the model can better approximate the true function) but increases variance (it fits noise as if it were signal). Conversely, a simpler model has lower variance but higher bias. The goal is to choose a complexity that minimises total expected error, which often lies somewhere in between. This principle is central to model selection, regularisation, and understanding why cross-validation is necessary.