Mathematics
Time Series Analysis and Autoregressive Models
Quick fact
An autoregressive model of order 1 (AR(1)) can be written as Xt = φ·Xt−1 + εt, where Xt is the value at time t, Xt−1 is the previous value, φ is a coefficient that controls the strength of the memory, and εt is random noise. If φ is between -1 and 1, the series is stable and mean-reverting, making it possible to forecast future values with measurable uncertainty.
Why this is interesting
Stock prices, daily temperatures, and even your heart rate have a hidden property: they can predict their own near future. But how can we turn that intuition into a mathematical model that actually forecasts?
Read the full explanation
Understanding Time Series Analysis and Autoregressive Models
Imagine you are watching a child swing back and forth. You instinctively know that where the swing is now tells you a lot about where it was a second ago and where it will be next. Time series analysis is about applying this intuition to data that is ordered in time. An autoregressive model is a formal way to say: 'The current value depends linearly on the previous values, plus some random jitter.' Think of a simple example: a lake's water level. Today's level is likely close to yesterday's level, but not exactly — rain, evaporation, and usage introduce random changes. We can model the level, Lt, as a fraction 'a' of yesterday's level, plus a random shock 'e'. That gives us Lt = a·Lt−1 + e. This is an AR(1) model. The coefficient 'a' tells us how strongly yesterday's level influences today. If a is near 1, the level is persistent — a high day tends to be followed by another high day. If a is 0, today's level is just random noise, unrelated to the past. The 'autoregressive' part means the model is regressing the variable on itself — using its own past as the predictor. The order of the model, denoted 'p', tells us how many past values we use. An AR(2) model, for example, would also include the value from two days ago. This idea is powerful because it captures the momentum or persistence so common in real-world data.
A deeper explanation
The mechanism behind autoregressive models rests on the concept of linear dependence across time. Formally, an AR(p) model is written as: Xt = c + φ₁Xt−1 + φ₂Xt−2 + ... + φₚXt−p + εt where c is a constant, φ₁...φₚ are coefficients, and εt is white noise (zero mean, constant variance, uncorrelated over time). The equation says that the value at time t is a weighted sum of the p most recent values, plus a random shock. The important insight is that this model can be solved recursively: given a seed value and a sequence of random shocks, we can generate the entire series. This makes it a stochastic process. The behavior of the series depends critically on the coefficients. For an AR(1) model, if |φ| < 1, the series is stationary — its mean and variance are constant over time, and shocks decay exponentially. If φ = 1, the series becomes a random walk, which is non-stationary and can drift without bound. This condition is a boundary point: the model's long-term behavior changes drastically. Why does this matter? Because stationarity is a requirement for many statistical inference tools. If a series is non-stationary, standard regression techniques can lead to spurious results. The autoregressive framework gives us a way to check and achieve stationarity, for instance by differencing the data. Moreover, AR models are not just a theoretical curiosity. They are the 'AR' in ARIMA, one of the most widely used forecasting methods. They also appear in economics (e.g., modelling GDP growth), in signal processing (e.g., spectral analysis), and in many other fields where we need to understand and predict sequences. The key takeaway is that temporal dependence is a structure we can model with a simple linear equation, and that understanding this structure is the first step to making good forecasts.