Forecasting: regression and forecast error measures
Least-squares trend and causal regression, correlation, and forecast error measures (bias, MAD, MSE, MAPE, tracking signal) with the MAD-sigma link.
Drafted with Aria, reviewed by the AiCanCode.org team. Spotted an error? Use Give Feedback at the bottom of the page.
Why it matters
When demand has a clear trend, or depends on a measurable driver such as advertising spend or the number of vehicles sold, averaging methods lag and regression gives a better forecast. Whatever method is used, a planner must also measure how wrong the forecasts are, decide whether the method is biased, and size safety stock from the error – which is what the error measures do.
Key ideas
Linear regression (least squares). A straight line y = a + b·x is fitted to n past pairs (x, y) so that the sum of squared vertical deviations is a minimum.
- Trend projection (time-series regression): x is the period number (1, 2, 3 …) and y is demand. The forecast for a future period is obtained by putting its period number in the line.
- Causal (associative) regression: x is a driver (price, promotion, construction permits) and y is demand. You must be able to forecast or know x for the future period.
- b is the slope (change in demand per unit change in x); a is the intercept (value of y at x = 0).
- The fitted line always passes through the point (x̄, ȳ).
- Assumptions: the relationship is linear; the past pattern continues; extrapolation far beyond the data range is risky.
Correlation and fit. The correlation coefficient r (−1 to +1) measures the strength and direction of the linear relationship. The coefficient of determination r² is the fraction of the variation in y explained by x. r close to ±1 means a strong linear relation; r near 0 means no linear relation (there may still be a curved one).
Coding time. When x values are equally spaced periods, you can code them so that Σx = 0 (for odd n: …, −2, −1, 0, 1, 2, …). Then b = Σxy/Σx² and a = ȳ, which saves arithmetic.
Forecast error. Error in period t is e(t) = D(t) − F(t), actual minus forecast. With this convention a positive error means the forecast was too low.
- Bias / mean forecast error (MFE) = Σe / n. A good method has bias near zero; positive bias means persistent under-forecasting.
- Running sum of forecast errors (RSFE) = Σe; it grows steadily when the method is biased.
- Mean absolute deviation (MAD) = Σ|e| / n, in units of demand. Easy to explain; used for tracking signals.
- Mean squared error (MSE) = Σe² / n, in units². Penalises large errors heavily; RMSE = √MSE is in units.
- Mean absolute percentage error (MAPE) = (100/n) Σ|e|/D, in %. Scale-free, so items of very different volume can be compared; undefined when any actual demand is zero and distorted when demand is very small.
- Tracking signal (TS) = RSFE / MAD, in "number of MADs". It is plotted each period; if it moves outside the control limits (commonly ±4, firms use ±3 to ±8) the method is biased and should be revised.
- For normally distributed errors, σ ≈ 1.25 × MAD (more precisely √(π/2) ≈ 1.2533). This links forecast error to safety stock.
Some books define error as F − D. Then signs of bias, RSFE and TS reverse; MAD, MSE and MAPE do not change. Always state the convention.
Formulas
y = a + b · x
- y = forecast demand (units), x = period number or causal variable, a = intercept (units), b = slope (units per unit of x).
b = (n·Σxy − Σx·Σy) / (n·Σx² − (Σx)²)
a = ȳ − b · x̄
- n = number of data pairs, x̄ = Σx/n, ȳ = Σy/n. Least-squares estimates.
r = (n·Σxy − Σx·Σy) / √{[n·Σx² − (Σx)²] · [n·Σy² − (Σy)²]}
- r = correlation coefficient (dimensionless); r² = coefficient of determination.
e(t) = D(t) − F(t)
MAD = Σ|e(t)| / n
MSE = Σe(t)² / n
MAPE = (100 / n) · Σ |e(t)| / D(t)
TS = RSFE / MAD = Σe(t) / MAD
- D = actual demand (units), F = forecast (units). MAD in units, MSE in units², MAPE in %, TS dimensionless.
σ ≈ 1.25 · MAD
- Standard deviation of forecast error, for normally distributed errors.
Worked examples
Example 1 (standard) – trend projection by least squares. Annual demand (thousand units) for years 1–5: 20, 24, 27, 33, 36. Forecast year 6.
- n = 5, Σx = 15, Σy = 140, Σxy = 1×20 + 2×24 + 3×27 + 4×33 + 5×36 = 461, Σx² = 55.
- b = (5 × 461 − 15 × 140)/(5 × 55 − 15²) = (2305 − 2100)/(275 − 225) = 205/50 = 4.1 thousand units per year.
- a = ȳ − b·x̄ = 28 − 4.1 × 3 = 15.7 thousand units.
- Forecast: y(6) = 15.7 + 4.1 × 6 = 40.3 thousand units.
- Check: r = 0.994, r² = 0.989, so the straight line explains about 99 % of the variation.
- Shortcut with coded time (x = −2, −1, 0, 1, 2): Σxy = −40 − 24 + 0 + 33 + 72 = 41, Σx² = 10, b = 4.1, a = ȳ = 28, and year 6 is x = 3: 28 + 4.1 × 3 = 40.3. Same answer.
Example 2 (GATE level) – error measures and tracking signal. Six months of actual demand D and forecast F (units): D: 100, 95, 110, 120, 105, 118 F: 102, 100, 104, 108, 110, 112
- Errors e = D − F: −2, −5, +6, +12, −5, +6.
- RSFE = Σe = −2 − 5 + 6 + 12 − 5 + 6 = 12 units; bias = 12/6 = 2 units per month (slight under-forecasting).
- MAD = (2 + 5 + 6 + 12 + 5 + 6)/6 = 36/6 = 6.0 units.
- MSE = (4 + 25 + 36 + 144 + 25 + 36)/6 = 270/6 = 45.0 units²; RMSE = 6.71 units.
- MAPE = (100/6)(2/100 + 5/95 + 6/110 + 12/120 + 5/105 + 6/118) = (100/6)(0.3256) = 5.43 %.
- Tracking signal = RSFE/MAD = 12/6 = 2.0, inside ±4, so the method is not judged biased.
- Estimated σ of forecast error ≈ 1.25 × 6.0 = 7.5 units.
Common mistakes
- Using y = a + bx with the regression of x on y, or swapping Σx² and (Σx)².
- Forgetting that a forecast by causal regression needs the future value of x.
- Mixing the sign conventions e = D − F and e = F − D within one problem; the TS sign depends on it.
- Computing MAD from signed errors (positives and negatives cancel). Use absolute values; only RSFE and bias use signed errors.
- Dividing by the forecast instead of the actual in MAPE.
- Reporting MSE in units instead of units², or comparing MSE of one method with MAD of another.
- Treating r near zero as "no relationship" when the data are clearly curved.
For GATE PI
Expect a small data table and a request for the least-squares slope, intercept or next-period forecast; or a set of actuals and forecasts with a request for MAD, MSE, MAPE, bias or the tracking signal. Conceptual items ask which measure penalises large errors, what a drifting tracking signal means, or the relation between MAD and σ. Practise filling the Σx, Σy, Σxy, Σx² table quickly and using coded time to save arithmetic.
Quick check
- For a fitted line, ȳ = 50, x̄ = 4 and b = 3. What is a?
- Errors for four periods are +3, −1, +4, −2 units. Find MAD and the tracking signal.
- Which error measure is expressed in units squared?
- Why is MAPE unsuitable for an item whose demand is zero in some months?
Answers: 1. a = 50 − 3 × 4 = 38. 2. MAD = 10/4 = 2.5 units; TS = 4/2.5 = 1.6. 3. MSE. 4. Each term divides by the actual demand, which is zero, so MAPE is undefined.
Interview questions
All Production Planning and Operations Management interview questionsTry answering each one aloud before you open it.
1.What is regression analysis in the context of forecasting?Concept
Regression analysis is a statistical method used to model the relationship between a dependent variable and one or more independent variables. In forecasting, it helps predict future values by analyzing past data trends and relationships.
2.Explain the difference between simple linear regression and multiple regression.Concept
Simple linear regression involves two variables: one independent variable and one dependent variable. It models the relationship as a straight line. Multiple regression involves more than one independent variable, allowing for a more complex model that can capture the influence of multiple factors on the dependent variable.
3.What are forecast error measures, and why are they important?Concept
Forecast error measures quantify the accuracy of a forecast by comparing predicted values to actual outcomes. They are important because they help assess the reliability of forecasting models and guide improvements. Common measures include Mean Absolute Error (MAE), Mean Squared Error (MSE), and Mean Absolute Percentage Error (MAPE).
4.Why is Mean Absolute Percentage Error (MAPE) often preferred over other error measures?Application
MAPE expresses the average absolute error as a percentage of actual demand, so it is scale-free: a planner can compare forecast accuracy for a fast-moving item selling 10,000 units and a slow one selling 50 units, which MAD or MSE cannot do. It is also easy to explain to managers. Its weaknesses are that it is undefined when any actual demand is zero and becomes very large when actuals are small, so for intermittent demand MAD or a weighted MAPE is used instead.
5.What happens if multicollinearity is present in a multiple regression model?Application
Multicollinearity occurs when independent variables in a regression model are highly correlated. It can lead to unreliable coefficient estimates, making it difficult to determine the effect of each variable. This can reduce the model's predictive power and make it challenging to interpret the results.
6.Explain how you would use regression analysis to forecast sales for a new product.Application
To forecast sales for a new product using regression analysis, you would first gather historical data on similar products, including factors like price, marketing spend, and economic indicators. Then, you would build a regression model to identify relationships between these factors and sales. Finally, you would use the model to predict sales for the new product based on its specific attributes.
7.Calculate the Mean Absolute Error (MAE) given the following forecasted and actual values: Forecasted: [100, 150, 200], Actual: [110, 140, 210].Numerical
- Calculate the absolute errors: |100 - 110| = 10, |150 - 140| = 10, |200 - 210| = 10.
- Sum the absolute errors: 10 + 10 + 10 = 30.
- Divide by the number of observations: 30 / 3 = 10. The Mean Absolute Error (MAE) is 10.
8.When would you choose MSE rather than MAD to compare forecasting methods?Application
Choose MSE (or RMSE) when large errors are disproportionately costly, for example when a big shortfall stops an assembly line, because squaring the errors penalises them heavily. MAD weights all errors linearly, is in the same units as demand and is less sensitive to a single outlier, so it is preferred for tracking signals and day-to-day monitoring. If the data contain occasional one-off spikes, MSE can be dominated by them, so look at both.
9.If a regression model consistently underestimates actual values, what might this indicate about the model?Application
If a regression model consistently underestimates actual values, it may indicate that the model is biased or missing important variables that influence the dependent variable. It could also suggest that the model's assumptions are not fully met, requiring a reevaluation of the model structure or inclusion of additional data.
Finished this topic? Mark it so your progress, study plan and readiness keep up.
Stuck on something here?