Machine learning

Independent research · 2024

Forecasting
U.S. GDP

Four models, 64 years of quarterly data, and a 26-quarter test period that happens to contain the pandemic. The random forest posts the lowest error. Most of that error comes from two quarters, and what is left says something different about the result than the ranking does.

Holdout · 2017 Q4 to 2024 Q1 Quarterly change · $bn
257
Quarters, 1960 to 2024
4
Model families compared
288
Best holdout MAE, $bn
42%
Of that error, from 2 quarters

In plain terms

Gross domestic product is the standard headline measure of how much an economy produced in a given period, and forecasts of it feed into decisions that get made before the official figure arrives. It is also, by definition, the sum of four things: what households consume, what businesses invest, what government spends, and what the country sells abroad minus what it buys. This project asks whether those four components, tracked quarterly since 1960, carry enough structure to predict the next move in GDP. It pulls five series from FRED, the Federal Reserve's public data service, runs the standard checks that decide which models are even applicable, and then puts four of them on the same 26-quarter test period: two classical time-series models from econometrics and two general-purpose machine learning models. It was written in Python with pandas, statsmodels and scikit-learn. The random forest records the lowest error, which is the result the original write-up reported — and the sections below take that number apart to show what it is and is not measuring, because the test period contains the sharpest two quarters in the history of the series.

01 The question

Four numbers add up to GDP. Do they also predict it?

GDP is not an estimate of anything mysterious. It is an accounting identity: consumption plus investment plus government spending plus net exports. The four components are published alongside it, on the same quarterly schedule, from the same national accounts.

So the identity itself is not a forecast — it holds exactly, after the fact, and knowing three components tells you the fourth only once all four are measured. The question is whether the recent history of those components carries information about where the total goes next.

One decision shapes everything downstream. The target here is not the level of GDP but the quarterly change in it. Levels grow relentlessly, so predicting a level is mostly predicting a trend and any model looks impressive. Predicting the change is the part that is genuinely hard.

The test period runs from 2017 Q4 to 2024 Q1, the last 26 quarters of the dataset, held out in chronological order. That window includes 2020, which turns out to matter more than any modelling choice made anywhere else in the project.

GDP at time t equals consumption plus investment plus government spending plus net exports.

The five series pulled from FRED are GDP itself, personal consumption expenditures, gross private domestic investment, net exports, and total government expenditure. Everything is nominal and quarterly, inner-joined into 257 aligned rows from 1960 Q1 to 2024 Q1.

02 The data

Diagnostics before modelling

Two things are wrong with the raw series

Most time-series models will not work on data that trends, and no model gives interpretable coefficients when the inputs move together. Both problems are present here, and only one of them gets fixed.

The series are not stationary

An augmented Dickey–Fuller test asks whether a series has a stable mean and variance. Small p-values mean stationary. Every raw series comes back at essentially 1.0 — as nonstationary as the test can report, which is what you would expect of anything that grows for sixty years.

The fix is differencing: model the change from one quarter to the next instead of the level itself.

Delta y at time t equals y at time t minus y at time t minus one.

After one difference every series clears the test comfortably. That is also the moment the target quietly becomes the quarterly change, and every error figure on this page is in those units.

The predictors are nearly the same variable

A variance inflation factor measures how much of one predictor is explained by the others. Anything above 5 is usually called a problem. Consumption comes back at 170.

Switch the panel between the raw and differenced data to see what differencing fixes and what it leaves alone.

Series

Stationarity ADF p-value, lower is stationary

    Collinearity variance inflation factor, raw levels

      Differencing does not appear here because the notebook measures inflation on the levels. The components share information by construction, and no transformation of them changes that.

      This is what steers the model choice. With predictors that collinear, a linear model's coefficients stop meaning anything individually, even when its forecasts are fine. Tree models split on one variable at a time and do not have coefficients to corrupt, which is the stated reason the notebook reaches for a random forest.

      One more check, and one thing it did not clear

      A Granger causality test asks whether past values of one series improve a forecast of another. The notebook runs it per component and keeps a lag for each. Consumption is the awkward one: its best p-value sits just outside the usual 5% threshold at lag 6, and the notebook says so and keeps it anyway. That is a judgement call, and it is worth knowing it was made.

      03 Four models

      Two classical, two machine learning

      What each one assumes

      The four were not picked to be exotic. With 257 quarters there is not enough data to justify anything large, so the comparison is between models that write their assumptions down and models that make very few.

      1. Vector autoregression

        Classical · needs stationary input

        Every series is regressed on the recent past of every series, itself included. It captures how the components lead and lag one another, and it requires the differencing from the previous section to be valid at all. Fitted with up to 14 lags.

        Delta y at t equals a constant plus a sum of matrices times lagged differences plus an error term.
      2. Vector error correction

        Classical · works on the levels

        Built for series that wander individually but stay tied together in the long run, which is exactly what an accounting identity produces. A Johansen test finds three such relationships, so the model is fitted on the raw levels with rank 3. It predicts levels, not changes, so its error is on a different scale from the other three and cannot be ranked against them.

        Delta y at t equals Pi times the lagged level plus a sum of Gamma matrices times lagged differences plus an error term.
      3. Random forest

        Machine learning · no assumptions about form

        Many decision trees, each fitted to a resampled version of the data, averaged. It splits on one predictor at a time, so the collinearity from Section 02 does not hurt it, and it is hard to overfit by accident. 100 trees, fixed seed.

      4. K nearest neighbours

        Machine learning · no fitting at all

        For a new quarter, find the most similar quarters in the training data and average what happened next. With five predictors and a thousand-odd observations the ratio is favourable. The notebook sweeps k from 1 to 9 and takes k = 2.

      All four use the components measured in the same quarter as the GDP figure being predicted. In a live setting those numbers are not available yet, so this is closer to nowcasting than to genuine forecasting. Section 07 comes back to it.

      04 The forecasts

      Interactive

      All four on the same 26 quarters

      Every value here is read from the notebook's output. Switch between the quarterly change, which is what three of the models actually predict, and the accumulated level, which is what the change adds up to over six years.

      Model
      View

      Quarter 2024 Q1 drag the chart to inspect
      Actual — quarterly change · $bn
      Model — 100 trees
      Miss — model minus actual
      Reported in the notebook, 26-quarter holdout
      Model Target MAE MAPE Note
      Random forest Quarterly change287.9744.31% Lowest of the three comparable models
      KNN, k = 2 Quarterly change303.4547.72% Within 6% of the random forest
      VAR, 14 lags Quarterly change328.3957.72% The classical baseline
      VECM, rank 3 GDP level1,482.53— Different target, not comparable

      MAPE divides by the actual value, and quarterly GDP changes pass through zero and go negative in this window. The percentages above are what the notebook printed, but they are not a stable way to compare these models and nothing on this page rests on them.

      05 Where the error is

      Interactive

      Two quarters carry two fifths of it

      A mean absolute error is an average, and averages hide their own composition. Here is the same 287.97, broken out by quarter.

      MAE equals one over n times the sum of absolute differences between actual and predicted changes.

      Each bar is one quarter's absolute miss. Use the switch to drop 2020 Q2 and Q3 — the collapse and the rebound — and watch what happens to the number and to the ranking.

      Model
      Absolute miss per quarter, in billions of dollars, for the selected model across the holdout.
      MAE shown 287.97 over 26 quarters
      Worst quarter 2020 Q2 miss of 1,631
      Ranking RF · KNN · VAR best to worst on this subset

      What that means, and what it does not

      In 2020 Q2 GDP fell by $1,793bn in a single quarter. The random forest predicted a fall of $162bn. One quarter later it rose by $1,734bn and the model predicted a rise of $234bn. Those two quarters alone contribute 41.8% of the model's total error over six years. The same is true of the other two: 41.1% for VAR, 41.0% for KNN.

      None of this is a flaw in the models. Nothing in sixty years of data suggested a quarter like 2020 Q2 was possible, and a model that had somehow predicted it would have been wrong about everything else. The point is narrower: the headline MAE is largely a measurement of how badly each model handled one event, not of how well it tracks a normal quarter.

      The ranking survives. Drop the two quarters and the random forest still leads, at 181.5 against 194.0 and 209.5. That is worth stating plainly, because it is the part of the original conclusion that holds up.

      06 Is it any good

      The comparison the notebook did not run

      Better than what?

      A model can be the best of four and still not be doing much. The way to find out is to put it against predictions that require no model at all.

      All on the same 26 quarters, in billions of dollars
      Prediction MAE Uses only past data What it is
      Always predict zero 482.38Yes Assume GDP does not move
      Training average, +83.26 411.93Yes The mean change over 1960 to 2017
      Last quarter's change 367.98Yes Assume the last move repeats
      Random forest 287.97Yes The model from Section 03
      Best fixed number, +345.7 270.21No Chosen after seeing the answers

      The first three are legitimate competitors: each could have been produced in 2017 with no knowledge of what followed. The random forest beats all of them, comfortably. That is a real result and the project earned it.

      The last row is not a legitimate competitor. It is the single constant that would have minimised error over the test period, which nobody could have known in advance. It is in the table because of what it reveals: no constant could have scored better than 270, and the model scored 288. Whatever the random forest is doing, it is not beating the best a flat line could manage.

      The models barely move

      That result has a visible cause. Over the holdout, actual quarterly changes range from −$1,793bn to +$1,735bn, with a standard deviation of 536. The random forest's predictions range from −$162bn to +$297bn, standard deviation 97. The VAR is flatter still, at 45.

      Actual −1,793 to +1,735
      Random forest −162 to +297
      KNN −66 to +264
      VAR +84 to +287

      Each bar shows the range each series covers, on a common scale. The models occupy a narrow band near the middle: they have learned roughly how much GDP grows in a typical quarter and they mostly predict that, every quarter.

      Which is a defensible thing for a model to do when the alternative is guessing at shocks. But it means the honest description of the result is not "the random forest forecasts GDP well". It is: the random forest learned the typical quarterly growth rate more precisely than a long-run average would give you, and it does not predict deviations from it.

      The accumulated view makes the same point. Add up the random forest's predicted changes across the holdout and you get a GDP level of $24.27tn at 2024 Q1 against an actual $28.26tn. Over six years the model recovers slightly more than half of the real rise, because the quarters it underestimates never get corrected — small errors in a differenced model compound when you put the levels back together.

      The predicted level at time T equals the last training level plus the sum of predicted changes.
      07 What I would change

      Concretely

      Four things, in order of how much they would matter

      1. Use the data that existed at the time

        FRED serves revised figures. The 2020 Q2 number in this dataset is not the number anyone had in 2020 Q2 — GDP estimates are revised for years afterward. A real-time test needs vintage data, where each quarter is the estimate as first published. This is the change most likely to move the results.

      2. Stop using same-quarter components

        Every model here reads consumption, investment, government spending and net exports from the quarter it is predicting. Those are published alongside GDP, not before it. Lagging every predictor by one quarter would make this a forecast rather than a reconstruction, and would almost certainly make the errors larger.

      3. Validate on a rolling origin

        A single 26-quarter holdout that happens to contain the pandemic gives one number with a very wide margin around it. Refitting at each step and forecasting one quarter ahead, repeatedly, would give a distribution of errors and show whether the gap between the random forest and KNN — 288 against 303 — means anything at all.

      4. Report the baselines alongside the models

        Section 06 is the check that should have been in the original notebook. It costs four lines of code and it changes what the headline number means.

      A fifth possibility, from the original write-up, was to smooth the 2020 to 2024 window before fitting. That would improve the metrics and I would not do it: the pandemic quarters are the honest part of this test, and removing them optimises for the score rather than the question.

      08 Scope

      Read this before quoting the number

      What this is, and what it is not

      A clean model comparison

      Four families on identical data, an identical chronological split and an identical metric. The ranking holds with and without the pandemic quarters.

      The machine learning models won

      Both beat both classical models, and the notebook's reasoning for why — too few observations for VAR and VECM, collinearity that trees tolerate — is sound.

      It is nowcasting, not forecasting

      The predictors come from the same quarter as the target. Nothing here tells you what GDP will be next quarter, only what it was consistent with once the components were known.

      Revised data, not real-time data

      Every figure is the current revised estimate. Testing against numbers that were not available when the prediction would have been made is a known way to overstate accuracy.

      One split, one number

      No rolling validation, no confidence interval, no repeat runs. The 15-point gap between the random forest and KNN is well inside the range a different split could produce.

      Not a GDP forecast

      This is an archive of a July 2024 student project, and the data ends at 2024 Q1. It is a study of forecasting methods, not a current view on the economy.

      Source

      github.com/Dronmong/GDP-Forecast — the notebook. It needs a free FRED API key to run.

      Forecasting US GDP using Machine Learning and Mathematics — the original write-up, published in Towards Data Science, July 2024.

      Data

      Federal Reserve Bank of St. Louis, FRED. Series GDP, PCE, GPDI, NETEXP and W068RCQ027SBEA, quarterly, nominal.

      References

      Hyndman & Athanasopoulos, and the two texts the notebook cites: Forecasting Time Series Data (Apress) for the classical models and An Introduction to Statistical Learning for the machine learning ones.

      Acknowledgement

      This page was written with the assistance of Claude, an AI system made by Anthropic. Every forecast value shown is read from the notebook's own output; the baseline and error-decomposition figures in Sections 05 and 06 were computed from those same values and are new to this page.

      Continue exploring

      Read the notebook, the original article, or a project where the metric was the hard part.