nep-for New Economics Papers
on Forecasting
Issue of 2026–09–07
eighteen papers chosen by
Malte Knüppel, Deutsche Bundesbank


  1. Sequentially valid inference for probabilistic inflation forecasts By Amadeo Grob; Maurizio Daniele; Johanna Ziegel
  2. A Design Concept of Forecasting Software for Normalized Vector Autoregressions with Fat Tails and Stochastic Volatility By Fei Shang; Xiaolei Wang; Tomasz Wo\'zniak
  3. Structural forecast analysis By Davide Brignone; Michele Piffer
  4. Conditional projection methods for large-scale Bayesian VARs By Niko Hauzenberger; Michael Pfarrhofer
  5. Measures of Macroeconomic Shocks and Uncertainty for Asia-Pacific Economies By Iris Claus; Leo Krippner
  6. Anchors aweigh? The effect of communicating forecast uncertainty By Michael McMahon; Matthew Naylor; Ryan Rholes; Peter Rickards
  7. Developing a Standard Strategy for Time Series Forecasting Integrating Statistical and Machine Learning Techniques Using a Meta-Model Approach and its Application in Generating External Debt Projections By Joshua Elias Fred A. Suero
  8. An Open Benchmark for Evaluating Time Series Forecasting Methods across Financial Markets By Jeremy Bejarano; Viren Desai; Kausthub Keshava; Arsh Kumar; Zixiao Wang; Vincent Hanyang Xu; Yangge Xu
  9. Cross-Sectional Heterogeneity in LSTM Networks for Financial Time Series By Julius D\"obelt
  10. Experimenting with Large Language Models for Inflation Forecasting in Colombia By Aaron L. Garavito-Acosta; Edgar Caicedo-Garcia; Wilmer Martinez-Rivera; Juan J. Ospina-Tejeiro
  11. Quantile-Covariance Three-Pass Regression Filter By Pedro Isaac Chavez-Lopez; Tae-Hwy Lee
  12. Forecasting the Cost of a Basic Basket of Goods: a comparative analysis using machine learning models and online prices By João Anderson da Silva Felix; Michel Alexandre; Cássio da Nóbrega Besarria
  13. From Financial Sentiment Classification to Return Predictability: A QLoRA Benchmark of Large Language Models By Fusheng Luo
  14. Calibration-Induced Degeneracy in LLM Financial Forecasting: An Audit-Trailed Case Study on Next-Day Market Risk By Arin Mohanty
  15. Return Predictability, Expectations, and Investment: Experimental Evidence By Andries, Marianne; Bianchi, Milo; Huynh, Karen; Pouget, Sebastien
  16. Developing a house price-at-risk framework for the UK By Tihana Škrinjarić
  17. Inflation Expectations of Households: Adaptive, Rational, or Sticky? Evidence from an Emerging Market Economy​​ By Eddie Boy L. Fuentes; Mary Kryslette C. Bunyi; Cherrie R. Mapa
  18. Uncertainty, Anchoring, and Expectations Formation: Experimental Evidence​ By Benjamin E. Radoc, Jr.; Sarah Lynne S. Daway-Ducanes

  1. By: Amadeo Grob; Maurizio Daniele; Johanna Ziegel
    Abstract: Traditional statistical tests are poorly suited for the sequential evaluation of probabilistic forecast calibration. We address this limitation in macroeconomic forecasting by applying a new sequential testing method based on e-values. The e-value-based methodology enables anytime-valid inference. It allows practitioners to test against calibration continuously without invalidating statistical guarantees. To illustrate the framework's practical value, we apply it to probabilistic inflation forecasts for the United States, the Euro Area, and Switzerland. Our analysis shows that the sequential approach gives detailed insights into the timing and nature of forecast misspecification. We find these diagnostics are particularly insightful during major structural breaks. During these events, we find evidence against calibration that static, full-sample tests often miss. Therefore, this work shows that e-value-based tests are a practical method for the evaluation of forecast calibration in empirical macroeconomics.
    Date: 2026–08
    URL: https://d.repec.org/n?u=RePEc:arx:papers:2608.23064
  2. By: Fei Shang; Xiaolei Wang; Tomasz Wo\'zniak
    Abstract: We present a suite of R packages for macroeconomic forecasting that leverages advanced Bayesian, structural, multivariate, dynamic, hierarchical, non-linear, and non-Gaussian models. The suite enables both structural and predictive analyses, and is adapted to time series data across various types, dimensions, and sampling frequencies. Each additional feature increases computational complexity. To address this challenge, our software design incorporates a carefully curated selection of models, efficient algorithms implemented in C++, advanced econometric and numerical methods, robust handling of complex input and output objects, and standardised workflows. This approach combines the computational efficiency of C++ with the convenience of working with data in R. We demonstrate that our packages facilitate original research contributions in forecasting, as illustrated by our example in which vector autoregressions with non-centred stochastic volatility enhance density and point predictions relative to models with centred stochastic volatility.
    Date: 2026–08
    URL: https://d.repec.org/n?u=RePEc:arx:papers:2608.28087
  3. By: Davide Brignone (Bank of England); Michele Piffer (Bank of England)
    Abstract: This paper shows how the structural representation of a vector autoregressive (VAR) model can support forecast analysis. We offer a unified framework that formalises how the structural form of the model can help form a narrative for two key statistics in real-time VAR forecasting: the forecast errors relative to the outturn of the data, and the consequent revisions of the forecast. To illustrate the method developed, we conduct a stylised real-time exercise on the UK, focusing on the inflation surge that followed the pandemic. We show that the inflation forecast produced by a four-variable VAR model was revised upwards not only due to contractionary supply-side shocks, but also due to a mix of expansionary demand-side shocks, and a revision in the past shocks.
    Keywords: VAR modelling;forecasting;structural shocks;decomposition
    JEL: C32 E52
    Date: 2026–01–09
    URL: https://d.repec.org/n?u=RePEc:boe:boeewp:023287
  4. By: Niko Hauzenberger; Michael Pfarrhofer
    Abstract: We develop fast methods for conditional forecasting and structural scenario analysis with high-dimensional Bayesian vector autoregressions (VARs). Our general framework features a factor structure on the reduced-form errors, which enables fast and order-invariant equation-by-equation estimation; suitably identified factors admit a structural interpretation. The scenarios are defined through separate distributional restrictions on observables, structural shocks and idiosyncratic components. The computational cost of our proposed algorithm is cubic only in the number of restrictions, while the dimension of the forecasted system enters linearly. In our application with $33$ macroeconomic and financial variables and ten set-identified structural shocks for the US, we compute counterfactual predictions for oil price scenarios in the context of the 2026 closure of the Strait of Hormuz. The same oil price path is consistent with outcomes ranging from a mostly benign episode to pronounced stagflation, depending on which structural and idiosyncratic shocks are allowed to deliver it.
    Date: 2026–07
    URL: https://d.repec.org/n?u=RePEc:arx:papers:2607.29215
  5. By: Iris Claus; Leo Krippner
    Abstract: We develop standardized time-series measures of shocks and uncertainty for output growth and inflation across 14 Asia-Pacific economies using data from Consensus Economics survey of professional economic forecasters. They are based on changes in mean forecasts and standard deviations, adjusted for intra-year patterns that arise from the annual-average percent changes basis of the data. As an example of how our shock measures may be used, we apply the vector autoregression method of sign restrictions to decompose output growth and inflation shocks into fundamental demand and supply shocks. Those shocks and our uncertainty measures align well with major economic events such as the Asian Financial Crisis, the COVID-19 pandemic, and geopolitical conflicts. We show that our uncertainty measures most closely reflect the concept of macroeconomic uncertainty, with apparent differences to established economic policy uncertainty measures across all comparable economies, but a close relationship with two well-known macroeconomic uncertainty measures produced only for the United States. Hence, extending the measures of macroeconomic uncertainty to all 14 economies in our analysis, along with the shocks for those economies, creates a valuable dataset for economic policy setting and empirical research.
    Keywords: macroeconomic shocks; uncertainty; Asia-Pacific; Consensus Economics survey of professional economic forecasters; demand and supply shocks; vector autoregression; sign restrictions
    Date: 2026–08–21
    URL: https://d.repec.org/n?u=RePEc:imf:imfwpa:2026/174
  6. By: Michael McMahon (University of Oxford, CEPR); Matthew Naylor (Bank of England, University of Oxford); Ryan Rholes (University of Mississippi); Peter Rickards (Reserve Bank of Australia, University of Oxford)
    Abstract: We examine how central banks can effectively communicate forecast uncertainty in a two-part experimental study. Part I tests how different visual media – fan charts, dot plots, box-and-whisker plots, speedometers, and ranges – communicate uncertainty to both the general public and expert audiences. We find that fan charts are well understood and perform best at jointly conveying both expectations and uncertainty. Part II implements a novel dynamic information experiment with 1, 600 UK participants across four stages, examining the effects of uncertainty communication on expectations and uncertainty perceptions over time. We find that while point forecasts anchor expectations marginally more than fan charts initially, forecast errors significantly de-anchor expectations, particularly for ‘unlucky’ errors that move inflation away from target. Critically, fan charts materially mitigate this de-anchoring, acting as an ‘insurance policy’ that helps protect central bank reputation. We also document that the public consistently underestimates the degree of uncertainty, and that communicating uncertainty via fan charts helps the public learn more realistic uncertainty perceptions. Our findings have important implications for central bank communication strategies.
    Keywords: Central bank communication;forecast uncertainty;fan charts;expectations;anchoring
    JEL: C91 D83 E52 E58
    Date: 2026–07–17
    URL: https://d.repec.org/n?u=RePEc:boe:boeewp:023318
  7. By: Joshua Elias Fred A. Suero (Bangko Sentral ng Pilipinas)
    Abstract: The pursuit of accurate forecasting and the rise of machine learning have led to the development of various forecasting models, making model selection increasingly difficult. This paper aims to address this challenge by developing a standard strategy for time series forecasting using a meta-model approach. A meta-model is constructed by combining individual forecasts from leading statistical and machine learning models, with Philippine external debt as the target variable. The baseline meta-model combines the following individual forecasts using ordinary least squares (OLS): (a) random walk with drift, (b) autoregressive integrated moving average (ARIMA), (c) exponential smoothing with error, trend, and seasonal components (ETS), (d) Holt-Winters, (e) multiple aggregation prediction algorithm (MAPA), (f) temporal hierarchical forecasting (THieF), (g) theta model, (h) Prophet, (i) neural network autoregression (NNAR), (j) long short-term memory (LSTM), (k) gradient boosting machine (GBM), (l) extreme gradient boosting (XGBoost), (m) random forest, (n) support vector regression (SVR), and (o) dynamic linear models (DLM). Alternative meta-models were also evaluated. The best-performing meta-model, which employs the Least Absolute Shrinkage and Selection Operator (LASSO), performs remarkably well, achieving a mean absolute percentage error (MAPE) of 3.0 percent in the validation set. The proposed standard strategy is versatile, with the potential to forecast other financial and economic variables. Additionally, forecasts generated by the meta-model can serve as a valuable benchmark, whether compared to current forecasting practices or in cases where no forecasting methodology exists.
    JEL: C22 C53 C61 F34
    Date: 2025–04
    URL: https://d.repec.org/n?u=RePEc:bhd:dpaper:202503
  8. By: Jeremy Bejarano; Viren Desai; Kausthub Keshava; Arsh Kumar; Zixiao Wang; Vincent Hanyang Xu; Yangge Xu
    Abstract: An open benchmark that holds financial data fixed across forecasting methods, revealing where machine learning sharpens forecasts of financial stress (Working Paper no. 26-05).
    Date: 2026–08–25
    URL: https://d.repec.org/n?u=RePEc:ofr:wpaper:26-05
  9. By: Julius D\"obelt
    Abstract: Predicting financial asset returns remains one of the most difficult challenges in empirical finance, driven by the low signal-to-noise ratio and the semi-strong form of market efficiency. While deep learning models, especially LSTM networks, have shown promise in capturing temporal dependencies, standard architectures often struggle to account for the cross-sectional heterogeneity of asset returns. This paper proposes a novel architectural extension to the basic LSTM model designed to improve both predictive accuracy and model interpretability. The framework integrates macro-financial covariates to capture broader economic signals and learnable sector embeddings to encompass heterogeneity by sector. The trading strategy involves constructing a long-short portfolio based on daily directional forecasts for each S&P 500 constituent, targeting stocks expected to under- or outperform the cross-sectional median return of the S&P 500. Model Performance is evaluated against three competitive benchmarks: a basic LSTM, a Random Forest model and a traditional market buy-and-hold strategy. The empirical results demonstrate that the LSTM with sector embeddings outperforms all benchmarks across key risk and return metrics. By utilizing sector embeddings, the model explicitly incorporates cross-sectional heterogeneity, allowing it to adapt to varying industry dynamics within the market. To address the black-box nature of deep learning, I use latent space visualizations to analyse how the model differentiates between sectors, providing insights into the internal representation of the sectors in the LSTM. The impact of the sector information can be quantified using a novel contribution metric by inspecting the weights of the LSTM. The predictive signal is driven by a short-term reversal factor and an industry momentum factor.
    Date: 2026–08
    URL: https://d.repec.org/n?u=RePEc:arx:papers:2608.05755
  10. By: Aaron L. Garavito-Acosta (Central Bank of Colombia); Edgar Caicedo-Garcia (Central Bank of Colombia); Wilmer Martinez-Rivera (Central Bank of Colombia); Juan J. Ospina-Tejeiro (Central Bank of Colombia)
    Abstract: This paper develops and applies a standardized framework for forecasting year-on-year inflation in Colombia using large language models (LLMs). We conduct six sequential experiments in which the information set available to the models is progressively expanded by incorporating historical macroeconomic data, contextual indicators, explicit economic structure, and contemporaneous information retrieved through web search. Each configuration is executed daily and generates 24-month inflation forecasts in real time rather than retrospectively, together with qualitative explanations of the forecasts and, in the more advanced configurations, assessments of the shocks affecting inflation. This real-time design mitigates look-ahead bias and produces genuine forecast vintages. The framework is implemented as a programmatic pipeline in Python that queries the OpenAI and Google APIs, executes predefined experiment-specific prompts, and automatically processes and stores the model responses, while applying forecast validation and revision procedures in the more advanced configurations. The results show that richer information environments produce less monotonic inflation paths that remain above the 3% target over the forecast horizon and are more consistent with the contemporaneous domestic and external shocks affecting inflation. The qualitative analysis also shows that the models consistently identify relevant inflation drivers and their interactions. These richer forecast paths are broadly consistent with the pattern observed in survey-based inflation expectations. A formal evaluation of forecast accuracy will be conducted as additional real-time forecast vintages become available.
    Keywords: Large language models; Inflation forecasting; Colombia; Prompt design
    JEL: C53 E01 E31 E37 F31 E23
    Date: 2026–08–25
    URL: https://d.repec.org/n?u=RePEc:gii:giihei:heidwp23-2026
  11. By: Pedro Isaac Chavez-Lopez (Bank of Mexico); Tae-Hwy Lee (Department of Economics, University of California Riverside)
    Abstract: We develop the Quantile-Covariance Three-Pass Regression Filter (Qcov3PRF), a supervised factor model for quantile regression that exploits quantile-covariance (qcov). This method extracts latent factors from a high-dimensional set of predictors to forecast conditional quantiles of a response variable. Unlike Partial Quantile Regression (PQR), Qcov3PRF identifies multiple relevant factors for the target conditional quantiles by qcov. We establish that the resulting forecasts are consistent and asymptotically normal as both the time-series and cross-sectional dimensions diverge. Monte Carlo evidence supports the theoretical results and indicates favorable finite-sample performance. An empirical application to Growth-at-Risk forecasting further demonstrates the advantages of Qcov3PRF over competing alternatives.
    Keywords: Factor models; quantile-covariance; quantile regression; Qcov3PRF; PQR; Growth-at-Risk
    JEL: C13 C22 C53
    Date: 2026–08
    URL: https://d.repec.org/n?u=RePEc:ucr:wpaper:202605
  12. By: João Anderson da Silva Felix; Michel Alexandre; Cássio da Nóbrega Besarria
    Abstract: Online prices can be used to successfully estimate and predict inflation indices represented by the cost of a basket of goods. Nevertheless, machine learning algorithms can deliver better inflation index forecasts than a simple weighted sum of prices due to their ability to handle complex relationships between predictive variables. Using data from five Brazilian state capitals (São Paulo, Porto Alegre, Rio de Janeiro, Goiânia, and Fortaleza) from February 2024 to May 2025, we attempt to predict the cost of a basket of goods using the online prices of the goods that make up such baskets as predictive features. We employ four machine learning models (k-NN, XGBoost, Random Forest, and ridge regression) and a forecast combination technique (Voting Regressor). We also verified the robustness of the machine learning models in situations where it was not possible to obtain online prices for all products. Machine learning models outperform the simple weighted sum of prices in forecasting the overall cost of the food basket, whether all the variables are available or not.
    Date: 2026–08
    URL: https://d.repec.org/n?u=RePEc:bcb:wpaper:650
  13. By: Fusheng Luo
    Abstract: Financial sentiment classifiers are commonly evaluated against human labels, but strong linguistic performance does not necessarily imply economically useful return predictability. This study separates these questions through two experiments. First, we construct a unified three-class benchmark from five financial text datasets and compare TF--IDF Naive Bayes, off-the-shelf FinBERT and Financial-RoBERTa encoders, zero-shot Qwen2.5-7B, and QLoRA-adapted Qwen2.5-7B, LLaMA3-8B, and Mistral-7B models. Mistral-7B achieves the best test accuracy (0.8840) and macro-F1 (0.8771), while QLoRA raises Qwen2.5's macro-F1 from 0.7274 to 0.8615. An inverse-frequency class-weighted loss does not improve Qwen2.5. Second, we evaluate economic validity on a temporally separate 2019 Benzinga sample containing 10, 637 unique headlines and 13, 115 headline--stock observations for a fixed S\&P~100 universe. Model probabilities are converted into continuous sentiment scores, aggregated by stock and signal date, and aligned with next-session returns over one-, two-, three-, and five-day horizons. All seven downstream models produce positive but small mean rank information coefficients at the one-day horizon; the largest is 0.0143 for FinBERT. None of the 28 model--horizon tests remains significant after Newey--West inference and false-discovery-rate correction. Portfolio results likewise fail to establish a robust advantage for the best-performing classifiers. The findings show that QLoRA is effective for financial sentiment adaptation, while also documenting a clear gap between classification accuracy and tradable cross-sectional signals.
    Date: 2026–08
    URL: https://d.repec.org/n?u=RePEc:arx:papers:2608.04200
  14. By: Arin Mohanty
    Abstract: Costly LLM features matter only if calibration lets them affect the forecast. We document a failure of this link in a next-day risk study of two broad-market funds. Full-history scoring preceded the 2022 calibration. Calibration then set all four LLM weights to zero. The 856 later scores therefore could not affect the evaluation. We call this calibration-induced degeneracy. Allowing signed weights reactivated all four mappings. None improved forecasts after familywise correction. By contrast, a near-zero-cost headline count reduced SPY variance-forecast loss by 0.001720 (95 percent familywise interval: [0.000719, 0.002830]). The cheap baseline is therefore a critical diagnostic. We propose a calibration-viability checkpoint. Fit the mapping, perturb the feature over prespecified calibration values, and require a meaningful forecast response before acquiring holdout features. The check uses no holdout outcomes. Here, it would have stopped the paid full-history inference phase.
    Date: 2026–08
    URL: https://d.repec.org/n?u=RePEc:arx:papers:2608.20304
  15. By: Andries, Marianne; Bianchi, Milo; Huynh, Karen; Pouget, Sebastien
    Abstract: In an investment experiment, we show variations in information affect belief and decision behaviors within the information-beliefs-decisions chain. Subjects observe the time series of a risky asset and a signal that, in random rounds, helps predict returns. When they perceive the signal as useless, subjects form extrapolative forecasts, and their investment decisions underreact to their beliefs. When they perceive the signal as predictive, the same subjects rationally use it in their forecasts, they no longer extrapolate, and they rely significantly more on their forecasts when making risk allocations. Analyzing investments without observing forecasts and information sets leads to erroneous interpretations.
    Date: 2024–07
    URL: https://d.repec.org/n?u=RePEc:cpr:ceprdp:19239
  16. By: Tihana Škrinjarić (Bank of England)
    Abstract: This paper develops a house price-at-risk framework for the UK. The model allows me to track and decompose different parts of the distribution of house price growth. The analysis covers both the national level, and nine English regions, along with Wales, Scotland, and Northern Ireland. I employ a comprehensive set of variables and indicators that could help to explain house price dynamics. My main findings are that since the 1970s, the most important predictors for the tail of the distribution have been transaction growth, changes in mortgage rate, credit to GDP gap, and financial stress. I utilise several forecasting horizons and demonstrate that this framework can be applied to forecast downside risks to house price growth and the probability of negative growth up to two years ahead. At the regional level, the analysis reveals considerable variation in the estimated coefficients for mortgage interest rates, with supply-inelastic regions showing higher values than other areas. Finally, I find that an increase in the housing supply in most regions is associated with subsequent easing of price pressures in regional markets.
    Keywords: House price dynamics;financial stability;quantile regression;sub-national house price growth
    JEL: C22 E32 E44 E58 G01 G28
    Date: 2026–06–26
    URL: https://d.repec.org/n?u=RePEc:boe:boeewp:023315
  17. By: Eddie Boy L. Fuentes (Bangko Sentral ng Pilipinas); Mary Kryslette C. Bunyi (Bangko Sentral ng Pilipinas); Cherrie R. Mapa (Bangko Sentral ng Pilipinas)
    Abstract: Understanding the formation of inflation expectations is crucial for monetary policymaking, given that these expectations shape inflation dynamics. Using Bangko Sentral ng Pilipinas (BSP) Consumer Expectations Survey (CES) data, we test whether Filipino households’ inflation expectations can be characterized as adaptive, rational, or sticky. Empirical evidence indicates that households are partly adaptive and not fully rational. However, the survey data are more consistent with the sticky information model, which posits that agents are generally inattentive and infrequently update their forecasts. Results from an epidemiological model, validated using Bayesian posterior estimation, suggest that half of the households retain lagged, previous--quarter expectations; a third follow “rational†professional forecasts; and a fifth update adaptively based on the latest inflation data. We also estimate a state--dependent model, which shows that agents update forecasts more frequently during inflationary periods.
    JEL: D10 D84 E31 E58
    Date: 2025–11
    URL: https://d.repec.org/n?u=RePEc:bhd:dpaper:202510
  18. By: Benjamin E. Radoc, Jr. (Bangko Sentral ng Pilipinas); Sarah Lynne S. Daway-Ducanes (University of the Philippines School of Economics)
    Abstract: The important role of expectations in intertemporal decision making has long been recognized but disagreement among economists on how expectations are formed persists. We conducted an online learning to forecast experiment to determine the impact of uncertainty (in terms of market volatility and number of players) and anchoring (or a non-binding target price band) on the quality of price forecasts. We find that forecast errors are significantly higher in more volatile markets; lower in the presence of a non-binding price bandwidth, suggesting an anchoring effect on expectations; and lower in later markets, confirming the importance of learning and e xperience. Employing a two-step system generalized method of moments, we further confirm these results, and also find that players make systematic forecast errors, in contrast to what is predicted by rational expectations hypothesis.
    JEL: C91 E31 E71
    Date: 2025–01
    URL: https://d.repec.org/n?u=RePEc:bhd:dpaper:202504

This nep-for issue is ©2026 by Malte Knüppel. It is provided as is without any express or implied warranty. It may be freely redistributed in whole or in part for any purpose. If distributed in part, please include this notice.
General information on the NEP project can be found at https://nep.repec.org. For comments please write to the director of NEP, Marco Novarese at <director@nep.repec.org>. Put “NEP” in the subject, otherwise your mail may be rejected.
NEP’s infrastructure is sponsored by the Griffith Business School of Griffith University in Australia.