|
on Forecasting |
| By: | Charisios Grivas; Mikkel Mandrup; Orimar Sauri |
| Abstract: | The paper considers the problem of variable selection for forecasting electricity spot prices. High-dimensional methods such as LASSO and Elastic Net are widely used for this purpose, and while they exhibit strong predictive performance, their tendency to select over-parameterized models raises questions about interpretability. We evaluate the performance of six variable selection procedures, includingthe recently proposed Boosting Multiple Testing (BMT) method, using an extensive dataset from six regional electricity markets. We assess their performance in terms of both out-of-sample forecasting ac-curacy and model parsimony. We find that, although LASSO and Elastic Net achieve similar accuracy and outperform most screening alternatives, BMT matches their forecasting performance while using less than one-tenth as many variables. Our results reveal that BMT offers researchers and practitioners a substantially more interpretable and computationally efficient alternative to shrinkage methods, without any loss of forecasting accuracy. These findings suggest that the over-parameterization typically associated with regularization methods is not a necessary price for predictive accuracy in electricity price forecasting. |
| Date: | 2026–08 |
| URL: | https://d.repec.org/n?u=RePEc:arx:papers:2608.09213 |
| By: | Jeron Tan Kang |
| Abstract: | This paper compares direct-yield and factor-based approaches to U.S. Treasury yield curve forecasting using a common high-dimensional macroeconomic information set. Forecasts are evaluated on monthly zero-coupon yields over the 2015-2025 out-of-sample period. Gains over the random walk are concentrated at short maturities and in slope forecasts, and decline with the forecast horizon. Direct-yield models perform best for slope forecasts and are relatively stronger at short horizons, while factor-based models become more competitive at longer horizons. Macroeconomic predictors provide clear incremental predictive power, strongest for slope-related movements. A trading simulation reinforces that macro-augmented models perform best in slope trades. The simulation also highlights a gap between statistical and economic performance, as the random walk is a strong benchmark under statistical loss but performs poorly as a trading signal. |
| Date: | 2026–07 |
| URL: | https://d.repec.org/n?u=RePEc:arx:papers:2608.07536 |
| By: | Thomas R. Cook; Mariia Dzholos; Johannes Matschke |
| Abstract: | Trade flows are volatile and notoriously difficult to predict. This paper introduces trimmed import and export growth rates for the United States to help improve trade forecasts. We calculate the growth rates as a trimmed mean that strips out subcomponents with less predictive power. By separating the signal from the noise, our trimmed growth rates provide a better perspective on where aggregate imports or exports are heading, particularly at horizons four to 12 months ahead. Consequently, out-of-sample forecasts based on our trimmed growth rates reduce forecast errors by up to 15 percent relative to forecasts from aggregate import or export growth. These trimmed forecasts are simple to compute, outperform both naive forecasts and models that rely on a richer information set, and are particularly useful during episodes with heightened volatility. We update the trimmed growth rates monthly and make them publicly available. |
| Keywords: | imports and exports; forecasting; signal extraction |
| JEL: | F17 C22 |
| Date: | 2026–08–18 |
| URL: | https://d.repec.org/n?u=RePEc:fip:fedkrw:103656 |
| By: | Rishab Ghosh; Vinay Devarakonda |
| Abstract: | Large language models (LLMs) are increasingly used as components of agentic systems that observe, plan, and act. In finance, even "assistive" systems become decision-relevant once their outputs are used to size trades or allocate risk. A key failure mode is the confidence--competence gap: a model that is only slightly better than chance but consistently overconfident will, under typical bet-sizing rules, generate negative long-run growth. Existing benchmarks emphasize semantic understanding or point accuracy, but do not directly test probabilistic calibration under the temporal constraints and non-stationarity that define real markets. We introduce FinBench, a benchmark designed to evaluate calibration and uncertainty quality for financial forecasting in a setting that is (i) strictly time-gated to avoid look-ahead bias and (ii) evaluated with strictly proper scoring rules that penalize hallucinated confidence. FinBench tasks require models to output (a) a probability of positive return and (b) an 80% prediction interval for realized log return; evaluation uses the Brier score and the Winkler interval score, along with skill scores against hard baselines. This paper describes the benchmark specification and reports a small pilot run (one trading day; three liquid tickers; 33 forecasts) as a sanity check of the pipeline. The pilot illustrates how calibration-sensitive metrics distinguish between "confident but fragile" behavior and uncertainty-aware forecasting. |
| Date: | 2026–06 |
| URL: | https://d.repec.org/n?u=RePEc:arx:papers:2607.16229 |
| By: | Ulrich Hounyo; Zhendong Li |
| Abstract: | Factor-MIDAS regressions forecast a low-frequency target by extracting common factors from a large panel of high-frequency predictors via principal component analysis (PCA). While PCA mitigates the curse of dimensionality, it relies on factor pervasiveness, an assumption often violated when factors are weak, as is common in macro-financial forecasting. We propose SsPCA-MIDAS, which integrates supervised scaled PCA (SsPCA) into the mixed-data sampling framework. We establish consistency and asymptotic normality under weak factors, permitting inference on the prediction target. Simulations show that SsPCA-MIDAS outperforms competing PCA-based and supervised methods, especially when weak factors are prevalent. Applying machine-learning techniques such as boosting to the cleaner factors it extracts yields further gains. An extensive application to U.S. macro-financial forecasting shows that SsPCA-MIDAS selects economically meaningful predictors and improves forecasts of GDP, inflation, unemployment, asset prices, and volatility. |
| Date: | 2026–08 |
| URL: | https://d.repec.org/n?u=RePEc:arx:papers:2608.12589 |
| By: | Lifeng Hao; Shaolin Ji |
| Abstract: | Implied volatility surface forecasting is essential for option valuation, hedging, and risk management, but remains difficult because future surfaces are stochastic while pricing inputs must satisfy static no-arbitrage shape restrictions. We propose a decoupled generative refinement framework for IVS forecasting as an operational risk surface modeling problem. The first stage uses a conditional diffusion model to learn the conditional distribution of future surfaces. The generated ensemble captures predictive distributional variation, and its median provides a robust representative surface for subsequent refinement. The second stage introduces a Surface Aware Attention Module (SAAM), a cross sectional refinement operator that improves fit to market observations and staticno-arbitrage diagnostics for the representative surface. This design separates distribution learning from surface refinement, allowing the diffusion model to capture stochastic market dynamics while SAAM controls static no-arbitrage residual violations on the final surface. We evaluate the framework on CSI 300 index options from June 2020 to September 2024 under daily and minute level forecasting protocols. The diffusion stage improves forecasting accuracy and produces predictive intervals that vary across moneyness, maturity, and sampling frequency. The refinement stage improves fitting accuracy against market observations and reduces measured static no-arbitrage residual violations, with stronger gains at the minute level. Attention diagnostics suggest that SAAM performs adaptive cross sectional refinement rather than fixed local smoothing |
| Date: | 2026–07 |
| URL: | https://d.repec.org/n?u=RePEc:arx:papers:2607.29220 |
| By: | Ekkehardt Bauer; Dirk Holl\"ander; David Scholz; Linus Wolff; Christoph Ostermair; Kyrillus Aiad; Joachim Hasebrook |
| Abstract: | This study focuses on developing an AI-supported prototype for multiperspective interest rate forecasting that combines classical econometric models with modern artificial intel-ligence methods. Tested in a major European bank, the system enables more precise and flexible prediction of interest rate developments, supporting strategic decision-making in Asset-Liability Management (ALM). It integrates topic modeling, sentiment analysis, econometric forecasting, and market-based analyses within an interactive platform. Leveraging AI to analyze large volumes of financial documents and market data enables the identification of monetary policy trends and sentiment signals at an early stage. The core econometric model is a Bayesian vector autoregression (BVAR) that enables simulation-based scenario analyses to evaluate economic developments from multiple perspectives. The system's innovation lies in its integration of several forecasting approaches that consolidate previously separate information sources and present them transparently and interpretably. Financial analysts and risk managers thus gain a better basis for making decisions, allowing them to assess interest rate risks more accurately and manage market movements more proactively. While the prototype demonstrates how AI can transform interest rate management in banking, further development is required to optimize real-time data integration and regulatory compliance. Even at this stage, the study shows that multi-perspective, AI-driven forecasting provides substantial added value for banks by increasing transparency, strengthening evidence-based decision-making, and improving risk management. |
| Date: | 2026–08 |
| URL: | https://d.repec.org/n?u=RePEc:arx:papers:2608.12424 |
| By: | Aditya Dutta |
| Abstract: | Production forecasting systems retrain models regularly, but a retrained candidate does not necessarily outperform a continuously maintained incumbent that has continued to learn. We introduce Shadow Before Swap (SBS), a deployment policy that warm-refits a challenger off the serving path, evaluates it against the maintained incumbent on the same next week of delayed labels, and promotes it only after a fixed paired negative-log-likelihood (NLL) advantage. In historical replay over two nonoverlapping Binance episodes spanning 48 UTC weeks, three seeds, eight underlyings, and two perpetual-futures contract types, SBS reduces NLL by 0.1472% relative to calendar replacement, 0.0755% relative to schedule-matched automatic promotion, and 0.0428% relative to continuous maintenance. The corresponding episode-stratified four-week block intervals are 0.1139%-0.1754%, 0.0521%-0.0980%, and 0.0301%-0.0554%, respectively. SBS promotes 114 of 528 challengers, reducing deployed model changes by 78.4% while improving the serving trajectory. The effect remains directionally consistent across seeds, trial budgets, promotion margins, an earlier 20-asset panel, and a topology-matched supervised objective. SBS thus provides a practical deployment policy that improves probabilistic forecasts while limiting consequential model-state transitions. |
| Date: | 2026–07 |
| URL: | https://d.repec.org/n?u=RePEc:arx:papers:2607.28577 |
| By: | Kwon, Alexander; Maliar, Lilia |
| Abstract: | We demonstrate that machine learning substantially improves predictions of individual decisions about retirement and Social Security (SS) claims. When predicting the number of people receiving SS, we achieve an error of less than 1%, while the benchmark model employed by the Social Security Administration (SSA) results in a greater than 4% error, and in forecasting SS claiming decisions, we attain an error of 0.2%, while the benchmark exceeding 2%. Based on averages, we show that a 3% difference in prediction amounts to 39.6 billion dollars annually. The set of important variables selected by our model significantly differs from that of the SSA model. We use Shapley values to evaluate the non-linear contributions of the selected variables to predictive outcomes. |
| Keywords: | retirement and Social Security |
| JEL: | C53 H55 J14 J26 |
| Date: | 2024–07 |
| URL: | https://d.repec.org/n?u=RePEc:cpr:ceprdp:19198 |
| By: | Amin Izadyar |
| Abstract: | I revisit the exchange rate disconnect puzzle, first documented by Meese and Rogoff (1983), using generative artificial intelligence (AI) to forecast currency returns based on economic fundamentals. Using ChatGPT and DeepSeek, I analyze a comprehensive dataset of economic data releases for major currency pairs and measure the fundamental strength of each currency. These AI-powered fundamentals exhibit significant cross-sectional predictive power. A simple trading strategy that goes long currencies with strong fundamentals and short currencies with weak fundamentals generates a Sharpe ratio exceeding 0.7 per annum. The excess returns of this strategy remain significant after controlling for traditional currency factors. To mitigate concerns of look-ahead bias, I run multiple exercises to ensure that predictability stems from AI reasoning rather than memorization. Finally, I explore the potential sources of predictability and find evidence that the Taylor rule framework, generally used by central banks to set interest rates, is a key mechanism connecting exchange rates to economic fundamentals. |
| Date: | 2026–08 |
| URL: | https://d.repec.org/n?u=RePEc:arx:papers:2608.00761 |
| By: | Kyungsub Lee; Kennedy Titus Kayaki |
| Abstract: | This paper proposes a GARCH-type volatility model in which level-and-slope updates of a latent power-law kernel generate state-dependent decay of past shocks within a two-dimensional Markov state. We derive a joint Foster--Lyapunov condition and establish positive Harris recurrence and uniqueness of the invariant distribution. Simulations show substantial low-frequency persistence in log-squared innovations, especially near the diagnostic stability boundary. Empirically, the model captures a substantial portion of observed volatility persistence and delivers competitive out-of-sample forecast accuracy using only a two-dimensional Markov state. |
| Date: | 2026–07 |
| URL: | https://d.repec.org/n?u=RePEc:arx:papers:2607.25189 |
| By: | De Polis, Andrea; Galvão, Ana Beatriz; Petrella, Ivan |
| Abstract: | Policymakers increasingly rely on real-time measures of economic activity to inform decisions, yet official statistics are typically available only with substantial delay. This paper develops a methodology to extract information from high-frequency indicators in order to produce weekly estimates of official monthly statistics in real time. Building on real-time tracking and nowcasting models, our approach addresses two challenges inherent in alternative indicators: the absence of seasonal adjustment and the prevalence of outliers. We incorporate seasonal components directly into the model, allowing the seasonal structure of low-frequency data to inform high frequency proxies, and employ fat-tailed distributions to mitigate the influence of large, infrequent shocks. Applying our methodology to UK data, we track retail sales, monthly GDP, and vacancies using proxies such as debit card spending (Revolut) and online job advertisements. Our results show that weekly estimates improve real-time prediction of official releases and highlight the usefulness of high-frequency alternative data. |
| Keywords: | real-time tracker; mixed-frequency models; seasonality; state-space models; high-frequency data |
| JEL: | C32 C53 E01 E37 |
| Date: | 2025–11–28 |
| URL: | https://d.repec.org/n?u=RePEc:eoe:escoed:escoe-dp-2025-19 |
| By: | Arnaud Garnier (LHEEA - Laboratoire de recherche en Hydrodynamique, Énergétique et Environnement Atmosphérique - CNRS - Centre National de la Recherche Scientifique - Nantes Univ - ECN - NANTES UNIVERSITÉ - École Centrale de Nantes - Nantes Univ - Nantes Université); Pierre Marty (LHEEA - Laboratoire de recherche en Hydrodynamique, Énergétique et Environnement Atmosphérique - CNRS - Centre National de la Recherche Scientifique - Nantes Univ - ECN - NANTES UNIVERSITÉ - École Centrale de Nantes - Nantes Univ - Nantes Université); Rodica Loisel (LEMNA - Laboratoire d'économie et de management de Nantes Atlantique - Nantes Univ - IAE Nantes - Nantes Université - Institut d'Administration des Entreprises - Nantes - Nantes Université - pôle Sociétés - Nantes Univ - Nantes Université) |
| Abstract: | The pathway towards decarbonisation of shipping is unclear, as many technical, economic, and regulatory challenges remain. This study builds a bottom-up model to forecast the merchant fleet vessel composition and CO2 emissions by 2050. A 35, 000 vessel fleet is modelled based on technical and operational data, on the population pyramid and historical fleet evolution triggered by trade demand. The emissions forecast in a ‘no-action' scenario shows that even low-growth traffic scenarios will largely deviate from the carbon neutrality objectives. It highlights fleet heterogeneity as a key point in understanding and considering global decarbonisation strategy. Fleet renewal analysis revealed technical and planning issues due to the tendency towards larger vessels and high building rates up to 2000 vessels per year from 2040 onwards. Alternatively, retrofitting could significantly contribute to carbon neutrality, concerning up to 40% of the shipping tonnage if the strategy of decarbonisation is not integrated early in shipyard industry planning. |
| Keywords: | Energy consumption model, Shipping decarbonisation, Bottom-up approach, Fleet renewal inertia, Traffic demand scenarios, Emission forecast, AIS data |
| Date: | 2026–09 |
| URL: | https://d.repec.org/n?u=RePEc:hal:journl:hal-05692446 |
| By: | Peter Cotton |
| Abstract: | Conformal prediction has been touted as a more formal, rigorous approach to adding uncertainty to a forecast. The sole objective of this note is to point out that rigor cuts both ways in the case of residual pooling, the technique used in the vast majority of conformal prediction applications. The fact that unconditional guarantee of coverage is provided is not in question, but we make clear, we believe for the first time, that there is an opposing guarantee too: a permanent gambit of logarithmic-score regret which no amount of data or tuning can subsequently reduce. We give the exact size of the sacrifice, and a financial reading of it as the growth rate of an oracle adversary betting against odds set by someone using residual pooling. |
| Date: | 2026–06 |
| URL: | https://d.repec.org/n?u=RePEc:arx:papers:2608.07479 |