|
on Computational Economics |
| By: | Gabaix, Xavier; Koijen, Ralph; Richmond, Robert; Yogo, Motohiro |
| Abstract: | Firm characteristics, based on accounting and financial market data, are commonly used to represent firms in economics and finance. However, investors collectively use a much richer information set beyond firm characteristics, including sources of information that are not readily available to researchers. We show theoretically that portfolio holdings contain all relevant information for asset pricing, which can be recovered under empirically realistic conditions. Such guarantees do not exist for other data sources, such as accounting or text data. We build on recent advances in artificial intelligence (AI) and machine learning (ML) that represent unstructured data (e.g., text, audio, and images) by high-dimensional latent vectors called embeddings. Just as word embeddings leverage the document structure to represent words, asset embeddings leverage portfolio holdings to represent firms. Thus, this paper is a bridge from recent advances in AI and ML to economics and finance. We explore various methods to estimate asset embeddings, including recommender systems, shallow neural network models such as Word2Vec, and transformer models such as BERT. We evaluate the performance of these models on three benchmarks that can be evaluated using a single quarter of data: predicting relative valuations, explaining the comovement of stock returns, and predicting institutional portfolio decisions. We also estimate investor embeddings (i.e., representations of investors and their strategies), which are useful for investor classification, performance evaluation, and detecting crowded trades. We discuss other applications of asset embeddings, including generative portfolios, risk management, and stress testing. Finally, we develop a framework to give an economic narrative to a group of similar firms, by applying large language models to firm-level text data. |
| Keywords: | Artificial intelligence; Asset pricing; Machine learning; Transformer models |
| JEL: | C53 G12 G23 |
| Date: | 2025–03 |
| URL: | https://d.repec.org/n?u=RePEc:cpr:ceprdp:20082 |
| By: | Liao, Yuan; Ma, Xinjie; Neuhierl, Andreas; Schilling, Linda |
| Abstract: | Machine learning in asset pricing typically predicts expected returns as point estimates, ignoring uncertainty. We develop new methods to construct forecast confidence intervals for expected returns obtained from neural networks. We show that neural network forecasts of expected returns share the same asymptotic distribution as classic nonparametric methods, enabling a closed-form expression for their standard errors. We also propose a computationally feasible bootstrap to obtain the asymptotic distribution. We incorporate these forecast confidence intervals into an uncertainty-averse investment framework. This provides an economic rationale for shrinkage implementations of portfolio selection. Empirically, our methods improve out-of-sample performance. |
| JEL: | G12 C45 C58 |
| Date: | 2025–03 |
| URL: | https://d.repec.org/n?u=RePEc:cpr:ceprdp:20080 |
| By: | Fr\'ed\'eric Godin |
| Abstract: | The deep hedging framework of Buehler et al. (2019) trains a neural network policy, via Monte Carlo simulation of price paths and stochastic gradient descent, to minimize a risk measure applied to the terminal hedging error. In a recent stream of papers, my coauthors and I have referred to this technique as reinforcement learning (RL), a characterization that referees on several submissions have challenged on two grounds, among others: first, that because feedback is generated only at the terminal date, with no intermediate reward signal, the method cannot constitute genuine RL; and second, that the absence of a value function, a Bellman equation, temporal-difference (TD) learning, and an explicit exploration mechanism disqualifies the method from the RL category altogether, so that it should instead be labeled a neural-network method for stochastic optimal control. I argue that both objections rest on an unduly narrow, TD-centric reading of what constitutes RL. Once RL is understood, as it is in the standard references of the field, to include Monte Carlo policy-gradient methods and direct (actor-only) policy search as first-class members, the deep hedging algorithm of Buehler et al. (2019) falls squarely within the RL umbrella. |
| Date: | 2026–07 |
| URL: | https://d.repec.org/n?u=RePEc:arx:papers:2607.13353 |
| By: | Hauzenberger, Niko; Huber, Florian; Klieber, Karin; Marcellino, Massimiliano |
| Abstract: | We propose a method to learn the nonlinear impulse responses to structural shocks using neural networks, and apply it to uncover the effects of US financial shocks. The results reveal substantial asymmetries with respect to the sign of the shock. Adverse financial shocks have powerful effects on the US economy, while benign shocks trigger much smaller reactions. Instead, with respect to the size of the shocks, we find no discernible asymmetries. |
| Keywords: | Bayesian neural networks |
| JEL: | C11 C30 C45 E3 E44 |
| Date: | 2025–02 |
| URL: | https://d.repec.org/n?u=RePEc:cpr:ceprdp:19964 |
| By: | Anamol Khadka; Milan Arjel; Ayush Lataula; Aayam Dhakal; Prajun Trital; Mingmar Sherpa; Biman Rimal |
| Abstract: | This study examines the dynamic relationship between the global oil prices and Nepal Stock Exchange (NEPSE) using an integrated approach which combines traditional econometric techniques with machine learning and explainable AI techniques. For this, Daily data of International Oil prices and NEPSE index is analyzed from approximately thirteen years (June 2013 to June 2026) using Granger causality, EGARCH(1, 1), and DCC-GARCH models to examine different properties like predictive relationships, asymmetric volatility behaviour, and time-varying correlations. To further supplement the econometric analysis, Machine Learning Models like Random Forest, LightGBM, and XGBoost algorithms were used to capture nonlinear relationships, along with explainable artificial intelligence techniques like SHAP values, Partial Dependence Plots, and Individual Conditional Expectation plots to further interpret the results of the model. The results from the econometric analysis showed a statistically significant unidirectional Granger causality from Brent crude oil to NEPSE with a four-day lag, high volatility persistence in both markets, and weak yet highly time-varying conditional correlations. Among the machine learning models, XGBoost achieves the best performance, and explainability analysis reveals that NEPSE own momentum and short-term volatility mainly influence its own behaviour and oil-related information serves as a minor, method-dependent contributor. The findings demonstrate that econometric and explainable machine learning approaches provide insights into the oil and equity market relationship in a way that each approach complements the result of one another. |
| Date: | 2026–07 |
| URL: | https://d.repec.org/n?u=RePEc:arx:papers:2607.11922 |
| By: | Catherine Chen; Chen Gao; Jonathon Hazell; Lihua Lei; Chen Lian |
| Abstract: | Does microeconomic heterogeneity help to forecast aggregate inflation in a non-stationary environment? We develop a scan test for whether one forecast outperforms another, over an interval with unknown starting point and duration. To exploit any occasional forecasting power that the scan test detects, we design an adaptive machine learning pipeline. We encode the distribution of price changes into a high-dimensional vector, which we combine with a gradient boosted trees algorithm. We then combine this micro forecast with other benchmark forecasts, using an adaptive algorithm that makes use of the micro forecast only when it performs well. We apply the pipeline to UK microdata, with four main results. First, the micro forecast outperforms a univariate benchmark, but only in the volatile period after 2020. Second, the scan test detects periods of micro outperformance, so the micro forecast enters the combined forecast. Third, the combined forecast performs comparably to the univariate benchmark before 2020 and better at every horizon after 2020. Fourth, the value of microdata for the combined forecast materializes after 2020. We conclude that microdata are valuable for forecasting aggregate inflation, but only after large shocks. |
| Date: | 2026–07 |
| URL: | https://d.repec.org/n?u=RePEc:arx:papers:2607.12345 |
| By: | Igor Halperin; Andrey Itkin |
| Abstract: | This paper introduces a dynamic portfolio optimization framework for large institutional investors using Scientific Physics-Informed Reinforcement Learning (SciPhyRL). Formulated in continuous time over an extended state space that includes explicit cumulative costs, the approach leverages offline historical data to learn optimal, distribution-aware strategies. A core innovation reduces the optimization challenge to solving an HJB equation by projecting it onto observed trajectories as a pathwise Hamilton-Jacobi equation. This is solved directly from data using PINN in a single offline sweep, eliminating the need for traditional value or policy iteration. To make the method effective at practical short horizons, the control variable is recast from a continuous trading rate to a discrete target holding. This ensures signal-implied positions are reached immediately, while execution costs are evaluated against a microstructure-grounded quadratic price impact model. Evaluated on a $14$-asset ETF universe using an engineered oracle signal, the learned Gibbs policy yields substantial out-of-sample Sharpe ratio improvements over static and myopic baselines. The results demonstrate that the proposed framework successfully translates known signal quality into a robust, multi-period, and cost-aware allocation mechanism with strictly controlled volatility and turnover. |
| Date: | 2026–07 |
| URL: | https://d.repec.org/n?u=RePEc:arx:papers:2607.15195 |
| By: | Beroud, Mohammed; Qiu, Feng; Wichmann, Bruno; Fan, Xiaoli |
| Abstract: | This paper develops a Machine Learning (ML) approach to quantify the impact of geospatial data quality on farmland valuation accuracy. The approach is useful when both the set of quality attributes and the predictor set are high-dimensional relative to the sample size, making standard non-market valuation methods difficult to implement. Conceptually, we draw on value of information (VoI) theory, which defines value as the difference in expected payoff under higherquality versus imperfect information. The VoI measure can be approximated as the difference between predicted farmland value under full-quality data and predicted value under incomplete data. We apply the approach to soil data and farmland values in the Canadian Prairies. Baseline and counterfactual values are predicted using a Random Forest model with hyperparameters tuned on spatially blocked folds and evaluated with blocked spatial cross-validation to limit spatial leakage. Counterfactual quality scenarios are implemented as random perturbations of soil features. Specifically, we simulate six scenarios: spatially localized bias, limited geographic coverage, coarse spatial resolution, measurement error, numeric rounding, and categorical misclassification. The results show that coarse spatial resolution generates the largest average valuation distortion (238.46 CAD/ha, 2006 CAD), followed by limited coverage (104.49 CAD/ha), while the remaining degradations have small effects. Quality–value curves traced over degradation intensities are nonlinear and concave, consistent with diminishing marginal returns to information improvements. The findings have policy implications for prioritizing investments in public geospatial data: budgets may yield higher returns by shifting from incremental precision upgrades toward improving spatial coverage and resolution. |
| Keywords: | Research Methods/ Statistical Methods |
| Date: | 2026 |
| URL: | https://d.repec.org/n?u=RePEc:ags:aaea26:404720 |
| By: | David Imhof; Thierry Madi\`es; Martin Huber |
| Abstract: | This paper analyzes the internal organization and economic effects of a bid-rigging cartel in the road construction sector of the Swiss canton of Ticino, active from 1999 to 2005. Using exceptionally rich documentary evidence, we reconstruct how cartel members coordinated bids and allocated contracts under a formal agreement known as the 'convention'. We show that, despite the absence of side payments, the cartel implemented a cost-based allocation mechanism that closely approximated the first-best collusive outcome. Regression and machine-learning analyses indicate that observable cost proxies systematically predict both winning bids and bid rankings. The evidence further suggests that cartel members strategically mimicked competitive bidding behavior, allowing them to evade standard econometric detection methods. Using double machine learning, we estimate average overcharges of at least 45\%, and potentially substantially higher, highlighting the significant financial harm caused by this sophisticated form of collusion. |
| Date: | 2026–06 |
| URL: | https://d.repec.org/n?u=RePEc:arx:papers:2606.30470 |
| By: | Hauzenberger, Niko; Marcellino, Massimiliano; Pfarrhofer, Michael; Stelzer, Anna |
| Abstract: | We develop Bayesian machine learning methods for mixed data sampling (MIDAS) regressions. This involves handling frequency mismatches and specifying functional relationships between many predictors and the dependent variable. We use Gaussian processes (GPs) and compress the input space with structured and unstructured MI-DAS variants. This yields several versions of GP-MIDAS with distinct properties and implications, which we evaluate in short-horizon now- and forecasting exercises with both simulated data and data on quarterly US output growth and inflation in the GDP deflator. Our proposed framework leverages macroeconomic Big Data in a computationally efficient way and offers gains in predictive accuracy along several dimensions. |
| JEL: | C11 C22 C53 E31 E37 |
| Date: | 2025–02 |
| URL: | https://d.repec.org/n?u=RePEc:cpr:ceprdp:19965 |
| By: | Maria Saveria Mavillonio; Stefano Borgioli; Caterina Giannetti; Chiara Ongari; Giampiero M. Gallo |
| Abstract: | Measuring sentiment from financial news is a central task in economics and finance, yet most existing indicators rely on dictionary-based approaches that infer sentiment from word counts and only partially capture context, negation, and semantic structure. This paper proposes a framework for constructing daily news mood indices using transformer-based language models and evaluates whether they better represent sentiment than dictionary-based alternatives. Using 143, 755 financial news articles from Factiva, we classify sentiment at the sentence level with FinBERT and aggregate these predictions into article-level and daily sentiment measures through alternative normalization schemes. We compare the resulting indices with benchmark measures based on Shapiro et al., 2022 and Barbaglia et al., 2025. A central contribution is the validation of alternative sentiment measures against human judgments. We conducted an incentivized annotation exercise in which 444 participants evaluated a validation subsample of 588 financial news articles. Consensus ratings from independent human evaluations serve as an external benchmark for assessing the quality of automated sentiment measures. Across correlation, regression, and classification exercises, transformer-based measures show stronger agreement with human judgments than vocabulary-based alternatives and perform substantially better in distinguishing positive, neutral, and negative articles. Overall, the results suggest that incorporating contextual information through transformer-based language models produces sentiment measures that more closely reflect human assessments of financial news. |
| Date: | 2026–07 |
| URL: | https://d.repec.org/n?u=RePEc:arx:papers:2607.13968 |
| By: | Andrii Babii; Luca Barbaglia; Eric Ghysels; Jonas Striaukas |
| Abstract: | This paper develops the asymptotic theory for high-dimensional panel data regressions in settings with cross-sectionally dependent errors driven by common shocks. We consider a factor-augmented sparse-group LASSO estimator that combines MIDAS aggregation with latent factors. The estimator can take advantage of the mixed-frequency group structure in the time-series dimension. Theory shows that it can outperform the standard LASSO estimator both for prediction and estimation while allowing for cross-sectional dependence. |
| Date: | 2026–07 |
| URL: | https://d.repec.org/n?u=RePEc:arx:papers:2607.06368 |