nep-big New Economics Papers
on Big Data
Issue of 2026–09–07
eighteen papers chosen by
Tom Coupé, University of Canterbury


  1. EcoFinBench – a natural language processing benchmark for economics and finance By Max Ahrens; Dragos Gorduza; Micheal McMahon
  2. Forecasting the Cost of a Basic Basket of Goods: a comparative analysis using machine learning models and online prices By João Anderson da Silva Felix; Michel Alexandre; Cássio da Nóbrega Besarria
  3. Pathways of Climate Variability, Agricultural Performance, and Conflict: A Machine Learning Approach to Complex Dependencies By Tulia Gattone; Donato Romano; Luca Tiberti
  4. Developing a Standard Strategy for Time Series Forecasting Integrating Statistical and Machine Learning Techniques Using a Meta-Model Approach and its Application in Generating External Debt Projections By Joshua Elias Fred A. Suero
  5. AI and Judicial Productivity: The Impact of MIDAS on the Courts of Fortaleza, Brazil By Pierri, Gastón; Fontenele, Marcelo; Nunes, Jose Luiz
  6. Do Deep Learning Methods Improve Financial Forecasts? By Jeremy Bejarano; Viren Desai; Kausthub Keshava; Arsh Kumar; Zixiao Wang; Vincent Hanyang Xu; Yangge Xu
  7. Reducing Basis Risk in Index Insurance Using Satellite-Based Vegetation Indices and Machine Learning By Ryu, Jaehyeon; Yu, Jisang
  8. An Open Benchmark for Evaluating Time Series Forecasting Methods across Financial Markets By Jeremy Bejarano; Viren Desai; Kausthub Keshava; Arsh Kumar; Zixiao Wang; Vincent Hanyang Xu; Yangge Xu
  9. Communicating Credit Risk with Large Language Models: Evaluation of Explanations from Standard and Alternative Data-Based Models By Sahab Zandi; Noah Kostesku; Christophe Mues; Mar\'ia \'Oskarsd\'ottir; Cristi\'an Bravo
  10. Neural Network Learning for Nonlinear Economies By Ashwin, Julian; Beaudry, Paul; Ellison, Martin
  11. Revealing economic facts: LLMs know more than they say By Marcus Buckmann; Quynh Anh Nguyen; Ed Hill
  12. Inflation attitudes of large language models By Nikoleta Anesti; Edward Hill; Andreas Joseph
  13. From Financial Sentiment Classification to Return Predictability: A QLoRA Benchmark of Large Language Models By Fusheng Luo
  14. Portfolio management with big data By Penaranda, Francisco; Sentana, Enrique
  15. Cross-Sectional Heterogeneity in LSTM Networks for Financial Time Series By Julius D\"obelt
  16. Estimating Heterogeneity in Travel Mode Choice Shifts with Causal Forests By Rishabh Singh Chauhan; Mahdi Ghadimi; Lishun Liu
  17. Financial advice behaviour: humans versus AI By Ylva Baeckström; Roman Matkovskyy
  18. Experimenting with Large Language Models for Inflation Forecasting in Colombia By Aaron L. Garavito-Acosta; Edgar Caicedo-Garcia; Wilmer Martinez-Rivera; Juan J. Ospina-Tejeiro

  1. By: Max Ahrens (Maihem.ai); Dragos Gorduza (Bank of England); Micheal McMahon (University of Oxford)
    Abstract: We introduce EcoFinBench, a natural language processing (NLP) benchmark suite for the domains of economics and finance. We comprehensively test a large array of NLP models across multiple domain-specific data sets for sentence classification. Specifically, we evaluate dictionary models, word count models, topic models, and modern transformer models. Furthermore, we introduce two new data sets to the research community. The Bluebook data set for text-only sentiment analysis in monetary policy, and the Greenbook data set for multimodal (text and numeric) sentiment analysis. We focus on data sets that require the models to work with relatively few data points and long average text lengths – typical characteristics of data sets in the economic and financial domain. We find that, dictionary models – still widely used as a default text analysis tool in economics and finance – underperform substantially across all evaluated data sets. From our findings, we conclude that given the underperformance of existing solutions in the multimodal domain, future modelling work is needed. With our benchmark suite we aim to lay the foundation for a systematic assessment on the most commonly used NLP models in economics and finance. To our knowledge, we are the first to provide such holistic benchmarking assessment for economics and finance.
    Keywords: Machine learning, natural language processing;artificial intelligence;benchmark.
    JEL: C45 Y10 C55 C88 G17
    Date: 2025–12–19
    URL: https://d.repec.org/n?u=RePEc:boe:boeewp:023285
  2. By: João Anderson da Silva Felix; Michel Alexandre; Cássio da Nóbrega Besarria
    Abstract: Online prices can be used to successfully estimate and predict inflation indices represented by the cost of a basket of goods. Nevertheless, machine learning algorithms can deliver better inflation index forecasts than a simple weighted sum of prices due to their ability to handle complex relationships between predictive variables. Using data from five Brazilian state capitals (São Paulo, Porto Alegre, Rio de Janeiro, Goiânia, and Fortaleza) from February 2024 to May 2025, we attempt to predict the cost of a basket of goods using the online prices of the goods that make up such baskets as predictive features. We employ four machine learning models (k-NN, XGBoost, Random Forest, and ridge regression) and a forecast combination technique (Voting Regressor). We also verified the robustness of the machine learning models in situations where it was not possible to obtain online prices for all products. Machine learning models outperform the simple weighted sum of prices in forecasting the overall cost of the food basket, whether all the variables are available or not.
    Date: 2026–08
    URL: https://d.repec.org/n?u=RePEc:bcb:wpaper:650
  3. By: Tulia Gattone; Donato Romano; Luca Tiberti
    Abstract: This study seeks to provide empirical validation of the conceptual framework set by Romano et al. (2025) in which climate variability influences conflict through agricultural and market-mediated channels. Climate variability, measured via the Standardized Precipitation-Evapotranspiration Index (SPEI), is linked to crop yields, crop commercialization, household consumption, and conflict outcomes. All models draw on socio-economic data from the World Bank LSMS-ISA project (waves 1–3, 2010–2016), which allows us to examine how changes in SPEI may translate into shifts in agricultural productivity, market participation, and household welfare. Our empirical strategy proceeds in three stages. We begin by testing the basic structure of the hypothesized relationships through a set of predictive machine learning (ML) models. The core specification relies on Artificial Neural Networks (ANNs), while Random Forest, Support Vector Machines, and Naive Bayes serve as complementary models to check whether the predictive power of the agricultural channel appears consistently across algorithms. In the second stage, we draw on wave 4 of the LSMS-ISA (2018 to 2019) as out-of-sample testing data and employ a stepwise ANN model. This design allows us to examine whether variation in crop yields mediates the effect of SPEI on conflict-related outcomes. In the final stage, we move from prediction to causal inference. Here, we apply Causal Forest, followed by Double ML as a robustness check, to identify heterogeneous effects and to assign substantive meaning to the relationships that earlier predictive models had revealed. Our results point toward a climate–conflict relationship shaped by nonlinearities, mediation through agricultural performance, and marked variation across local economic conditions.
    Keywords: Climate change, conflict, agrifood system, machine learning, artificial neural networks
    JEL: Q54 Q12 D74 C45 C53
    Date: 2026
    URL: https://d.repec.org/n?u=RePEc:frz:wpaper:wp2026_10.rdf
  4. By: Joshua Elias Fred A. Suero (Bangko Sentral ng Pilipinas)
    Abstract: The pursuit of accurate forecasting and the rise of machine learning have led to the development of various forecasting models, making model selection increasingly difficult. This paper aims to address this challenge by developing a standard strategy for time series forecasting using a meta-model approach. A meta-model is constructed by combining individual forecasts from leading statistical and machine learning models, with Philippine external debt as the target variable. The baseline meta-model combines the following individual forecasts using ordinary least squares (OLS): (a) random walk with drift, (b) autoregressive integrated moving average (ARIMA), (c) exponential smoothing with error, trend, and seasonal components (ETS), (d) Holt-Winters, (e) multiple aggregation prediction algorithm (MAPA), (f) temporal hierarchical forecasting (THieF), (g) theta model, (h) Prophet, (i) neural network autoregression (NNAR), (j) long short-term memory (LSTM), (k) gradient boosting machine (GBM), (l) extreme gradient boosting (XGBoost), (m) random forest, (n) support vector regression (SVR), and (o) dynamic linear models (DLM). Alternative meta-models were also evaluated. The best-performing meta-model, which employs the Least Absolute Shrinkage and Selection Operator (LASSO), performs remarkably well, achieving a mean absolute percentage error (MAPE) of 3.0 percent in the validation set. The proposed standard strategy is versatile, with the potential to forecast other financial and economic variables. Additionally, forecasts generated by the meta-model can serve as a valuable benchmark, whether compared to current forecasting practices or in cases where no forecasting methodology exists.
    JEL: C22 C53 C61 F34
    Date: 2025–04
    URL: https://d.repec.org/n?u=RePEc:bhd:dpaper:202503
  5. By: Pierri, Gastón; Fontenele, Marcelo; Nunes, Jose Luiz
    Abstract: This paper presents preliminary results from a pilot study conducted in the courts of Ceará, Brazil. The study evaluates the impact of introducing a tool that uses natural language processing and machine learning techniques to cluster judicial acts by textual similarity on clerk productivity, measured as the number of case files a clerk can produce in a day. Estimates indicate that treatment-group clerks produced approximately 10 more case files per day than control-group clerks, a statistically significant difference equivalent to a 37% increase relative to the control group mean. The results are robust to the exclusion of outlier observations and exceptionally productive clerks.
    Keywords: artificial intelligence;Judicial Productivity;Natural Language Processing;Court Administration;Public Sector Automation;machine learning;Field experiment;access to justice
    JEL: O33 H83 K40 C93 J24
    Date: 2026–07
    URL: https://d.repec.org/n?u=RePEc:idb:brikps:14700
  6. By: Jeremy Bejarano; Viren Desai; Kausthub Keshava; Arsh Kumar; Zixiao Wang; Vincent Hanyang Xu; Yangge Xu
    Abstract: The newest deep learning methods work best when the data contain rich, repeating structures.
    Date: 2026–08–25
    URL: https://d.repec.org/n?u=RePEc:ofr:ofrblg:26-10
  7. By: Ryu, Jaehyeon; Yu, Jisang
    Keywords: Risk and Uncertainty
    Date: 2026
    URL: https://d.repec.org/n?u=RePEc:ags:aaea26:404414
  8. By: Jeremy Bejarano; Viren Desai; Kausthub Keshava; Arsh Kumar; Zixiao Wang; Vincent Hanyang Xu; Yangge Xu
    Abstract: An open benchmark that holds financial data fixed across forecasting methods, revealing where machine learning sharpens forecasts of financial stress (Working Paper no. 26-05).
    Date: 2026–08–25
    URL: https://d.repec.org/n?u=RePEc:ofr:wpaper:26-05
  9. By: Sahab Zandi; Noah Kostesku; Christophe Mues; Mar\'ia \'Oskarsd\'ottir; Cristi\'an Bravo
    Abstract: Credit decisioning is a high-stakes task in which model outputs must be accurate and explainable to support compliant decisions. Although modern credit risk models such as eXtreme Gradient Boosting (XGBoost) and Graph Neural Networks (GNNs) improve predictive performance, their explanations are often too technical for stakeholders creating communication gaps that can shape approvals, denials, and fairness judgments. We examine whether Large Language Models (LLMs) can serve as explanation layers that translate post-hoc explanation artefacts into stakeholder-appropriate risk narratives. Using Freddie Mac single-family loan-level data, we develop three pipelines: standard tabular (XGBoost + SHAP), and two with alternative data, a pure network-based (GNN + GNNExplainer), and a bimodal one (combining tabular and network data). We generate narratives with three LLM configurations: a small fine-tuned LLM (Gemma 3 4B), a large fine-tuned LLM (DeepSeek R1 70B), and a zero-shot commercial LLM (Gemini 2.5). Explanation quality is evaluated through automated checks across all pipelines and a human study of bimodal explanations comparing credit risk professionals and non-professionals on eight decision-relevant dimensions. We have three main findings. First, the pipeline accounts for higher variance in evidence-grounding scores than the language model, meaning that the binding constraint on explanation quality is the evidence representation, not the model used. Second, the explanation narratives reliably name the influential factors but are less reliable when stating the direction of influence, which may be consequential for adverse-action communication. Finally, professionals apply stricter evidentiary standards than non-professionals. We discuss implications for the governance of risk models, including deployment considerations and the value of domain-aligned LLMs in regulated credit settings.
    Date: 2026–08
    URL: https://d.repec.org/n?u=RePEc:arx:papers:2608.17715
  10. By: Ashwin, Julian; Beaudry, Paul; Ellison, Martin
    Abstract: Neural networks offer a promising tool for the analysis of nonlinear economies. In this paper, we derive conditions for the global stability of nonlinear rational expectations equilibria under neural network learning. We demonstrate the applicability of the conditions in analytical and numerical examples where the nonlinearity is caused by monetary policy targeting a range, rather than a specific value, of inflation. If shock persistence is high or there is inertia in the structure of the economy, then the only rational expectations equilibria that are learnable may involve inflation spending long periods outside its target range. Neural network learning is also useful for solving and selecting between multiple equilibria and steady states in other settings, such as when there is a zero lower bound on the nominal interest rate.
    Keywords: Inflation targeting; Machine learning; Neural networks; Zero lower bound
    Date: 2024–07
    URL: https://d.repec.org/n?u=RePEc:cpr:ceprdp:19295
  11. By: Marcus Buckmann (Bank of England); Quynh Anh Nguyen (Bank of England); Ed Hill (Bank of England)
    Abstract: We investigate whether hidden states of large language models (LLMs) can be used to estimate and impute economic and financial statistics. Focusing on county-level (eg unemployment) and firm-level (eg total assets) variables, we show that a linear regression trained on the hidden states of open-source LLMs outperforms the models' own text outputs. This indicates that internal representations encode richer economic information than is revealed directly in generated responses. A learning curve analysis shows that, in many cases, only a few dozen labelled examples suffice for training. We further propose a transfer learning method that improves estimation accuracy without requiring any labelled data for the target variable. Finally, we demonstrate the practical utility of hidden states in data imputation and super-resolution tasks.
    Keywords: Large language models;embeddings;economic statistics;data imputation.
    JEL: C45 C81 C21
    Date: 2025–10–31
    URL: https://d.repec.org/n?u=RePEc:boe:boeewp:023272
  12. By: Nikoleta Anesti (Bank of England); Edward Hill (Bank of England); Andreas Joseph (Bank of England)
    Abstract: This paper investigates the ability of large language models (LLMs), primarily ‘GPT-3.5 Turbo’ (GPT), to form inflation perceptions and expectations based on macroeconomic price signals. We compare the LLM’s output to household survey data and official statistics, mimicking the information set and demographic characteristics of the Bank of England’s Inflation Attitudes Survey (IAS). Our quasi-experimental design exploits the timing of GPT’s training cut-off in September 2021 which means it has no knowledge of the subsequent UK inflation surge. This setting turns out to be crucial to track aggregate survey results and official statistics at short horizons. At a disaggregated level, GPT replicates key empirical regularities of households’ inflation perceptions, particularly for income, housing tenure, and social class. A novel Shapley value decomposition of LLM outputs suited for the synthetic survey setting provides well-defined insights into the drivers of model outputs linked to prompt content. We find that GPT demonstrates a heightened sensitivity to food inflation information like that of human respondents. However, we also find that it lacks a consistent model of consumer price inflation, eg by exhibiting unexplained kinks in component sensitivity. More generally, our approach could be used to evaluate the behaviour of LLMs for use in the social sciences, to compare different models, or to assist in survey design.
    Keywords: Large language models;inflation expectations;household surveys;Shapley values
    JEL: C8 C14 C45 C83 E31
    Date: 2026–06–19
    URL: https://d.repec.org/n?u=RePEc:boe:boeewp:023312
  13. By: Fusheng Luo
    Abstract: Financial sentiment classifiers are commonly evaluated against human labels, but strong linguistic performance does not necessarily imply economically useful return predictability. This study separates these questions through two experiments. First, we construct a unified three-class benchmark from five financial text datasets and compare TF--IDF Naive Bayes, off-the-shelf FinBERT and Financial-RoBERTa encoders, zero-shot Qwen2.5-7B, and QLoRA-adapted Qwen2.5-7B, LLaMA3-8B, and Mistral-7B models. Mistral-7B achieves the best test accuracy (0.8840) and macro-F1 (0.8771), while QLoRA raises Qwen2.5's macro-F1 from 0.7274 to 0.8615. An inverse-frequency class-weighted loss does not improve Qwen2.5. Second, we evaluate economic validity on a temporally separate 2019 Benzinga sample containing 10, 637 unique headlines and 13, 115 headline--stock observations for a fixed S\&P~100 universe. Model probabilities are converted into continuous sentiment scores, aggregated by stock and signal date, and aligned with next-session returns over one-, two-, three-, and five-day horizons. All seven downstream models produce positive but small mean rank information coefficients at the one-day horizon; the largest is 0.0143 for FinBERT. None of the 28 model--horizon tests remains significant after Newey--West inference and false-discovery-rate correction. Portfolio results likewise fail to establish a robust advantage for the best-performing classifiers. The findings show that QLoRA is effective for financial sentiment adaptation, while also documenting a clear gap between classification accuracy and tradable cross-sectional signals.
    Date: 2026–08
    URL: https://d.repec.org/n?u=RePEc:arx:papers:2608.04200
  14. By: Penaranda, Francisco; Sentana, Enrique
    Abstract: The purpose of this survey is to summarize the academic literature that studies some of the ways in which portfolio management has been affected in recent years by the availability of big datasets: many assets, many characteristics for each of them, many macro predictors, and various sources of unstructured data. Thus, we deliberately focus on applications rather than methods. We also include brief reviews of the financial theories underlying asset management, which provide the relevant background to assess the plethora of recent contributions to such an active research field.
    Keywords: Machine learning; Mean-variance analysis; Stochastic discount factors
    JEL: G11 G12 C55 G17
    Date: 2024–07
    URL: https://d.repec.org/n?u=RePEc:cpr:ceprdp:19314
  15. By: Julius D\"obelt
    Abstract: Predicting financial asset returns remains one of the most difficult challenges in empirical finance, driven by the low signal-to-noise ratio and the semi-strong form of market efficiency. While deep learning models, especially LSTM networks, have shown promise in capturing temporal dependencies, standard architectures often struggle to account for the cross-sectional heterogeneity of asset returns. This paper proposes a novel architectural extension to the basic LSTM model designed to improve both predictive accuracy and model interpretability. The framework integrates macro-financial covariates to capture broader economic signals and learnable sector embeddings to encompass heterogeneity by sector. The trading strategy involves constructing a long-short portfolio based on daily directional forecasts for each S&P 500 constituent, targeting stocks expected to under- or outperform the cross-sectional median return of the S&P 500. Model Performance is evaluated against three competitive benchmarks: a basic LSTM, a Random Forest model and a traditional market buy-and-hold strategy. The empirical results demonstrate that the LSTM with sector embeddings outperforms all benchmarks across key risk and return metrics. By utilizing sector embeddings, the model explicitly incorporates cross-sectional heterogeneity, allowing it to adapt to varying industry dynamics within the market. To address the black-box nature of deep learning, I use latent space visualizations to analyse how the model differentiates between sectors, providing insights into the internal representation of the sectors in the LSTM. The impact of the sector information can be quantified using a novel contribution metric by inspecting the weights of the LSTM. The predictive signal is driven by a short-term reversal factor and an industry momentum factor.
    Date: 2026–08
    URL: https://d.repec.org/n?u=RePEc:arx:papers:2608.05755
  16. By: Rishabh Singh Chauhan; Mahdi Ghadimi; Lishun Liu
    Abstract: Objectives: While causal analysis of travel behavior is an emerging field, estimating heterogeneity in mode choice through causal modeling remains unexplored. This study demonstrates the application of a novel causal method, causal forest, to quantify the heterogeneity in travel mode choice shifts caused by the COVID-19 pandemic. Methods: We applied causal forests, a non-parametric causal machine learning method, to 802, 935 trip records from the 2017 and 2022 waves of the National Household Travel Survey. The 2017 wave serves as the pre-pandemic control group, while the 2022 wave represents the treatment condition. Within the potential outcomes framework, we estimate average treatment effects (ATE), heterogeneous treatment effects (HTE), and conditional average treatment effects (CATE) across diverse socio-demographic groups and trip characteristics. Findings: Our results reveal an estimated ATE of a 1.86 percentage point (pp) increase in car-mode share, contrasted with decreases of 0.38 pp and 1.57 pp in public transit and walking, respectively. The largest increases in car use appeared for short-distance trips (one mile or less), households with annual incomes exceeding USD 200, 000, and female travelers. Novelty: This is one of the first applications of causal forests to travel mode choice, and the first to use causal machine learning to estimate the pandemic's causal effect on mode choice analysis. Practical Applications: This study discusses methodological advantages, inherent assumptions, and limitations of causal forests within the context of transportation planning. This methodology is applied to COVID-19 travel data to illustrate how causal heterogeneity analysis can offer a deeper understanding of changes in mode choice. These insights are valuable for planners and policymakers in making policies related to mode shifts under an intervention.
    Date: 2026–08
    URL: https://d.repec.org/n?u=RePEc:arx:papers:2608.04208
  17. By: Ylva Baeckström; Roman Matkovskyy (Rennes SB - Rennes School of Business)
    Abstract: Financial advice can attenuate underinvestment but is costly, biased, and skewed towards the wealthy. AI-powered co-advisors could help deliver more scalable and affordable advice. To understand how, our vignette-based survey experiment compares the portfolio recommendations made by professional human advisors with GenAI large language models (LLMs) under biased and unbiased prompts. We document human financial advice projection whereby human advisors strongly project their own portfolios onto their clients. AI financial advice projection is prompt and model family dependent: ChatGPT is the least biased, while strong Gemini-Biased projection collapses when removing advisor demographics. LLMs are systematically more conservative than professional human advisors, recommending portfolios with lower Sharpe ratios that deliver up to 18% lower 20-year terminal wealth. However, human advisory fees erode much of this excess gain, with a 20-year breakeven fee of 1.03% p.a. Our results have direct implications for financial regulators, the advice profession, and LLM developers seeking to deploy AI-generated financial advice.
    Keywords: Financial advice, Large language models, Artificial intelligence, Portfolio asset allocation
    Date: 2026–09
    URL: https://d.repec.org/n?u=RePEc:hal:journl:hal-05725514
  18. By: Aaron L. Garavito-Acosta (Central Bank of Colombia); Edgar Caicedo-Garcia (Central Bank of Colombia); Wilmer Martinez-Rivera (Central Bank of Colombia); Juan J. Ospina-Tejeiro (Central Bank of Colombia)
    Abstract: This paper develops and applies a standardized framework for forecasting year-on-year inflation in Colombia using large language models (LLMs). We conduct six sequential experiments in which the information set available to the models is progressively expanded by incorporating historical macroeconomic data, contextual indicators, explicit economic structure, and contemporaneous information retrieved through web search. Each configuration is executed daily and generates 24-month inflation forecasts in real time rather than retrospectively, together with qualitative explanations of the forecasts and, in the more advanced configurations, assessments of the shocks affecting inflation. This real-time design mitigates look-ahead bias and produces genuine forecast vintages. The framework is implemented as a programmatic pipeline in Python that queries the OpenAI and Google APIs, executes predefined experiment-specific prompts, and automatically processes and stores the model responses, while applying forecast validation and revision procedures in the more advanced configurations. The results show that richer information environments produce less monotonic inflation paths that remain above the 3% target over the forecast horizon and are more consistent with the contemporaneous domestic and external shocks affecting inflation. The qualitative analysis also shows that the models consistently identify relevant inflation drivers and their interactions. These richer forecast paths are broadly consistent with the pattern observed in survey-based inflation expectations. A formal evaluation of forecast accuracy will be conducted as additional real-time forecast vintages become available.
    Keywords: Large language models; Inflation forecasting; Colombia; Prompt design
    JEL: C53 E01 E31 E37 F31 E23
    Date: 2026–08–25
    URL: https://d.repec.org/n?u=RePEc:gii:giihei:heidwp23-2026

This nep-big issue is ©2026 by Tom Coupé. It is provided as is without any express or implied warranty. It may be freely redistributed in whole or in part for any purpose. If distributed in part, please include this notice.
General information on the NEP project can be found at https://nep.repec.org. For comments please write to the director of NEP, Marco Novarese at <director@nep.repec.org>. Put “NEP” in the subject, otherwise your mail may be rejected.
NEP’s infrastructure is sponsored by the Griffith Business School of Griffith University in Australia.