nep-cmp New Economics Papers
on Computational Economics
Issue of 2026–09–07
thirty-one papers chosen by
Stan Miles, Thompson Rivers University


  1. Communicating Credit Risk with Large Language Models: Evaluation of Explanations from Standard and Alternative Data-Based Models By Sahab Zandi; Noah Kostesku; Christophe Mues; Mar\'ia \'Oskarsd\'ottir; Cristi\'an Bravo
  2. Reinforcement Learning for Perpetual Futures Market Making By Martin Cekal
  3. Revealing economic facts: LLMs know more than they say By Marcus Buckmann; Quynh Anh Nguyen; Ed Hill
  4. Forecasting the Cost of a Basic Basket of Goods: a comparative analysis using machine learning models and online prices By João Anderson da Silva Felix; Michel Alexandre; Cássio da Nóbrega Besarria
  5. Neural Network Learning for Nonlinear Economies By Ashwin, Julian; Beaudry, Paul; Ellison, Martin
  6. An Open Benchmark for Evaluating Time Series Forecasting Methods across Financial Markets By Jeremy Bejarano; Viren Desai; Kausthub Keshava; Arsh Kumar; Zixiao Wang; Vincent Hanyang Xu; Yangge Xu
  7. Mapping the Evolution of Artificial Intelligence in Banking: A Systematic Bibliometric Review of Research Trends, Techniques, Applications, and Emerging Challenges By Ayoub Louhab; Abdelhak Yaacoubi; Mohamed Azouazi
  8. Cross-Sectional Heterogeneity in LSTM Networks for Financial Time Series By Julius D\"obelt
  9. Do Deep Learning Methods Improve Financial Forecasts? By Jeremy Bejarano; Viren Desai; Kausthub Keshava; Arsh Kumar; Zixiao Wang; Vincent Hanyang Xu; Yangge Xu
  10. Experimenting with Large Language Models for Inflation Forecasting in Colombia By Aaron L. Garavito-Acosta; Edgar Caicedo-Garcia; Wilmer Martinez-Rivera; Juan J. Ospina-Tejeiro
  11. Reducing Basis Risk in Index Insurance Using Satellite-Based Vegetation Indices and Machine Learning By Ryu, Jaehyeon; Yu, Jisang
  12. Developing a Standard Strategy for Time Series Forecasting Integrating Statistical and Machine Learning Techniques Using a Meta-Model Approach and its Application in Generating External Debt Projections By Joshua Elias Fred A. Suero
  13. Inference with AI-Generated Covariates By Junting Duan; Markus Pelger
  14. Multi-Level Market Making with Reinforcement Learning By Patrick Cheridito; Moritz Weiss
  15. What survives honest evaluation? Leakage-safe, search-aware assessment of LLM-driven trading strategy discovery By Eray Gen\c{c}ay
  16. From Financial Sentiment Classification to Return Predictability: A QLoRA Benchmark of Large Language Models By Fusheng Luo
  17. Predicting Debt Distress in Low-Income Countries By von Luckner, Clemens Graf; Horn, Sebastian; Kraay, Aart C.; Ramalho, Rita
  18. Portfolio management with big data By Penaranda, Francisco; Sentana, Enrique
  19. EcoFinBench – a natural language processing benchmark for economics and finance By Max Ahrens; Dragos Gorduza; Micheal McMahon
  20. Accounting for intra-household joint travel in agent-based transport simulations By Javaudin Lucas; Araldo Andrea; Coulombel Nicolas
  21. AI Agents and Prompt Engineering in Econometric Coding By Sebastian Galiani; Federico Ariel López; Raul A. Sosa
  22. Concentrated Liquidity Provision: a Reinforcement Learning Perspective By Georgios Chionas; Charalampos Kleitsikas; Stefanos Leonardos; Leandro S\'anchez-Betancourt; Carmine Ventre
  23. RetailAgent: Structured Adverse Timing in Self-Conditioned Multimodal LLM Trading Agents By Yupeng Zhang; Liuyuan Jiang; Hongyi Huang; Bingheng Li; Lisha Chen
  24. A Bayesian latent class reinforcement learning framework to capture adaptive, feedback-driven travel behaviour By Georges Sfeir; Stephane Hess; Thomas O. Hancock; Filipe Rodrigues; Jamal Amani Rad; Michiel Bliemer; Matthew Beck; Fayyaz Khan
  25. Inflation Unpacked: Breaking Down the Key Components Using a Neural Phillips Curve By Joan Christine S. Allon-Pineda
  26. Lead-Lag Relationships in Financial Markets: A Comparison of Multiple Clustering Algorithms By Ruichen Deng; Yichi Zhang
  27. Double/Debiased Machine Learning for Functional-Form-Robust Spatial Autoregression By Jieun Lee
  28. Training AI For When Humans Will Use It By Kevin A. Bryan; Joshua S. Gans
  29. How AI Prompts Can Teach Us About the Structure of Human Behavior By Matthew O. Jackson; Benjamin S. Manning; Yutong Xie; Walter Yuan; Qiaozhu Mei
  30. The devil in the DeTail: assessing state-contingent tail effects of a releasable macroprudential capital buffer using a parsimonious agent-based framework By Enrico Minnella; Ana Pereira; Eugen Tereanu
  31. A Simulation Model for Predicting Tramp Shipping Supply By Bläser, Nikolaj; Magnussen, Búgvi Benjamin; Fuentes, Gabriel; Reinhardt, Line; Lindén, Anders

  1. By: Sahab Zandi; Noah Kostesku; Christophe Mues; Mar\'ia \'Oskarsd\'ottir; Cristi\'an Bravo
    Abstract: Credit decisioning is a high-stakes task in which model outputs must be accurate and explainable to support compliant decisions. Although modern credit risk models such as eXtreme Gradient Boosting (XGBoost) and Graph Neural Networks (GNNs) improve predictive performance, their explanations are often too technical for stakeholders creating communication gaps that can shape approvals, denials, and fairness judgments. We examine whether Large Language Models (LLMs) can serve as explanation layers that translate post-hoc explanation artefacts into stakeholder-appropriate risk narratives. Using Freddie Mac single-family loan-level data, we develop three pipelines: standard tabular (XGBoost + SHAP), and two with alternative data, a pure network-based (GNN + GNNExplainer), and a bimodal one (combining tabular and network data). We generate narratives with three LLM configurations: a small fine-tuned LLM (Gemma 3 4B), a large fine-tuned LLM (DeepSeek R1 70B), and a zero-shot commercial LLM (Gemini 2.5). Explanation quality is evaluated through automated checks across all pipelines and a human study of bimodal explanations comparing credit risk professionals and non-professionals on eight decision-relevant dimensions. We have three main findings. First, the pipeline accounts for higher variance in evidence-grounding scores than the language model, meaning that the binding constraint on explanation quality is the evidence representation, not the model used. Second, the explanation narratives reliably name the influential factors but are less reliable when stating the direction of influence, which may be consequential for adverse-action communication. Finally, professionals apply stricter evidentiary standards than non-professionals. We discuss implications for the governance of risk models, including deployment considerations and the value of domain-aligned LLMs in regulated credit settings.
    Date: 2026–08
    URL: https://d.repec.org/n?u=RePEc:arx:papers:2608.17715
  2. By: Martin Cekal
    Date: 2026–01–22
    URL: https://d.repec.org/n?u=RePEc:prg:jnlwps:v:6:y:2026:id:6.004
  3. By: Marcus Buckmann (Bank of England); Quynh Anh Nguyen (Bank of England); Ed Hill (Bank of England)
    Abstract: We investigate whether hidden states of large language models (LLMs) can be used to estimate and impute economic and financial statistics. Focusing on county-level (eg unemployment) and firm-level (eg total assets) variables, we show that a linear regression trained on the hidden states of open-source LLMs outperforms the models' own text outputs. This indicates that internal representations encode richer economic information than is revealed directly in generated responses. A learning curve analysis shows that, in many cases, only a few dozen labelled examples suffice for training. We further propose a transfer learning method that improves estimation accuracy without requiring any labelled data for the target variable. Finally, we demonstrate the practical utility of hidden states in data imputation and super-resolution tasks.
    Keywords: Large language models;embeddings;economic statistics;data imputation.
    JEL: C45 C81 C21
    Date: 2025–10–31
    URL: https://d.repec.org/n?u=RePEc:boe:boeewp:023272
  4. By: João Anderson da Silva Felix; Michel Alexandre; Cássio da Nóbrega Besarria
    Abstract: Online prices can be used to successfully estimate and predict inflation indices represented by the cost of a basket of goods. Nevertheless, machine learning algorithms can deliver better inflation index forecasts than a simple weighted sum of prices due to their ability to handle complex relationships between predictive variables. Using data from five Brazilian state capitals (São Paulo, Porto Alegre, Rio de Janeiro, Goiânia, and Fortaleza) from February 2024 to May 2025, we attempt to predict the cost of a basket of goods using the online prices of the goods that make up such baskets as predictive features. We employ four machine learning models (k-NN, XGBoost, Random Forest, and ridge regression) and a forecast combination technique (Voting Regressor). We also verified the robustness of the machine learning models in situations where it was not possible to obtain online prices for all products. Machine learning models outperform the simple weighted sum of prices in forecasting the overall cost of the food basket, whether all the variables are available or not.
    Date: 2026–08
    URL: https://d.repec.org/n?u=RePEc:bcb:wpaper:650
  5. By: Ashwin, Julian; Beaudry, Paul; Ellison, Martin
    Abstract: Neural networks offer a promising tool for the analysis of nonlinear economies. In this paper, we derive conditions for the global stability of nonlinear rational expectations equilibria under neural network learning. We demonstrate the applicability of the conditions in analytical and numerical examples where the nonlinearity is caused by monetary policy targeting a range, rather than a specific value, of inflation. If shock persistence is high or there is inertia in the structure of the economy, then the only rational expectations equilibria that are learnable may involve inflation spending long periods outside its target range. Neural network learning is also useful for solving and selecting between multiple equilibria and steady states in other settings, such as when there is a zero lower bound on the nominal interest rate.
    Keywords: Inflation targeting; Machine learning; Neural networks; Zero lower bound
    Date: 2024–07
    URL: https://d.repec.org/n?u=RePEc:cpr:ceprdp:19295
  6. By: Jeremy Bejarano; Viren Desai; Kausthub Keshava; Arsh Kumar; Zixiao Wang; Vincent Hanyang Xu; Yangge Xu
    Abstract: An open benchmark that holds financial data fixed across forecasting methods, revealing where machine learning sharpens forecasts of financial stress (Working Paper no. 26-05).
    Date: 2026–08–25
    URL: https://d.repec.org/n?u=RePEc:ofr:wpaper:26-05
  7. By: Ayoub Louhab (FSJES AIN SEBAA, Hassan II University –Casablanca); Abdelhak Yaacoubi (FSJES AIN SEBAA, Hassan II University –Casablanca); Mohamed Azouazi (F.S. Ben M'sik, Casablanca)
    Abstract: Artificial intelligence (AI) development has revolutionized the financial industry, providing solutions for better fraud detection, credit risk assessment, customer relationship management, bank operations and regulation. Despite numerous publications dedicated to the topic in question, there are no bibliometric studies devoted to exploring the development, structure of knowledge bases and trends in research on AI applications in banking. Therefore, to fill the gap, this paper provides a systematic review of the existing literature concerning the application of AI in banking based on a dataset of 114 papers indexed in Scopus from 2020 to 2025. The systematic approach is applied to the research based on the methodology described in the PRISMA guidelines and such analytical tools as Bibliometrix, Biblioshiny and VOSviewer are used to assess the scientific output, key sources of information, Institutions and countries and thematic links via keyword co-occurrences. Thus, the annual increase rate of publication number in this area is estimated at 39.77% and involves 96 publication sources and 394 authors. Moreover, six topical issues related to the field under investigation are revealed, including fraud detection using machine learning techniques, credit and risk management, digital banking, cybersecurity, customer-focused applications and explainable AI. The overall findings from the findings indicate that the research is moving from the stage of developing individual algorithms to building a holistic approach towards developing intelligent banking systems that can be regulated. Through integrating bibliometric analysis with an extensive review on the use of AI in banking, this paper adds to the existing body of knowledge and identifies new areas for research in explainable AI, data governance, ethical implementation, and intelligent financial services.
    Keywords: VOSviewer, Bibliometrix. JEL Classification: G21, Bibliometric Analysis, Risk Management, Fraud Detection, Machine Learning, Banking, Artificial Intelligence, O33, G28, Artificial Intelligence Banking Machine Learning Fraud Detection Risk Management Bibliometric Analysis VOSviewer Bibliometrix. JEL Classification: G21
    Date: 2026–07–16
    URL: https://d.repec.org/n?u=RePEc:hal:journl:hal-05694871
  8. By: Julius D\"obelt
    Abstract: Predicting financial asset returns remains one of the most difficult challenges in empirical finance, driven by the low signal-to-noise ratio and the semi-strong form of market efficiency. While deep learning models, especially LSTM networks, have shown promise in capturing temporal dependencies, standard architectures often struggle to account for the cross-sectional heterogeneity of asset returns. This paper proposes a novel architectural extension to the basic LSTM model designed to improve both predictive accuracy and model interpretability. The framework integrates macro-financial covariates to capture broader economic signals and learnable sector embeddings to encompass heterogeneity by sector. The trading strategy involves constructing a long-short portfolio based on daily directional forecasts for each S&P 500 constituent, targeting stocks expected to under- or outperform the cross-sectional median return of the S&P 500. Model Performance is evaluated against three competitive benchmarks: a basic LSTM, a Random Forest model and a traditional market buy-and-hold strategy. The empirical results demonstrate that the LSTM with sector embeddings outperforms all benchmarks across key risk and return metrics. By utilizing sector embeddings, the model explicitly incorporates cross-sectional heterogeneity, allowing it to adapt to varying industry dynamics within the market. To address the black-box nature of deep learning, I use latent space visualizations to analyse how the model differentiates between sectors, providing insights into the internal representation of the sectors in the LSTM. The impact of the sector information can be quantified using a novel contribution metric by inspecting the weights of the LSTM. The predictive signal is driven by a short-term reversal factor and an industry momentum factor.
    Date: 2026–08
    URL: https://d.repec.org/n?u=RePEc:arx:papers:2608.05755
  9. By: Jeremy Bejarano; Viren Desai; Kausthub Keshava; Arsh Kumar; Zixiao Wang; Vincent Hanyang Xu; Yangge Xu
    Abstract: The newest deep learning methods work best when the data contain rich, repeating structures.
    Date: 2026–08–25
    URL: https://d.repec.org/n?u=RePEc:ofr:ofrblg:26-10
  10. By: Aaron L. Garavito-Acosta (Central Bank of Colombia); Edgar Caicedo-Garcia (Central Bank of Colombia); Wilmer Martinez-Rivera (Central Bank of Colombia); Juan J. Ospina-Tejeiro (Central Bank of Colombia)
    Abstract: This paper develops and applies a standardized framework for forecasting year-on-year inflation in Colombia using large language models (LLMs). We conduct six sequential experiments in which the information set available to the models is progressively expanded by incorporating historical macroeconomic data, contextual indicators, explicit economic structure, and contemporaneous information retrieved through web search. Each configuration is executed daily and generates 24-month inflation forecasts in real time rather than retrospectively, together with qualitative explanations of the forecasts and, in the more advanced configurations, assessments of the shocks affecting inflation. This real-time design mitigates look-ahead bias and produces genuine forecast vintages. The framework is implemented as a programmatic pipeline in Python that queries the OpenAI and Google APIs, executes predefined experiment-specific prompts, and automatically processes and stores the model responses, while applying forecast validation and revision procedures in the more advanced configurations. The results show that richer information environments produce less monotonic inflation paths that remain above the 3% target over the forecast horizon and are more consistent with the contemporaneous domestic and external shocks affecting inflation. The qualitative analysis also shows that the models consistently identify relevant inflation drivers and their interactions. These richer forecast paths are broadly consistent with the pattern observed in survey-based inflation expectations. A formal evaluation of forecast accuracy will be conducted as additional real-time forecast vintages become available.
    Keywords: Large language models; Inflation forecasting; Colombia; Prompt design
    JEL: C53 E01 E31 E37 F31 E23
    Date: 2026–08–25
    URL: https://d.repec.org/n?u=RePEc:gii:giihei:heidwp23-2026
  11. By: Ryu, Jaehyeon; Yu, Jisang
    Keywords: Risk and Uncertainty
    Date: 2026
    URL: https://d.repec.org/n?u=RePEc:ags:aaea26:404414
  12. By: Joshua Elias Fred A. Suero (Bangko Sentral ng Pilipinas)
    Abstract: The pursuit of accurate forecasting and the rise of machine learning have led to the development of various forecasting models, making model selection increasingly difficult. This paper aims to address this challenge by developing a standard strategy for time series forecasting using a meta-model approach. A meta-model is constructed by combining individual forecasts from leading statistical and machine learning models, with Philippine external debt as the target variable. The baseline meta-model combines the following individual forecasts using ordinary least squares (OLS): (a) random walk with drift, (b) autoregressive integrated moving average (ARIMA), (c) exponential smoothing with error, trend, and seasonal components (ETS), (d) Holt-Winters, (e) multiple aggregation prediction algorithm (MAPA), (f) temporal hierarchical forecasting (THieF), (g) theta model, (h) Prophet, (i) neural network autoregression (NNAR), (j) long short-term memory (LSTM), (k) gradient boosting machine (GBM), (l) extreme gradient boosting (XGBoost), (m) random forest, (n) support vector regression (SVR), and (o) dynamic linear models (DLM). Alternative meta-models were also evaluated. The best-performing meta-model, which employs the Least Absolute Shrinkage and Selection Operator (LASSO), performs remarkably well, achieving a mean absolute percentage error (MAPE) of 3.0 percent in the validation set. The proposed standard strategy is versatile, with the potential to forecast other financial and economic variables. Additionally, forecasts generated by the meta-model can serve as a valuable benchmark, whether compared to current forecasting practices or in cases where no forecasting methodology exists.
    JEL: C22 C53 C61 F34
    Date: 2025–04
    URL: https://d.repec.org/n?u=RePEc:bhd:dpaper:202503
  13. By: Junting Duan; Markus Pelger
    Abstract: Empirical researchers increasingly use large language models (LLMs) to extract structured features, such as sentiment scores, classifications, and expectations, from unstructured data and treat these generated features as observed covariates in downstream estimation. This practice can invalidate inference when systematic, input-dependent errors in generated features, such as hallucination and look-ahead bias, distort the downstream moment conditions. Even after correction, generated features remain noisy proxies whose error profiles differ across models and prompts. We introduce AI-Powered Inference (AI-PI), a method-of-moments framework for valid and efficient inference that combines three components: a moment-specific bias correction based on a small human-labeled calibration set; adaptive weights that optimally combine multiple model-prompt pairs; and an optimal calibration-set design that concentrates costly human labels where the generated features are least reliable. We establish consistency and asymptotic normality of the AI-PI estimator, allowing for data-adaptive labeling designs, cross-fitted LLM-pipeline tuning, and overidentified GMM. Simulations confirm substantial gains over naive LLM regressions and over debiasing without optimal weighting or labeling design. In an application to news-based sentiment and stock returns, AI-PI produces stable conclusions where naive analyses vary substantially across LLM and prompt choices, with a confidence interval roughly half as long as using the human-labeled data alone.
    JEL: C10 C13 C50 C55 C80 G12
    Date: 2026–07
    URL: https://d.repec.org/n?u=RePEc:nbr:nberwo:35481
  14. By: Patrick Cheridito; Moritz Weiss
    Abstract: We introduce a reinforcement learning framework for market making in a limit order book. Our algorithm aims to maximize trading revenue by dynamically submitting market and limit orders of varying sizes across multiple price levels while controlling inventory size. We use multivariate logistic-normal distributions to model order allocations and employ a deep-set encoder to aggregate features from variable-length order sets into a fixed-dimensional latent representation. Additionally, we incorporate potential-based reward shaping to accelerate learning without altering the optimal policy. We illustrate the performance of the method in three simulated market environments consisting of noise traders who submit random trades, tactical traders who respond to instantaneous volume imbalance, and strategic traders who trade in the direction of an exponentially weighted volume imbalance signal.
    Date: 2026–08
    URL: https://d.repec.org/n?u=RePEc:arx:papers:2608.18195
  15. By: Eray Gen\c{c}ay
    Abstract: Large language models (LLMs) are increasingly used to discover trading strategies, and much of the resulting literature shares a methodological weakness: many candidate strategies are generated, the best is reported, and neither look-ahead bias nor the intensity of the search behind the reported result is corrected for. We present a strategy-discovery system that makes both corrections structural rather than procedural. First, the agent can only act through registry-validated tools whose feature space excludes look-ahead by construction; we show that this guardrail is not redundant with statistical correction: a deliberately leaky oracle posting a Sharpe ratio of 35 survives Deflated Sharpe and probability-of-backtest-overfitting testing completely. Second, the system records every strategy evaluation its search performs and deflates all reported performance by that trial count, tracing how the best in-sample Sharpe ratio climbs with each trial while the deflation threshold, driven by the agent's own search, climbs faster. Across a 453-stock point-in-time US equity universe and a 39-ETF multi-asset universe with realistic transaction, impact, and borrow costs, honest evaluation certifies passive benchmarks (out-of-sample confidence intervals excluding zero), rejects every LLM-discovered strategy (across two frontier models, search budgets up to one hundred candidates, and five repeated runs), catching selection luck, predicted rank degradation, and out-of-sample collapse through complementary instruments, and evaluates a human trader's production rule system under identical instruments. The framework formalizes why pre-registered hypotheses earn lower evidential bars than brute search, and quantifies the sample sizes that credible certification of moderate edges actually requires.
    Date: 2026–08
    URL: https://d.repec.org/n?u=RePEc:arx:papers:2608.27734
  16. By: Fusheng Luo
    Abstract: Financial sentiment classifiers are commonly evaluated against human labels, but strong linguistic performance does not necessarily imply economically useful return predictability. This study separates these questions through two experiments. First, we construct a unified three-class benchmark from five financial text datasets and compare TF--IDF Naive Bayes, off-the-shelf FinBERT and Financial-RoBERTa encoders, zero-shot Qwen2.5-7B, and QLoRA-adapted Qwen2.5-7B, LLaMA3-8B, and Mistral-7B models. Mistral-7B achieves the best test accuracy (0.8840) and macro-F1 (0.8771), while QLoRA raises Qwen2.5's macro-F1 from 0.7274 to 0.8615. An inverse-frequency class-weighted loss does not improve Qwen2.5. Second, we evaluate economic validity on a temporally separate 2019 Benzinga sample containing 10, 637 unique headlines and 13, 115 headline--stock observations for a fixed S\&P~100 universe. Model probabilities are converted into continuous sentiment scores, aggregated by stock and signal date, and aligned with next-session returns over one-, two-, three-, and five-day horizons. All seven downstream models produce positive but small mean rank information coefficients at the one-day horizon; the largest is 0.0143 for FinBERT. None of the 28 model--horizon tests remains significant after Newey--West inference and false-discovery-rate correction. Portfolio results likewise fail to establish a robust advantage for the best-performing classifiers. The findings show that QLoRA is effective for financial sentiment adaptation, while also documenting a clear gap between classification accuracy and tradable cross-sectional signals.
    Date: 2026–08
    URL: https://d.repec.org/n?u=RePEc:arx:papers:2608.04200
  17. By: von Luckner, Clemens Graf; Horn, Sebastian; Kraay, Aart C.; Ramalho, Rita
    Abstract: This paper develops an empirical model to predict episodes of debt servicing difficulties (“debt distress”) in low-income countries, with three main contributions to the existing literature. First, it develops more refined measures of external debt distress episodes that allow timing the onset of distress episodes with increased precision. Second, it develops a systematic algorithm to comprehensively assess the out-of-sample predictive performance of more than 550, 000 candidate binary prediction models using J-K-fold cross-validation. Third, it tests whether more sophisticated machine learning algorithms can outperform simple probit models. The paper finds that simple single-equation probit models have better predictive power for debt distress than more sophisticated algorithms and are comparable in terms of predictive performance to important policy benchmarks such as the IMF and World Bank debt sustainability framework for low-income countries.
    Date: 2026–09–02
    URL: https://d.repec.org/n?u=RePEc:wbk:wbrwps:11441
  18. By: Penaranda, Francisco; Sentana, Enrique
    Abstract: The purpose of this survey is to summarize the academic literature that studies some of the ways in which portfolio management has been affected in recent years by the availability of big datasets: many assets, many characteristics for each of them, many macro predictors, and various sources of unstructured data. Thus, we deliberately focus on applications rather than methods. We also include brief reviews of the financial theories underlying asset management, which provide the relevant background to assess the plethora of recent contributions to such an active research field.
    Keywords: Machine learning; Mean-variance analysis; Stochastic discount factors
    JEL: G11 G12 C55 G17
    Date: 2024–07
    URL: https://d.repec.org/n?u=RePEc:cpr:ceprdp:19314
  19. By: Max Ahrens (Maihem.ai); Dragos Gorduza (Bank of England); Micheal McMahon (University of Oxford)
    Abstract: We introduce EcoFinBench, a natural language processing (NLP) benchmark suite for the domains of economics and finance. We comprehensively test a large array of NLP models across multiple domain-specific data sets for sentence classification. Specifically, we evaluate dictionary models, word count models, topic models, and modern transformer models. Furthermore, we introduce two new data sets to the research community. The Bluebook data set for text-only sentiment analysis in monetary policy, and the Greenbook data set for multimodal (text and numeric) sentiment analysis. We focus on data sets that require the models to work with relatively few data points and long average text lengths – typical characteristics of data sets in the economic and financial domain. We find that, dictionary models – still widely used as a default text analysis tool in economics and finance – underperform substantially across all evaluated data sets. From our findings, we conclude that given the underperformance of existing solutions in the multimodal domain, future modelling work is needed. With our benchmark suite we aim to lay the foundation for a systematic assessment on the most commonly used NLP models in economics and finance. To our knowledge, we are the first to provide such holistic benchmarking assessment for economics and finance.
    Keywords: Machine learning, natural language processing;artificial intelligence;benchmark.
    JEL: C45 Y10 C55 C88 G17
    Date: 2025–12–19
    URL: https://d.repec.org/n?u=RePEc:boe:boeewp:023285
  20. By: Javaudin Lucas; Araldo Andrea; Coulombel Nicolas
    Abstract: Intra-household joint home-based tours - trips in which household members depart together, engage in shared activities, and return together - represent a significant share of daily travel, yet are systematically ignored in transport simulations. Conflating joint and solo tours within a single mode choice framework introduces bias in preference parameter estimates. This paper proposes a three-step methodology to integrate joint tours in agent-based transport models: a Random Forest classifier to identify joint tours, a Multinomial Logit model estimating mode choice specific to joint tours, and a Penalized Logistic Regression for driver/passenger assignment. Applied to the Paris region using household travel survey data, the methodology successfully replicates observed joint tour shares and mode distributions in a synthetic population. The proposed framework enables more reliable evaluation of policies whose impacts differ between joint and solo travel, such as HOV lanes or family transit fare discounts.
    Date: 2026–08
    URL: https://d.repec.org/n?u=RePEc:arx:papers:2608.18657
  21. By: Sebastian Galiani; Federico Ariel López; Raul A. Sosa
    Abstract: We study how large language models write code for econometric analysis. We compare three dimensions of AI-assisted coding: statistical software (Stata, R, or Python), prompting (zero-shot versus few-shot), and the degree of agency, from a chatbot that writes a single script to an agent that executes and revises its own code. On a benchmark of applied econometric and statistical tasks, moving from the chatbot to the constrained agent raises task success from 74 to 96 percent, at about eight additional cents per run. For Claude Sonnet 4.6 and GPT-5.4 through Codex, few-shot prompting improves the chatbot far more than the constrained agent, indicating that prompting and agency act as substitutes. For these models, differences across statistical software are sizeable under the chatbot but largely disappear under the constrained agent.
    JEL: C18 C87
    Date: 2026–08
    URL: https://d.repec.org/n?u=RePEc:nbr:nberwo:35588
  22. By: Georgios Chionas; Charalampos Kleitsikas; Stefanos Leonardos; Leandro S\'anchez-Betancourt; Carmine Ventre
    Abstract: Automated market makers (AMMs) are a cornerstone of decentralised finance (DeFi). Constant product markets with concentrated liquidity, such as UniswapV3, are now a well-established design. In these markets, liquidity providers (LPs) face a sequential decision problem: they must decide when to rebalance their positions and which price ranges to allocate capital to as market conditions evolve. We formulate dynamic liquidity provision as a stochastic impulse control problem and use reinforcement learning (RL) to solve it, focusing on providing interpretable solutions. We show that learned policies exhibit rich state-dependent behaviour, allocating liquidity according to mispricing, rebalancing costs, uncertainty, inventory exposure, and heterogeneous risk preferences. These behaviours help compress the left tail of the Profit and Loss (PnL) distribution and avoid catastrophic outcomes under high uncertainty. Finally, we benchmark the RL agents against baseline and sophisticated agents from the AMM microstructure literature and analyse their performance.
    Date: 2026–08
    URL: https://d.repec.org/n?u=RePEc:arx:papers:2608.19389
  23. By: Yupeng Zhang; Liuyuan Jiang; Hongyi Huang; Bingheng Li; Lisha Chen
    Abstract: In financial markets, a sequential policy that reacts systematically to price movements may become predictable to other market participants. This paper studies whether large language model (LLM) agents exhibit such directional structure through RetailAgent, an experimental framework in which an LLM observes anonymized intraday equity price histories and permitted state, then repeatedly chooses long (hold the stock) or flat (stay out) before the subsequent interval return is revealed. We compare returns during long and flat intervals along the same stock's intraday path after removing the overall fraction of long decisions. This exposure-matched measure reveals persistent negative timing across modality, horizon, state, and model family. Shuffling saved action sequences substantially attenuates the effect, showing that alignment between actions and subsequent returns drives the negative score. Feeding self-authored memories into decisions further increases policy persistence, while timing becomes more negative among stock-days on which the agent uses both actions. These results reveal stable, recoverable directional structure in sequential LLM financial decisions and a behavioral signal for studying how another participant could respond to a predictable policy.
    Date: 2026–08
    URL: https://d.repec.org/n?u=RePEc:arx:papers:2608.28399
  24. By: Georges Sfeir; Stephane Hess; Thomas O. Hancock; Filipe Rodrigues; Jamal Amani Rad; Michiel Bliemer; Matthew Beck; Fayyaz Khan
    Abstract: Many travel decisions involve a degree of experience formation, where individuals learn their preferences over time. At the same time, there is extensive scope for heterogeneity across individual travellers, both in their underlying preferences and in how these evolve. The present paper puts forward a Latent Class Reinforcement Learning (LCRL) model that allows analysts to capture both of these phenomena. We apply the model to a driving simulator dataset and estimate the parameters through Variational Bayes. We identify three distinct classes of individuals that differ markedly in how they adapt their preferences: the first displays context-dependent preferences with context-specific exploitative tendencies; the second follows a persistent exploitative strategy regardless of context; and the third engages in an exploratory strategy combined with context-specific preferences.
    Date: 2025–12
    URL: https://d.repec.org/n?u=RePEc:arx:papers:2512.14713
  25. By: Joan Christine S. Allon-Pineda (Bangko Sentral ng Pilipinas)
    Abstract: Identifying the sources of inflation is a complex yet crucial endeavor for effective monetary policymaking. A significant challenge arises from the fact that key elements of inflation dynamics, such as the output gap and inflation expectations, are often unobserved. This study implements a Hemisphere Neural Network whose peculiar architecture allows the estimation of the unobserved states within an augmented New Keynesian Phillips Curve. In the context of Philippine inflation, the estimated latent states effectively capture macroeconomic concepts such as real activity, inflation expectations, and commodity price dynamics, as evidenced by their correlation with input variables and alignment with major economic events. Long-run expectations is found to remain steady between 3.5-4.5 percent, while commodity prices account for most spikes in realized inflation. The model’s estimated output gap is aligned with existing measures from the Bangko Sentral ng Pilipinas, while the estimated inflation expectations align with short- to medium-term expectations from businesses and professional forecasters. Finally, the research offers significant insights into inflation dynamics and provides an analytical tool for monitoring policy-relevant inflationary pressures.
    JEL: C45 E31 E32 E52
    Date: 2025–04
    URL: https://d.repec.org/n?u=RePEc:bhd:dpaper:202505
  26. By: Ruichen Deng; Yichi Zhang
    Abstract: Lead-lag relationships are widely used in financial time series, and many clustering algorithms based on them have been developed. The traditional DTW-KMedoids algorithm performs well both on the synthetic dataset and the real financial dataset. However, there are still several limitations to these algorithms: low efficiency caused by high time complexity, poor mathematical properties from DTW distance, the clustering effect is sensitive to the number of clusters. To solve the problems above and improve the performance, this paper introduces three clustering algorithms: MiniRocket-KMeans, KShape, Ensemble algorithm (a combination of KShape and DTW-KMedoids) and compares their performance on synthetic and real stock datasets with DTW-KMedoids algorithm under the same trade strategy. In addition, this paper also finds the best number of clusters by maximizing the silhouette coefficient in each clustering algorithm to improve the stability of the experiment results. Our main conclusions are as follows: MiniRocket-KMeans performs best under the lead strategy, achieving a Sharpe ratio of 0.866 with a maximum drawdown controlled at -63.9\%; the ensemble algorithm exhibits excellent stability; the robustness is significantly improved after finding the best number of clusters; the p-values of the hypothesis test on the Sharpe ratio of all strategies are 0.0, verifying the statistical validity of the lead-lag trading strategy. Finally, future improvement directions such as customized lead-lag matrices and optimized ensemble voting mechanisms are proposed.
    Date: 2026–08
    URL: https://d.repec.org/n?u=RePEc:arx:papers:2608.24703
  27. By: Jieun Lee
    Abstract: Spatial autoregressive inference is typically conditional on the spatial weights matrix, W, even though the underlying interaction structure is often unknown and empirical conclusions can be sensitive to its specification. This paper develops double/debiased machine learning inference for low-dimensional SAR parameters when the spatial interaction operator is learned flexibly from potentially endogenous characteristics. Within a maintained admissible support, interaction strength is generated by an unknown function of geographic and socioeconomic characteristics, making inference robust to functional form specification of the weights within that support. Endogeneity in the characteristics generating W is addressed through a nonlinear control function based on locally relevant first-stage residual information. Because the learned operator enters both the spatial lag and spatially transformed instruments, treating the estimated W as known generally leaves a first-order generated-W effect. I construct an operator-orthogonal SAR-IV/GMM score that removes this leading sensitivity and combine it with buffered spatial cross-fitting that separates evaluation-score footprints from nuisance-training observations. Under near epoch dependence on a spatially mixing innovation field and target-relevant nuisance rate and regularity conditions, the estimator is asymptotically linear and root-n normal. Monte Carlo simulations show improved finite-sample inference relative to nonorthogonal alternatives when the interaction function is misspecified, weight generating characteristics are endogenous, and observations are spatially dependent. In a U.S. application, diabetes estimates vary with the choice of W, showing the sensitivity of SAR inference to the interaction structure. Even for the same learned W, results differ across inferential methods, highlighting the importance of inference when W is learned.
    Date: 2026–08
    URL: https://d.repec.org/n?u=RePEc:arx:papers:2608.22706
  28. By: Kevin A. Bryan; Joshua S. Gans
    Abstract: AI predicts; humans use its predictions to make decisions. These predictions are combined with human verification and analysis, queries to other statistical models, and so on. The economic value of an AI, therefore, depends on how it interacts with the surrounding decision environment. We describe the value of AI as part of this ``composite experiment'' where AI makes a coarse prediction of the state of the world, show what this means for optimal model training via a geometric argument, explain why optimal training can be discontinuous in economic variables, and study how heterogeneous users or monopoly model trainers affect these results. In particular, maximizing the unconditional accuracy of AI predictions is generally suboptimal.
    Date: 2026–08
    URL: https://d.repec.org/n?u=RePEc:arx:papers:2608.12538
  29. By: Matthew O. Jackson; Benjamin S. Manning; Yutong Xie; Walter Yuan; Qiaozhu Mei
    Abstract: We introduce a general, easy-to-implement AI-based method for studying the structure and complexity of human behavior. We assign a large language model a ``type vector'' and then prompt it to choose actions across settings in which we observe human choices. For instance, the type vector (2, 4) becomes ``You are a player characterized by the following profile: 2 out of 5 in Altruism, 4 out of 5 in Risk Aversion, '' after which it is prompted to make choices. We vary the dimensions (e.g., Altruism, Fairness, Trust, $\dots$) and values (e.g., 1--5) to minimize distance to human choices. Applying the method to 119, 147 decisions made by 78, 657 subjects from more than 35 countries across 10 classic economic game roles, we find that human behavior can be closely matched using three dimensions: Risk Aversion, Strategic Sophistication, and Trust. Moreover, the types needed to fit individuals across games cluster into fewer than a dozen groups, and can predict behavior in held-out games with different rules and available actions. The results suggest that behavior across diverse settings can be approximated by a low-dimensional, portable representation, supporting the possibility of general yet parsimonious theories across the behavioral sciences. More broadly, the method can provide insights into the structure of many human behaviors.
    Date: 2026–08
    URL: https://d.repec.org/n?u=RePEc:arx:papers:2608.18265
  30. By: Enrico Minnella (Bank of England); Ana Pereira (Bank of England); Eugen Tereanu (European Central Bank and Joint Vienna Institute)
    Abstract: This paper develops an agent-based framework (DeTail) to assess the state-contingent tail effects of releasable macroprudential capital buffers. The model features heterogeneous firms, households, and banks, and a single central bank, all interacting in a fully integrated, stock-flow consistent framework which generates endogenous credit cycles. Using this approach, we evaluate how time-varying capital requirements affect the time-varying distributions of credit growth, firm and household default rates, and bank losses along the credit cycle. Policy experiments show that releasing capital buffers during economic downturns preserves credit supply by improving risky (lower-tail) credit outcomes, reduces both households and firms defaults, and supports macro-financial resilience by limiting tail bank losses. At the same time, capital buffer accumulation during upturns imposes minimal costs and does not significantly constrain lending. These findings support the active use of releasable buffers to mitigate systemic risk and smooth credit cycles without weakening the banking system.
    Keywords: Agent-based modelling;macroprudential policy;macro-financial linkages;credit cycles;bank resilience;tail risk;state-dependent effects.
    JEL: C63 E44 E58 G28
    Date: 2026–07–24
    URL: https://d.repec.org/n?u=RePEc:boe:boeewp:023320
  31. By: Bläser, Nikolaj (Dept. of People and Technology, Roskilde University); Magnussen, Búgvi Benjamin (Dept. of People and Technology, Roskilde University); Fuentes, Gabriel (Dept. of Business and Management Science, Norwegian School of Economics); Reinhardt, Line (Dept. of People and Technology, Roskilde University); Lindén, Anders (Research Department, TORM A/S)
    Abstract: Tramp shipping is a key part of the maritime industry which operates mainly in the spot market relying on voyage-by-voyage contracting, which forces them to reposition frequently in search of favourable cargoes. Market dynamics therefore emerge from how regional cargo demand aligns with the shifting distribution of available vessels. Forming multi-month forecasts of this evolving relationship between demand and supply is essential for market participants seeking to respond to rapidly changing market conditions. The supply side of the relationship remains relatively unexplored, particularly addressing vessels reposition and evolvement of regional availability over time. Bridging this gap, this paper introduces a simulation-based framework that models behaviour at the individual-vessel level and generates forward-looking forecasts of regional tramp-hipping supply over a 90-day horizon. The regional tramp shipping vessel supply prediction is presented through a mathematical formulation and an agent-based framework in which each vessel acts as an autonomous agent responding to market conditions is developed. To this end Neural network-based stochastic estimators of vessel behaviour are produced from historical data and used to simulate vessel-level decisions, yielding coherent forecasts of regional vessel supply. The framework is evaluated on the clean petroleum products market using datasets spanning the period 2020-01-01 to 2024-06-30. The results are compared with a regression benchmark relying on macro-economic variables, and the developed framework show to achieve higher supply prediction accuracy in 23 of 24 region-(vessel-segment) combinations, reducing average mean absolute percentage error from 13.01% to 4.79%.
    Keywords: Tramp shipping; Fleet simulation; Vessel supply prediction; Supply demand dynamics; Stochastic modelling
    JEL: C44 C63 R40
    Date: 2026–08–24
    URL: https://d.repec.org/n?u=RePEc:hhs:nhhfms:2026_010

This nep-cmp issue is ©2026 by Stan Miles. It is provided as is without any express or implied warranty. It may be freely redistributed in whole or in part for any purpose. If distributed in part, please include this notice.
General information on the NEP project can be found at https://nep.repec.org. For comments please write to the director of NEP, Marco Novarese at <director@nep.repec.org>. Put “NEP” in the subject, otherwise your mail may be rejected.
NEP’s infrastructure is sponsored by the Griffith Business School of Griffith University in Australia.