|
on Computational Economics |
| By: | Giorgos Iacovides; Wuyang Zhou; Danilo Mandic |
| Abstract: | Recent advances in Generative AI have substantially improved financial sentiment analysis through post-trained financial large language models (LLMs). However, existing approaches remain confined to a market-agnostic, supervised learning paradigm that relies on limited, static and human-annotated datasets, and thus are incapable of adapting to evolving market conditions. To address this limitation, we introduce FinSMART, the first market-aligned reinforcement learning framework for financial sentiment analysis, which directly optimizes sentiment signals using realized market outcomes. To deal with the noisy, non-stationary, and multifactorial nature of financial markets, FinSMART incorporates a signal extraction pipeline that combines market-aware data filtering with a discrete asymmetric trading reward, enabling stable reinforcement learning from economically meaningful market feedback. Experimental results demonstrate that FinSMART significantly outperforms existing state-of-the-art methods in profitability, risk-adjusted performance, and sentiment signal quality, improving cumulative trading returns by 220% over the strongest baseline. Uniquely, the FinSMART framework naturally supports market-aware retraining, at any point in time, by replacing costly manual annotation with newly observed financial articles and their realized market outcomes. Such a retraining strategy enables the model to continuously adapt to changing market dynamics, resulting in consistent performance gains over its static counterpart. These findings demonstrate the practical applicability of market-aligned reinforcement learning and highlight its potential as a next-generation paradigm for developing adaptive financial LLMs. |
| Date: | 2026–07 |
| URL: | https://d.repec.org/n?u=RePEc:arx:papers:2607.28127 |
| By: | Ash, Elliott; Hansen, Stephen; Muvdi, Yabra |
| Abstract: | This chapter explores the transformative impact of large language models (LLMs) on text analysis in economics. We trace the evolution from traditional methods like bag-of-words to advanced models such as BERT and GPT, highlighting how these models address limitations in understanding context and allowing higher-order reasoning. Although LLMs are complex, costly, and lacking in transparency, they are powerful tools for research, such as measuring sentiment or predicting metadata associated with documents. |
| Keywords: | Large Language Models; Transformer models; Text as data; Unstructured Data |
| JEL: | C18 C45 C55 |
| Date: | 2024–09 |
| URL: | https://d.repec.org/n?u=RePEc:cpr:ceprdp:19479 |
| By: | Mukashov, Askar; Kim, Soonho; Fang, Peixun; Diao, Xinshen; Thurlow, James; Proctor, Joshua; Rennison, Alan |
| Abstract: | The demand for high-quality, rapid economic analysis to navigate complex issues faced by many low- and middle-income countries has led to the development of detailed structural simulation models, such as Computable General Equilibrium (CGE) models. Policy analysis with such models requires deep knowledge of their structure and applicability to the policy issues at hand. Policymakers in these settings often lack access to the expertise required for articulating, analyzing, and interpreting the relevant causal chains captured by the models. Attempting to circumvent these barriers by submitting complex economic questions directly to off-the-shelf large language models (LLMs) introduces severe analytical risks, including hallucinations and insufficient expert guidance. To resolve this limitation, we developed and empirically evaluated an agentic AI assistant called RIAPA-AI that integrates LLMs with a CGE model. We evaluated the performance of RIAPA-AI against expert human CGE modelers and a general-purpose plain LLM baseline across samples of complex economic scenarios, utilizing an independent panel of senior economists to grade the outputs. Our statistical analysis reveals no statistically significant difference in analytical accuracy between RIAPA-AI and human experts, while the AI accelerates reproducible policy analysis from weeks to minutes. Furthermore, by operating without manual processing limits, RIAPA-AI eliminates the 6.7% error rate observed among human modelers. Conversely, the general-purpose plain LLM exhibits profound failure rates, failing to achieve policy-ready scores in over 60% of depth evaluations. Without an underlying CGE model acting as a bounding force to reflect economic structural constraints, the general-purpose plain LLM defaults to linear economic assumptions and inserts unmodeled socio-political narratives. Crucially, by explicitly restricting the AI's narrative interpretation solely to deterministic numerical outputs, RIAPA-AI mitigates the risk of unverified assumptions and logic hallucinations. We conclude that by deploying an agentic AI assistant that layers a generative AI over a formal CGE model, RIAPA-AI successfully delivers sensible, rigorous, and rapid policy analysis. |
| Keywords: | artificial intelligence; machine learning; modelling; computable general equilibrium models; large language models; agent-based models; econometric models |
| Date: | 2026–06–26 |
| URL: | https://d.repec.org/n?u=RePEc:fpr:ifprid:183521 |
| By: | Qihui Chen; Ka Yan Cheng; Zheng Fang |
| Abstract: | We develop a general framework of identification and estimation for automatic debiased machine learning (DML) where the parameter of interest $\theta_0$ is identified by a moment condition involving a nuisance $\gamma_0$ that may be high dimensional. We establish conditions under which the Riesz representer $\alpha_0$, which is at the core of DML, is identified, and show that the identification occurs precisely when $\alpha_0$ uniquely optimizes a quadratic functional. This characterization enables us to develop a general estimation procedure for $\alpha_0$ that allows for generic $\gamma_0$ including those defined by models with endogeneity and encompasses both classical sieves and modern architectures such as deep neural networks. To improve estimation precision and mitigate the curse of dimensionality, we incorporate shape constraints on $\gamma_0$ by embedding them into a possibly nonlinear parameter space. We illustrate our estimation procedure through simulations and empirical applications. |
| Date: | 2026–07 |
| URL: | https://d.repec.org/n?u=RePEc:arx:papers:2607.24472 |
| By: | Hauzenberger, Niko; Huber, Florian; Klieber, Karin; Marcellino, Massimiliano |
| Abstract: | Macroeconomic data is characterized by a limited number of observations (small T), many time series (big K) but also by featuring temporal dependence. Neural networks, by contrast, are designed for datasets with millions of observations and covariates. In this paper, we develop Bayesian neural networks (BNNs) that are well-suited for handling datasets commonly used for macroeconomic analysis in policy institutions. Our approach avoids extensive specification searches through a novel mixture specification for the activation function that appropriately selects the form of nonlinearities. Shrinkage priors are used to prune the network and force irrelevant neurons to zero. To cope with heteroskedasticity, the BNN is augmented with a stochastic volatility model for the error term. We illustrate how the model can be used in a policy institution through simulations and by showing that BNNs produce more accurate point and density forecasts compared to other machine learning methods. |
| Keywords: | Bayesian neural networks; Model selection; Shrinkage priors; Macro forecasting |
| JEL: | C11 C30 C45 C53 E3 E44 |
| Date: | 2024–08 |
| URL: | https://d.repec.org/n?u=RePEc:cpr:ceprdp:19381 |
| By: | Zane Shen; Xinli Xu; Guangyi Zhang; Jialong Chen; Jinsong Zhou; Cong Chen; Guibao Shen; Dongyu Yan; Luozhou Wang; Zhen Yang |
| Abstract: | Parent-order execution is a core problem in algorithmic trading, where the goal is to split a large order into smaller orders while reducing execution costs. Existing approaches either rely on pre-specified market assumptions that may not hold in practice, or require task-specific training that limits adaptability to new settings. To overcome these limitations, we present the first systematic study of large language models (LLMs) for parent-order execution. This extends the use of LLMs in finance from what to trade to how to execute. We propose PACE (Plan-Ahead Controlled Execution), a hierarchical framework that decomposes parent-order execution into long-horizon planning and short-horizon execution, requiring neither explicit market assumptions nor task-specific training. Experiments on Shenzhen Stock Exchange Level-1 data show that PACE outperforms TWAP, Almgren-Chriss, and learning-based baselines, exceeding the strongest baseline by 0.65 bps. Behavioral analysis reveals that LLMs make execution decisions differently from human investors: higher model confidence predicts better performance rather than worse returns, and the model trades earlier rather than procrastinating toward the deadline. These findings suggest that LLMs can complement human traders in execution decisions. |
| Date: | 2026–07 |
| URL: | https://d.repec.org/n?u=RePEc:arx:papers:2607.28410 |
| By: | Guillaume Coqueret; Joan Llull; Florian Oswald; Christophe P\'erignon; Christoph Scheuch; Lars Vilhuber |
| Abstract: | Large language models (LLMs) are increasingly used to generate data for research. Typical use cases are classifications, annotations, information extraction, and generation of numerical scores. Unlike conventional measurements, LLM outputs can vary across repeated requests even when the prompt and apparent model settings remain unchanged. This variation arises from deliberate sampling, silent model updates, numerical rounding, or expert routing. Setting a dedicated temperature parameter to zero removes deliberate sampling when that option is available, but it does not eliminate the other sources of randomness. Exact reproduction is therefore generally not possible when using proprietary application programming interfaces. Local execution of open-weight models offers greater control, but reproducibility still depends on the complete hardware and software stack. We illustrate these issues through sentiment classifications of corporate filings and examine their consequences for downstream regression results. We then propose a reporting standard for articles and replication packages, as well as guidance for data editors and authors. Together, these findings and recommendations establish that LLM outputs should be treated as draws from a distribution rather than as fixed measurements. |
| Date: | 2026–07 |
| URL: | https://d.repec.org/n?u=RePEc:arx:papers:2607.24372 |
| By: | Uehara, Masatoshi; Shi, Chengchun; Kallus, Nathan |
| Abstract: | Reinforcement learning (RL) is one of the most vibrant research frontiers in machine learning and has been recently applied to solve a number of challenging problems. In this paper, we primarily focus on off-policy evaluation (OPE), one of the most fundamental topics in RL. In recent years, a number of OPE methods have been developed in the statistics and computer science literature. We provide a discussion on the efficiency bound of OPE, some of the existing state-of-the-art OPE methods, their statistical properties and some other related research directions that are currently actively explored. |
| Keywords: | off-policy evaluation;semiparametric methods;causal inference;dynamic treatment regime;offline reinforcement learning;contextual bandits |
| JEL: | C1 |
| Date: | 2026–08–31 |
| URL: | https://d.repec.org/n?u=RePEc:ehl:lserod:127940 |
| By: | Yuya Shimizu |
| Abstract: | Unstructured data, such as images and text, are increasingly used in empirical economics. Since training machine-learning models on unstructured data is costly, economists often use off-the-shelf pre-trained deep learning models developed by computer scientists to extract embeddings, which are then used as covariates in target economic analyses. Despite the popularity of this practice, its theoretical foundations remain limited. There are two main difficulties. First, the pre-trained model is usually trained on a different dataset and for a different task. Consequently, it is unclear when such a model can be used reliably for the target task. Second, the embedding function is subject to an identification problem, which makes it difficult to analyze the estimation error of the embedding function and its effect on the target task. In this paper, we provide sufficient conditions to overcome these difficulties and derive the convergence rate of machine learning models with pre-trained embeddings. We illustrate the theory through double machine learning applications for estimating parameters of interest, such as partially linear regression with unstructured controls, price elasticity in demand estimation considering the product quality measured by images and text, missing data imputation with unstructured data, and the average treatment effect with unstructured confounders. |
| Date: | 2026–07 |
| URL: | https://d.repec.org/n?u=RePEc:arx:papers:2607.17378 |
| By: | Travis L. Johnson; Jiannan Jiang; Soumyabrata Chaudhuri; Yihao Chen; Lauren Falvey; Donal O'Cofaigh |
| Abstract: | Specialist training beats generalist scale when forecasting financial statements. To our knowledge, no prior work jointly forecasts complete financial statements beyond one year, yet in a discounted-cash-flow valuation most firm value sits past that window. We release ProForma-20Q, a reproducible benchmark for forecasting 78 statement line items 1-20 quarters ahead, for anonymized firms, from past statements and an industry code, scored by change-space $R^2$. On it, Forma, a transformer that reads statements as sets of (account, quarter, value) tuples and maximizes a masked-tuple Gaussian likelihood, beats every competitor we field: classical machine learning, chained gradient boosting, a zero-shot time-series foundation model, and frontier large language models. Its lead widens with horizon, where valuation needs accuracy most, and its Gaussian predictive intervals never under-cover. Forma's forecasts nearly satisfy accounting identities; exact coherence is recoverable at no statistically significant accuracy cost. Its tuple interface supports scenario analysis without retraining, and we show that pinning future revenue paths sharpens the rest of the statement. |
| Date: | 2026–08 |
| URL: | https://d.repec.org/n?u=RePEc:arx:papers:2608.11327 |
| By: | Wu, Tianyi; Wang, Tengyao; Samworth, Richard J. |
| Abstract: | In the context of multivariate nonparametric regression with missing covariates, we propose Pattern Embedded Neural Networks (PENNs), which can be applied in conjunction with any existing imputation technique. In addition to a neural net work trained on the imputed data, PENNs pass the vectors of observation indicators through a second neural network to provide a compact representation. The outputs are then combined in a third neural network to produce final predictions. Our main theoretical result exploits an assumption that the observation patterns can be partitioned into cells on which the Bayes regression function behaves similarly, and belongs to a compositional H¨older class. It provides a finite-sample excess risk bound that holds for an arbitrary missingness mechanism, and in combination with a complementary minimax lower bound, demonstrates that our PENN estimator attains in typical cases the minimax rate of convergence as if the cells of the par tition were known in advance, up to a poly-logarithmic factor in the sample size. Numerical experiments on simulated, semi-synthetic and real data confirm that the PENN estimator consistently improves, often dramatically, on standard neural net works without pattern embedding. Code to reproduce our experiments, as well as a tutorial on how to apply our method, is publicly available. |
| Keywords: | deep learning;missing data;nonparametric regression |
| JEL: | C1 |
| Date: | 2026–07–29 |
| URL: | https://d.repec.org/n?u=RePEc:ehl:lserod:139005 |
| By: | Rametta, Jack T.; Fuller, Sam (Harvard University) |
| Abstract: | Are random forests, the workhorse of supervised machine learning methods in the social sciences, still “good enough” versus new methods that tout dramatic performance benefits? In this article we present a large, diverse Monte Carlo study and existing real-world data benchmarks to compare tree-based methods with TabPFN, a new pretrained transformer foundation model that performs in-context learning over millions of synthetic datasets designed for tabular data. We compare tree-based methods with TabPFN in the popular R-learner framework for conditional average treatment effect estimation. This allows us to assess both predictive model performance and resulting gains for downstream inference. First, our simulations suggest that TabPFN does outperform random forest, achieving near-oracle results for conditional effect estimation. TabPFN’s improvements manifest in the most difficult simulation setups, where the data generating process is complex and there are fewer observations. Second, in real-world data analyses TabPFN performs well, outperforming random forest in some cases, especially as sample dwindles and the number of predictive covariates increases. Our results suggest that tree-based methods are still well suited for social science data, but TabPFN specifically, and the prior data fitted network approach generally, is a strong competitor worthy of consideration. |
| Date: | 2026–07–30 |
| URL: | https://d.repec.org/n?u=RePEc:osf:socarx:g29xc_v1 |
| By: | Ebrahimi Kahou, Mahdi; Fernández-Villaverde, Jesús; Gomez Cardona, Sebastian; Perla, Jesse; Rosa, Jan |
| Abstract: | In the long run, we are all dead. Nonetheless, when studying the short-run dynamics of economic models, it is crucial to consider boundary conditions that govern long-run, forward-looking behavior, such as transversality conditions. We demonstrate that machine learning (ML) can automatically satisfy these conditions due to its inherent inductive bias toward finding flat solutions to functional equations. This characteristic enables ML algorithms to solve for transition dynamics, ensuring that long-run boundary conditions are approximately met. ML can even select the correct equilibria in cases of steady-state multiplicity. Additionally, the inductive bias provides a foundation for modeling forward-looking behavioral agents with self-consistent expectations. |
| JEL: | C1 E1 |
| Date: | 2024–08 |
| URL: | https://d.repec.org/n?u=RePEc:cpr:ceprdp:19386 |
| By: | Keller, Wolfgang; Shiue, Carol; Yan, Sen |
| Abstract: | Primary historical sources are often by-passed for secondary sources due to high human costs of accessing and extracting primary information–especially in lower-resource settings. We propose a supervised machine-learning approach to the natural language processing of Chinese historical data. An application to identifying different forms of social unrest in the Veritable Records of the Qing Dynasty shows that approach cuts dramatically down the cost of using primary source data at the same time when it is free from human bias, reproducible, and flexible enough to address particular questions. External evidence on triggers of unrest also suggests that the computer-based approach is no less successful in identifying social unrest than human researchers are. |
| Keywords: | Natural language processing |
| JEL: | N45 C8 |
| Date: | 2024–09 |
| URL: | https://d.repec.org/n?u=RePEc:cpr:ceprdp:19517 |
| By: | Arishi Orra; Himanshu Choudhary; Manoj Thakur |
| Abstract: | Reinforcement learning has gained increasing attention as a data-driven approach for stock trading. However, learning a policy that is both profitable and stable remains challenging due to non-stationary market behaviour and noisy reward signals. Auxiliary tasks are often used to improve representation learning and stabilize training, yet they are usually designed manually and depend heavily on prior assumptions about targets and prediction horizons. Such fixed designs may not remain suitable across changing market regimes. In this work, we propose a self-supervised framework that automatically discovers auxiliary tasks to support reinforcement learning for stock trading. The auxiliary tasks are formulated as General Value Functions so that their predictions enrich the learned state representation and assist policy optimization. The framework consists of two networks. The main network learns the trading policy along with the auxiliary predictions, while the secondary network generates the definitions of auxiliary tasks through learned cumulants and discount factors. These tasks are updated using a meta gradient mechanism that accounts for their long-term impact on trading performance and improves training stability. We evaluate the proposed approach across four major equity indices: DJI, FTSE, Sensex, and TAIEX. The empirical results demonstrate that automatically discovered auxiliary tasks lead to more robust learning and improved trading performance compared to existing baselines. |
| Date: | 2026–08 |
| URL: | https://d.repec.org/n?u=RePEc:arx:papers:2608.15841 |
| By: | Hongyu Lin; Yulin Chen; Yuanrong Wang; Antonio Briola; Tomaso Aste |
| Abstract: | Using neural networks for stock return prediction typically requires choices about depth and hidden-layer width that are difficult to connect to financial interpretation. We study an alternative: estimate dependence among firm characteristics with a Maximally Filtered Clique Forest (MFCF), then map its clique structure to a Homological Neural Network (HNN). The MFCF maximum clique size K is the only parameter controlling architectural complexity, and it has a clear graphical meaning: it bounds the number of characteristics in each maximal clique and hence the highest interaction order the network can represent. The filtered graph then fixes the neural network's depth, layer widths, and sparse connections before training, in place of a separately chosen depth and width sequence. We apply two HNN variants to annual out-of-sample forecasts of U.S. stock excess returns from 1987 to 2016 using 94 firm characteristics. The HNN models match a three-hidden-layer benchmark on pooled predictive accuracy, rank the cross-section more accurately, and use roughly 80 times fewer parameters than a fully connected network with the same induced layer widths. Two structural ablations indicate that both the sparse connectivity and the estimated grouping of characteristics contribute to the ranking advantage, and both effects remain significant after correcting for multiple testing. These findings show that HNNs offer a practical and interpretable way to incorporate estimated dependence among firm characteristics into neural architecture design. |
| Date: | 2026–08 |
| URL: | https://d.repec.org/n?u=RePEc:arx:papers:2608.14323 |
| By: | Esteban Sánchez-Gómez (Economic Division, Central Bank of Costa Rica) |
| Abstract: | This paper evaluates the performance of machine learning (ML) methods for forecasting year-over-year inflation in Costa Rica using monthly data from 2012-2025 and compares their performance against standard benchmarks within a rolling out-of-sample framework. ML techniques are particularly useful for capturing nonlinearities and complex interactions between inflation and a broad set of macroeconomic covariates. The results show that nonlinear ensemble methods such as XGBoost and BART provide the strongest gains at short horizons, while linear shrinkage methods are more competitive at longer horizons. ***Resumen: Este documento evalúa el desempeño de los métodos de aprendizaje automático (ML) para pronosticar la inflación interanual en Costa Rica, utilizando datos mensuales de 2012 a 2025 y comparándolos con estándares de referencia dentro de un esquema de muestra móvil fuera de muestra. Las técnicas de ML son especialmente útiles para captar no linealidades e interacciones complejas entre la inflación y un amplio conjunto de variables macroeconómicas. Los resultados muestran que los métodos de ensamblaje no lineal como XGBoost y BART presentan los mayores beneficios en horizontes cortos, mientras que los métodos de reducción lineal son más competitivos en horizontes largos. |
| Keywords: | Inflation, Macroeconomic Forecasting, Machine Learning, Monetary Policy, Hybrid Models, inflación, Pronósticos, Aprendizaje automático, Política Monetaria, Modelos Híbridos. |
| JEL: | C53 C33 |
| Date: | 2026–08 |
| URL: | https://d.repec.org/n?u=RePEc:apk:doctra:2606 |
| By: | Christian Terwiesch; Lennart Meincke; Karan Girotra; Ethan Mollick; Gideon Nave; Karl T. Ulrich |
| Abstract: | This research examines how well large language models, or LLMs, generate new product ideas for college students priced under $50. Across a series of studies, we identify key strengths and weaknesses of using LLMs for product innovation. Our first study shows that LLM-generated product ideas have higher average quality than human ideas, based on purchase intent, and are 7 times more likely to rank in the top 10%. Our second study shows that this AI-induced creativity boost is not explained by the LLM's more persuasive pitching skills. Our third and fourth studies identify a weakness of using LLMs for brainstorming: AI-generated ideas are less novel at the idea level and less diverse at the set level. In our fifth study, we analyze prior LLM-based creativity studies and find consistently lower idea diversity across all of them, demonstrating the generalizability of these findings. Our sixth and seventh studies investigate techniques to mitigate this diversity loss. We compare LLMs from different vendors and versions and find that more recent models generate more diverse ideas, though they still fall short of human-level diversity. We also demonstrate techniques that increase idea diversity almost to the level of human idea generation: pooling ideas across vendors; prompt engineering, including Chain-of-Thought prompting and injecting heterogeneous personas or constraints; and creative agents that broadly explore the solution landscape to restore diversity. Finally, in our eighth study, we show that exploiting the near-zero marginal cost of AI idea generation by scaling the number of ideas steadily improves coverage of the idea space, approaching human-level coverage. We conclude by presenting actionable recommendations for innovation managers who want to identify better new product ideas with the help of LLMs. |
| Date: | 2026–07 |
| URL: | https://d.repec.org/n?u=RePEc:arx:papers:2607.27553 |
| By: | Bohren, Noah; Hakimov, Rustamdjan; Lalive, Rafael |
| Abstract: | Generative artificial intelligence (AI) has made substantial progress, but some capabilities of AI are not well understood. This study compares the ability of AI to a representative population of US adults in creative and strategic tasks. The creative ideas produced by AI chatbots are rated more creative than those created by humans. Moreover, ChatGPT is substantially more creative than humans, while Bard lags behind. Augmenting humans with AI improves human creativity, albeit not as much as ideas created by ChatGPT alone. Competition from AI does not significantly reduce the creativity of men, but it decreases the creativity of women. Humans who rate the text cannot discriminate well between ideas created by AI or other humans but assign lower scores to the responses they believe to be AI-generated. As for strategic capabilities, while ChatGPT shows a clear ability to adjust its moves in a strategic game to the play of the opponent, humans are, on average, more successful in this adaptation. |
| Keywords: | ChatGPT; Creativity; Experiment |
| JEL: | I24 J24 D91 C90 |
| Date: | 2024–09 |
| URL: | https://d.repec.org/n?u=RePEc:cpr:ceprdp:19507 |
| By: | Junyi Ye; Gargi Vijay Borde |
| Abstract: | Financial volatility is regime dependent, yet incorporating regime information into neural networks can also destabilize training. This paper asks where such information should enter a neural cross-sectional volatility forecasting model. We study five-day realized-volatility forecasts for 1, 027 U.S. equities using a rolling walk-forward evaluation framework in which information, model capacity, hyperparameter tuning, and random seeds are matched across architectures. We propose RG-ResMoE, a regime-gated residual mixture-of-experts architecture in which regime information is used only for expert routing rather than for direct forecasting. The base predictor models volatility from stock features, while a gating network uses regime state variables to route residual corrections. RG-ResMoE consistently outperforms a capacity-matched MLP in both forecasting accuracy and training stability in the main U.S. study. Similar gains are observed on an independent Japanese panel. The integration pathway is decisive: appending the same regime variables directly to the forecasting input degrades both predictive performance and training stability, whereas restricting them to the routing gate improves accuracy and Value-at-Risk calibration. Hard routing consistently underperforms soft routing. The results suggest that, in compact neural volatility forecasting models, the primary value of mixture-of-experts models lies less in increasing model capacity than in controlling how nonstationary regime information influences prediction. |
| Date: | 2026–08 |
| URL: | https://d.repec.org/n?u=RePEc:arx:papers:2608.12251 |
| By: | Bruno Bouchard (CEREMADE); Lucas Gnecco Heredia (LAMSADE); Ludovic Moreau (CEREMADE); Kim-Anh Pham (CEREMADE) |
| Abstract: | As in Bouchard et al. (2010) and Bouchard and Nutz (2014), we study a utility maximization problem with expectation constraint. We first consider a uniformly elliptic case in which the endogenous state boundary associated with the constraint in expectation is proved to be smooth. This allows one to derive a proper Dirichlet condition for the value function of the optimal control problem on this boundary. We then propose a new truncation argument in the martingale representation of the expectation constraint. This leads to an approximating sequence of auxiliary systems of PDEs for which comparison holds. Convergence to the initial optimal control problem is proved. In the degenerate case, we propose another approximation which consists in adding a small noise term to recover uniformly ellipticity. Convergence is also proved. To the best of our knowledge, it is the first time that a full analysis is performed for such control problems, so as to open the doors to the use of numerical schemes. Numerical resolution in a toy example is performed using neural networks. It is complemented by an estimation of the numerical error, also performed by using a neural network approach. |
| Date: | 2026–07 |
| URL: | https://d.repec.org/n?u=RePEc:arx:papers:2607.24114 |
| By: | Yijia Xiao; Rujun Han; Yanfei Chen; Zifeng Wang; Ke Jiang; Zhongying CuiZhu; Vishy Tirumalashetty; Wei Wang; Burak Gokturk; Tomas Pfister; Chen-Yu Lee |
| Abstract: | Powered by advances in LLMs and autonomous agents, deep research has become one of the most widely adopted agentic products. However, most deep research systems write general-purpose reports, which are inadequate for financial deep research. Financial research demands specialized knowledge to analyze historical patterns and forecast upcoming events. Automating financial deep research therefore requires both a layered harness to drive the research agent and a verifiable, point-in-time benchmark that prevents leakage of future information. We present FinanceHarness, a harness that runs finance-oriented tools and practitioner-guided workflows, automating financial deep research end to end: environment and data construction, the agent execution loop, and reward modeling. We further propose FinanceGym, comprising thesis-driven research questions and rubrics that combine pre-cutoff and post-cutoff criteria. Professional expert validation yields an 82% pass rate. With the same open-weight backbone, FinanceHarness improves the overall rubric score from 25.3% to 32.4%, demonstrating the effectiveness of our specialized harness design. However, even pairing FinanceHarness with the most cutting edge LLM (e.g. Opus-5), the FinanceGym score is below 45%, showing that it is a challenging benchmark for financial deep research. Leaderboard is available at: https://financegym.github.io/ and FinanceHarness code is available at: https://github.com/Yijia-Xiao/FinanceHar ness. |
| Date: | 2026–07 |
| URL: | https://d.repec.org/n?u=RePEc:arx:papers:2607.27853 |
| By: | Kasun Dewage; Suranadi De Silva; Shankhadeep Mondal |
| Abstract: | Foundation models for time series forecasting demonstrate impressive zero-shot generalization but often underperform on specialized domains such as high-frequency finance. We present a comprehensive study of hybrid neural-classical correction for adapting frozen TimesFM (200M parameters) to stock return prediction during the volatile opening trading hour. We compare two neural correction architectures - AttnCorrect (multi-head self-attention, approximately 471K parameters) and GatedLinear (low-rank bilinear projection with gating, approximately 49K parameters) - each augmented with Random Forest residual learning. Through systematic ablation across 10 major technology stocks (NVDA, MSFT, AAPL, GOOG, GOOGL, AMZN, META, AVGO, TSLA, NFLX) spanning 2 million data points, we reveal critical insights: (1) The hybrid neural-classical approach achieves 0.597 pooled correlation and 6.4x mean per-day correlation improvement over frozen TimesFM; (2) Classical residual learning (Random Forest) provides the largest single-component contribution, matching or exceeding the neural correction component; (3) Simpler neural architectures surprisingly outperform complex ones when classical residual learning is removed; (4) Self-attention provides the largest neural-only contribution. GatedLinear+RF achieves best overall performance with 9x fewer neural parameters than AttnCorrect+RF. We report three complementary correlation metrics - mean per-day, cross-day cumulative, and pooled - to provide a complete picture of predictive quality. Our results provide practical guidance: effective foundation model adaptation requires careful integration of neural and classical components, with classical methods playing a crucial complementary role. |
| Date: | 2026–08 |
| URL: | https://d.repec.org/n?u=RePEc:arx:papers:2608.08825 |
| By: | Howard Su; Huan-Hsin Tseng; Chi-Sheng Chen; Lance Bai |
| Abstract: | Solving high-dimensional parabolic partial differential equations (PDEs) is important in engineering, physics, and stochastic control. Deep BSDE methods reformulate semilinear PDEs as backward stochastic differential equations and admit a model-based reinforcement learning interpretation, where trajectories are generated from known stochastic dynamics while a trainable model learns the gradient-related control process. We propose a Quantum Transformer BSDE solver based on Multi-Layer Fully-Connected Variational Quantum Circuits (FC-VQC). The method treats the normalized state trajectory as time--coordinate tokens and applies causal self-attention to learn interactions in the adapted BSDE gradient process. All trainable model parameters are contained within the FC-VQC embedding, projection, feed-forward, and decoder modules, while attention and structural operations remain classical and parameter-free. Experiments on three d=36 PDE benchmarks show that QTransformer consistently improves over the non-attentive FC-VQC baseline and outperforms the classical Transformer at compact hidden widths, while the wider classical Transformer achieves the best overall accuracy. These results demonstrate that combining causal attention with FC-VQC provides an effective quantum architecture for high-dimensional BSDE trajectory learning. |
| Date: | 2026–07 |
| URL: | https://d.repec.org/n?u=RePEc:arx:papers:2607.25162 |
| By: | Christian Bongiorno; Efstratios Manolakis; Rosario Nunzio Mantegna |
| Abstract: | This paper introduces a compact reformulation of a modular end-to-end neural network for global minimum-variance portfolio optimization that decouples model complexity from both look-back window length and universe size. A five-parameter hyperbolic weighted moving average combined with a saturating exponential replaces the original 2, 400-parameter lag-transformation layer, and a bidirectional gated-recurrent-unit eigencleaning module together with a streamlined marginal-volatility network reduce total learnable parameters from 39, 586 to just 2, 175. In out-of-sample tests against state-of-the-art nonlinear-shrinkage and risk-parity benchmarks, the compact network attains the lowest realized portfolio variance without compromising expected return. Under long-only constraints, the variance reduction supports substantially higher leverage while maintaining comparable drawdown control. Validation in a high-fidelity trading simulator that incorporates realistic margin-call dynamics confirms enhanced over-leverage resilience. These findings demonstrate that end-to-end variance-minimization architectures can achieve substantial parameter efficiency and robust capital-efficiency gains without sacrificing risk-adjusted performance. |
| Date: | 2026–07 |
| URL: | https://d.repec.org/n?u=RePEc:arx:papers:2607.23068 |
| By: | Mayer, Thierry; Rapoport, Hillel; Umana-Dajud, Camilo |
| Abstract: | Using provisions to ease the movement of business visitors in trade agreements, we show that removing barriers to the movement of business people promotes trade. We document the increasing complexity of Free Trade Agreements and develop an algorithm that combines machine learning and text analysis techniques to examine the content of FTAs. We use the algorithm to determine which FTAs include provisions to facilitate the movement of business people and whether these are included in dispute settlement mechanisms. We show that provisions facilitating business travel are effective in promoting them and eventually increase bilateral trade flows. The paper provides (indirect) evidence of the role of face-to-face interaction on aggregate bilateral trade flows. |
| Keywords: | Migration; Machine learning; Text analysis |
| JEL: | F13 F22 F23 |
| Date: | 2024–09 |
| URL: | https://d.repec.org/n?u=RePEc:cpr:ceprdp:19463 |
| By: | Alireza Kargarzadeh; Nariman Khaledian; Navid Parvini; Arman Khaledian |
| Abstract: | Large language models can extract richer signals from financial news than fixed sentiment lexicons, and recent work has explored feeding such signals into portfolio construction. We study an uncertainty-aware construction that feeds model-predicted risk -- decomposed into aleatoric and epistemic components -- directly into the covariance matrix of portfolio allocators, rather than treating portfolio risk as fixed or adjusting only expected returns. We evaluate the pipeline on Russell 2000 equities under three stock-selection regimes: a pure-alpha trigger that isolates abnormal stock moves not explained by macro indicators, a pure-beta trigger that captures macro-indicator moves before the stock itself fires, and a beta trigger in which both channels agree. Across the full holding-period grid, the separated pure-alpha and pure-beta legs usually dominate the beta intersection on Sharpe and return. Two horizons are especially informative. At one day, pure beta can work under low and moderate transaction costs because it captures immediate lead-lag spillovers from liquid macro and sector indicators into exposed small-cap stocks, but this advantage disappears at 100 bps when turnover and microstructure noise dominate. At 40 days, pure beta works for a different reason: slower macro repricing overtakes the firm-specific pure-alpha channel. The strongest conservative row is pure beta with GPT-4o mini sentiment, a Student-t target, a 40-day holding period, and risk parity allocation, reaching Sharpe 2.33 at 100 bps. The results suggest that stock-selection regime and allocator choice matter at least as much as the sentiment model, and that separating firm-specific and macro-exposure triggers is more informative than requiring both to fire simultaneously. |
| Date: | 2026–08 |
| URL: | https://d.repec.org/n?u=RePEc:arx:papers:2608.12283 |
| By: | Binzhi Chen; Annalivia Polselli; Paul S. Clarke |
| Abstract: | Factor structures are central to empirical work in economics and finance, and are usually used to model time-varying unobserved heterogeneity through interactive fixed effects (IFE). Existing IFE estimators rest on low-dimensional and linear specifications in the covariates, assumptions which are increasingly restrictive in applications drawing on rich datasets with controls of unknown functional form. This paper develops a Double Machine Learning estimator for the high-dimensional partially linear panel model with interactive fixed effects (panel DML-IFE). The method combines projection-based defactorisation of the data, in the spirit of Common Correlated Effects (CCE), with a Neyman-orthogonal score function and cross-fitting procedure, and accommodates low-rank factor structures in outcomes and treatments alongside high-dimensional, potentially nonlinear covariate effects estimated by machine learning algorithms. Monte Carlo simulations show that panel DML-IFE outperforms conventional IFE estimator outside the correctly-specified linear case, with bias reduction driven primarily by the time and covariate dimensions. An empirical application to U.S. stock returns shows that several effects documented under linear specifications lose statistical significance once high-dimensional nonlinear confounding and the presence of IFE are jointly accounted for. |
| Date: | 2026–08 |
| URL: | https://d.repec.org/n?u=RePEc:arx:papers:2608.01137 |
| By: | Liexin Cheng; Xue Cheng; Shuaiqiang Liu; Cornelis W. Oosterlee |
| Abstract: | Automated code generation is becoming an important tool in quantitative finance, where large language models can generate option pricing implementations directly from mathematical model specifications. Validating such implementations, however, requires considerably more than conventional software testing: numerical pricing methods must remain mathematically consistent, numerically stable, and reliable across a wide range of model parameters. We introduce RIDGE, an autonomous validation framework in which generated pricing implementations are subjected to structured no-arbitrage tests, stress tests, benchmark comparisons, and consistency checks. Validation evidence is interpreted diagnostically, while the resulting knowledge is accumulated in a repository and reused across models and successive validation iterations. This enables systematic refinement of both the pricing implementation and the validation methodology. The framework is applied to five stochastic volatility models. Across these studies, all detected implementation defects are removed and, in two cases, the validation process reveals methodological limitations and motivates the development of alternative numerical methods. The supplementary material is available in the GitHub repository: https://github.com/ShQiangLiu/ridge. |
| Date: | 2026–07 |
| URL: | https://d.repec.org/n?u=RePEc:arx:papers:2607.25199 |