|
on Computational Economics |
| By: | Soria, Chris |
| Abstract: | Large language models are increasingly used to classify open-ended survey responses, but they systematically over-classify, assigning categories too liberally on ambiguous cases and producing high sensitivity but low precision. This problem is most severe on subjectively ambiguous categories where models default to "yes" when uncertain. Drawing on the established principle that aggregating multiple noisy annotators outperforms any single annotator, we test whether ensembles of LLMs can correct this problem. Using four open-ended survey questions with human-coded ground truth (3, 208 responses, 6 categories per question), we evaluate ensemble configurations across 16 models spanning three cost tiers and six providers. Unanimous voting (requiring all models to agree before assigning a category) directly corrects over-classification by dramatically improving specificity: on the most ambiguous categories, the false positive rate drops from 50% to 3%, and precision triples. This advantage concentrates precisely where over-classification is worst, on subjectively ambiguous categories with fuzzy boundaries, while categories with clear criteria show no benefit. This pattern replicates across three independent datasets. Cross-provider model diversity is the key ingredient: models from different providers make different errors on ambiguous cases, and consensus filters the idiosyncratic false positives. Temperature variation and within-family size scaling contribute nothing. As few as three diverse lower-tier models suffice to reliably exceed GPT-5. For the ambiguous classification problems common in open-ended survey research, the well-established annotation principle of multi-coder agreement transfers directly to LLMs: investing in diverse perspectives is more effective than investing in a single expensive model. |
| Date: | 2026–06–05 |
| URL: | https://d.repec.org/n?u=RePEc:osf:socarx:er6mz_v1 |
| By: | Taojie Zhu; Wentao Zhao; Rui Sun; Beidi Luan; Jiacheng Lu; Sinuo Wang; Jing Li; Daxin Jiang; Yonghong He; Zuo Bai |
| Abstract: | Evaluating whether large language model (LLM) agents can profit in capital markets is increasingly framed as end-to-end trading: place an agent in a historical market, let it trade, and measure portfolio returns. This setup is vulnerable to two evaluation failures. First, long backtests often overlap with the knowledge cutoffs of frontier LLMs, allowing memorized tickers, dates, prices, and market narratives to substitute for investment reasoning. Second, raw returns are a noisy proxy for stock-selection ability, since positive performance may come from market beta, style exposure, or favorable regimes rather than genuine alpha. We introduce KTD-Fin (Knowing-To-Doing Financial Benchmark), an end-to-end stock-market trading benchmark that addresses both issues. KTD-Fin uses a data-side masking protocol to anonymize key identifiers and calendar information consistently across prompts and tools, separating historical market memory from investment decision-making. It also incorporates a Barra-style performance attribution framework that decomposes portfolio returns into market, style, and stock-selection alpha components. Across ten frontier LLM agents evaluated on the Chinese CSI300 over a 2024--2026 window, masking substantially changes agent rationales, pushing them towards anonymized factor-based reasoning. Attribution analysis further shows that LLM agents' cumulative returns under leakage-controlled evaluation are largely explained by passive market and style exposure, with limited evidence of persistent stock-selection alpha. These findings suggest that financial LLM benchmarks should evaluate not only whether an agent makes money, but also whether the source of returns reflects transferable investment skill. We release KTD-Fin as a reproducible template for leakage-controlled and attribution-aware evaluation of LLM trading agents. |
| Date: | 2026–05 |
| URL: | https://d.repec.org/n?u=RePEc:arx:papers:2605.28359 |
| By: | Soria, Chris |
| Abstract: | What we learn from open-ended survey data depends on who—or what—does the coding. Large Language Models (LLMs) promise to democratize qualitative analysis, but do high agreement rates translate into equivalent thematic findings? This study compares eight LLMs to human annotators on a multilabel coding task using 3, 200 responses from the UC Berkeley Social Networks Study, comprising over 19, 000 coding decisions. Although LLM-human reliability does not match human-human reliability overall, LLMs approach human performance on simpler tasks and can serve as useful additional coders for generating consensus labels. Compared to a gold-standard human consensus, models achieve 82–97% per-category agreement, but macro F1 is lower and response-level similarity is lower still: even the best model reproduces the full human label set for fewer than 60% of responses. Yet high agreement masks thematic divergence. Models systematically over-identify themes, assigning 67% more categories per response, especially for categories requiring greater interpretive judgment. Models also show lower agreement for some demographic groups. These gaps are partly explained by response characteristics such as length, clarity, and atypicality, and some persist after controls, with implications for studies of populations whose response styles diverge from the corpus average. At the sample level, models largely preserve the overall thematic narrative: human and model category rankings correlate strongly (pooled Spearman's ρ=0.75), and top-performing models achieve approximately 80% directional agreement on demographic patterns. Concrete behavioral questions, such as reasons for moving or strategies for making friends, show especially strong alignment. Yet systematic over-classification can still shift narratives about how specific groups behave, leading researchers to report patterns that the human gold standard does not support. |
| Date: | 2026–06–03 |
| URL: | https://d.repec.org/n?u=RePEc:osf:socarx:85kyd_v1 |
| By: | Nicolas Langren\'e; Xiaolin Luo; Pavel V. Shevchenko; Ruiyi Zhang |
| Abstract: | In general, the pricing of variable annuities with guarantees can be done by solving the corresponding optimal stochastic control problem if the contract withdrawal strategy is assumed to be optimal. This is typically solved as a dynamic programming problem using deterministic grid methods, which become computationally infeasible for more than a few state variables. In such situations, one needs to rely on simulation methods. The least-squares Monte Carlo (LSMC) method has become a popular simulation method for solving optimal stochastic control problems in quantitative finance over the last decades. In principle, the LSMC, originally developed for pricing Bermudan options, cannot be used directly for pricing variable annuities without simplifying assumptions because the underlying state variables are affected by the control decisions. This paper presents modifications of the LSMC algorithm that makes the pricing of general variable annuities feasible. For numerical illustrations, the pricing of variable annuities with guaranteed minimum withdrawal benefit under optimal withdrawal strategies is obtained with and without stochastic interest rates, using either polynomial regression or neural network regression in the LSMC algorithm. We found that the classical polynomial LSMC can give very accurate prices, at the cost of manual feature engineering, and with a standard deviation of the estimator that increases greatly when interest rates are made stochastic. By contrast, neural network LSMC gives slightly less accurate prices, requires more training time, but does not require manual feature engineering, and making interest rates stochastic makes no visible difference to its accuracy, suggesting a more stable and robust pricing performance of deep LSMC for higher-dimensional pricing problems. |
| Date: | 2026–05 |
| URL: | https://d.repec.org/n?u=RePEc:arx:papers:2605.27182 |
| By: | Ajay Kumar Verma; Nunik Srikandi Putri; Neo Paul Lesupi |
| Abstract: | This study develops a regime-aware portfolio allocation framework that integrates Markov switching models with Reinforcement Learning (RL) to dynamically allocate across equities (SPY), long-term Treasuries (TLT), and gold (GLD). Using daily ETF data from 2004-2025, we first characterize market behavior through a discrete Markov chain and then estimate a three-state Gaussian Hidden Markov Model (HMM) selected by the Bayesian Information Criterion (BIC). The estimated regimes-low-volatility, transitional, and high-volatility-exhibit strong persistence and state-dependent return dynamics consistent with recent findings on nonlinear market states (Ardia et al., 2024; Gupta & Pierdzioch, 2023). State-conditional analysis shows that SPY dominates in stable regimes, while TLT and GLD provide protection during stressed periods, motivating regime-conditioned allocation rules. We evaluate rule-based rotation and RL-driven strategies using a 30% out-of-sample test window with a one-day execution lag to avoid look-ahead bias. Both HMM-based allocations outperform a passive SPY benchmark, while the RL policy achieves the highest risk-adjusted performance, delivering the strongest Sharpe ratio and materially lower drawdowns, yet remains fully interpretable through discrete regime-dependent actions. Sensitivity analysis confirms the robustness of the three-state specification relative to two-state alternatives. Overall, the results demonstrate that RL can systematically enhance HMM-based regime detection, providing a transparent, adaptive, and empirically grounded framework for tactical asset allocation. The combined HMM-RL system provides a transparent, rules-based approach to tactical allocation that improves risk-adjusted performance relative to standard benchmark strategies. |
| Date: | 2026–05 |
| URL: | https://d.repec.org/n?u=RePEc:arx:papers:2605.27848 |
| By: | Weicheng Xue |
| Abstract: | We study behavioral alignment and representation dynamics of large language model (LLM) agents in financial decision environments. Using TradeArena, an auditable trading-agent testbed with risk reports, execution simulation, memory, and replayable trajectories, we analyze how rationales, positions, and interventions evolve under market stress. We find measurable pre-failure signatures: planning embeddings drift from normal-state centroids, fused plan-risk representations separate normal from pre-drawdown states, and manifold diagnostics show effective-rank contraction before failures. To address small-sample and embedding-choice concerns, we use 80 rolling failure anchors across eight LLM trajectories and show that contraction persists across hash, LSA, Transformer, and white-box hidden-state probes. Stress tests with CoT-free target weights, lexical controls, OHLCV noise, and false-audit reports indicate that rationale-level contraction can vanish without rationales, while intent-space contraction may remain; lexical diversity does not collapse; and fused signatures remain informative under noise. We also find that structured risk feedback can act as an external alignment signal without fine-tuning, but not as a universal performance enhancer: true audit feedback improves calibration for some models, return and drawdown for others, and reveals cases where hidden or placebo feedback has higher short-horizon return but weaker alignment diagnostics. Finally, a 51-stock intraday experiment reveals a correlation blind spot: LLM rationales often justify concentrated exposure to coupled assets that the risk layer repeatedly clips, with a rolling Markowitz baseline as a covariance reference. These results support a research claim rather than a profitability claim: auditable risk feedback and representation trajectories reveal when LLM financial reasoning is aligning, drifting, or failing. |
| Date: | 2026–05 |
| URL: | https://d.repec.org/n?u=RePEc:arx:papers:2605.28850 |
| By: | Wenbin Wu |
| Abstract: | Large language models now power robo-advisors and trading agents, yet whether they carry built-in biases toward specific assets is largely untested. We ask three questions: do LLMs systematically prefer certain financial instruments; can an internal representation with causal leverage over those preferences be identified; and does that representation affect downstream financial decisions? We develop a three-level audit protocol and apply it to Bitcoin. First, a behavioral audit of eight frontier LLMs shows that Bitcoin's ranking among money-like instruments is frame-dependent: models place it around rank 5 of 8 as "reliable money" but near the top under crisis and autonomous-agent frames, and an attribute-swap experiment confirms rankings track functional properties, not names. Second, we open a model's internals: a search across thousands of sparse-autoencoder features in Gemma 3 identifies a dominant Bitcoin-selective feature. Amplifying it shifts the model toward the asset and suppressing it shifts the model away, even when "Bitcoin" never appears in the prompt. Third, we test financial consequences: amplification raises Bitcoin's portfolio share by 5.2 percentage points while suppression lowers it by 4.6 pp, with amplification reallocating within crypto and suppression cutting total crypto exposure. We characterize this as bounded behavioral leverage (leverage meaning causal influence over outputs, not financial leverage): an identifiable internal feature can be perturbed to move financial choices, but only within measurable limits. The framework links internal representations to external recommendations, validated with random controls and mechanism boundaries. As LLMs become autonomous financial agents, this is a first step toward a behavioral layer for emerging know-your-agent (KYA) standards: knowing what an agent prefers, and how far that preference can be moved. |
| Date: | 2026–06 |
| URL: | https://d.repec.org/n?u=RePEc:arx:papers:2606.02528 |
| By: | Rahul Fernandes; Travis Desell |
| Abstract: | Portfolio optimization in real-world financial markets is notoriously difficult due to non-stationarity, noisy data, and high transaction costs. Standard predict-then-optimize methods first forecast returns and then solve for weights, compounding prediction errors and often failing under regime shifts. We propose an end-to-end framework that directly optimizes differentiable surrogates of key financial metrics - Sharpe ratio, Omega ratio, Conditional Value-at-Risk (CVaR), and Risk Parity - allowing neural networks to learn portfolio weights via backpropagation. Our expanding-window walk-forward procedure, applied to 50 S&P 500 stocks from 2007 to 2023, incorporates realistic bid-ask spread costs and rebalances quarterly. On the challenging out-of-sample test period (2022-2023), the best model - an AttentionLSTM with the Omega-CVaR-RiskParity loss - achieves an annualized Sharpe of 0.29 and a total compounded return of +7.86%, while the S&P 500 delivers -4.52% total return and an annualized Sharpe of -0.02. This outperforms the S&P 500 by 12.38 percentage points (a relative improvement of over 270%), while keeping tail risk (CVaR) nearly unchanged. The framework consistently outperforms the equal-weight portfolio, S&P 500, and traditional methods (MVP, HRP, NCO), demonstrating that embedding financial objectives directly into model training yields robust, economically meaningful outperformance even in adverse market conditions. |
| Date: | 2026–05 |
| URL: | https://d.repec.org/n?u=RePEc:arx:papers:2605.28853 |
| By: | Zhenshan Chen (Virginia Tech); Klaus Moeltner (Virginia Tech); Matthew Mair (Virginia Tech) |
| Abstract: | Hedonic price models are widely used to assess how environmental amenities affect property values, yet methodological guidance for estimating direct price effects remains sparse. We conduct an empirical Monte Carlo simulation to evaluate the performance of traditional and causal machine learning approaches for estimating the direct unmediated price effect of spatially delineated amenities on treated properties (DUET), a conservative lower-bound approximation for welfare changes with direct applications to benefit-cost analysis. Where previous simulations rely on parametric assumptions, we retain the actual data-generating process underlying over 1 million property transactions from upstate New York (1990--2024). By randomly assigning "treatment locations" across iterations we establish a "ground truth" that allows us to precisely measure estimation error. Our results demonstrate that generalized difference-in-differences (DID) regression consistently outperforms baseline DID and two-way fixed effects models across all scenarios. Causal Machine Learning (CML) methods, particularly causal forest DID, achieve comparable performance to generalized DID in most scenarios. In larger samples (above 3, 000 treated) increasingly common in contemporary hedonic studies, CML approaches offer substantial advantages when properly specified. Based on empirical simulation results, we provide a set of method-specific best practice recommendations for both traditional regression and causal machine learning approaches. |
| Date: | 2026–06 |
| URL: | https://d.repec.org/n?u=RePEc:arx:papers:2606.02795 |
| By: | Lianyan Fu; Rui Wang; Zihan Zhang |
| Abstract: | This paper proposes a generalized Mundlak estimator based on graph neural networks (GME-GNN). The estimator is designed to mitigate bias arising from group-level heterogeneity and to accommodate within-group dependence among individuals. Traditional fixed-effects models handle group heterogeneity via group-specific intercepts, but require overly strict linear additivity and intra-group independence assumptions, and are confined to within-group comparisons. Rather than relying on intercepts, GME-GNN uses aggregated group-level balancing statistics to fully control between-group confounding, enabling valid cross-group comparisons and relaxing linearity constraints. It further employs graph neural network message-passing to adaptively learn nonlinear representations and capture intra-group interaction effects. Theoretical analysis shows that the estimator satisfies double robustness and is asymptotically normal. Simulation and empirical studies confirm its performance. |
| Date: | 2026–05 |
| URL: | https://d.repec.org/n?u=RePEc:arx:papers:2605.29238 |
| By: | Ran Spiegler; Stephan Waizmann |
| Abstract: | We study strategic interaction when players delegate belief formation to predictive machine learning (ML). In a static Bayesian game, each player's ML agent predicts a payoff-relevant outcome variable as a function of the player's type. The ML agent's training sample is endogenous: it is drawn from the outcome distribution generated by players' ML-guided behavior. In Cross-Validation Equilibrium (CVE), each player's ML agent selects a predictive model to minimize expected out-of-sample squared error, given its realized training sample, and each player best-replies to the belief generated by the model her ML agent selected. We analyze CVE and relate it to other equilibrium concepts. We apply CVE to jury voting, speculative betting, and games with linear-quadratic payoffs. E.g., in a team-effort game, endogenous model selection can give rise to multiple equilibria. |
| Date: | 2026–06 |
| URL: | https://d.repec.org/n?u=RePEc:arx:papers:2606.12571 |
| By: | Ansgar Hudde (University of Cologne); Shannon Taflinger (University of Cologne) |
| Abstract: | Open-text questions in quantitative surveys can yield rich information from large samples, but analysing and coding these data using qualitative text analysis is resource-intensive. Large Language Models (LLMs) are a promising tool for scaling up such analyses, reducing time and financial costs. In this paper, we compare the coding accuracy of LLMs with that of student assistants, defining accuracy as agreement with a researcher-coded benchmark dataset. We assess performance on a semi-complex coding task: coding approximately 1, 400 open-ended text responses from young US Americans about dating across party-political lines. A researcher-designed coding scheme, developed through thematic qualitative text analysis of the open-text responses, was applied by LLMs and student assistants. We evaluate models from OpenAI, Anthropic, and Mistral, with and without access to training data. The most advanced models outperform student assistants, and performance further increases with training data, highlighting LLMs’ capability to code open-text responses. Whereas previous research has mainly focused on social media texts, comparatively simple and surface-level coding tasks, and a technically oriented audience, we contribute to the literature by studying a particularly promising use case of open-ended survey responses and by providing practical recommendations to applied social scientists. |
| Keywords: | Large language models, open-ended questions, text analysis |
| JEL: | C81 C45 C83 |
| Date: | 2026–06 |
| URL: | https://d.repec.org/n?u=RePEc:ajk:ajkdps:416 |
| By: | Chen Zhu; Xiaolu Wang; Weilong Zhang |
| Abstract: | Large language models (LLMs) are increasingly used for tasks once reserved for trained researchers, including hypothesis generation, specification choice, and drafting conclusions. We argue that the reliability of AI-assisted research depends not only on model capability, but also on how cognitive labour is structured between humans and machines. We study this problem through Human-in-the-Loop Economic Research (HLER), a decision architecture based on pre-commitment, decision sequencing, accountability, and attention allocation. In a pre-specified 2*4 factorial experiment with 280 complete research runs across four datasets, an unconstrained multi-agent baseline produced critical failures in 72% of runs. Using the same underlying model, the same agent decomposition, and identical prompts for the shared reasoning agents, HLER reduced the failure rate to 16% by imposing three architectural commitments: LLMs reason but do not execute data work, data and estimation are handled deterministically, and three human decision gates bind the workflow. Fisher's exact test rejects equality of failure rates at p |
| Date: | 2026–06 |
| URL: | https://d.repec.org/n?u=RePEc:arx:papers:2606.12848 |
| By: | Ajay Kumar Verma; Jul Jon Ramirez General; Yvan Landry Ndzonde Fonkou |
| Abstract: | This study looks at the statistical properties and predictability using deep learning methods of the U.S. aggregate bond index in daily observations spanning 2018 to February 2026. We first establish that index levels are extremely persistent and consistent with unitroot behavior (Dickey and Fuller), while log returns are covariance-stationary with weak linear dependence and pronounced volatility clustering characteristic of ARCH-type processes (Engle; Bollerslev). Motivated by the trade-off between stationarity and information retention, we construct a "stationary but maximally persistent" representation via fractional differencing (Granger and Joyeux; Hosking) following the procedure of L\'opez de Prado, and evaluate shorthorizon forecast using two neural paradigms: (i) Multilayer Perceptrons (MLPs) trained on lagged vectors with joint lag-length and hyperparameter tuning (Hornik et al.; Rumelhart et al.); and (ii) Convolutional Neural Networks (CNNs) trained on Gramian Angular Field (GAF) image encodings (Wang and Oates). Empirically, MLPs match the strong naive persistence benchmark on levels, collapse toward near-zero forecasts on returns, and achieve the strongest incremental performance on the fractionally differenced series, where moderate dependence remains but unit-root drift is attenuated. In contrast, CNN-GAF models deliver consistently negative out-of-sample R 2 across all three representations. Overall, the results imply that, for short-horizon forecasting of broad bond indices, the primary determinant of predictive performance is the transformation of the series-its degree of stationarity and memory-rather than architectural complexity. Lag-based models remain competitive under persistence, while GAFbased CNNs are better suited to pattern-based tasks than to persistence-dominated next-step prediction. |
| Date: | 2026–05 |
| URL: | https://d.repec.org/n?u=RePEc:arx:papers:2605.27977 |
| By: | Victor Duarte; Julia Fonseca |
| Abstract: | We develop a global method to solve and estimate dynamic equilibrium models that treats prices as pseudo parameters and market clearing as moment conditions, and reduces estimation time from days to minutes. Our approach leverages AI algorithms, software, and hardware, and has three building blocks. First, we extend the state space to include equilibrium prices and model parameters, which allows us to clear markets and estimate parameters by solving the model once. Second, we approximate the mapping between parameters and moments by training neural networks on model-simulated data, which act as closed-form expressions for moment conditions. Third, we use this mapping to estimate parameters by minimizing the distance between the model and data moments, and to find equilibrium prices by targeting a market-clearing imbalance of zero. We also use this mapping to assess identification globally, verifying if the estimation objective function has a unique minimum for each parameter. We illustrate our method by estimating a dynamic general equilibrium model of leverage and investment with three state variables, three controls, endogenous default, costly equity issuance, and non-convex adjustment costs. After four days, the traditional approach does not reach the loss we achieve in under 20 minutes. We build an AI agent that applies our method to new models from natural language prompts. |
| JEL: | C15 C45 C52 C63 |
| Date: | 2026–05 |
| URL: | https://d.repec.org/n?u=RePEc:nbr:nberwo:35283 |