nep-cmp New Economics Papers
on Computational Economics
Issue of 2026–08–31
twenty-six papers chosen by
Stan Miles, Thompson Rivers University


  1. AlphaZeroBeta: Deep Reinforcement Learning for Market-Neutral Portfolios By Boris Belyakov
  2. Store-Level Food Healthfulness Measurements Using Scanner Data and Machine Learning Imputation By Yang, Bixuan; Kropp, Jaclyn; Grigsby-Calage, Charles; Sowell, Samantha; Mullally, Conner; Volpe, Ricky; Byrne, Anne
  3. Photonic Quantum Computing vs. Classical Solvers in Constrained Factor Portfolio Optimization By Nirvik Sahoo; Chyng Wen Tee; Paul Robert Griffin
  4. Yield Curve Prediction with Machine Learning: Forecasting Approaches and the Role of Macroeconomic Predictors By Jeron Tan Kang
  5. Predicting Retirement and Social Security Claiming Decisions using Machine Learning By Kwon, Alexander; Maliar, Lilia
  6. FinBench: Time-Gated Calibration and Uncertainty Benchmarking for Agentic Financial Forecasting By Rishab Ghosh; Vinay Devarakonda
  7. Data-Driven Measures of High-Frequency Trading By Gbenga Ibikunle; Ben Moews; Dmitriy Muravyev; Khaladdin Rzayev
  8. Learning Optimal Dynamic Matching via Graph Neural Networks By Genta Okada; Shunya Noda; Junpei Komiyama; Akira Matsushita
  9. Machine Learning for Estimating Catastrophic Health Spending in Disaster-Affected, Data-Scarce Settings By Himaz, Rozana; Salmanidou, Dimitra; Ghaffarian, Saman
  10. Defining Current and Expected Financial Constraints using AI: Reinterpreting the Cash Flow Sensitivity of Cash By Cho, Rachel; Görtz, Christoph; McGowan, Danny; Schröder, Max
  11. Integrating Non-traditional Data and AI into Central Banking: A Canadian Perspective By James Chapman; Ajit Desai; Maryam Haghighi; James (Jim) C. MacGee
  12. Reinforcement Learning-Guided NSGA-II Enhanced with Gray Relational Coefficient for Multi-Objective Optimization: Application to NASDAQ Portfolio Optimization By Zhiyuan Wang; Qinxu Ding; Ding Ding; Siying Zhu; Jing Ren; Yue Wang; Chong Hui Tan
  13. Mastering Stochastic OLG Models in Continuous Time By Yves Achdou; Johannes Brumm; Lukas Frank
  14. Learning under Opponent Unawareness in Linear-Quadratic Stochastic Games By Dantong Chu; Xuefeng Gao; Yufei Zhang
  15. Predicting Startup Exit from Textual Descriptors - A Computational Linguistics Framework By Alberto M. G. Saruggia; Sebastien Germano
  16. Can Open-Weight Models Compete on Financial Text Comprehension? By Jan Sp\"orer
  17. AgonAlpha: Autonomous Alpha Discovery via Prompt Economy and Scalable Agentic Search By Weicheng Ye; Youran Sun; Xingyu Ren; Shunyao Yu; Chugang Yi; Haizhao Yang
  18. From 4-Digit Codes to Social Class: Benchmarking Supervised Classification, Semantic Retrieval, and LLM Reranking for Multilingual Occupational Coding and Downstream Stratification (Technical report: model design, evaluation, and derived-measure accuracy) By YILMAZ, Erdem; Kruithof, Elias H.; Verhaeghe, Pieter-Paul
  19. When LLM Agents Negotiate: Private Information and Dynamic Bargaining in Supply Chains By Chen Liang; Fasheng Xu
  20. The Innate Economic Preferences of Language Models By Joy Buchanan; Joshua Foster
  21. Dynamic Coalition Formation and Communication Pricing in Skill-Based Agentic AI Systems By Mojtaba Eslami
  22. AI-Driven Multiscenario Interest Rate Forecasting: A Proof of Concept for Banking Asset Management By Ekkehardt Bauer; Dirk Holl\"ander; David Scholz; Linus Wolff; Christoph Ostermair; Kyrillus Aiad; Joachim Hasebrook
  23. Macro News in Market Moves: Classifying News through Asset Co-movements By Bruno Feunou; Jean-Sébastien Fontaine; Rishi Vala
  24. Simulation-Driven Analysis of Warehouse Operations in Multi-Level Robotic Mobile Fulfillment Systems to Support Decision-Making By Wenzel, Julia
  25. Optimal risk for pension funds: the sustainability of the UK Universities pension scheme By Miles, David; Sefton, James
  26. An Analytic COS Method for Compound Option Valuation By Zhipeng Huang; Cornelis W. Oosterlee

  1. By: Boris Belyakov
    Abstract: Market-neutral portfolios aim to generate consistent returns while offsetting systematic market risk. Traditional approaches based on factor models or convex optimization often underperform during market regime shifts or when structural assumptions break down. We propose AlphaZeroBeta, a deep reinforcement learning framework designed to deliver benchmark-relative alpha (excess returns) with near-zero beta (market neutrality). AlphaZeroBeta combines a composite reward function that balances risk-adjusted excess return, benchmark correlation, and transaction costs with a CNN-GRU policy trained end-to-end via Recurrent PPO and evaluated through a rolling walk-forward protocol. Backtests covering 2014-2024 across seven equity indices show that the model achieves higher Sharpe ratios than the baselines while maintaining near-zero benchmark correlations and competitive drawdowns.
    Date: 2026–07
    URL: https://d.repec.org/n?u=RePEc:arx:papers:2607.18001
  2. By: Yang, Bixuan; Kropp, Jaclyn; Grigsby-Calage, Charles; Sowell, Samantha; Mullally, Conner; Volpe, Ricky; Byrne, Anne
    Keywords: Food Consumption/Nutrition/Food Safety
    Date: 2026
    URL: https://d.repec.org/n?u=RePEc:ags:aaea26:404588
  3. By: Nirvik Sahoo; Chyng Wen Tee; Paul Robert Griffin
    Abstract: The authors present a rigorous empirical evaluation of three distinct optimization paradigms for institutional factor portfolio construction: an entropy-based photonic quantum annealer (Dirac-3, Quantum Computing Inc.), a commercial mixed-integer programming solver (Gurobi), and a model-free deep reinforcement learning agent (SAC). Evaluating these pipelines on the Jensen-Kelly-Pedersen 13-factor equity library across 164 months test window, we implement a full factorial penalty sweep comprising 48 hyperparameter configurations that govern return, volatility, and skewness trade-offs. Our findings demonstrate that while photonic hardware can locate superior risk-return topologies within a narrow operating range, classical mixed-integer programming remains superior for risk-constrained mandates requiring tight tail-risk control and cross-seed stability. Furthermore, we document structural failure modes in reinforcement learning factor allocators under unanchored higher-moment shaping. We translate these empirical results into actionable, mandate-specific guidelines for quantitative portfolio managers deploying advanced optimization engines.
    Date: 2026–08
    URL: https://d.repec.org/n?u=RePEc:arx:papers:2608.14134
  4. By: Jeron Tan Kang
    Abstract: This paper compares direct-yield and factor-based approaches to U.S. Treasury yield curve forecasting using a common high-dimensional macroeconomic information set. Forecasts are evaluated on monthly zero-coupon yields over the 2015-2025 out-of-sample period. Gains over the random walk are concentrated at short maturities and in slope forecasts, and decline with the forecast horizon. Direct-yield models perform best for slope forecasts and are relatively stronger at short horizons, while factor-based models become more competitive at longer horizons. Macroeconomic predictors provide clear incremental predictive power, strongest for slope-related movements. A trading simulation reinforces that macro-augmented models perform best in slope trades. The simulation also highlights a gap between statistical and economic performance, as the random walk is a strong benchmark under statistical loss but performs poorly as a trading signal.
    Date: 2026–07
    URL: https://d.repec.org/n?u=RePEc:arx:papers:2608.07536
  5. By: Kwon, Alexander; Maliar, Lilia
    Abstract: We demonstrate that machine learning substantially improves predictions of individual decisions about retirement and Social Security (SS) claims. When predicting the number of people receiving SS, we achieve an error of less than 1%, while the benchmark model employed by the Social Security Administration (SSA) results in a greater than 4% error, and in forecasting SS claiming decisions, we attain an error of 0.2%, while the benchmark exceeding 2%. Based on averages, we show that a 3% difference in prediction amounts to 39.6 billion dollars annually. The set of important variables selected by our model significantly differs from that of the SSA model. We use Shapley values to evaluate the non-linear contributions of the selected variables to predictive outcomes.
    Keywords: retirement and Social Security
    JEL: C53 H55 J14 J26
    Date: 2024–07
    URL: https://d.repec.org/n?u=RePEc:cpr:ceprdp:19198
  6. By: Rishab Ghosh; Vinay Devarakonda
    Abstract: Large language models (LLMs) are increasingly used as components of agentic systems that observe, plan, and act. In finance, even "assistive" systems become decision-relevant once their outputs are used to size trades or allocate risk. A key failure mode is the confidence--competence gap: a model that is only slightly better than chance but consistently overconfident will, under typical bet-sizing rules, generate negative long-run growth. Existing benchmarks emphasize semantic understanding or point accuracy, but do not directly test probabilistic calibration under the temporal constraints and non-stationarity that define real markets. We introduce FinBench, a benchmark designed to evaluate calibration and uncertainty quality for financial forecasting in a setting that is (i) strictly time-gated to avoid look-ahead bias and (ii) evaluated with strictly proper scoring rules that penalize hallucinated confidence. FinBench tasks require models to output (a) a probability of positive return and (b) an 80% prediction interval for realized log return; evaluation uses the Brier score and the Winkler interval score, along with skill scores against hard baselines. This paper describes the benchmark specification and reports a small pilot run (one trading day; three liquid tickers; 33 forecasts) as a sanity check of the pipeline. The pilot illustrates how calibration-sensitive metrics distinguish between "confident but fragile" behavior and uncertainty-aware forecasting.
    Date: 2026–06
    URL: https://d.repec.org/n?u=RePEc:arx:papers:2607.16229
  7. By: Gbenga Ibikunle; Ben Moews; Dmitriy Muravyev; Khaladdin Rzayev
    Abstract: We introduce data-driven measures of high-frequency trading (HFT) that distinguish between liquidity-supplying and liquidity-demanding strategies. We train machine learning models on a proprietary dataset with observed HFT activity, then apply these models to public intraday data to generate HFT measures across all U.S. stocks during 2010-2023. Our measures outperform conventional proxies, which struggle to capture the temporal dynamics of HFT. Consistent with theory, our measures respond to a quasi-exogenous speed bump introduction and a data feed upgrade. The measures help uncover the differential impact of HFT on information acquisition. Liquidity-supplying HFT improves price informativeness around earnings announcements, while liquidity-demanding HFT impedes it.
    Date: 2026–08
    URL: https://d.repec.org/n?u=RePEc:arx:papers:2608.00858
  8. By: Genta Okada; Shunya Noda; Junpei Komiyama; Akira Matsushita
    Abstract: Dynamic matching markets require decisions about whom to match and when: matching now yields value but removes participants who may create better future opportunities. We develop a value-based reinforcement-learning framework for this problem on finite, evolving weighted graphs. We study an infinite-horizon continuous-time model with stochastic arrivals, node-type transitions, edge realizations, and exogenous exits. We prove an event-time reduction: without loss of optimality, the planner acts immediately after each exogenous event and then waits for the next one. We further show that the optimal edge-wise $Q$-function is characterized by a single continuation-value function on post-decision residual graphs, reducing the learned object from state-action values to graph values. Exact action selection still requires combinatorial matching optimization; we approximate the value with a graph neural network, train it by temporal-difference learning, and use it in a forward-greedy matching heuristic. In a binary-type benchmark, the learned policy substantially outperforms immediate and threshold-greedy rules by preserving common nodes for rare arrivals of valuable matches while forming lower-value matches only in thick pools. In a kidney paired donation benchmark, it performs similarly to immediate greedy when exits are unpredictable, recovers the logic of patient matching when warnings are reliable, and outperforms the better of Immediate Greedy and Patient Greedy across intermediate warning probabilities. These results show that residual-graph value learning yields state-dependent dynamic matching policies that adapt to realized connectivity and exit information.
    Date: 2026–07
    URL: https://d.repec.org/n?u=RePEc:arx:papers:2607.28925
  9. By: Himaz, Rozana; Salmanidou, Dimitra; Ghaffarian, Saman
    Abstract: Natural hazard events can increase out-of-pocket health costs and push vulnerable households into poverty. Mitigation measures require understanding changes in health spending patterns using pre- and post-event data, but such data are often unavailable in disaster-affected settings. This represents a fundamental measurement challenge: the absence of pre-event baseline data makes it impossible to construct the counterfactual quantities needed for welfare analysis. To address this measurement problem, we develop a hybrid machine learning approach to estimate unobserved household health spending using longitudinal survey data from Indonesia. We first develop a model around the 2006 Yogyakarta earthquake, for which complete data are available. The model learns spending patterns across income, hazard intensity, and other characteristics, achieving >70% accuracy in a noisy and complex domain. After testing the model for transportability, we apply it to post-2004 Indian Ocean tsunami survey data in Indonesia, to predict plausible baseline health spending. These predictions are used to evaluate the impact of the tsunami on health spending to reveal that without targeted aid, catastrophic health spending would have increased from 4.5% to 29.4% and that moderately damaged households experienced more cost increases than heavily damaged ones. By combining artificial intelligence with 2 household survey data, our framework is a proof-of-concept, for addressing data gaps in official economic statistics, demonstrating how machine learning can enable counterfactual welfare measurement where conventional data collection is absent or incomplete.
    Keywords: Natural hazards; catastrophic health spending; disaster risk reduction; tsunami; earthquake; Indonesia; machine learning
    JEL: C45 C51 C52 C53 I19 O13 Q54
    Date: 2026–03–23
    URL: https://d.repec.org/n?u=RePEc:eoe:escoed:escoe-dp-2026-05
  10. By: Cho, Rachel; Görtz, Christoph; McGowan, Danny; Schröder, Max
    Abstract: We propose a new approach to identify firm-level financial constraints by applying artificial intelligence to text of 10-K filings by U.S. public firms from 1993 to 2021. Leveraging transformer-based natural language processing, our model captures contextual and semantic nuances often missed by traditional text classification techniques, enabling more accurate detection of financial constraints. A key contribution is to differentiate between constraints that affect firms presently and those anticipated in the future. These two types of constraints are associated with distinctly different financial profiles: while firms expecting future constraints tend to accumulate cash preemptively, currently constrained firms exhibit reduced liquidity and higher leverage. We show that only firms anticipating financial constraints exhibit significant cash flow sensitivity of cash, whereas currently constrained and unconstrained firms do not. This calls for a narrower interpretation of this widely used cash-based constraints measure, as it may conflate distinct firm types – unconstrained and currently constrained – and fail to capture all financially constrained firms. Our findings underscore the critical role of constraint timing in shaping corporate financial behavior.
    Keywords: Financial Constraints; Artificial Intelligence; Expectations; Cash; Cash Flow; Corporate Finance Behavior
    JEL: D92 G31 G32
    Date: 2025–09–18
    URL: https://d.repec.org/n?u=RePEc:eoe:escoed:escoe-dp-2025-11
  11. By: James Chapman; Ajit Desai; Maryam Haghighi; James (Jim) C. MacGee
    Abstract: Rapid advances in artificial intelligence (AI)—including machine learning, natural language processing, and generative AI—are expanding the ability to extract meaningful insights from non-traditional data sources such as text, speeches, images, and real-time transactions, thereby strengthening policy analysis and operational decision-making. These tools also enable more sophisticated analytical approaches to the study of economic dynamics while creating opportunities to improve efficiency across institutional processes and operations. This paper documents the growing use of non-traditional data and AI at the Bank of Canada and their contribution to deeper insight and operational effectiveness. The experience highlights critical considerations for accelerating the responsible integration of AI into central banking functions, including evolving ways of working and career paths, fostering a robust ecosystem for innovation, and addressing emerging risks. A successful AI strategy must balance innovation with trust, transparency, security, reproducibility, sound model governance, data residency, and effective operational risk management.
    Keywords: Financial system; Financial stability and systemic risk; Monetary policy; Monetary policy tools and implementation; Money and payments; Payment and financial market infrastructures
    JEL: C45 C55 C88 L23 M15 O33
    Date: 2026–05
    URL: https://d.repec.org/n?u=RePEc:bca:bocsap:26-17
  12. By: Zhiyuan Wang; Qinxu Ding; Ding Ding; Siying Zhu; Jing Ren; Yue Wang; Chong Hui Tan
    Abstract: In modern financial markets, decision-makers increasingly rely on quantitative methods to navigate complex trade-offs among multiple, often conflicting objectives. This paper addresses constrained multi-objective optimization (MOO) with an application to portfolio optimization for minimizing risk and maximizing return. To address existing gaps, we propose a novel reinforcement learning (RL)-guided non-dominated sorting genetic algorithm II (NSGA-II) enhanced with gray relational coefficients (GRC), termed RL-NSGA-II-GRC, which combines an RL agent controller and GRC-based selection to improve convergence and diversity of Pareto fronts. The agent adapts evolutionary parameters online using metrics of hypervolume, feasibility, and diversity, while the GRC tournament operator ranks parents via a unified score considering dominance rank, crowding distance, and proximity to ideal reference. We evaluate the framework on the Kursawe and CONSTR benchmarks and a NASDAQ portfolio application. On the benchmarks, RL-NSGA-II-GRC achieves convergence improvements of about 5.8% and 4.4% over NSGA-II, while preserving well-distributed non-dominated solutions. In the portfolio application, it produces a smooth, densely populated efficient frontier supporting identification of the maximum Sharpe ratio portfolio (annualized Sharpe =1.92) and utility-optimal portfolios for different risk-aversion levels. The main contributions are three-fold: 1) we propose an RL-NSGA-II-GRC method integrating an RL agent into the evolutionary framework to adaptively control parameters via generational feedback; 2) we design a GRC-enhanced binary tournament operator providing a comprehensive indicator to guide the search toward the Pareto front; 3) we demonstrate, on benchmark MOO and a NASDAQ case study, that the method delivers improved convergence and well-populated frontiers supporting actionable insights.
    Date: 2026–04
    URL: https://d.repec.org/n?u=RePEc:arx:papers:2607.16194
  13. By: Yves Achdou; Johannes Brumm; Lukas Frank
    Abstract: We propose a comprehensive framework for solving overlapping-generations (OLG) models in continuous time with both idiosyncratic and aggregate risk. Our general characterization of equilibrium through the master equation operates on the joint distribution over the continuous idiosyncratic states, age and wealth. Our computational strategy is to take a finite-dimensional representation of this distribution as an input of a neural net which in turn outputs a finite-difference representation of the (conditional) value function. This idea can be applied generally to heterogeneous agent models with aggregate risk, and we call it finite-difference neural operator. Our method combines advantages from modern neural nets and traditional finite-difference methods: It is grid-free in the high-dimensional distribution, and retains control on boundary conditions in low-dimensional state variables. Moreover, our method is able to enforce shape constraints. We showcase its flexibility by solving a continuous-time OLG model with aggregate risk alone where we characterize the distribution by its supporting function; and to an OLG model with both types of risk.
    Date: 2026–08
    URL: https://d.repec.org/n?u=RePEc:arx:papers:2608.11134
  14. By: Dantong Chu; Xuefeng Gao; Yufei Zhang
    Abstract: As firms increasingly deploy machine learning for strategic decision-making, understanding algorithmic interactions has become central to operations research and economics. This paper studies learning in infinite-horizon, nonzero-sum linear-quadratic stochastic games under a radically uncoupled information structure, where players are either unaware of opponents or strategically oblivious, observing only a common state and their own action history. Under this minimal information, we analyze an asynchronous decentralized learning process in which each player independently runs a single-agent $\epsilon$-greedy iterated least-squares algorithm. We prove that, despite being unable to identify the system parameters, players' learning dynamics converge almost surely to the complete-information Nash equilibrium and characterize the convergence rate. We then apply the framework to a dynamic Cournot competition with sticky prices. Numerical experiments validate the theoretical results and show that learning under limited information reduces firm profits under both low and high price stickiness, while total surplus declines and market concentration increases when price stickiness is high. Publicly revealing aggregate market output substantially accelerates convergence and mitigates these welfare losses.
    Date: 2026–08
    URL: https://d.repec.org/n?u=RePEc:arx:papers:2608.08268
  15. By: Alberto M. G. Saruggia; Sebastien Germano
    Abstract: This study shows that textual descriptors alone can predict early-stage startup success, defined as Exit, without relying on contextual, financial, or human capital variables. Using venture capital-curated datasets covering 7, 419 startups over 20 years, the research isolates text-based framing variables and engineers 850 features through startup narrative mapping. Data subsets and vector embeddings are evaluated for statistical significance, followed by supervised machine learning experiments across six models. LightGBM achieved the highest predictive performance (F1 = 0.48), while textual descriptors alone achieved F1 = 0.30, confirming the standalone predictive value of founder narratives. Feature analysis shows that optimized densities of hyping markers, including adjectives, jargon, and buzzwords, are associated with higher Exit probability, whereas excessive statement or name length reduces it. The study also introduces a quantifiable Hyping Score for venture capital applications, demonstrating that startup framing provides measurable signals for predicting Exit under conditions of high information asymmetry.
    Date: 2026–07
    URL: https://d.repec.org/n?u=RePEc:arx:papers:2608.00045
  16. By: Jan Sp\"orer
    Abstract: Open-weight language models from Chinese AI labs caught up on benchmarks relative to proprietary frontier models in recent months. Yet their reliability on real-world financial tasks remains largely untested. We updated the Financial Touchstone benchmark, which now has 2, 967 question context-answer triplets across 495 international annual reports. We also apply a new set of models on the benchmark, expanding coverage from eleven to twenty models across ten providers, including recent open-weight models such as GLM 4.7, GLM 5, Kimi K2.6, and DeepSeek V3.2, as well as Alibaba's proprietary flagship Qwen3-Max. Anthropic's Claude Opus 4.6 achieves the highest accuracy (88.4%), while Google's Gemini 2.5 Pro maintains the lowest hallucination rate (0.08%). Notably, the open-weight Kimi K2.6 ranks third in accuracy, and the non-reasoning models GLM 5 and Mistral 3 rank fourth and fifth, challenging the assumption that reasoning architectures or proprietary weights are a prerequisite for strong financial comprehension. Information retrieval remains the primary bottleneck, accounting for 48.9% of all failures. We also document a new finding: geopolitical content filters in Chinese models refuse legitimate financial questions (0.08% of attempts), sometimes without clear reason, and the refusal behavior depends on the access route as much as on the model. The complete dataset and evaluation framework are publicly available.
    Date: 2026–08
    URL: https://d.repec.org/n?u=RePEc:arx:papers:2608.08634
  17. By: Weicheng Ye; Youran Sun; Xingyu Ren; Shunyao Yu; Chugang Yi; Haizhao Yang
    Abstract: Language models can propose many plausible trading factors, but an autonomous research system must also allocate its evaluation budget, verify its own evidence, and preserve how each candidate was produced. We present AgonAlpha, an architecture that searches over frozen research artifacts---hypotheses, executable expressions, platform evidence, rationales, and review status---rather than formulas alone. To our knowledge, AgonAlpha is the first alpha-mining system to combine verified artifact search, a fresh-context adversarial reviewer with re-execution and veto authority, and pending-aware parallel budget allocation, together with a complete public evidence trail. Independent deployments on WorldQuant BRAIN produced SPECTACULAR-grade alphas across five users and six model backends, with Fitness reaching 9.50 and Sharpe reaching 3.48, while retaining prompt-to-expression provenance for every submission.
    Date: 2026–08
    URL: https://d.repec.org/n?u=RePEc:arx:papers:2608.11250
  18. By: YILMAZ, Erdem; Kruithof, Elias H.; Verhaeghe, Pieter-Paul
    Abstract: This report documents the design, implementation, and evaluation of a fully local pipeline that codes multilingual free-text occupation responses into the International Standard Classification of Occupations (ISCO-08) and derives from those codes the social-class and status measures (e.g., ESeC, ISEI, Oesch) on which much of stratification research depends. This coding is costly, slow, and only moderately consistent when done by hand. The pipeline is local (hence GDPR-friendly), reproducible, and runs on consumer-grade hardware. We benchmark and compare three approaches to the coding step: (M1) a supervised multilingual transformer classifier; (M2) semantic retrieval against a 361k-document multilingual occupation index; and (M3) the supervised classifier with an added large-language-model reranking step. On held-out European Social Survey (ESS) data the supervised classifier reaches 59.2% exact four-digit accuracy, comparable to human inter-coder agreement for this intrinsically ambiguous task, which is itself low at the four-digit level (reported by past research to be ≈42–53% between individual coders; κ ≈ 0.5 between agencies). More importantly for stratification research, given that coding errors are mostly near-misses that coarser schemes absorb, the sociologically relevant aggregate measures are recovered far more accurately, namely 80.4% for the 9-class ESeC (87.8% for its 5-class collapse), 80.5% for the 7-group ESeG, 77.9% for the 16-category Oesch scheme, and a Pearson correlation of 0.87 for the ISEI socio-economic index. The same class- and status-level accuracy holds when the model is trained and tested on canonical occupation text, so it is not an artefact of the ESS sample. Retrieval alone (M2) is substantially weaker, and contrary to the intuitive expectation, adding an LLM reranker (M3) reduces accuracy on noisy multilingual survey responses (−3.7 pp), while on clean canonical titles its effect is approximately neutral, i.e., +0.5 pp over the backbone it reranks, with no gain over the best supervised model. Instead, the supervised model’s (M1) well-calibrated confidence (ECE = 0.038) supports a human-in-the-loop workflow in which only genuinely uncertain cases are escalated to human coders. We report bootstrap confidence intervals and paired significance tests for every comparison, a train/test leakage audit, and a decomposition of geographic variation showing that the cross-country accuracy spread (in the ESS corpus) reflects response style and occupational-title ambiguity rather than language resources, whereas a genuine resource effect appears only across languages on canonical text in the normative corpus. The pipeline runs entirely offline on a single consumer-grade workstation, making automated, auditable occupation coding feasible without sending respondent data to external services.
    Date: 2026–07–25
    URL: https://d.repec.org/n?u=RePEc:osf:socarx:xv2gf_v1
  19. By: Chen Liang; Fasheng Xu
    Abstract: As LLM agents move from decision support to autonomous procurement, firms need to know whether delegated negotiators create value, divide it predictably, and avoid money-losing contracts. We study this in a canonical supply chain bargaining problem: a buyer with private demand information negotiates a quantity-payment contract with an uninformed seller. We benchmark nine LLMs from OpenAI, Google, and Alibaba against a validated Perfect Bayesian Equilibrium across 9, 840 LLM-to-LLM negotiations. First, capability governs value creation. Agents agree in 98.9% of negotiations and capture 95.4% of first-best surplus undiscounted, but average 2.98 rounds against the benchmark's 1.25, and this delay erodes 21-34% of surplus. Capability also governs reliability: baseline models accept individually irrational contracts in 19.2% of cases, versus 0.0-0.6% at mid-tier and flagship, making automated profit verification the binding guardrail below that threshold. Second, surplus capture is relational. Provider identity predicts who captures surplus better than capability rank: self-play buyer shares average 40% for OpenAI, 50% for Google, and 70% for Alibaba's Qwen, an ordering that survives restricted communication and no discounting. Reversing which provider sells moves the division by 7-18 percentage points, and the capable Qwen flagship is the weakest cross-family seller: vendor choice is a first-order distributional decision. Third, the prompt is a strategic lever. Delegation separates the principal's economic patience from the agent's prompted strategic patience, a free deployment choice that is the single strongest driver of surplus division (90% of explained variance). Together these establish an equilibrium-referenced audit of AI agents along three dimensions: discounted efficiency, distributional profile, and operational reliability.
    Date: 2026–07
    URL: https://d.repec.org/n?u=RePEc:arx:papers:2608.07538
  20. By: Joy Buchanan; Joshua Foster
    Abstract: Language models increasingly settle real resource tradeoffs on behalf of principals yet their economic preferences remain unobserved. We demonstrate their generation rule is isomorphic to the random utility model of discrete choice. This allows internal logit scores to structurally identify preferences. Estimating risk attitudes across twelve models in a portfolio task reveals universal but heterogeneous risk aversion. Although models reject strictly dominated options, their elicited preferences fail invariance tests and violate the independence of irrelevant alternatives across varying experimental prompts. Finally, fine tuning establishes that a principal can explicitly engineer a target risk attitude.
    Date: 2026–07
    URL: https://d.repec.org/n?u=RePEc:arx:papers:2607.26288
  21. By: Mojtaba Eslami
    Abstract: Modern agentic AI systems combine multiple large language model agents with heterogeneous skills, yet most architectures either fix communication in advance or allow full broadcast. Both can be inefficient because token cost, latency, redundancy, and error propagation increase with the number of active agents and communication links. We model agent selection and communication as a cooperative game with task-conditioned net utility $U(C\mid x)=V(C\mid x)-\sum_{i\in C}c_i$, separating coalition-level costs from agent activation costs. We propose a marginal-value activation rule and greedy router, extend the model to optimize communication edges with per-edge costs, and use estimated Shapley values to predict which agents are worth contacting before and during execution. We connect the problem to submodular maximization and prove two limited guarantees: a curvature-refined bound for a monotone, cardinality-constrained special case, and a tight $1/2$-approximation, with a correction for signed objectives, for an unconstrained non-monotone case via double greedy. Neither guarantee applies directly to the main router, which remains a heuristic. We also prove a Shapley-submodularity sandwich bound linking the error of marginal-value routing to a per-agent diminishing-returns quantity. In synthetic experiments, greedy routing achieves $99.5%$ of brute-force-optimal utility while activating $1.96$ of $8$ agents on average, compared with $38.8%$ for full broadcast. Performance is robust to activation cost and redundancy weight but falls to $66%$ under strong violations of submodularity or noisy value estimates. We distinguish the framework from Shapley pricing, hedonic coalition formation, and communication-graph pruning, and propose evaluation on real multi-agent LLM benchmarks.
    Date: 2026–07
    URL: https://d.repec.org/n?u=RePEc:arx:papers:2608.07532
  22. By: Ekkehardt Bauer; Dirk Holl\"ander; David Scholz; Linus Wolff; Christoph Ostermair; Kyrillus Aiad; Joachim Hasebrook
    Abstract: This study focuses on developing an AI-supported prototype for multiperspective interest rate forecasting that combines classical econometric models with modern artificial intel-ligence methods. Tested in a major European bank, the system enables more precise and flexible prediction of interest rate developments, supporting strategic decision-making in Asset-Liability Management (ALM). It integrates topic modeling, sentiment analysis, econometric forecasting, and market-based analyses within an interactive platform. Leveraging AI to analyze large volumes of financial documents and market data enables the identification of monetary policy trends and sentiment signals at an early stage. The core econometric model is a Bayesian vector autoregression (BVAR) that enables simulation-based scenario analyses to evaluate economic developments from multiple perspectives. The system's innovation lies in its integration of several forecasting approaches that consolidate previously separate information sources and present them transparently and interpretably. Financial analysts and risk managers thus gain a better basis for making decisions, allowing them to assess interest rate risks more accurately and manage market movements more proactively. While the prototype demonstrates how AI can transform interest rate management in banking, further development is required to optimize real-time data integration and regulatory compliance. Even at this stage, the study shows that multi-perspective, AI-driven forecasting provides substantial added value for banks by increasing transparency, strengthening evidence-based decision-making, and improving risk management.
    Date: 2026–08
    URL: https://d.repec.org/n?u=RePEc:arx:papers:2608.12424
  23. By: Bruno Feunou; Jean-Sébastien Fontaine; Rishi Vala
    Abstract: This paper introduces CLONE (Classification Of News), a method that decomposes asset price movements into four types of macroeconomic news—aggregate demand, productivity, inflation, and monetary policy—based on joint changes in prices of stocks, bonds, and inflation swaps. CLONE’s simplicity and forward-looking focus enable the identification of real-time economic signals that are critical for understanding market behavior and guiding policy decisions. We show that from 2004 to 2024 aggregate demand news historically dominated daily variation in asset prices, while inflation and monetary policy news have gained importance since 2021. We validate our method against sign-restricted VAR models and apply it to major U.S. macroeconomic data releases, providing insights into how market participants interpret and react to forward-looking information. We discuss several benefits of our approach relative to the standard sign restriction method.
    Keywords: Models and tools; Econometric, statistical and computational methods; Monetary policy; Monetary policy framework and transmission
    JEL: E32 E44 G12 G14
    Date: 2026–03
    URL: https://d.repec.org/n?u=RePEc:bca:bocsap:26-7
  24. By: Wenzel, Julia
    Abstract: In recent years, Robotic Mobile Fulfillment System (RMFS) as part of a third generation of automated parts-to-picker systems has emerged. These systems were developed to meet the growing demands for flexibility and scalability in modern warehouses, which predominantly handle e-commerce orders. By employing mobile robots, RMFSs efficiently process smaller, customized orders, manage a wide range of storage items, and handle the increasing order lines in e-commerce. Despite their advantages, RMFSs have a notable drawback compared to second-generation order picking systems, such as automated storage and retrieval systems, in terms of space utilization. This deficit is crucial for logistics managers and often limits the practical adoption of RMFSs. The use of multi-level RMFSs, e.g., as mezzanine structures, offers the potential to improve space efficiency. However, this approach remains unexplored and is only partially utilized. Multi-level RMFSs pose additional challenges for logistics managers, resulting in decision problems at strategic, tactical, and operational levels. These decision problems include designing the layout of each level, allocating items across multiple levels, and assigning pods and robots within each level. As a result, planning and operating such systems is significantly more complex compared to single-level warehouses. This thesis examines these decision problems and the performance of multi-level RMFSs through three research questions. The resulting studies address, first, optimal algorithms for assigning orders and pods in multi-level RMFSs (Publication 1); second, the analysis of interactions between individual system parameters of a multi-level RMFS and their impact on system performance (Publication 2); and third, the comparison of multi-level RMFSs with traditional order picking systems (Publication 3). The findings provide valuable insights for optimizing RMFSs in dynamic warehouse environments for e-commerce and establish a foundation for future developments in automated warehouse logistics. In order to do this, first a heuristic is created to solve the combined planning problem. Next, key system factors that greatly improve performance and the best design of an RMFS are identified. Finally, the performance and cost-effectiveness of RMFS are compared to traditional order picking systems using numbers.
    Date: 2026–06–10
    URL: https://d.repec.org/n?u=RePEc:dar:wpaper:161017
  25. By: Miles, David; Sefton, James
    Abstract: We use stochastic simulations to analyse the probability distribution of outcomes for the UK University pension scheme (the USS) in the light of conflicting claims about its sustainability. We use the results to draw wider conclusions about the nature of defined benefit (DB) pension schemes and whether they bring benefits to members based on risk sharing. We find that a substantial investment in riskier assets (equities) makes the average outcome one in which the scheme is comfortably able to pay accrued benefits. But the risk of having far fewer funds than needed to pay existing pension promises is significant and the chances of large deficits is substantial. The ambiguity about how pension fund surpluses or deficits would be allocated between scheme members and the scheme sponsor (for the USS that is Universities) means that agreement on the optimal portfolio allocation for the scheme's funds is not likely. Among scheme members of different ages and different attitudes towards risk agreement on what are acceptable trade-offs between risk and return on assets is unlikely. We contrast this with the position for defined contribution pensions.
    Keywords: Pensions
    JEL: G11 G50 G22
    Date: 2024–07
    URL: https://d.repec.org/n?u=RePEc:cpr:ceprdp:19254
  26. By: Zhipeng Huang; Cornelis W. Oosterlee
    Abstract: We develop an analytic Fourier cosine (COS) method for the valuation of compound options. By deriving closed-form expressions for the cosine coefficients at all compound stages, the proposed method eliminates the need for numerical quadrature in intermediate exercise stages while retaining the convergence properties of the underlying COS approximation. The formulation extends to multi-stage compound structures and a broader class of payoffs, and remains applicable to a wide class of stochastic models characterized by known characteristic functions, including jump-diffusion dynamics. Numerical experiments demonstrate improved computational efficiency compared with quadrature-based implementations while maintaining high accuracy. Applications to staged real-option problems further illustrate the flexibility of the method in handling nested decision structures under different uncertainty dynamics.
    Date: 2026–07
    URL: https://d.repec.org/n?u=RePEc:arx:papers:2607.25599

This nep-cmp issue is ©2026 by Stan Miles. It is provided as is without any express or implied warranty. It may be freely redistributed in whole or in part for any purpose. If distributed in part, please include this notice.
General information on the NEP project can be found at https://nep.repec.org. For comments please write to the director of NEP, Marco Novarese at <director@nep.repec.org>. Put “NEP” in the subject, otherwise your mail may be rejected.
NEP’s infrastructure is sponsored by the Griffith Business School of Griffith University in Australia.