nep-ain New Economics Papers
on Artificial Intelligence
Issue of 2026–09–07
twenty-one papers chosen by
Ben Greiner, Wirtschaftsuniversität Wien


  1. AI worsens climate change, integrated assessment shows By Huiying Ye; Richard S. J. Tol; Fangzhi Wang
  2. Training AI For When Humans Will Use It By Kevin A. Bryan; Joshua S. Gans
  3. How AI Prompts Can Teach Us About the Structure of Human Behavior By Matthew O. Jackson; Benjamin S. Manning; Yutong Xie; Walter Yuan; Qiaozhu Mei
  4. Governing Delegation to Generative Artificial Intelligence: Human Direction, Work-Related Orientation, and Modes of Use By Jorge F\'abrega
  5. Optimal Liability Design for Medical AI By Rui Mao; Tingliang Huang; Houcai Shen
  6. Staged Access and Liability for Dual-Use Artificial Intelligence By Joshua S. Gans
  7. AI Agents and Prompt Engineering in Econometric Coding By Sebastian Galiani; Federico Ariel López; Raul A. Sosa
  8. Aging Economies and AI Adoption: Firm-Level Evidence from the World Bank Enterprise Surveys By Ha Minh Nguyen
  9. AI Adoption, Regional Productivity, and Inflation Evidence from Korea and Implications for Monetary Policy By Cyn-Young Park; Kwanho Shin
  10. Artificial Intelligence and Labor Market Adjustment in Türkiye : Evidence from LinkedIn Data By Fatima, Freeha; Ozen, Efsan Nas; Raju, Dhushyanth
  11. Stranded credentials: how a skill-signaling market absorbed generative AI By Song Yao
  12. CentaurBench: Benchmarking LLM Capabilities on Augmenting vs. Automating Real-World Work Tasks By Pattaraphon Kenny Wongchamcharoen; Kris Gulati; Min Min Fong; Abhishek Nagaraj
  13. Monitor Digital Working Society – Brief report on the fourth wave of surveys: AI is becoming increasingly embedded in everyday work: Rising usage, greater degrees of autonomy and growing concerns. By Marcinkowski, Frank; Keller, Birte; Lünich, Marco; Flaßhoff, Florian Golo
  14. Governments as Adopters and Regulators of AI: A Challenge for Democracy? By Trein, Philipp; Maggetti, Martino
  15. Financial advice behaviour: humans versus AI By Ylva Baeckström; Roman Matkovskyy
  16. Communicating Credit Risk with Large Language Models: Evaluation of Explanations from Standard and Alternative Data-Based Models By Sahab Zandi; Noah Kostesku; Christophe Mues; Mar\'ia \'Oskarsd\'ottir; Cristi\'an Bravo
  17. Benchmark Mineability and the Financing of AI Innovation By Alex Chan
  18. From Financial Sentiment Classification to Return Predictability: A QLoRA Benchmark of Large Language Models By Fusheng Luo
  19. Inference with AI-Generated Covariates By Junting Duan; Markus Pelger
  20. The Measurement Revolution? Credible Measurement and Inference in the Age of AI By Melissa Dell; Ashesh Rambachan
  21. Designing AI-Augmented Peer Review By Joshua S. Gans

  1. By: Huiying Ye; Richard S. J. Tol; Fangzhi Wang
    Abstract: Artificial intelligence (AI) interacts with climate in various ways, while a unified analytical framework of this intricate interplay is lacking. To align AI investment with climate policy, we propose such a framework integrating AI's impact on emissions, output, and climate damages into the DICE model. We distinguish between ICT-like and Industrial Revolution (IR)-like AI prospects. Calibrated to the best available evidence, we find that AI development is net polluting. Under current low abatement, ICT-like AI adds 0.1 degree C to 2100 warming, while IR-like AI adds 0.8 degree C. The associated climate costs offset roughly one-fifth and one-quarter of AI's economic gains, respectively. Meeting the 2 degree C target saves the optimal ICT(IR)-like AI investment rate by 2100 from 3.3% (5.1%) under the low-abatement scenario to 3.7% (12.7%), indicating that mitigation is complementary to AI development. We further show that the investment trade-off between AI and abatement is driven primarily by AI's economic prospects, not by its emissions footprint.
    Date: 2026–08
    URL: https://d.repec.org/n?u=RePEc:arx:papers:2608.24670
  2. By: Kevin A. Bryan; Joshua S. Gans
    Abstract: AI predicts; humans use its predictions to make decisions. These predictions are combined with human verification and analysis, queries to other statistical models, and so on. The economic value of an AI, therefore, depends on how it interacts with the surrounding decision environment. We describe the value of AI as part of this ``composite experiment'' where AI makes a coarse prediction of the state of the world, show what this means for optimal model training via a geometric argument, explain why optimal training can be discontinuous in economic variables, and study how heterogeneous users or monopoly model trainers affect these results. In particular, maximizing the unconditional accuracy of AI predictions is generally suboptimal.
    Date: 2026–08
    URL: https://d.repec.org/n?u=RePEc:arx:papers:2608.12538
  3. By: Matthew O. Jackson; Benjamin S. Manning; Yutong Xie; Walter Yuan; Qiaozhu Mei
    Abstract: We introduce a general, easy-to-implement AI-based method for studying the structure and complexity of human behavior. We assign a large language model a ``type vector'' and then prompt it to choose actions across settings in which we observe human choices. For instance, the type vector (2, 4) becomes ``You are a player characterized by the following profile: 2 out of 5 in Altruism, 4 out of 5 in Risk Aversion, '' after which it is prompted to make choices. We vary the dimensions (e.g., Altruism, Fairness, Trust, $\dots$) and values (e.g., 1--5) to minimize distance to human choices. Applying the method to 119, 147 decisions made by 78, 657 subjects from more than 35 countries across 10 classic economic game roles, we find that human behavior can be closely matched using three dimensions: Risk Aversion, Strategic Sophistication, and Trust. Moreover, the types needed to fit individuals across games cluster into fewer than a dozen groups, and can predict behavior in held-out games with different rules and available actions. The results suggest that behavior across diverse settings can be approximated by a low-dimensional, portable representation, supporting the possibility of general yet parsimonious theories across the behavioral sciences. More broadly, the method can provide insights into the structure of many human behaviors.
    Date: 2026–08
    URL: https://d.repec.org/n?u=RePEc:arx:papers:2608.18265
  4. By: Jorge F\'abrega
    Abstract: Delegating cognitive operations to generative artificial intelligence redistributes execution and raises a governance problem: where human direction of the task remains. We distinguish two routes. Specified delegation places that direction before execution, through instructions, constraints, or criteria that delimit the task. Iterative coproduction places it during production, through interventions that correct or redirect provisional outputs. To examine both routes, we use aggregate monthly cells from the Anthropic Economic Index for April and May 2026. The AEI distinguishes two modes of use: 1P API, which corresponds to direct traffic through Anthropic's API, and Claude.ai, which combines activity from Chat and Cowork. On this basis, we test whether a stronger work-related orientation of human-AI interaction is associated with more specified delegation within each mode and whether the increase in the iterative profile is greater in Claude.ai than in 1P API. The main analysis uses level-0 O*NET tasks and estimates how both profiles change when an eligible record reallocates ten percentage points from personal use to work-related use. The iterative comparison is restricted to 1, 411 node-month pairs observed and eligible in both modes. Specified delegation increases by 2.76 points in 1P API (95% CI: [2.30, 3.22]) and by 1.45 in Claude.ai (95% CI: [0.93, 1.97]). On the common support, iterative coproduction changes by-0.30 points in 1P API and by 0.15 in Claude.ai, yielding a between-mode difference of 0.45 points (95% CI: [0.15, 0.75]). These findings show that work-related orien tation is associated with stronger traces of prior human direction and that the observable iterative response varies across modes of use. The article shifts attention from how much the AI executes to when human direction leaves observable traces.
    Date: 2026–08
    URL: https://d.repec.org/n?u=RePEc:arx:papers:2608.17624
  5. By: Rui Mao; Tingliang Huang; Houcai Shen
    Abstract: Artificial intelligence (AI) is increasingly integrated into medical decision-making, yet its liability implications remain complex, particularly when physicians differ in diagnostic skills and their quality is unobservable. This paper develops a principal-agent model in which a social planner designs medical liability to regulate a physician with private quality information who chooses between a standard treatment, a personalized judgment-based treatment, or following an imperfect AI recommendation. Our analysis yields several novel insights. First, we show that the optimal mechanism under asymmetric information is surprisingly simple: a uniform, one-size-fits-all liability level for all physician types who deviate from the standard of care. Despite physician heterogeneity, this simple policy often achieves the full-information first-best outcome, particularly when standard care is reliable or AI is highly accurate. Second, the relationship between AI accuracy and optimal liability is non-monotonic. Contrary to common intuition, better AI does not always imply more relaxed liability. As AI accuracy increases, the optimal liability either decreases monotonically or follows an inverted-U pattern, depending on the uncertainty of the standard treatment. Third, asymmetric information does not universally reduce social welfare. Welfare loss arises only when standard care is unreliable and AI accuracy is too low; even then, its magnitude follows an inverted U-shape, initially increasing as AI complicates the regulatory problem, but declining as more accurate AI helps mitigate it. Finally, we find that information asymmetry is a double-edged sword in the presence of AI, and greater transparency does not benefit all stakeholders equally.
    Date: 2026–08
    URL: https://d.repec.org/n?u=RePEc:arx:papers:2608.03114
  6. By: Joshua S. Gans
    Abstract: How should regulators combine staged access to a dual-use AI model with developer liability? I model a defender and an adversary searching for the same software flaws. After release, liability induces more defensive search but also makes the adversary search harder, limiting its effect on harm. Exclusive access removes this strategic response, so staged access and liability are complements in protection. Defensive effort rises towards the release date, making additional delay progressively less productive. Optimal evaluation windows are therefore bounded, and their response to greater harm saturates. Because defensive search uses real resources, the efficient liability rate can be below full internalisation of harm. That rate generally cannot induce the developer to choose the regulator's preferred release date, leaving a distinct role for a timing mandate.
    JEL: D74 K13 L51 O33
    Date: 2026–08
    URL: https://d.repec.org/n?u=RePEc:nbr:nberwo:35586
  7. By: Sebastian Galiani; Federico Ariel López; Raul A. Sosa
    Abstract: We study how large language models write code for econometric analysis. We compare three dimensions of AI-assisted coding: statistical software (Stata, R, or Python), prompting (zero-shot versus few-shot), and the degree of agency, from a chatbot that writes a single script to an agent that executes and revises its own code. On a benchmark of applied econometric and statistical tasks, moving from the chatbot to the constrained agent raises task success from 74 to 96 percent, at about eight additional cents per run. For Claude Sonnet 4.6 and GPT-5.4 through Codex, few-shot prompting improves the chatbot far more than the constrained agent, indicating that prompting and agency act as substitutes. For these models, differences across statistical software are sizeable under the chatbot but largely disappear under the constrained agent.
    JEL: C18 C87
    Date: 2026–08
    URL: https://d.repec.org/n?u=RePEc:nbr:nberwo:35588
  8. By: Ha Minh Nguyen
    Abstract: How do demographic trends shape the adoption of AI and automation technologies? This paper provides the first large-scale cross-country firm-level test of the demographic–automation hypothesis using World Bank Enterprise Surveys data covering 89, 380 firms across 144 countries from 2022 to 2025. I classify adopters by applying a large language model to firms’ open-ended process innovation descriptions, identifying 1, 656 AI and automation adopters (1.9 percent of the sample). A ten-percentage-point increase in the old-age dependency ratio raises process adoption probability by approximately 0.6 percentage points, after accounting for countries’ income levels, digital infrastructure, firm size and sector, and broad regional and time differences. The result is robust across specifications and supported by an instrumental variable strategy based on predetermined demographic cohort structure. Heterogeneity analysis shows the effect concentrates in manufacturing, large firms, and developing economies for the broad adoption measure; restricting to firms with explicit references to AI reverses the sector pattern, with services firms significantly more likely to adopt than manufacturing firms, pointing to distinct sectoral profiles for software-based AI and hardware-based automation. Aging also predicts firms’ development of AI-enabled products across both manufacturing and services. The results indicate that demographic aging shapes AI and automation adoption through both process and product innovation channels: firms substitute technology for increasingly scarce and costly labor in production, and separately develop AI-enabled products for labor-constrained customers.
    Keywords: Aging; automation; artificial intelligence; firm-level; technology adoption; labor-saving technology; World Bank Enterprise Surveys; demographics; labor substitution
    Date: 2026–08–21
    URL: https://d.repec.org/n?u=RePEc:imf:imfwpa:2026/176
  9. By: Cyn-Young Park (The South East Asian Central Banks (SEACEN) Research and Training Centre); Kwanho Shin (Korea University)
    Abstract: This paper examines whether the early diffusion of artificial intelligence (AI) is visible in productivity and price outcomes relevant to monetary policy. We combine firm-level information on AI adoption from Korea’s Survey of Business Activities with annual industry- and region-level data for 2017–2023. We construct value-added- and employment-weighted measures of AI intensity and use their 2019 values as predetermined measures of initial AI intensity. Both measures strongly predict the cross-sectional distribution of AI intensity in 2023. We then estimate reduced-form panel regressions that compare 2023 outcomes across industries and regions with different initial levels of AI intensity, controlling for unit and year fixed effects. We find no systematic evidence that more AI-intensive industries or regions experienced stronger output or labour-productivity growth in 2023. Industry-level price effects are also statistically insignificant and vary across price measures. At the regional level, however, employment-weighted AI intensity is positively associated with overall consumer price inflation, while restaurant price inflation is higher under both measures of AI intensity. These findings suggest that the supply-side benefits of AI had not yet become visible in aggregate productivity by 2023, whereas inflationary pressures may have emerged in some locally determined consumer services. This pattern is consistent with demand responding before productivity gains are fully realised, although our empirical design does not identify the underlying mechanism. The findings have important implications for monetary policy: during the early stages of AI diffusion, central banks should not assume that anticipated productivity gains will immediately expand effective supply or alleviate inflationary pressures. We discuss the implications of this transitional asymmetry for central banks in Asian economies, where rapid AI adoption may coincide with persistent supply constraints and sector-specific price pressures.
    Keywords: Artificial Intelligence (AI), AI Adoption, Productivity Growth, Inflation, Monetary Policy and Central Banking
    JEL: E31 E52 O33 O47
    Date: 2026–09
    URL: https://d.repec.org/n?u=RePEc:sea:wpaper:wp63
  10. By: Fatima, Freeha; Ozen, Efsan Nas; Raju, Dhushyanth
    Abstract: This paper examines how artificial intelligence (AI) is reshaping Türkiye’s labor market by documenting patterns in skill supply, employer demand, and labor market adjustment using high-frequency digital labor market indicators from LinkedIn. The analysis focuses on the mechanisms through which AI-related change is associated with shifts in skills, hiring, occupational mobility, exposure to generative AI, and international migration. The evidence shows a relatively broad presence of foundational digital and AI literacy skills across sectors and demographic groups, alongside a persistent and increasing concentration of advanced AI engineering talent within a narrow set of occupations and industries. Measured skill penetration follows non-monotonic patterns over time, while frontier AI talent accumulates steadily, indicating a divergence between the breadth and depth of AI capability. Entry into AI roles often follows strongly path-dependent pathways, and employer demand signals for technical and AI-adjacent capabilities are only partially reflected in realized hiring, with no sustained positive divergence in AI-related hiring relative to overall labor demand. Potential exposure to generative AI varies systematically across sectors and demographic groups, with the balance between task augmentation and disruption differing across sectors rather than uniformly favoring one over the other. International migration emerges as a salient adjustment margin for highly specialized AI talent, operating alongside domestic reallocation mechanisms and influencing the availability of frontier skills within the domestic labor market. These patterns indicate that the central challenge associated with AI in Türkiye’s labor market lies not in whether AI-related capabilities will spread, but in how reallocation unfolds across skills, occupations, and workers over time. The findings highlight the role of skill formation systems, hiring and credentialing practices, occupational structures, and cross-border mobility in shaping the trajectory of labor market adju stment. The analysis also illustrates how digital labor market data can complement traditional sources by providing timely evidence on emerging skills, evolving demand, and early adjustment dynamics in middle-income economies navigating the AI transition.
    Date: 2026–04–01
    URL: https://d.repec.org/n?u=RePEc:wbk:hdnspu:209921
  11. By: Song Yao
    Abstract: Generative AI can now perform many tasks that credentialing institutions count on to assess skill. During the AI era, do credentials retain their signaling value for subsequent performance? Mostly, yes. We audit the 2010-2026 archive of Kaggle, the largest data science competition platform, which ran two evaluation formats concurrently: upload-competitions, which directly score entrants' predictions computed on published data, and code-competitions, which score predictions by executing entrants' code on hidden data. Across 444, 698 participations, competition medals predict subsequent leaderboard performance almost entirely in the first year after being earned, in both formats. Fresh medals retained most of their signaling value through the AI transition; credential stocks are only as informative as their replenishment. Although upload-competition medal stocks lost 82% of their informativeness, institutional stranding explains half to three quarters of the loss: upload-competitions had exited for reasons predating AI, and their frozen medal stock aged out under the pre-existing decay pattern. Old upload-competition medals look more valuable only in isolation, by proxying for the rest of the holder's record (e.g., experience). The measured changes are institutional rather than personal: an AI-like working style predicts performance similarly in both formats. The platform's official credential tiers, based on lifetime medal counts, discard 13-16% of the medals' information; an index weighting recent medals more heavily, built on pre-AI-era data alone, outperforms the official tiers in predicting AI-era performance. In conclusion, credentials are informative, perishable, institution-bound, and interdependent; sustaining their value under AI is a high-stakes, socio-economic problem of institutional design.
    Date: 2026–08
    URL: https://d.repec.org/n?u=RePEc:arx:papers:2608.17111
  12. By: Pattaraphon Kenny Wongchamcharoen; Kris Gulati; Min Min Fong; Abhishek Nagaraj
    Abstract: Most LLM benchmarks rank models on their ability to automate work tasks. In practice, however, models are often used to assist other (human or LLM) agents. The question that drives model selection is therefore not only which model produces the best output, but which model most improves the work of another (weaker) agent. We introduce a unified framework that evaluates the capability of models to automate and augment another agent's performance. Across seven economically grounded real-world tasks, an assistant model writes assistance text for a standardized lower-capacity worker model, which produces the deliverable. In automation mode, the assistant produces the output directly. Outputs are scored through blind pairwise comparisons by an LLM judge panel with task-specific rubrics, replicated across ten runs. Rankings across the two regimes are only modestly correlated, and the automation winner loses augmentation on five of seven tasks. Assistance is not reliably positive. The unaided worker outranks every assisted condition on three tasks, and only one model's guidance beats no guidance on average. These results suggest that automation ability is an incomplete proxy for assistance quality, motivating benchmarks that evaluate models according to the roles they play in human-AI and multi-agent systems.
    Date: 2026–08
    URL: https://d.repec.org/n?u=RePEc:arx:papers:2608.18554
  13. By: Marcinkowski, Frank; Keller, Birte; Lünich, Marco; Flaßhoff, Florian Golo
    Abstract: This brief report presents key findings from the last of four representative survey waves of the Monitor Digital Working Society, conducted as part of the research project Opinion Monitor Artificial Intelligence 3.0. The analysis is based on responses from 1, 607 participants surveyed in July 2026, including 915 employed and 692 non-employed respondents. At the end of the survey series, interest in AI has returned to the elevated level observed in summer 2025 following an interim decline, while respondents’ self-assessed knowledge remains stable at a higher level than in the first wave. At the same time, assessments of the societal consequences of AI have become more critical: fewer respondents believe that the benefits outweigh the risks, while risk-oriented and ambivalent assessments have become more prevalent. In the workplace, the risk–benefit assessment of AI for respondents’ own occupations remains stable and slightly benefit-oriented. Evaluations of specific working conditions have become somewhat more positive, while personal job insecurity has increased noticeably. Most strikingly, occupational AI use has risen to the highest level recorded across the survey series. Growth is particularly pronounced for generative, recommendation-based, analytical, and decision-related functions. Employees also report greater freedom of use, decision-making authority, and opportunities to participate in shaping AI use, although a majority still have no influence over the introduction of AI in their workplace. Overall, the findings point to an ongoing normalization of AI: as its use becomes more widespread, both its concrete benefits and associated scope for action, as well as potential risks for employment and society, are becoming more visible and are assessed in increasingly differentiated ways.
    Date: 2026–08–13
    URL: https://d.repec.org/n?u=RePEc:osf:socarx:3byft_v1
  14. By: Trein, Philipp (University of Lausanne); Maggetti, Martino
    Abstract: This paper explores the dual role of governments as both adopters and regulators of artificial intelligence (AI), focusing on the implications for democratic governance. As AI technologies become increasingly embedded in public administration – from policymaking to service delivery – they offer opportunities for improving efficiency and personalization but also raise concerns about transparency, accountability, and fairness. To deal with these questions, this paper examines how AI is framed as a policy problem, the regulatory approaches adopted in different political systems, and the politicization of AI governance. It also analyzes the democratic risks posed by algorithmic decision-making, polarization, and corporate concentration of power, while highlighting the potential of AI to enhance democratic quality through improved public services, inclusive discourse, and citizen engagement. The paper concludes by identifying key areas for future research, including legitimacy, trust, regulatory design, and equity in AI governance.
    Date: 2026–08–15
    URL: https://d.repec.org/n?u=RePEc:osf:socarx:rcqbt_v1
  15. By: Ylva Baeckström; Roman Matkovskyy (Rennes SB - Rennes School of Business)
    Abstract: Financial advice can attenuate underinvestment but is costly, biased, and skewed towards the wealthy. AI-powered co-advisors could help deliver more scalable and affordable advice. To understand how, our vignette-based survey experiment compares the portfolio recommendations made by professional human advisors with GenAI large language models (LLMs) under biased and unbiased prompts. We document human financial advice projection whereby human advisors strongly project their own portfolios onto their clients. AI financial advice projection is prompt and model family dependent: ChatGPT is the least biased, while strong Gemini-Biased projection collapses when removing advisor demographics. LLMs are systematically more conservative than professional human advisors, recommending portfolios with lower Sharpe ratios that deliver up to 18% lower 20-year terminal wealth. However, human advisory fees erode much of this excess gain, with a 20-year breakeven fee of 1.03% p.a. Our results have direct implications for financial regulators, the advice profession, and LLM developers seeking to deploy AI-generated financial advice.
    Keywords: Financial advice, Large language models, Artificial intelligence, Portfolio asset allocation
    Date: 2026–09
    URL: https://d.repec.org/n?u=RePEc:hal:journl:hal-05725514
  16. By: Sahab Zandi; Noah Kostesku; Christophe Mues; Mar\'ia \'Oskarsd\'ottir; Cristi\'an Bravo
    Abstract: Credit decisioning is a high-stakes task in which model outputs must be accurate and explainable to support compliant decisions. Although modern credit risk models such as eXtreme Gradient Boosting (XGBoost) and Graph Neural Networks (GNNs) improve predictive performance, their explanations are often too technical for stakeholders creating communication gaps that can shape approvals, denials, and fairness judgments. We examine whether Large Language Models (LLMs) can serve as explanation layers that translate post-hoc explanation artefacts into stakeholder-appropriate risk narratives. Using Freddie Mac single-family loan-level data, we develop three pipelines: standard tabular (XGBoost + SHAP), and two with alternative data, a pure network-based (GNN + GNNExplainer), and a bimodal one (combining tabular and network data). We generate narratives with three LLM configurations: a small fine-tuned LLM (Gemma 3 4B), a large fine-tuned LLM (DeepSeek R1 70B), and a zero-shot commercial LLM (Gemini 2.5). Explanation quality is evaluated through automated checks across all pipelines and a human study of bimodal explanations comparing credit risk professionals and non-professionals on eight decision-relevant dimensions. We have three main findings. First, the pipeline accounts for higher variance in evidence-grounding scores than the language model, meaning that the binding constraint on explanation quality is the evidence representation, not the model used. Second, the explanation narratives reliably name the influential factors but are less reliable when stating the direction of influence, which may be consequential for adverse-action communication. Finally, professionals apply stricter evidentiary standards than non-professionals. We discuss implications for the governance of risk models, including deployment considerations and the value of domain-aligned LLMs in regulated credit settings.
    Date: 2026–08
    URL: https://d.repec.org/n?u=RePEc:arx:papers:2608.17715
  17. By: Alex Chan
    Abstract: Public AI benchmarks steer research and allocate investments. They are therefore market designs. Public examples can reveal the process behind a private final test, while a finite public score cannot cover a broad task space inherent to general intelligence. I show how both gaps become profitable when scores move capital and how targeted effort erodes the signal used by later investors. The market design lesson is to separate development from certification: publish practice tasks, but choose the investment-consequential generator after the submitted system's evaluation policy is fixed.
    JEL: D40 D47 D82 D83 D86 O3 O32
    Date: 2026–08
    URL: https://d.repec.org/n?u=RePEc:nbr:nberwo:35641
  18. By: Fusheng Luo
    Abstract: Financial sentiment classifiers are commonly evaluated against human labels, but strong linguistic performance does not necessarily imply economically useful return predictability. This study separates these questions through two experiments. First, we construct a unified three-class benchmark from five financial text datasets and compare TF--IDF Naive Bayes, off-the-shelf FinBERT and Financial-RoBERTa encoders, zero-shot Qwen2.5-7B, and QLoRA-adapted Qwen2.5-7B, LLaMA3-8B, and Mistral-7B models. Mistral-7B achieves the best test accuracy (0.8840) and macro-F1 (0.8771), while QLoRA raises Qwen2.5's macro-F1 from 0.7274 to 0.8615. An inverse-frequency class-weighted loss does not improve Qwen2.5. Second, we evaluate economic validity on a temporally separate 2019 Benzinga sample containing 10, 637 unique headlines and 13, 115 headline--stock observations for a fixed S\&P~100 universe. Model probabilities are converted into continuous sentiment scores, aggregated by stock and signal date, and aligned with next-session returns over one-, two-, three-, and five-day horizons. All seven downstream models produce positive but small mean rank information coefficients at the one-day horizon; the largest is 0.0143 for FinBERT. None of the 28 model--horizon tests remains significant after Newey--West inference and false-discovery-rate correction. Portfolio results likewise fail to establish a robust advantage for the best-performing classifiers. The findings show that QLoRA is effective for financial sentiment adaptation, while also documenting a clear gap between classification accuracy and tradable cross-sectional signals.
    Date: 2026–08
    URL: https://d.repec.org/n?u=RePEc:arx:papers:2608.04200
  19. By: Junting Duan; Markus Pelger
    Abstract: Empirical researchers increasingly use large language models (LLMs) to extract structured features, such as sentiment scores, classifications, and expectations, from unstructured data and treat these generated features as observed covariates in downstream estimation. This practice can invalidate inference when systematic, input-dependent errors in generated features, such as hallucination and look-ahead bias, distort the downstream moment conditions. Even after correction, generated features remain noisy proxies whose error profiles differ across models and prompts. We introduce AI-Powered Inference (AI-PI), a method-of-moments framework for valid and efficient inference that combines three components: a moment-specific bias correction based on a small human-labeled calibration set; adaptive weights that optimally combine multiple model-prompt pairs; and an optimal calibration-set design that concentrates costly human labels where the generated features are least reliable. We establish consistency and asymptotic normality of the AI-PI estimator, allowing for data-adaptive labeling designs, cross-fitted LLM-pipeline tuning, and overidentified GMM. Simulations confirm substantial gains over naive LLM regressions and over debiasing without optimal weighting or labeling design. In an application to news-based sentiment and stock returns, AI-PI produces stable conclusions where naive analyses vary substantially across LLM and prompt choices, with a confidence interval roughly half as long as using the human-labeled data alone.
    JEL: C10 C13 C50 C55 C80 G12
    Date: 2026–07
    URL: https://d.repec.org/n?u=RePEc:nbr:nberwo:35481
  20. By: Melissa Dell; Ashesh Rambachan
    Abstract: Artificial intelligence (AI) is transforming measurement in economics. AI models convert unstructured data, such as text and images, into structured variables at low cost, making previously prohibitive measurement feasible at scale. This shifts the bottleneck from finding any scalable measure of a phenomenon to choosing among many plausible ones, which may support different empirical conclusions. This review provides guidance for navigating that shift. We describe three stages at which AI enters the measurement pipeline---discovery, construct definition, and observation---and what each demands of researchers. We argue that credible inference with AI-generated variables requires appropriately designed validation: anchoring measurement to explicit criteria, rather than informal claims that a proxy is reasonable. We then examine how validation samples support valid inference even when AI predictions are arbitrarily biased, and what can be done when a random validation sample is unavailable.
    Date: 2026–08
    URL: https://d.repec.org/n?u=RePEc:arx:papers:2608.23524
  21. By: Joshua S. Gans
    Abstract: AI-generated assessments of manuscripts could improve the quality of peer review, but sharing them with reviewers might decrease the information that editors possess about the manuscript. Human peer reviews add value because reviewers add what the AI assessment reveals. When both reviewers see the same assessment and make similar errors, they can replicate each other’s mistakes, coincidentally investigate related areas, and share the same blind spots. If AI assessment improves the quality of reviewers’ work while they investigate the manuscript, giving the report to only one reviewer can better combine AI assistance with independent, human investigation. If some reviewers already use AI-generated assessments on their own, changing the journal policy to allow sharing the assessment can improve peer review if the change takes some of the burden of investigation off private tools that make similar errors, though giving the assessment to only one reviewer may still outperform giving it to both or neither. Even a report delivered after the reviews are submitted can still shift what reviewers investigate if they know in advance which issues it will check. Journal policy will depend on the AI’s effects on reviewer effort and the questions they investigate, whether reviewers are willing and able to comply with the policy, the monitoring costs, and the journal’s ability to protect the confidentiality of reviewer identities and emails. Journals can evaluate competing policies without knowing the true quality of the manuscripts they review by randomly varying who receives the assessment and then comparing reviewer disagreement across different levels of report sharing.
    JEL: D82 D83 L86 O33
    Date: 2026–08
    URL: https://d.repec.org/n?u=RePEc:nbr:nberwo:35688

This nep-ain issue is ©2026 by Ben Greiner. It is provided as is without any express or implied warranty. It may be freely redistributed in whole or in part for any purpose. If distributed in part, please include this notice.
General information on the NEP project can be found at https://nep.repec.org. For comments please write to the director of NEP, Marco Novarese at <director@nep.repec.org>. Put “NEP” in the subject, otherwise your mail may be rejected.
NEP’s infrastructure is sponsored by the Griffith Business School of Griffith University in Australia.