nep-ain New Economics Papers
on Artificial Intelligence
Issue of 2026–09–21
ten papers chosen by
Ben Greiner, Wirtschaftsuniversität Wien


  1. The Utility of AI Tools in Auditing Adherence to Pre-Analysis Plans By Jeffrey Clemens; Anwita Mahajan
  2. Normative boundaries of AI in scientific work: Evidence from PhD researchers By Francesco Angelini; Johan Lyrvall
  3. Reproducibility is not construct validity: LLM measurement of institutionally situated communication By Veronika Batzdorfer; Carlo Romano Marcello Alessandro Santagiustina
  4. The Role of AI in Online Reviews By Valeria Lermana; Oren Rigbi; Yaniv Dover
  5. Access to Live AI Advice and Behavior Under Risk: An Incentivized Experiment By Paul Althaus; Leon Houf; Christiane Schwieren
  6. Voice AI in Firms: A Natural Field Experiment on Automated Job Interviews By Brian Jabarian; Luca Henkel
  7. AI Innovation and Firm Performance in the Medical Device Industry By Fazliddin Shermatov; Stephane Robin; Aldo Geuna
  8. Are AI Risks Priced in the U.S. Stock Market? Evidence from Financial News Factors By Yanhui Shen
  9. One advisor for the whole world? Cross-country evidence on financial advice from large language models By Bäckman, Claes
  10. CIFQA: A Deterministic Tool-Grounded Multi-Agent LLM Framework for Financial Query Answering By Kunjesh Parekh; Anil Kumar Tiwari; Divya Saxena

  1. By: Jeffrey Clemens; Anwita Mahajan
    Abstract: Pre-analysis plans (PAPs) can improve research reproducibility by reducing researchers’ degrees of freedom, but their value depends on adherence. We argue that large language models (LLMs) can provide an efficient, scalable, and systematic way for authors and reviewers to assess adherence to PAPs. In an application to our own research, an LLM systematically identifies precommitted design choices, evaluates deviations, and diagnoses gaps in pre-specification, substantially reducing the human labor required for these tasks. However, variability in audit output across LLMs underscores the continued importance of human judgment. We discuss implications for best practices in AI-assisted PAP auditing.
    JEL: C18 C80 D82
    Date: 2026–09
    URL: https://d.repec.org/n?u=RePEc:nbr:nberwo:35719
  2. By: Francesco Angelini; Johan Lyrvall
    Abstract: Artificial intelligence (AI) is increasingly embedded in scientific work, but researchers may not evaluate its use uniformly across research tasks. This study examines task-specific attitudes towards AI among an international, self-selected sample of 3, 785 PhD students in STEM and medical and health sciences who participated in Nature's Graduate Survey 2025. We analyse respondents' comfort with using AI for writing a research article, collecting and analysing data, designing experiments, tracking scientific literature, and summarising it. Latent class analysis identifies four distinct attitudinal profiles. The dominant profile reflects a "division of labour, " in which AI is widely accepted for literature-related tasks but resisted in activities closely associated with intellectual contribution, such as writing, data analysis, and experimental design. A "status quo" profile is broadly uncomfortable across tasks, an "all-purpose" profile is broadly comfortable, and an "undecided" profile expresses substantial uncertainty. These patterns suggest that attitudes towards AI in research are organised less around a simple acceptance-rejection divide than around task-specific boundaries, likely concerning delegation, authorship, and responsibility. Because the survey measures comfort rather than legitimacy, the profiles are best interpreted as attitudinal configurations with a normative dimension. The findings highlight the importance of task-specific approaches to AI governance, doctoral training, disclosure, and research evaluation.
    Date: 2026–08
    URL: https://d.repec.org/n?u=RePEc:arx:papers:2608.25678
  3. By: Veronika Batzdorfer (KIT); Carlo Romano Marcello Alessandro Santagiustina (ALMAnaCH, m\'edialab, Sciences Po)
    Abstract: High annotation reproducibility does not necessarily imply that an LLM-inferred measure captures the construct it is intended to measure. We test this distinction using a dataset from the European Commission's AI Act consultation, linking structured survey responses to free-text consultation submissions from the same stakeholders. LLM annotations of consultation submissions are highly reproducible (intraclass correlations > 0.99), yet show limited convergence with survey-reported measures of the nominal construct they were intended to approximate. Divergence between survey-and LLM-inferred text-based measures varies systematically across stakeholder groups: business associations express greater concern about AI risks in text-based consultations than in survey responses ({\=g} = +1.0), whereas public authorities and several nonbusiness groups show smaller or negative divergences. Divergences between scores suggest positive spatial autocorrelation across European countries (Moran's I = 0.347, p = 0.036), indicating that stakeholders from neighboring countries tend toward more similar text-based stances towards AI safety concerns. Despite divergence, survey-reported concerns remain strongly associated with support for explainability across all divergence levels. These results demonstrate that LLM annotation reproducibility can coexist with poor construct correspondence and motivate validation procedures that distinguish reproducibility, construct validity, and communication context variation when LLMs are used as measurement instruments.
    Date: 2026–09
    URL: https://d.repec.org/n?u=RePEc:arx:papers:2609.19866
  4. By: Valeria Lermana; Oren Rigbi; Yaniv Dover
    Abstract: The rapid adoption of large language models (LLMs) creates new opportunities for strategic content generation on online platforms, including potentially harmful forms of manipulation that may undermine platform effectiveness and reshape platform dynamics. However, measuring such activity is difficult because AI-generated content is rarely directly observable. We introduce an empirical approach that leverages discrete LLM supply shocks - abrupt changes in model prices and capabilities, and contrasts verified with non-verified reviews to identify changes in platform activity associated with generative AI supply improvements. We apply this approach to more than 13 million reviews from Trustpilot, one of the leading online platforms for business reviews. A robust finding is that following LLM supply shocks, unverified reviews shift toward greater negativity: more 1-stars, fewer 5-stars, and lower ratings, with effects driven primarily by new model releases and concentrated among firms with the lowest and highest review volumes, suggesting that strategic AI use may reshape platform competition dynamics. We further find that LLM supply shocks trigger short, concentrated bursts of review activity. Together, these findings suggest that generative AI is already reshaping how reputation and competition operate on online platforms.
    Keywords: generative AI, large language models, online reviews, digital platforms, user-generated content
    JEL: L86 O33 L15 M31
    Date: 2026
    URL: https://d.repec.org/n?u=RePEc:ces:ceswps:_12960
  5. By: Paul Althaus; Leon Houf; Christiane Schwieren
    Abstract: Generative AI has become an everyday advisor, and the systems people consult are live and interactive, not pre-scripted. We ask whether access to such a system changes behavior under risk. In an incentivized experiment (N = 158), participants made lottery choices with an optional decision aid presented as a conventional pre-written tool, a live one-shot AI, or a live interactive AI they could query, with information format held equivalent across conditions. Risk preferences are elicited via DOSE. We find no evidence that access to a live AI advisor changes risk aversion.
    Date: 2026–09
    URL: https://d.repec.org/n?u=RePEc:arx:papers:2609.07358
  6. By: Brian Jabarian; Luca Henkel
    Abstract: We study AI agents as information-collection technologies: automated systems that elicit decision-relevant signals from humans through live interactions. We test how such AI automation impacts information collection and organizational outcomes using a natural field experiment with 70, 000 applicants applying for real jobs. Applicants were randomly assigned to be interviewed by either human recruiters or AI voice agents. Afterward, human recruiters evaluate the interviews and make hiring decisions. Applicants interviewed by AI agents are 12% more likely to receive job offers, and these gains translate into higher job starts and worker retention, with no decline in the productivity of hired workers. Analyzing interview transcripts reveals that AI voice agents achieve controlled variance: their interviews are more structured and consistent while remaining responsive to individual applicants, which is associated with more hiring-relevant information collected. Our results suggest that a key advantage of AI automation lies in environments where information collection is delegated across many human workers and repeated such that variance in task execution becomes noise in decision-relevant signals, which AI compresses through adaptive standardization.
    Keywords: artificial intelligence, interviews, hiring, organizational design, field experiment
    JEL: C93 J24 M15 M51 O33
    Date: 2026
    URL: https://d.repec.org/n?u=RePEc:ces:ceswps:_12984
  7. By: Fazliddin Shermatov; Stephane Robin; Aldo Geuna
    Abstract: Whether artificial intelligence pays off for the firms that build it into their products is hard to establish, because AI innovation is itself hard to observe. The medical technology sector is a rare exception: an AI-enabled device must obtain clearance from a national health authority before it can reach a patient, leaving a dated, firm-attributable record of AI innovation output that can be observed directly rather than proxied. We exploit this setting with a three-stage recursive model estimated on a novel firm-level dataset linking FDA premarket clearances, USPTO patents, Scopus publications, and Orbis financials, tracing the full innovation chain from external collaboration through AI device introduction to firm performance. We find that external AI research collaboration is a robust driver of AI device introduction across firm sizes and estimators, with a larger effect for small firms, consistent with external knowledge ties substituting for limited internal R&D capacity. Decomposing by partner type, the effect is largest for industry and clinical collaborations and smallest for academic ties, consistent with the former being closer to the regulatory and commercialisation process. Firms that bring AI devices to market display higher labour productivity, an effect robust for small firms and the full sample that holds under both sequential and joint maximum-likelihood estimation and accumulates across successive device introductions. Effects on profit margins are present but weaker and do not survive all specifications, a pattern consistent with competitive entry eroding pricing power as AI devices diffuse through the sector.
    Date: 2026–09
    URL: https://d.repec.org/n?u=RePEc:arx:papers:2609.08485
  8. By: Yanhui Shen
    Abstract: This paper asks whether firms' exposures to news about different types of AI risk are priced in U.S. stock returns. Using AI and risk keywords, I identify 7, 787 Wall Street Journal articles from January 2016 to December 2025. I combine latent Dirichlet allocation (LDA) with the Domain Taxonomy in the MIT AI Risk Repository to construct four news-based systematic risk factors. I estimate betas to factor innovations and test pricing with univariate portfolio analysis and Fama-MacBeth regressions. Only the taxonomy-mapped Misinformation factor (D3) is robustly priced. Its high-minus-low beta portfolio earns monthly alphas of 0.49%-0.57%, and the estimated D3 price of risk is positive and statistically significant across beta-estimation windows, conventional factor and industry controls, alternative innovation models, and the pre-ChatGPT subsample. The other factors are not reliably priced, indicating that AI-risk pricing is domain-specific.
    Date: 2026–08
    URL: https://d.repec.org/n?u=RePEc:arx:papers:2609.05485
  9. By: Bäckman, Claes
    Abstract: Households around the world increasingly take portfolio advice from the same handful of large language models. Does that advice adjust to what local circumstances warrant? I pose the same portfolio problem to leading models across twenty-one countries, each in its own language. Advice is nearly uniform across countries, even though the explanations invoke local context. What moves the advice instead is the model a household consults, not the country it lives in. The models follow retail financial advice rather than what academic finance would prescribe, and fail to incorporate the household's balance sheet even when it is stated. While automated advice once held out the promise of tailoring guidance to everyone, large language models encode the same advice for everyone.
    Keywords: large language models, financial advice, household finance, portfolio choice, robo-advising, cross-country variation
    JEL: G11 G41 G51 D14 O33
    Date: 2026
    URL: https://d.repec.org/n?u=RePEc:zbw:safewp:343061
  10. By: Kunjesh Parekh; Anil Kumar Tiwari; Divya Saxena
    Abstract: Calculation-intensive financial question answering requires exact reasoning over structured rates, temporal conditions, numerical formulas, and rule-based constraints. Although Large Language Models (LLMs) perform strongly on natural language tasks, they often produce numerically incorrect yet plausible answers when solving multi-step financial calculations. To address this limitation, we introduce CIFQA (Calculation-Intensive Financial Query Answering), a deterministic tool-grounded multi-agent LLM framework for financial question answering. CIFQA separates language understanding from numerical execution by assigning specialized agents to query interpretation, routing, parameter extraction, computation planning, and response generation, while deterministic Python-based tools perform financial calculations and rule application. We instantiate CIFQA for fixed deposit query answering and evaluate it on a curated benchmark of fixed deposit queries. CIFQA achieves 95.54% accuracy on calculation-intensive queries and 90.87% overall accuracy, substantially outperforming direct LLM baselines even when provided with complete formulas, rate cards, and benchmark instructions. Ablation studies show that deterministic components such as exact rate lookup, tenure computation, rolling-year adjustment, and premature-withdrawal logic are critical contributors to performance. Notably, a 17B open-source backbone operating within CIFQA outperforms substantially larger frontier models evaluated with the same financial information, demonstrating that architectural design is a more important determinant of numerical reliability than model scale. While evaluated on fixed deposit queries, CIFQA provides a generalizable framework for calculation-intensive financial reasoning tasks.
    Date: 2026–06
    URL: https://d.repec.org/n?u=RePEc:arx:papers:2608.26114

This nep-ain issue is ©2026 by Ben Greiner. It is provided as is without any express or implied warranty. It may be freely redistributed in whole or in part for any purpose. If distributed in part, please include this notice.
General information on the NEP project can be found at https://nep.repec.org. For comments please write to the director of NEP, Marco Novarese at <director@nep.repec.org>. Put “NEP” in the subject, otherwise your mail may be rejected.
NEP’s infrastructure is sponsored by the Griffith Business School of Griffith University in Australia.