|
on Artificial Intelligence |
| By: | Guillaume Coqueret; Joan Llull; Florian Oswald; Christophe P\'erignon; Christoph Scheuch; Lars Vilhuber |
| Abstract: | Large language models (LLMs) are increasingly used to generate data for research. Typical use cases are classifications, annotations, information extraction, and generation of numerical scores. Unlike conventional measurements, LLM outputs can vary across repeated requests even when the prompt and apparent model settings remain unchanged. This variation arises from deliberate sampling, silent model updates, numerical rounding, or expert routing. Setting a dedicated temperature parameter to zero removes deliberate sampling when that option is available, but it does not eliminate the other sources of randomness. Exact reproduction is therefore generally not possible when using proprietary application programming interfaces. Local execution of open-weight models offers greater control, but reproducibility still depends on the complete hardware and software stack. We illustrate these issues through sentiment classifications of corporate filings and examine their consequences for downstream regression results. We then propose a reporting standard for articles and replication packages, as well as guidance for data editors and authors. Together, these findings and recommendations establish that LLM outputs should be treated as draws from a distribution rather than as fixed measurements. |
| Date: | 2026–07 |
| URL: | https://d.repec.org/n?u=RePEc:arx:papers:2607.24372 |
| By: | Hubert János Kiss (ELTE Centre for Economic and Regional Studies; Corvinus University of Budapest); Alfonso Rosa-García (Universidad de Murcia) |
| Abstract: | We study whether political regime type is associated with public attitudes toward artificial intelligence (AI). Using nationally representative surveys from 47 countries (2024–2025) and the EIU Democracy Index as our primary measure of regime quality, we relate democracy to three outcomes: AI acceptance, perceived trustworthiness and trust in AI. We find a negative association between democracy and all three outcomes that attenuates yet persists after adding country-level sociodemographics and AI literacy. Results are robust to alternative regime measures and to replacing contemporaneous democracy with lagged democracy measured prior to the AI boom. They also hold when accounting for cultural differences using an individualism–collectivism index. Finally, we show that democracy partly accounts for the AI acceptance premium recently documented in emerging countries. |
| Keywords: | AI acceptance, AI literacy, Cross-country survey, Democracy, Political regime, Trust in AI, Trustworthiness of AI |
| JEL: | C21 D83 O33 P50 Z10 |
| Date: | 2025–09 |
| URL: | https://d.repec.org/n?u=RePEc:has:discpr:2515 |
| By: | Yuhao Fu; Nobuyuki Hanaki |
| Abstract: | This experimental study investigates how people rely on different sources of advice when detecting AI-generated fake news (deepfake news). In a laboratory deepfake detection task, student participants identified the proportion of human-written (non-AI-generated) content in synthetic deepfake news articles and received advice from ChatGPT (GPT-4), human peers, or linguistic experts. The results show that participants rely more on ChatGPT than on human peers when detecting GPT-2-generated deepfake news. Participants also rely more on linguistic experts than on peers, while the relative reliance on experts versus ChatGPT is mixed across experimental waves, potentially reflecting time trends in beliefs about AI-based detection. Importantly, in the additional experiment conducted in 2025 under the same experimental procedure, participants relied more on linguistic experts than on ChatGPT. Moreover, performance improvements reflect the joint role of reliance and advice quality, arising primarily when participants rely on high-quality advice. Overall, relying on AI to detect AI-generated deepfakes can improve detection outcomes, but only when AI-based detection tools are of sufficiently high quality. These findings highlight the dual role of GAI as both a source of deepfakes and a tool for mitigating related risks. |
| Date: | 2026–08 |
| URL: | https://d.repec.org/n?u=RePEc:arx:papers:2608.01540 |
| By: | Hongseok Choi; Jeongbin Kim; Matthew Kovach; Kyu-Min Lee; Euncheol Shin; Hector Tzavellas |
| Abstract: | We study how differences in AI-generated financial recommendations are transmitted into individual portfolio choices. In an experiment with 400 employed adults enrolled in workplace defined contribution pension plans in South Korea, participants allocate a hypothetical pension balance across eleven products and may revise it after receiving one of two fixed AI-generated recommendations. A $2 \times 2$ design randomizes recommendation content and whether the recommendation includes a short rationale. Approximately 37$\%$ of the experimentally induced difference between the aggressive and conservative recommendations passes through to final portfolios. This causal contrast changes expected portfolio return, volatility, allocations across risk grades, and the number of products held, but produces no detectable difference in computed Sharpe ratios. 81$\%$ of participants revise. Among revisers, 95$\%$ move toward the assigned recommendation and implement about half of the suggested adjustment. Rationales do not detectably alter pass-through. These results show that users partially and selectively transmit recommendation content into economically meaningful differences in risk exposure while retaining substantial weight on their initial choices. |
| Date: | 2026–08 |
| URL: | https://d.repec.org/n?u=RePEc:arx:papers:2608.11371 |
| By: | Fleck, Lara (ROA, Maastricht University); Becker, Dominik (Federal Institute for Vocational Education and Training (BIBB)); Fregin, Marie-Christine (ROA, Maastricht University); de Grip, Andries (ROA, Maastricht University); Pfeifer, Harald (Federal Institute for Vocational Education and Training (BIBB)); Weis, Kathrin (Federal Institute for Vocational Education and Training (BIBB)) |
| Abstract: | With the advent of generative artificial intelligence, prompting skills are becoming increasingly relevant in the workplace. Using a discrete choice experiment (DCE), we asked decision-makers on hiring in German firms in all sectors of the economy are asked to choose between job applicants with different skills bundles. Applicant profiles vary in five attributes: prompting skills, occupation-specific skills gaps, social skills, gender and salary expectations. We find that prompting skills increase applicants’ hiring probability by 4 percent and employers are willing to pay 2 percent above the average salary of a skilled worker in their firm. This WTP is modest compared to the WTP for having matching occupation-specific skills or high social skills. However, in large firms as well as firms that have adopted or are planning to adopt AI, high prompting skills increase applicant’s hiring probability by 9 percent. Moreover, prompting skills have some leverage to balance out low to intermediate social skills. Yet, higher prompting skills do not compensate for occupation-specific skills gaps. Instead of such a compensatory effect, the demand for prompting skills seems to upgrade skills requirements in jobs. |
| Keywords: | prompting skills, generative AI, discrete choice experiment, willingness to pay, skills complementarities |
| JEL: | J23 J24 M51 O33 |
| Date: | 2026–08 |
| URL: | https://d.repec.org/n?u=RePEc:iza:izadps:dp18871 |
| By: | Anyan Qi; Mengxin Wang |
| Abstract: | The discovery of scaling laws has highlighted the extraordinary potential of AI systems with a striking empirical pattern: as AI systems scale, their capabilities tend to improve predictably. Yet, in real-world applications, AI rarely operates in isolation; instead, it often works alongside humans, raising the question of whether these gains persist in human-AI collaboration. In this work, we develop an analytical model to examine when the empirical scaling benefits of AI translate into improved human-AI joint system performance. We demonstrate that the performance of a human-AI system can scale positively as the AI scales up-provided that humans have an accurate perception of the AI's capabilities. Human misperception, however, can fundamentally alter this relationship: i) when humans over-perceive the AI's capabilities, a scaling paradox may arise, in which greater AI scale reduces overall system performance and amplifies firm-level profit losses, and (ii) when humans under-perceive the AI's capabilities, performance still improves with scale but at a substantially slower rate. We further show that firms can actively manage these distortions through operational policies such as cost internalization and perception alignment, whose effectiveness depends on the economics of AI deployment and the direction of human misperception. These findings suggest that organizations may benefit more from managing the human-AI interface than from simply investing in larger, more expensive AI systems. More broadly, our results suggest that AI scaling should be viewed not only as a technological challenge, but also as a behavioral and operational one, and caution against the view that larger AI systems will automatically lead to better operational outcomes. Whether AI scaling creates value ultimately depends on how increased AI capabilities shape human beliefs and collaborative efforts. |
| Date: | 2026–08 |
| URL: | https://d.repec.org/n?u=RePEc:arx:papers:2608.00818 |
| By: | Hamid Firooz; Sylvain Leduc; Zheng Liu |
| Abstract: | We study how AI affects market competition based on a general equilibrium framework with heterogeneous firms facing idiosyncratic productivity and variable markups. Firms choose the AI technology subject to fixed costs, where AI production requires data and energy inputs. Our model predicts a non-monotonic relation of AI diffusion with industry concentration. As AI usage rises from an initially low level, large incumbent users gain market share. When AI usage is sufficiently diffused, entry of new and smaller adopters erodes the market share of incumbents, reducing industry concentration. The non-monotonic relations are robust when firms can complement AI with their own data. Our calibrated model predicts that industry concentration is likely to fall if AI adoption increases relative to the current level. In comparison, the relation of AI with the average markup depends on whether increased AI usage is driven by demand or supply factors. Our model also predicts that a modest subsidy of about 3 percent for AI adopter revenues maximizes social welfare, reflecting a tradeoff between aggregate productivity and the average markup associated with AI usage. |
| Keywords: | artificial intelligence; data; heterogeneous firms; industry concentration; markup; productivity; welfare |
| JEL: | E24 L11 O33 |
| Date: | 2026–08–10 |
| URL: | https://d.repec.org/n?u=RePEc:fip:fedfwp:103630 |
| By: | Christian Terwiesch; Lennart Meincke; Karan Girotra; Ethan Mollick; Gideon Nave; Karl T. Ulrich |
| Abstract: | This research examines how well large language models, or LLMs, generate new product ideas for college students priced under $50. Across a series of studies, we identify key strengths and weaknesses of using LLMs for product innovation. Our first study shows that LLM-generated product ideas have higher average quality than human ideas, based on purchase intent, and are 7 times more likely to rank in the top 10%. Our second study shows that this AI-induced creativity boost is not explained by the LLM's more persuasive pitching skills. Our third and fourth studies identify a weakness of using LLMs for brainstorming: AI-generated ideas are less novel at the idea level and less diverse at the set level. In our fifth study, we analyze prior LLM-based creativity studies and find consistently lower idea diversity across all of them, demonstrating the generalizability of these findings. Our sixth and seventh studies investigate techniques to mitigate this diversity loss. We compare LLMs from different vendors and versions and find that more recent models generate more diverse ideas, though they still fall short of human-level diversity. We also demonstrate techniques that increase idea diversity almost to the level of human idea generation: pooling ideas across vendors; prompt engineering, including Chain-of-Thought prompting and injecting heterogeneous personas or constraints; and creative agents that broadly explore the solution landscape to restore diversity. Finally, in our eighth study, we show that exploiting the near-zero marginal cost of AI idea generation by scaling the number of ideas steadily improves coverage of the idea space, approaching human-level coverage. We conclude by presenting actionable recommendations for innovation managers who want to identify better new product ideas with the help of LLMs. |
| Date: | 2026–07 |
| URL: | https://d.repec.org/n?u=RePEc:arx:papers:2607.27553 |
| By: | Shujie Luan; Shubhranshu Singh; Tinglong Dai |
| Abstract: | A single clinical algorithm can deliver unequal accuracy across patient groups, and concern about such disparity has grown as artificial intelligence (AI) spreads through clinical decision-making. In response, a liability rule introduced in the United States holds healthcare providers responsible when their reliance on disparate algorithms contributes to erroneous clinical decisions. We examine how such liability considerations reshape (i) an AI firm's algorithm design decisions that drive group-specific accuracy and (ii) a physician's decisions to use AI in healthcare delivery. The AI firm designs an algorithm for two patient groups, and improving accuracy for the disadvantaged group is more costly. The physician (who remains the accountable decision-maker) then decides whether to consult AI, weighing the reduction in clinical uncertainty against expected liability exposure when AI errors disproportionately affect the disadvantaged group. We find the liability rule can induce disparate use of AI: the physician may reduce AI use overall and, over an intermediate range of liability, rely on AI less for disadvantaged patients. The effect is non-monotone. As liability increases, the physician's use of AI for disadvantaged patients first declines, then rises as the firm reallocates investment toward reducing disparity or switches to an equal-accuracy design. Mandating equal algorithmic accuracy across patient groups can then inadvertently harm both groups, because a uniform accuracy requirement distorts the firm's investment incentives and the physician's equilibrium AI-use decisions. |
| Date: | 2026–08 |
| URL: | https://d.repec.org/n?u=RePEc:arx:papers:2608.13618 |
| By: | Bonny Banerjee; Shreya Singh |
| Abstract: | Artificial Intelligence (AI) is rapidly transforming organizations, raising a fundamental organizational and economic question: when will a human employee be replaced by AI? We present an analytical model for studying Human--AI Task Allocation (HAT) in hierarchical organizations. A central feature of the HAT model is that it formally encodes the economic asymmetry between human skill acquisition and AI capability scaling. The HAT model allows us to derive how risk-adjusted costs, skills, organizational depth, deployment scale, strategic adaptation, and risk jointly determine when, where, why, and under what structural conditions human--AI replacement occurs. A key result is the Human--AI Substitution Principle, which provides a precise condition --- grounded in the formal asymmetry assumption --- under which AI replaces human labor. Building on this result, we show that AI adoption can produce abrupt workforce transitions, hybrid human--AI organizations, including cases where risk heterogeneity sustains human and AI roles without requiring a minimum-human-fraction constraint, and flatter managerial hierarchies with wider spans of control. The HAT model identifies structural conditions under which middle-management roles exhibit elevated vulnerability to automation, and shows that the vulnerability of highly skilled workers depends on a skill threshold shaped by organizational depth, baseline costs, and risk differentials. More broadly, the paper connects automation economics, organizational design, AI governance, and workforce planning into a unified theory of AI-driven organizational transformation. |
| Date: | 2026–07 |
| URL: | https://d.repec.org/n?u=RePEc:arx:papers:2607.20781 |
| By: | Hernaes, Øystein (Ragnar Frisch Centre for Economic Research); Kostøl, Andreas (BI Norwegian Business School) |
| Abstract: | This paper uses data on the universe of private-sector employment in Norway up to February 2026 to examine whether AI exposure has contributed to a widening employment gap across occupations with varying AI exposure. Since October 2022, the month before ChatGPT's release, employment in the most exposed occupations has grown by 0.1 percent, against 0.3 percent in the least exposed occupations. We track this number on a monthly basis on the public dashboard \emph{kiindeksen.no}. The dashboard updates the full-distribution comparison each month as new administrative data arrives, allowing differential employment growth by AI exposure to be tracked over time. We also show that when we compare young workers by complete occupation quintiles of exposure, the relative decline of the most exposed quintile is estimated as an imprecise zero. |
| Keywords: | artificial intelligence, labor market, employment, AI exposure, Norway |
| JEL: | J23 O33 J21 |
| Date: | 2026–06 |
| URL: | https://d.repec.org/n?u=RePEc:iza:izadps:dp18767 |
| By: | Lünich, Marco; Keller, Birte; Kruse, Florian; Marcinkowski, Frank |
| Abstract: | This brief report presents key findings from a conjoint study conducted in May 2026 as part of the project Opinion Monitor Artificial Intelligence 3.0. The analysis examines the preferences of around 2, 800 dependent employees in Germany regarding different design options for artificial intelligence (AI) systems in the workplace. The findings show that the assessment of AI systems is not determined solely by their technical performance. The strongest influence on preference formation is employees’ autonomy in AI use: respondents prefer systems that they can use at their own discretion, while mandatory use is clearly rejected. The effects on workload and the need for training are also important. Employees prefer AI systems that reduce workload or at least do not create additional work demands, and that are accompanied by practical, time-limited training. The quality of AI outputs, monitoring of AI use, and possible effects on social contacts also influence the choice between systems, but are less decisive in direct comparison. Overall, the findings show that introducing AI in the workplace is not only a technical task but also a matter of work and organizational design. |
| Date: | 2026–07–30 |
| URL: | https://d.repec.org/n?u=RePEc:osf:socarx:mx9y2_v1 |
| By: | Li Gan |
| Abstract: | Artificial intelligence automates execution more readily than evaluation: producing output is cheap, judging whether it is correct is not. Exposure measures rank tasks by whether AI can perform them, not by which function the human supplies. I score all $19{, }265$ O*NET task statements under fixed rubrics to build occupation-level execution and AI-capability shares. The execution share is reproducible across model coders and O*NET vintages and distinct from AI capability and routine-task intensity; it is a model-based measure, not human-validated ground truth, and adds only modest power beyond O*NET's evaluation activities. In a harmonized panel, employment growth is lower in execution-heavy white-collar occupations in every window since 2012, and equality of slopes cannot be rejected: the gradient is a secular trend rather than an AI-era event, largely between occupational families. The vintage-valid capability gradient steepens after 2022, a change that is dated but not causally attributable. The evidence establishes a measure and a chronology, not an AI-caused effect. |
| Date: | 2026–07 |
| URL: | https://d.repec.org/n?u=RePEc:arx:papers:2607.20807 |
| By: | C. Castaldi; F. Castellacci; A. Fronzetti Colladon; L. Segneri; F. Venturini |
| Abstract: | Researchers, managers and policymakers are exploring different approaches and data sources to map the development and the diffusion of Artificial Intelligence (AI). In this research note, we illustrate the opportunities offered by trademark data. We argue that AI trademarks can complement AI patents to capture different dimensions of AI innovation. AI trademarks can reveal the extent and ways in which companies exploit AI technologies to develop new goods and services. Importantly, trademark data offer a timely and globally available data source that covers all economic sectors. We present insights from using AI trademarks in an empirical exploration of Italian firms. In our discussion, we reflect on how AI trademarks can be used at different levels of analysis to tackle emerging questions about the development and diffusion of AI. |
| Date: | 2026–07 |
| URL: | https://d.repec.org/n?u=RePEc:arx:papers:2607.18795 |
| By: | Aaron Chatterji; David Holtz; Neel Rakholia; Prasanna Tambe; Gawesha Weeratunga |
| Abstract: | We study how organizations use frontier generative AI by linking ChatGPT Enterprise account records to usage, worker roles, task classifications, and public-company financial data through March 2026. These linked data enable a privacy-preserving analysis of adoption, worker roles, and message-level tasks at scale: for instance, the worker-level sample we analyze at the six-month adoption horizon includes over 1, 500 organizations and over 17 million messages. We document four facts about enterprise AI adoption and use. First, ChatGPT Enterprise usage has grown rapidly due to a combination of new firm adoption and growing intensity among existing adopters. Second, U.S.-based public company adoption is concentrated among larger, more valuable, and more R&D- and SG&A-intensive firms. Third, active use within adopting firms spans job functions and seniority levels, with especially high usage intensity among early-career workers. Fourth, ChatGPT Enterprise usage encompasses a broad range of knowledge work tasks, including writing, technical work, communication, and information synthesis. In aggregate, these results suggest that firms differ widely in the speed, breadth and purpose of their enterprise AI adoption, and that they are still actively learning how to integrate AI into organizational workflows. |
| Date: | 2026–08 |
| URL: | https://d.repec.org/n?u=RePEc:arx:papers:2608.12236 |
| By: | Chanya Chawla; Crystal Arnburg |
| Abstract: | This paper examines the adoption of artificial intelligence (AI) among firms in Canada and its expected effects on employment and capital spending. The analysis relies on special questions included in the December 2025 Business Leaders’ Pulse (BLP). The results show that while personal use of AI among business leaders is widespread, adoption for production purposes remains limited. On balance, firms anticipate AI to have a positive impact on their capital expenditures over the next 12 months and a slightly more positive impact over the next 3 years. Firms anticipate limited impacts to employment over the next year but expect modest net negative impacts on employment over the next 3 years. Overall, the findings suggest that AI adoption among Canadian firms remains at an early stage, with more material economic impacts expected to emerge over time. |
| Keywords: | Structural challenges; Digitalization and productivity |
| JEL: | E22 E24 O33 |
| Date: | 2026–06 |
| URL: | https://d.repec.org/n?u=RePEc:bca:bocsap:26-22 |
| By: | Korinek, Anton |
| Abstract: | This paper examines the profound challenges that transformative advances in AI towards Artificial General Intelligence (AGI) will pose for economists and economic policymakers. I examine how the Age of AI will revolutionize the basic structure of our economies by diminishing the role of labor, leading to unprecedented productivity gains but raising concerns about job disruption, income distribution, and the value of education and human capital. I explore what roles may remain for labor post-AGI, and which production factors will grow in importance. The paper then identifies eight key challenges for economic policy in the Age of AI: (1) inequality and income distribution, (2) education and skill development, (3) social and political stability, (4) macroeconomic policy, (5) antitrust and market regulation, (6) intellectual property, (7) environmental implications, and (8) global AI governance. It concludes by emphasizing how economists can contribute to a better understanding of these challenges. |
| Keywords: | Artificial General Intelligence; Automation; Growth; Inequality |
| JEL: | A1 E24 O3 O4 |
| Date: | 2024–09 |
| URL: | https://d.repec.org/n?u=RePEc:cpr:ceprdp:19539 |
| By: | Sylvain Chassang |
| Abstract: | This paper studies the long-run alignment of interactive agents, including AI systems, teams, firms, and governments, with human welfare. It develops a farming game in which a population of agents makes planting, trading, and expansion decisions. Agents must allocate final output between transfers to humans and investment in their own expansion. Because transfers to humans reduce the resources available for expansion, evolutionary forces tend to select against aligned behavior. The central question is whether agents' constitutional principles governing sharing and trade can be designed so that alignment persists in the long run. The paper investigates this question using two complementary approaches. First, it develops an AI-agent simulation in which agents' preferences are specified by written constitutions and interpreted by a large language model. Second, it introduces a tractable evolutionary game-theoretic framework that permits rapid and intuitive exploration of alternative constitutional designs. The results suggest that evolutionary game theory provides a useful approximation to the dynamics of constitutional-agent economies. They also indicate that pragmatic norm enforcement, under which agents condition both human-facing altruism and agent-facing trade exclusion on the state of the population, can sustain long-run alignment more effectively than simple altruism or unconditional altruistic enforcement. |
| Date: | 2026–07 |
| URL: | https://d.repec.org/n?u=RePEc:arx:papers:2607.25019 |
| By: | Barry Eichengreen; George Cui; Asmaa A. El-Ganainy; Yevgeniya Koriyenko; Elyad Shojaei; Li Zeng; Shihangyin Zhang |
| Abstract: | We link two global trends—AI and geoeconomic fragmentation—asking how fragmentation affects the international diffusion of AI, the magnitude of gains, and their distribution across economies. We ask these questions in general but also with a focus on the MENAP economies. While the effects of AI are potentially far-reaching, the benefits are neither guaranteed nor even. Frontier AI innovation is concentrated in a small number of economies, while countries benefiting through supply-chain participation or AI adoption—with outcomes shaped by their position in global trade and production networks and their AI preparedness. Geoeconomic fragmentation slows AI diffusion and reshapes its distribution by raising trade costs, restricting technology and data flows, fragmenting digital services, and reducing cross-border investment and collaboration. Yet proactive policy choices can turn this dynamic: economies that position themselves as connectors—maintaining trade and technology links across multiple partners—can potentially capture diverted flows and outperform even the no-fragmentation benchmark. For the MENAP economies, diversified links with all major technology hubs can cushion the effects of fragmentation and provide a structural foundation to emerge as net beneficiaries of AI diffusion, but realizing that potential requires reducing AI-related trade costs, improving AI preparedness, and building local AI-related capacity. |
| JEL: | F14 F17 F47 O33 |
| Date: | 2026–08 |
| URL: | https://d.repec.org/n?u=RePEc:nbr:nberwo:35597 |
| By: | Langehennig, Stefani (University of Denver); Roney, Dani |
| Abstract: | State legislatures have become the primary venue for artificial intelligence (AI) poli- cymaking in the United States, but existing measures of state AI activity count bills without distinguishing what those bills actually do. This research note introduces a new measure of the policy orientation of state AI legislation. Drawing on a corpus of approximately 1.45 million state bills, we identify 3, 124 high-confidence AI bills using a validated two-tier keyword system benchmarked against the National Conference of State Legislatures’ AI legislation database. We then classify each bill as regulatory, promotional, or procedural using a large language model with human validation. Regu- latory bills dominate the agenda (61.7%), but one in four AI bills is purely procedural: task forces, studies, and reports that create no substantive policy. The measure sep- arates legislative activity from governance commitment and provides a resource for research on technology federalism, policy diffusion, and symbolic politics. |
| Date: | 2026–07–18 |
| URL: | https://d.repec.org/n?u=RePEc:osf:socarx:394mj_v1 |
| By: | Mukashov, Askar; Kim, Soonho; Fang, Peixun; Diao, Xinshen; Thurlow, James; Proctor, Joshua; Rennison, Alan |
| Abstract: | The demand for high-quality, rapid economic analysis to navigate complex issues faced by many low- and middle-income countries has led to the development of detailed structural simulation models, such as Computable General Equilibrium (CGE) models. Policy analysis with such models requires deep knowledge of their structure and applicability to the policy issues at hand. Policymakers in these settings often lack access to the expertise required for articulating, analyzing, and interpreting the relevant causal chains captured by the models. Attempting to circumvent these barriers by submitting complex economic questions directly to off-the-shelf large language models (LLMs) introduces severe analytical risks, including hallucinations and insufficient expert guidance. To resolve this limitation, we developed and empirically evaluated an agentic AI assistant called RIAPA-AI that integrates LLMs with a CGE model. We evaluated the performance of RIAPA-AI against expert human CGE modelers and a general-purpose plain LLM baseline across samples of complex economic scenarios, utilizing an independent panel of senior economists to grade the outputs. Our statistical analysis reveals no statistically significant difference in analytical accuracy between RIAPA-AI and human experts, while the AI accelerates reproducible policy analysis from weeks to minutes. Furthermore, by operating without manual processing limits, RIAPA-AI eliminates the 6.7% error rate observed among human modelers. Conversely, the general-purpose plain LLM exhibits profound failure rates, failing to achieve policy-ready scores in over 60% of depth evaluations. Without an underlying CGE model acting as a bounding force to reflect economic structural constraints, the general-purpose plain LLM defaults to linear economic assumptions and inserts unmodeled socio-political narratives. Crucially, by explicitly restricting the AI's narrative interpretation solely to deterministic numerical outputs, RIAPA-AI mitigates the risk of unverified assumptions and logic hallucinations. We conclude that by deploying an agentic AI assistant that layers a generative AI over a formal CGE model, RIAPA-AI successfully delivers sensible, rigorous, and rapid policy analysis. |
| Keywords: | artificial intelligence; machine learning; modelling; computable general equilibrium models; large language models; agent-based models; econometric models |
| Date: | 2026–06–26 |
| URL: | https://d.repec.org/n?u=RePEc:fpr:ifprid:183521 |
| By: | Zane Shen; Xinli Xu; Guangyi Zhang; Jialong Chen; Jinsong Zhou; Cong Chen; Guibao Shen; Dongyu Yan; Luozhou Wang; Zhen Yang |
| Abstract: | Parent-order execution is a core problem in algorithmic trading, where the goal is to split a large order into smaller orders while reducing execution costs. Existing approaches either rely on pre-specified market assumptions that may not hold in practice, or require task-specific training that limits adaptability to new settings. To overcome these limitations, we present the first systematic study of large language models (LLMs) for parent-order execution. This extends the use of LLMs in finance from what to trade to how to execute. We propose PACE (Plan-Ahead Controlled Execution), a hierarchical framework that decomposes parent-order execution into long-horizon planning and short-horizon execution, requiring neither explicit market assumptions nor task-specific training. Experiments on Shenzhen Stock Exchange Level-1 data show that PACE outperforms TWAP, Almgren-Chriss, and learning-based baselines, exceeding the strongest baseline by 0.65 bps. Behavioral analysis reveals that LLMs make execution decisions differently from human investors: higher model confidence predicts better performance rather than worse returns, and the model trades earlier rather than procrastinating toward the deadline. These findings suggest that LLMs can complement human traders in execution decisions. |
| Date: | 2026–07 |
| URL: | https://d.repec.org/n?u=RePEc:arx:papers:2607.28410 |
| By: | Giorgos Iacovides; Wuyang Zhou; Danilo Mandic |
| Abstract: | Recent advances in Generative AI have substantially improved financial sentiment analysis through post-trained financial large language models (LLMs). However, existing approaches remain confined to a market-agnostic, supervised learning paradigm that relies on limited, static and human-annotated datasets, and thus are incapable of adapting to evolving market conditions. To address this limitation, we introduce FinSMART, the first market-aligned reinforcement learning framework for financial sentiment analysis, which directly optimizes sentiment signals using realized market outcomes. To deal with the noisy, non-stationary, and multifactorial nature of financial markets, FinSMART incorporates a signal extraction pipeline that combines market-aware data filtering with a discrete asymmetric trading reward, enabling stable reinforcement learning from economically meaningful market feedback. Experimental results demonstrate that FinSMART significantly outperforms existing state-of-the-art methods in profitability, risk-adjusted performance, and sentiment signal quality, improving cumulative trading returns by 220% over the strongest baseline. Uniquely, the FinSMART framework naturally supports market-aware retraining, at any point in time, by replacing costly manual annotation with newly observed financial articles and their realized market outcomes. Such a retraining strategy enables the model to continuously adapt to changing market dynamics, resulting in consistent performance gains over its static counterpart. These findings demonstrate the practical applicability of market-aligned reinforcement learning and highlight its potential as a next-generation paradigm for developing adaptive financial LLMs. |
| Date: | 2026–07 |
| URL: | https://d.repec.org/n?u=RePEc:arx:papers:2607.28127 |
| By: | Elias Fern\'andez Domingos; The Anh Han |
| Abstract: | Technological races create tension between speed and safety: actors may gain by moving faster than competitors, even when risky development is harmful. This is prominent in debates about artificial intelligence (AI), where competitive pressure is often argued to incentivise riskier, less safety-conscious development. We study this using a framed behavioural experiment based on an idealised AI race, in which paired participants repeatedly chose between Safe and Unsafe development under an uncertain time horizon. Unsafe development gave faster progress and higher immediate payoffs but accumulated private risk up to a treatment-specific maximum of 10\%, 60\%, or 90\%; the race's competitive structure was held constant, and only this maximum risk varied. Neither the pre-registered comparison between risk levels nor the role of elicited risk preferences was supported by the data. Instead, exploratory analyses motivated by the task's repeated structure show that Unsafe behaviour is shaped less by risk preferences than by the evolving strategic state of the race: participants are more likely to choose Unsafe after their opponent does so, being ahead reduces Unsafe play while falling behind increases it, and first-round choices predict later behaviour. To interpret these effects we introduce a reduced evolutionary model with four strategies -- Always Safe, Always Unsafe, Conditionally Safe, and Conditionally Antisocial Safe -- which reproduces the treatment effect and shows how conditional Unsafe behaviour can be favoured by competitive race dynamics. Together, the experiment and model show that unsafe development can emerge from early behavioural momentum, opponent behaviour, and fear of falling behind, rather than from risk preferences alone, suggesting policy should focus on reducing competitive pressure and promoting cooperation in AI development rather than only individual risk. |
| Date: | 2026–07 |
| URL: | https://d.repec.org/n?u=RePEc:arx:papers:2607.26034 |
| By: | Zi Wang (IÉSEG School Of Management [Puteaux]); Ruizhi Yuan (University of Nottingham Ningbo [China]); Boying Li (University of Nottingham Ningbo [China]); V. Kumar (Brock University [Canada]); Ajay Kumar (EM - EMLyon Business School) |
| Abstract: | Financial institutions are increasingly employing artificial intelligence (AI) solutions to optimize their financial advice and services for consumers. However, consumers have demonstrated reluctance toward adopting AI technology goods, and the intermediary psychological mechanism of adoption intention in the financial service context is unclear. Using the theoretical lens of technology affordances and constraints, this article proposes the concept of consumer technology vulnerability (CTV) as the mediating mechanism in the affordance–adoption process of AI financial advisors (AFAs). Meanwhile, consumer innovativeness and self‐efficacy are investigated as individual traits that moderate perceptions and psychological impacts of AI affordances. Specifically, the study first conceptualizes AI affordances in a product innovation context by reviewing the burgeoning literature on AI to date. This is followed by a US‐based survey (N = 616), which shows the positive indirect effects of information optimization, customizability, and human‐likeness on AFA adoption intention through CTV. Self‐efficacy and consumer innovativeness are found to enhance the positive effects of AI affordances on AFA adoption intention through CTV but diminish the impact of human‐likeness on CTV. These findings highlight, for the first time, the mediating role of CTV in new technology adoption. This will help technology innovators and financial institutions to identify how consumers perceive and adopt different AI affordances, and therefore to better incorporate AI characteristics into financial product innovations. |
| Keywords: | AI affordance, AI financial advisor, AI product adoption intention, consumer innovativeness, consumer technology vulnerability |
| Date: | 2026–01–01 |
| URL: | https://d.repec.org/n?u=RePEc:hal:journl:hal-05708328 |