|
on Big Data |
| By: | Christopher W. Karvetski; Sheldon S. Huang; Simas Ku\v{c}inskas; Nadja Flechner; Jingyu Hu; Philip Tetlock; Ezra Karger |
| Abstract: | Decision-makers routinely rely on expert judgments accompanied by written explanations, yet explanation quality is difficult to measure at scale. Forecasting tournaments offer a natural testing ground: probabilistic judgments are paired with natural-language rationales and scored against realized outcomes. We introduce Explanation Quality Markers (EQMs), a set of sixty theory-guided reasoning patterns scored by large language models (LLMs). In a pre-registered analysis of over 55, 000 forecast-rationale pairs from a multiyear forecasting tournament, EQMs predict accuracy at both the forecast and forecaster levels, consistently outperforming pre-LLM text-analysis methods. More than 90% of statistically significant pattern-level EQM-accuracy correlations match our directional hypotheses. The signal is asymmetric: EQMs identify likely underperformers more reliably than they distinguish the very best forecasters. Benchmarked against traditional indicators of forecasting skill, EQMs are the strongest predictor at the forecast level and competitive at the forecaster level, though weaker than prior accuracy. Human ratings of rationale quality are less consistently correlated with accuracy and place disproportionate weight on rationale length. Results transfer to an independent forecasting study. EQMs provide a scalable, interpretable method for extracting judgment-relevant information from written explanations. |
| Date: | 2026–06 |
| URL: | https://d.repec.org/n?u=RePEc:arx:papers:2606.30987 |
| By: | Alam, M. Jahangir; Boyle, Shane; Li, Huiyu; Sekhposyan, Tatevik |
| Abstract: | Recent research suggests that generic large language models (LLMs) can match the accuracy of traditional methods when forecasting macroeconomic variables in pseudo out-of-sample settings generated via prompts. This paper assesses the out-of-sample forecasting accuracy of LLMs by eliciting real-time forecasts of U.S. inflation from ChatGPT. We find that out-of-sample predictions are largely inaccurate and stale, even though forecasts generated in pseudo out-of-sample environments are comparable to existing benchmarks. Our results underscore the importance of out-of-sample benchmarking for LLM predictions. |
| Date: | 2026–01 |
| URL: | https://d.repec.org/n?u=RePEc:cpr:ceprdp:21057 |
| By: | Hannes Wallimann; C\'edric Br\"utsch; Martin Huber |
| Abstract: | Fare evasion generates substantial revenue losses for public transport operators and is typically combated through fare inspections, yet little is known about how the mode of inspection-uniformed versus plainclothes-affects detection efficiency. Using a unique dataset of 21, 727 inspection records from PostAuto, the largest regional bus operator in Switzerland, we apply causal machine learning to estimate the causal effect of inspector visibility on inspection efficiency, defined as detected fare evaders per inspection hour. Our results indicate that plainclothes inspections are, on average, significantly more effective than uniformed inspections, with an estimated average treatment effect of -0.173 incidents per hour, corresponding to a relative reduction of approximately 26%. Heterogeneity analyses find no evidence of systematic effect variation across contextual characteristics, suggesting that the superiority of plainclothes inspections is robust and pervasive across the PostAuto network. When applying optimal policy learning (based on policy trees) to optimally target subgroups by one or the other treatment depending on relative effectiveness, plainclothes inspections are recommended for the large majority of contexts (83.3%), with uniformed inspections suggested only for lines characterised by a below-median share of foreign residents and above-median population size. |
| Date: | 2026–06 |
| URL: | https://d.repec.org/n?u=RePEc:arx:papers:2606.24181 |
| By: | Andreas Ferrara |
| Abstract: | Large language models (LLMs) are lowering the entry barriers to working with exciting data sources that used to require strong data science skills, such as handwritten ledgers, text, images, or sound recordings. This guide provides an introduction for researchers who are new to LLMs. It sets out a step-by-step workflow for turning a research idea into working code and data, and describes the four main ways of interacting with an LLM: the chat window, editor-integrated assistants, agentic coding tools, and the API. It then works through the decisions a practitioner meets in sequence, beginning with whether an LLM is the right tool and whether the data are allowed to be sent to one, then how to select models, write prompts, manage context limits, and control costs, and finally how to validate, reproduce, document, and correct LLM-generated measures in regression settings. A review of recent research shows how these tools already extract, link, harmonize, and classify historical data at scale. Four worked examples with replication files illustrate the use of LLMs. They classify emotions in paintings, link census records without names, measure newspaper salience and sentiment around the 1882 Chinese Exclusion Act, and score the emotional delivery of Franklin D. Roosevelt's wartime speeches. The guide also condenses the workflow, the best-practice recommendations, and the preparation of replication packages into summary tables and checklists to aid applied economists. |
| JEL: | C55 C8 N0 |
| Date: | 2026–06 |
| URL: | https://d.repec.org/n?u=RePEc:nbr:nberwo:35374 |
| By: | Shujie Li (Paderborn University); Yuanhua Feng (Paderborn University) |
| Abstract: | Macroeconomic time series forecasting is crucial for guiding government policy decisions, business strategies, and understanding economic trends. However, predicting macroeconomic variables remains a significant challenge. The complexity of economic systems, insufficient data, high levels of volatility complicate the task of accurate forecasting. To enhance forecasting accuracy, we propose two novel models to capture both linear and nonlinear dynamics. First, we generalize the random walk model by incorporating a drift term, which is estimated using a simple neural network model. Second, a hybrid model is introduced to combine local linear regression and the neural network model. Additionally, we adopt other models from Fritz et al. (2024) for combination. These models are combined using a simple averaging method. Our results demonstrate that the newly proposed neural network-based models produce the lowest average MASE. Additionally, model combination is an effective strategy for enhancing the performance of GDP forecasting in most countries and it is less risky than relying on a single model. |
| Keywords: | nonparametric approaches, combination of forecasting, NNAR, Random Walk |
| JEL: | C14 C51 |
| Date: | 2026–03 |
| URL: | https://d.repec.org/n?u=RePEc:pdn:ciepap:172 |
| By: | Born, Benjamin; Lamersdorf, Nora; Schuster, Jana-Lynn; Steffen, Sascha |
| Abstract: | Using modern natural language processing, we construct a high-frequency inflation expectations index from German-language tweets. This index closely tracks realized inflation and aligns even more closely with household survey expectations. It also improves short-run forecasts relative to standard benchmarks. In response to monetary policy tightening, the index declines within about a week, with the effects concentrated in tweets by private individuals and during the recent period of elevated inflation. Using 117 million online transactions from German retailers, we show that higher inflation expectations are followed by lower household spending on discretionary goods. By linking these shifts in demand to stock returns, we find that, during periods of elevated inflation, firms operating in discretionary sectors experience significantly lower stock returns when inflation expectations rise. Thus, our Twitter-based index provides market participants and policymakers with a timely tool to monitor inflation sentiment and its economic consequences. |
| Keywords: | Inflation expectations; Social media; Large Language Models; Nlp; Household consumption; Stock returns; Monetary policy |
| JEL: | E31 D84 E58 C45 C81 |
| Date: | 2025–12 |
| URL: | https://d.repec.org/n?u=RePEc:cpr:ceprdp:20977 |
| By: | Paul X. McCarthy; Rasika Amarasiri; Xian Gong |
| Abstract: | Universities, funders, investors, and policy agencies often need to identify research with translational relevance before patents, licenses, startups, or industry collaborations are visible. This study introduces the Translation Readiness Index (TRI), a text-based measure evaluating a publication's semantic similarity to papers that appear in high-confidence patent-paper pairs. Using 20, 610 publications from OpenAlex, including 9, 431 publications from the Reliance on Science patent-paper pairs data and 11, 179 matched comparison publications, we created paper-level 768-dimensional semantic embeddings from titles and abstracts with SPECTER2. After evaluating four machine learning classifiers, XGBoost achieved the highest ROC-AUC (0.77). We define TRI as the model-estimated probability that a publication belongs to the patent-paper-paired class. Linguistic analysis revealed that patent-paired publications more often use an invention-oriented framing, distinct from the observational language of the comparison group. External validation across University of Western Australia (UWA) publications and leading global universities demonstrated positive associations between high TRI scores and independent translational indicators. TRI provides a text-based method for identifying translation-ready research, though it should be interpreted as a measure of semantic proximity to patented science rather than a direct measure of realized commercialization. |
| Date: | 2026–06 |
| URL: | https://d.repec.org/n?u=RePEc:arx:papers:2606.31102 |
| By: | Sergio A. Correia; Stephan Luck; Emil Verner |
| Abstract: | Banking crises are commonly associated with bank runs and banking panics, yet our empirical understanding of bank runs is constrained by a lack of bank-level data. In a new paper, we use large language models (LLMs) to extract information on bank runs from millions of digitized historical newspaper pages, creating the most comprehensive database of bank runs in U.S. history. Every bank run episode that we identify is documented on a companion website where users can browse and examine individual episodes, and read the original newspaper articles. In this post, we describe how we built this dataset and discuss what its basic features reveal. |
| Keywords: | bank runs; banking crises; bank failures; deposit insurance; liquidity; solvency; artificial intelligence (AI) |
| JEL: | G01 |
| Date: | 2026–07–07 |
| URL: | https://d.repec.org/n?u=RePEc:fip:fednls:103501 |
| By: | Masood Tadi; Milan Fičura; Jiří Witzany |
| Abstract: | We study natural gas storage valuation under a stochastic futures term structure using deep reinforcement learning (DRL). The storage problem is formulated as a continuous-state, continuous-action Markov Decision Process and solved using the Deep Deterministic Policy Gradient (DDPG) algorithm with Prioritized Experience Replay (PER) buffer and a constraint-aware policy network. We benchmark the approach against intrinsic and rolling intrinsic strategies and find that DRL consistently outperforms intrinsic valuation and achieves competitive performance relative to rolling intrinsic in markets with jumps and seasonality. The results show that DRL provides a practical valuation framework that captures additional extrinsic value under realistic market dynamics and operational constraints. |
| Keywords: | Natural Gas Storage, Rolling Intrinsic Valuation, Deep Reinforcement Learning |
| Date: | 2026–06–12 |
| URL: | https://d.repec.org/n?u=RePEc:prg:jnlwps:v:6:y:2026:id:6.003 |
| By: | Ziwen Zu |
| Abstract: | Large language models (LLMs), a prominent form of artificial intelligence (AI), are becoming everyday interfaces for political questions, but most exchanges are dyadic rather than audiencefacing. This paper asks whether AI conversation functions as a new arena for political expression or as a conversational intermediary for routine political demand. Using 4.30 million humanAI conversations from three large public datasets, we apply two validated classifiers to user messages, identifying political content, use case, and expressed ideology. Political content appears in 3.9% of conversations, varies sharply by platform publicness and conversation depth, and is mostly practical: users ask for information, draft text, and process documents far more often than they state opinions. A regression-discontinuity-in-time design around the 2024 U.S. presidential result call shows that the call changed the expressive subset: among U.S. users, stance-taking, affective language, and ideological extremity rose; comparable conversations elsewhere did not. AI conversation is less a public square than a conversational political intermediary, absorbing routine demand and becoming expressive when major events make political stakes explicit. |
| Date: | 2026–07 |
| URL: | https://d.repec.org/n?u=RePEc:arx:papers:2607.00551 |
| By: | Kwon, Byeungchun; Park, Taejin; Rungcharoenkitkul, Phurichai; Smets, Frank |
| Abstract: | Macroeconomic indicators provide quantitative signals that must be pieced together and interpreted by economists. We propose a reversed approach of parsing press narratives directly using Large Language Models (LLM) to recover growth and inflation sentiment indices. A key advantage of this LLM-based approach is the ability to decompose aggregate sentiment into its drivers, readily enabling an interpretation of macroeconomic dynamics. Our sentiment indices track hard-data counterparts closely, providing an accurate, near real-time picture of the macroeconomy. Their components–demand, supply, and deeper structural forces–are intuitive and consistent with prior model-based studies. Incorporating sentiment indices improves the forecasting performance of simple statistical models, pointing to information unspanned by traditional data. |
| JEL: | E30 E44 E60 C55 C82 |
| Date: | 2025–11 |
| URL: | https://d.repec.org/n?u=RePEc:cpr:ceprdp:20828 |
| By: | Cosmin Borsa; Michael Ludkovski |
| Abstract: | Simulation based solvers for optimal stopping problems must discretize the stopping decision. Under classical dynamic programming, a coarse exercise grid with only a few stopping opportunities can materially undervalue the optimal expected reward, whereas on a very fine grid, approximation errors accumulate through the backward recursion. To remove this limitation, we develop a new reinforcement-learning inspired algorithm that enables us to learn the exercise rule at arbitrarily fine time resolution. Our CARLOS (Continuous-time Adaptive Reinforcement Learning for Optimal Stopping) algorithm utilizes an aggregate deep neural network (ADNN) to learn a joint space-time decision boundary. Starting from a coarse time grid, we progressively increase the frequency of stopping opportunities, while in parallel training the ADNN to refine its timing-value estimates. We moreover design an adaptive sampling strategy that gradually concentrates training effort near the stopping boundary. Benchmarked results show that CARLOS delivers higher prices than existing Bermudan solvers, approaching the American upper bound, and achieves high computational efficiency relative to non-RL comparators. |
| Date: | 2026–06 |
| URL: | https://d.repec.org/n?u=RePEc:arx:papers:2606.17545 |
| By: | Riboni, Alessandro; Ruge-Murcia, Francisco; Tran, Linh |
| Abstract: | Natural language processing is used to extract information from FOMC transcripts and construct quantitative text-based measures of voiced policy stance, emotions, and collaboration. These measures are inputs in an econometric model of deliberation where members interact with one another across rounds of a meeting and over time across meetings. Evidence shows that members learn from one another during within-meeting deliberation and exert influence across meetings. Although emotional tone has limited effects on policy stances and decisions, it has strong predictive power for dissent behavior. |
| JEL: | D7 E5 |
| Date: | 2025–11 |
| URL: | https://d.repec.org/n?u=RePEc:cpr:ceprdp:20840 |
| By: | Foltyn, Richard; Olsson, Jonna |
| Abstract: | Do large language models (LLMs) provide gender-neutral financial advice? We answer this question by prompting 33 widely used LLMs from five vendors, varying only a single word in otherwise identical prompts: “man†versus “woman.†We find that women are advised to allocate 1.8 percentage points less to equity funds than men; this gap persists across vendors, model generations, and model complexity. Providing richer investor information attenuates but does not entirely eliminate the gender gap. Since even modest allocation differences imply persistent return differentials, algorithmic financial advice can shape wealth accumulation across demographic groups. |
| JEL: | C1 G11 J16 |
| Date: | 2026–03 |
| URL: | https://d.repec.org/n?u=RePEc:cpr:ceprdp:21323 |
| By: | Gorin, Clement; Combes, Pierre-Philippe; Duranton, Gilles; Gobillon, Laurent |
| Abstract: | We provide a highly granular account of land-use change across France, comparing two snapshots taken approximately 160 years apart. Built or paved land increased from about 0.7% of the territory in 1860 to 5.2% in 2020. Over the same period, cropland contracted by 20 percentage points to 42%, and meadows nearly disappeared. Only a small fraction of these declines is explained by land development; instead, the shares of forests and pastures roughly doubled. Although recent policies aim to limit land development, our results suggest that the spatial dispersion of developed land may pose a greater concern than its total share. |
| Keywords: | historical maps; Computer vision |
| JEL: | N53 Q15 R14 |
| Date: | 2025–12 |
| URL: | https://d.repec.org/n?u=RePEc:cpr:ceprdp:20946 |
| By: | Tianjia Dong; Nadav Kunievsky; James A. Evans |
| Abstract: | Large language models are increasingly deployed as autonomous decision makers, yet the behavioral mapping they exhibit can vary substantially across decision environments that are payoff-equivalent by construction-environments that share identical payoff-relevant structure but differ in surface presentation. This sensitivity renders suite-based evaluation fragile and raises a fundamental question of behavioral portability: how well does a behavioral mapping learned in one decision environment informative on another that preserves the same underlying incentive structure? We introduce a formal framework to measure this property. Our protocol fits an interpretable behavioral model on data pooled from a set of source environments and evaluates its out-of-sample predictive performance in a held-out target environment, benchmarking against an oracle trained directly on target data. Portability is quantified via a loss-agnostic measure that delivers worst-case bounds on the performance of the induced prediction-action mapping in the target environment. In controlled experiments spanning seven canonical economic decision problems, we document substantial and systematic portability losses, suggesting that behavioral characterizations of LLMs obtained in one decision environment cannot be assumed to transfer reliably to structurally equivalent alternatives. |
| Date: | 2026–06 |
| URL: | https://d.repec.org/n?u=RePEc:arx:papers:2606.22797 |
| By: | Matthieu Bunel; Elisabeth Tovar; Marie-Noëlle Lefebvre |
| Abstract: | This paper studies how large language models (LLMs) trade off moral norms against economic incentives in discriminatory hiring decisions. Bridging discrimination economics and the literature on the moral alignment in computer science, we submit 18 frontier and local LLMs to the factorial vignette experiment of a published human survey that manipulates the motive of discrimination (customer taste-based versus statistical), the cost of non-discrimination, and explicit moral injunctions. We extend this design with LLM-relevant factors: model and user personas, reasoning instructions, scenario realism, and conversational memory. In line with the literature, we find that LLMs align with humans in the direction of the effects manipulated in the survey; we also find important inter-model heterogeneity. Beyond, we contribute to the literature with, to the best of our knowledge, five novel results: 1) the market-oriented motive (customer-taste) overwhelmingly sways models in favour of discrimination, much more than what happens for human respondents; 2) models are more polarised than humans in response to moral injunctions; 3) classic prompt engineering interventions (user stated motives and model personas) have a weak impact on the models’ “moral compass”; 4) post-training alignement, not scale, shape inter-model heterogeneity, which means that de-biasing is possible but must be explicitly implemented by model providers and 5) memory effects suggest that moral permissivity in the models can be induced by conversational contextual effects. |
| Keywords: | Moral judgment on discrimination ; Artificial intelligence ; LLM audit |
| JEL: | D63 D91 C90 |
| Date: | 2026 |
| URL: | https://d.repec.org/n?u=RePEc:drm:wpaper:2026-16 |