nep-big New Economics Papers
on Big Data
Issue of 2026–08–31
seven papers chosen by
Tom Coupé, University of Canterbury


  1. Machine Learning for Estimating Catastrophic Health Spending in Disaster-Affected, Data-Scarce Settings By Himaz, Rozana; Salmanidou, Dimitra; Ghaffarian, Saman
  2. Predicting Startup Exit from Textual Descriptors - A Computational Linguistics Framework By Alberto M. G. Saruggia; Sebastien Germano
  3. Predicting Retirement and Social Security Claiming Decisions using Machine Learning By Kwon, Alexander; Maliar, Lilia
  4. Defining Current and Expected Financial Constraints using AI: Reinterpreting the Cash Flow Sensitivity of Cash By Cho, Rachel; Görtz, Christoph; McGowan, Danny; Schröder, Max
  5. Yield Curve Prediction with Machine Learning: Forecasting Approaches and the Role of Macroeconomic Predictors By Jeron Tan Kang
  6. Tracking Weekly Activity using New Data Sources By De Polis, Andrea; Galvão, Ana Beatriz; Petrella, Ivan
  7. When the Crowd Speaks: AI-based retail investor sentiment indicator with Reddit data By Wagner, Fabian; Metzler, Julian

  1. By: Himaz, Rozana; Salmanidou, Dimitra; Ghaffarian, Saman
    Abstract: Natural hazard events can increase out-of-pocket health costs and push vulnerable households into poverty. Mitigation measures require understanding changes in health spending patterns using pre- and post-event data, but such data are often unavailable in disaster-affected settings. This represents a fundamental measurement challenge: the absence of pre-event baseline data makes it impossible to construct the counterfactual quantities needed for welfare analysis. To address this measurement problem, we develop a hybrid machine learning approach to estimate unobserved household health spending using longitudinal survey data from Indonesia. We first develop a model around the 2006 Yogyakarta earthquake, for which complete data are available. The model learns spending patterns across income, hazard intensity, and other characteristics, achieving >70% accuracy in a noisy and complex domain. After testing the model for transportability, we apply it to post-2004 Indian Ocean tsunami survey data in Indonesia, to predict plausible baseline health spending. These predictions are used to evaluate the impact of the tsunami on health spending to reveal that without targeted aid, catastrophic health spending would have increased from 4.5% to 29.4% and that moderately damaged households experienced more cost increases than heavily damaged ones. By combining artificial intelligence with 2 household survey data, our framework is a proof-of-concept, for addressing data gaps in official economic statistics, demonstrating how machine learning can enable counterfactual welfare measurement where conventional data collection is absent or incomplete.
    Keywords: Natural hazards; catastrophic health spending; disaster risk reduction; tsunami; earthquake; Indonesia; machine learning
    JEL: C45 C51 C52 C53 I19 O13 Q54
    Date: 2026–03–23
    URL: https://d.repec.org/n?u=RePEc:eoe:escoed:escoe-dp-2026-05
  2. By: Alberto M. G. Saruggia; Sebastien Germano
    Abstract: This study shows that textual descriptors alone can predict early-stage startup success, defined as Exit, without relying on contextual, financial, or human capital variables. Using venture capital-curated datasets covering 7, 419 startups over 20 years, the research isolates text-based framing variables and engineers 850 features through startup narrative mapping. Data subsets and vector embeddings are evaluated for statistical significance, followed by supervised machine learning experiments across six models. LightGBM achieved the highest predictive performance (F1 = 0.48), while textual descriptors alone achieved F1 = 0.30, confirming the standalone predictive value of founder narratives. Feature analysis shows that optimized densities of hyping markers, including adjectives, jargon, and buzzwords, are associated with higher Exit probability, whereas excessive statement or name length reduces it. The study also introduces a quantifiable Hyping Score for venture capital applications, demonstrating that startup framing provides measurable signals for predicting Exit under conditions of high information asymmetry.
    Date: 2026–07
    URL: https://d.repec.org/n?u=RePEc:arx:papers:2608.00045
  3. By: Kwon, Alexander; Maliar, Lilia
    Abstract: We demonstrate that machine learning substantially improves predictions of individual decisions about retirement and Social Security (SS) claims. When predicting the number of people receiving SS, we achieve an error of less than 1%, while the benchmark model employed by the Social Security Administration (SSA) results in a greater than 4% error, and in forecasting SS claiming decisions, we attain an error of 0.2%, while the benchmark exceeding 2%. Based on averages, we show that a 3% difference in prediction amounts to 39.6 billion dollars annually. The set of important variables selected by our model significantly differs from that of the SSA model. We use Shapley values to evaluate the non-linear contributions of the selected variables to predictive outcomes.
    Keywords: retirement and Social Security
    JEL: C53 H55 J14 J26
    Date: 2024–07
    URL: https://d.repec.org/n?u=RePEc:cpr:ceprdp:19198
  4. By: Cho, Rachel; Görtz, Christoph; McGowan, Danny; Schröder, Max
    Abstract: We propose a new approach to identify firm-level financial constraints by applying artificial intelligence to text of 10-K filings by U.S. public firms from 1993 to 2021. Leveraging transformer-based natural language processing, our model captures contextual and semantic nuances often missed by traditional text classification techniques, enabling more accurate detection of financial constraints. A key contribution is to differentiate between constraints that affect firms presently and those anticipated in the future. These two types of constraints are associated with distinctly different financial profiles: while firms expecting future constraints tend to accumulate cash preemptively, currently constrained firms exhibit reduced liquidity and higher leverage. We show that only firms anticipating financial constraints exhibit significant cash flow sensitivity of cash, whereas currently constrained and unconstrained firms do not. This calls for a narrower interpretation of this widely used cash-based constraints measure, as it may conflate distinct firm types – unconstrained and currently constrained – and fail to capture all financially constrained firms. Our findings underscore the critical role of constraint timing in shaping corporate financial behavior.
    Keywords: Financial Constraints; Artificial Intelligence; Expectations; Cash; Cash Flow; Corporate Finance Behavior
    JEL: D92 G31 G32
    Date: 2025–09–18
    URL: https://d.repec.org/n?u=RePEc:eoe:escoed:escoe-dp-2025-11
  5. By: Jeron Tan Kang
    Abstract: This paper compares direct-yield and factor-based approaches to U.S. Treasury yield curve forecasting using a common high-dimensional macroeconomic information set. Forecasts are evaluated on monthly zero-coupon yields over the 2015-2025 out-of-sample period. Gains over the random walk are concentrated at short maturities and in slope forecasts, and decline with the forecast horizon. Direct-yield models perform best for slope forecasts and are relatively stronger at short horizons, while factor-based models become more competitive at longer horizons. Macroeconomic predictors provide clear incremental predictive power, strongest for slope-related movements. A trading simulation reinforces that macro-augmented models perform best in slope trades. The simulation also highlights a gap between statistical and economic performance, as the random walk is a strong benchmark under statistical loss but performs poorly as a trading signal.
    Date: 2026–07
    URL: https://d.repec.org/n?u=RePEc:arx:papers:2608.07536
  6. By: De Polis, Andrea; Galvão, Ana Beatriz; Petrella, Ivan
    Abstract: Policymakers increasingly rely on real-time measures of economic activity to inform decisions, yet official statistics are typically available only with substantial delay. This paper develops a methodology to extract information from high-frequency indicators in order to produce weekly estimates of official monthly statistics in real time. Building on real-time tracking and nowcasting models, our approach addresses two challenges inherent in alternative indicators: the absence of seasonal adjustment and the prevalence of outliers. We incorporate seasonal components directly into the model, allowing the seasonal structure of low-frequency data to inform high frequency proxies, and employ fat-tailed distributions to mitigate the influence of large, infrequent shocks. Applying our methodology to UK data, we track retail sales, monthly GDP, and vacancies using proxies such as debit card spending (Revolut) and online job advertisements. Our results show that weekly estimates improve real-time prediction of official releases and highlight the usefulness of high-frequency alternative data.
    Keywords: real-time tracker; mixed-frequency models; seasonality; state-space models; high-frequency data
    JEL: C32 C53 E01 E37
    Date: 2025–11–28
    URL: https://d.repec.org/n?u=RePEc:eoe:escoed:escoe-dp-2025-19
  7. By: Wagner, Fabian; Metzler, Julian
    Abstract: Retail investors increasingly discuss markets in real time on social media, yet these discussions remain difficult to measure systematically. This paper introduces the Reddit Retail Investor Sentiment Indicator (R-RISI), a high-frequency measure of retail investor sentiment based on Reddit posts from major investing and crypto-related subreddits between 2015 and April 2026. Using ChatGPT 5.1, we classify posts by sentiment and topic and show that large language models provide more accurate labels for informal social media language than widely used finance-specific models such as FinBERT, which tends to over-assign neutral sentiment. R-RISI is constructed by aggregating daily positive and negative posts, weighting them by user engagement and standardising the resulting series using a rolling one-year z-score methodology. The indicator can be flexibly built for broad asset classes, portfolios, individual securities or thematic groups, allowing sentiment to be tracked at a much higher granularity than traditional retail investor indicators. We show how sentiment and dominant discussion topics evolve over time within each of the extracted subreddits. R-RISI closely reflects major market developments and aligns with established sentiment indicators while providing higher frequency and more timely signals. Empirically, changes in R-RISI contain statistically significant short-term information for market prices. This relationship is particularly relevant during periods of market stress: changes in R-RISI matter more for next-day S&P 500 returns when markets are in decline. The indicator therefore offers a timely and granular tool for analysing retail investor behaviour and monitoring sentiment-driven market dynamics in an environment of growing retail investor participation. A regularly updated version of the indicator is available through an accompanying R-RISI online dashboard. JEL Classification: F1, G1, G4, G5
    Keywords: AI/LLM-based text classification, alternative data, behavioural finance, Reddit, retail investors, sentiment analysis, social media
    Date: 2026–08
    URL: https://d.repec.org/n?u=RePEc:ecb:ecbwps:20263276

This nep-big issue is ©2026 by Tom Coupé. It is provided as is without any express or implied warranty. It may be freely redistributed in whole or in part for any purpose. If distributed in part, please include this notice.
General information on the NEP project can be found at https://nep.repec.org. For comments please write to the director of NEP, Marco Novarese at <director@nep.repec.org>. Put “NEP” in the subject, otherwise your mail may be rejected.
NEP’s infrastructure is sponsored by the Griffith Business School of Griffith University in Australia.