nep-big New Economics Papers
on Big Data
Issue of 2026–07–27
seventeen papers chosen by
Tom Coupé, University of Canterbury


  1. Machine Learning and Liquidity Dynamics in European Stock Markets By Veni Arakelia; Guglielmo Maria Caporale; Mirto M. Gasparinatou; Menelaos Karanasos
  2. Estimating Supply Incrementality in Two-sided Marketplaces: A Causal Machine Learning Approach By Yufei Wu; Daniel Schmierer; Dan Zylberglejd
  3. Predicting Financial Market Stress with Machine Learning By Aldasoro, Inaki; Hördahl, Peter; Schrimpf, Andreas; Zhu, Sonya
  4. Leveraging Non-traditional Data for Macroeconomic Nowcasting: The Case of Morocco By Dina M Hamed
  5. Visual Bias in the Brexit Referendum: A Quantitative Analysis of Newspaper Images By Chung, Wanyu; Dai, Duiyi; Elliott, Robert
  6. Scalable Targeting of Social Protection: When Do Algorithms Out-Perform Surveys and Community Knowledge? By Aiken, Emily; Ashraf, Anik; Blumenstock, Joshua; Guiteras, Raymond; Mobarak, Ahmed
  7. Measuring Industrial Policy: A Text-Based Approach By Juhász, Réka; Lane, Nathaniel; Oehlsen, Emily; Perez, Veronica
  8. Treatment Targeting by Scaled Behavioral Measurement By Kevin Bauer; Andreas Grunewald; Florian Hett; Johanna Jagow; Maximilian Speicher
  9. Public Communication and Collusion: New Screening Tools for Competition Authorities By Duso, Tomaso; Harrington, Jr, Joseph E.; Kreuzberg, Carl; Sapi, Geza
  10. Macroeconomic Forecasting and Machine Learning By Chi, Ta-Chung; Fan, Ting-Han; Ghigliazza, Raffaele; Giannone, Domenico; Wang, Zixuan (Kevin)
  11. Breaking the Echo Chamber: Social Media Networks and Political Conflict By Menéndez, Luis; Montolio, Daniel; Mueller, Hannes; Slataper, Francesco
  12. Management Practices and Firm Performance during the Great Recession By Englmaier, Florian; Galdón Sánchez, José Enrique; Gil, Ricard; Kaiser, Michael; Strandt, Helene
  13. Evaluating predictors of medium-term job growth using machine learning By Sinha, Rishabh
  14. Capturing Heterogeneity: Machine Learning Approaches to Implied Volatility Forecasting By Hyung Joo Kim; Dong Hwan Oh
  15. How economics classifies itself: text-based JEL codes and their consistency By Garau, Alessio
  16. Not-quite-human tastes: the stylized omnivorousness of LLM survey surrogates By Xiangyu Ma; Mengmi Zhang; Shannon Ang; Minne Chen
  17. Macro Shocks and Firm-Level Response Heterogeneity By Davis, Steven; Hansen, Stephen; Seminario-Amez, Cristhian

  1. By: Veni Arakelia; Guglielmo Maria Caporale; Mirto M. Gasparinatou; Menelaos Karanasos
    Abstract: This paper examines the forecasting of liquidity dynamics in European stock markets by means of traditional econometric models and machine learning techniques. It uses daily data for the DAX, CAC 40, FTSE 100, FTSE MIB, and IBEX 35 over 2010–2026, liquidity being measured by the logarithmic Amihud illiquidity indicator. The empirical framework compares ARIMA models and a dynamic panel specification with Random Forest, Extreme Gradient Boosting (XGBoost), and Support Vector Regression (SVR) within a common rolling one-step-ahead forecasting framework. The results show that liquidity is highly persistent and that the dynamic panel model achieves the lowest forecast errors, although Diebold–Mariano tests indicate no significant predictive advantage over the leading machine learning models. SHAP analysis reveals that trading activity, lagged liquidity, and market uncertainty are the main determinants of liquidity forecasts. The findings highlight the complementary role of explainable machine learning in empirical finance.
    Keywords: liquidity dynamics, european stock markets, forecasting, econometric models, machine learning (ML), artificial intelligence (AI)
    JEL: C22 C33 C53 G17
    Date: 2026
    URL: https://d.repec.org/n?u=RePEc:ces:ceswps:_12829
  2. By: Yufei Wu; Daniel Schmierer; Dan Zylberglejd
    Abstract: In two-sided marketplaces with heterogeneous products, it is important to understand the causal relationship between additional supply and marketplace outcomes, such as the total quantity transacted or transaction value in the marketplace. This paper studies a causal machine learning approach to estimating this relationship across product segments. We use the Airbnb marketplace as an example, focusing on the impact of additional listing supply on total bookings, but the methodology applies to other two-sided marketplaces. Our approach combines double/debiased machine learning with a hierarchical Bayesian framework that leverages pre-existing knowledge as priors. We construct tractable and informative features for the model by leveraging measures of product segment similarity from the geospatial literature. We find that such a model provides plausible estimates of the marketplace returns to additional supply and strong out of sample performance.
    Date: 2026–06
    URL: https://d.repec.org/n?u=RePEc:arx:papers:2606.30999
  3. By: Aldasoro, Inaki; Hördahl, Peter; Schrimpf, Andreas; Zhu, Sonya
    Abstract: Using newly constructed market conditions indicators (MCIs) for three pivotal markets centered around the US dollar (Treasury, foreign exchange, and money markets), we demonstrate that tree-based machine learning (ML) models significantly outperform traditional time-series approaches in predicting the full distribution of future market stress. Through quantile regressions, we show that the random forest method achieves up to 27\% lower quantile loss than autoregressive benchmarks, particularly at longer horizons (up to 12 months). Shapley value analysis reveals that variables related to macro expectations and uncertainty — especially about the monetary policy stance — are important predictors of future tail realizations of market conditions. For individual market segments, the state of the global financial cycle, as well as liquidity conditions, also play important roles. These results highlight the value of ML in forecasting tail risks and identifying systemic vulnerabilities in real time, bridging the gap between high-frequency data and macroeconomic stability frameworks.
    Keywords: Shapley value
    JEL: G01 C53 G17 G12 G28
    Date: 2025–07
    URL: https://d.repec.org/n?u=RePEc:cpr:ceprdp:20439
  4. By: Dina M Hamed
    Abstract: Making informed policy decisions is contingent upon the availability of reliable and timely data. The use of non-traditional data has been shown to be a powerful tool for enabling policymakers to conduct robust nowcasting—the practice of estimating the current period’s economic indicator(s), ahead of official releases, using a wide range of macroeconomic and high-frequency data. This paper showcases how different types of non-traditional data, such as indices extracted from satellite imagery, Google Trends, and flight tracking information, can be leveraged to complement official statistics and monitor economic activity, and how these timely signals can be incorporated into nowcasting models to provide early estimates of key macroeconomic variables in Morocco. The approach is applied to agricultural gross value added, tourism revenues, and the unemployment rate. The results demonstrate that non-traditional data substantially improves nowcasting models by enhancing predictive accuracy and enabling the rapid generation of nowcast estimates prior to the release of official data.
    Keywords: Nowcasting; Macroeconomic Forecasting; Non-traditional data; Satellite Imagery; Google Trends; Tourism Revenues; Agriculture GVA; Unemployment Rate; Machine learning; Morocco
    Date: 2026–06–05
    URL: https://d.repec.org/n?u=RePEc:imf:imfwpa:2026/108
  5. By: Chung, Wanyu; Dai, Duiyi; Elliott, Robert
    Abstract: In this paper, we investigate whether, and to what extent, UK newspapers exhibited image-based bias in their portrayal of politicians during the 2016 Brexit referendum, potentially shaping public perceptions. We use computer vision and machine learning techniques to first, identify the faces of politicians and assess the emotional content conveyed through their expressions and second, to measure the contextual sentiment of the overall image, including elements such as background, objects, and color. Our findings reveal that tabloid newspapers displayed significant partisan bias. Specifically, pro-leave tabloids were more likely to depict pro-leave politicians with positive facial expressions and in more favorable visual contexts, while pro-remain politicians were portrayed more negatively. This bias was especially pronounced in front-page images and those featuring key political figures, while no comparable patterns were found in broadsheets. However, the visual bias diminished immediately after the referendum vote. Our scalable framework offers a systematic way to detect visual bias in political imagery, with broader applicability to media coverage of other electoral or policy events.
    Keywords: Visual media bias; image analysis; Machine learning; Brexit referendum
    JEL: L82 D91 D83
    Date: 2025–08
    URL: https://d.repec.org/n?u=RePEc:cpr:ceprdp:20524
  6. By: Aiken, Emily; Ashraf, Anik; Blumenstock, Joshua; Guiteras, Raymond; Mobarak, Ahmed
    Abstract: Innovations in big data and algorithms are enabling new approaches to target interventions at scale. We compare the accuracy of three different systems for identifying the poor to receive benefit transfers — proxy means-testing, nominations from community members, and an algorithmic approach using machine learning to predict poverty using mobile phone usage behavior— and study how their cost-effectiveness varies with the scale and scope of the program. We collect mobile phone records from all major telecom operators in Bangladesh and conduct community-based wealth rankings and detailed consumption surveys of 5, 000 households, to select 22, 000 poorest households for $300 transfers from 106, 000 listed households. While proxy-means testing is most accurate, algorithmic targeting becomes more cost-effective for national-scale programs where large numbers of households have to be screened. We explore the external validity of these insights using survey data and mobile phone records data from Togo, and cross-country information on benefit transfer programs from the World Bank.
    Keywords: Poverty; Development
    JEL: C55 I32 I38
    Date: 2025–06
    URL: https://d.repec.org/n?u=RePEc:cpr:ceprdp:20332
  7. By: Juhász, Réka; Lane, Nathaniel; Oehlsen, Emily; Perez, Veronica
    Abstract: Since the 18th century, policymakers have debated the merits of industrial policy (IP). Yet, economists lack basic facts about its use due to measurement challenges. We propose a new approach to IP measurement based on information contained in policy text. We show how off-the-shelf supervised machine learning tools can be used to categorize industrial policies at scale. Using this approach, we validate longstanding concerns with earlier approaches to measurement which conflate IP with other types of policy. We apply our methodology to a global database of commercial policy descriptions, and provide a first look at IP use at the country, industry, and year levels (2010-2022). The new data on IP suggest that i) IP is on the rise; ii) modern IP tends to use subsidies and export promotion measures as opposed to tariffs; iii) rich countries heavily dominate IP use; iv) IP tends to target sectors with an established comparative advantage, particularly in high-income countries.
    JEL: O25 L52 C38
    Date: 2025–06
    URL: https://d.repec.org/n?u=RePEc:cpr:ceprdp:20333
  8. By: Kevin Bauer; Andreas Grunewald; Florian Hett; Johanna Jagow; Maximilian Speicher
    Abstract: We study how behavioral economics and machine learning can jointly construct effective treatment-targeting rules. In a large field experiment at an online fashion retailer with approximately 500, 000 consumers, we test a loss-framed discount message. We elicit individual loss aversion in a nested incentivized behavioral measurement experiment (N=582) and use machine learning to impute it from digital footprints. Targeting based on scaled behavioral measurement yields statistically significant revenue gains and outperforms causal forests. The results show how scaling behavioral measurement can improve algorithmic treatment assignment relative to purely data-driven approaches, especially when pilot data are unavailable, noisy, or costly.
    Keywords: treatment targeting, behavioral measurement, machine learning
    JEL: C93 C55 D91 M31 L81
    Date: 2026
    URL: https://d.repec.org/n?u=RePEc:ces:ceswps:_12772
  9. By: Duso, Tomaso; Harrington, Jr, Joseph E.; Kreuzberg, Carl; Sapi, Geza
    Abstract: Competition authorities increasingly rely on economic screening tools to identify markets where firms deviate from competitive norms. Traditional screening methods assume that collusion occurs through secret agreements. However, recent research highlights that firms can use public announcements to coordinate decisions, reducing competition while avoiding detection. We propose a novel approach to screening for collusion in public corporate statements. Using natural language processing, we analyze more than 300, 000 earnings call transcripts issued worldwide between 2004 and 2022. By identifying expressions commonly associated with collusion, our method provides competition authorities with a tool to detect potentially anticompetitive behavior in public communications. Our approach can extend beyond earnings calls to other sources, such as news articles, trade press, and industry reports. Our method informed the European Commission’s 2024 unannounced inspections in the car tire sector, prompted by concerns over price coordination through public communication.
    Keywords: Communication; Collusion; Screening
    JEL: C23 D22 L1 L4 L64
    Date: 2025–07
    URL: https://d.repec.org/n?u=RePEc:cpr:ceprdp:20490
  10. By: Chi, Ta-Chung; Fan, Ting-Han; Ghigliazza, Raffaele; Giannone, Domenico; Wang, Zixuan (Kevin)
    Abstract: We forecast the full conditional distribution of macroeconomic outcomes by systematically integrating three key principles: using high-dimensional data with appropriate regularization, adopting rigorous out-of-sample validation procedures, and incorporating nonlinearities. By exploiting the rich information embedded in a large set of macroeconomic and financial predictors, we produce accurate predictions of the entire profile of macroeconomic risk in real time. Our findings show that regularization via shrinkage is essential to control model complexity, while introducing nonlinearities yields limited improvements in predictive accuracy. Out-of-sample validation plays a critical role in selecting model architecture and preventing overfitting.
    Keywords: Regularization
    JEL: C22 C52 C53 C55
    Date: 2025–10
    URL: https://d.repec.org/n?u=RePEc:cpr:ceprdp:20727
  11. By: Menéndez, Luis; Montolio, Daniel; Mueller, Hannes; Slataper, Francesco
    Abstract: This article exploits data from a political conflict between language groups to show how political events can rapidly redefine how these groups interact on social media. Leveraging on a unique dataset of 26 million retweets by 120 000 Catalan- and Spanish-speaking Twitter users, we estimate individual exposure to tweets with a network-based model. We then compare two shocks in the same region and year: the Barcelona terror attack and the Catalan independence referendum of 2017. The referendum, and related police violence, triggered a sharp, symmetric jump in retweeting across language groups. The terror attack, by contrast, did not lead to a similar realignment.
    Keywords: Social Networks; Polarization
    JEL: D74 C55 C45
    Date: 2025–08
    URL: https://d.repec.org/n?u=RePEc:cpr:ceprdp:20559
  12. By: Englmaier, Florian; Galdón Sánchez, José Enrique; Gil, Ricard; Kaiser, Michael; Strandt, Helene
    Abstract: This paper examines how management practices affect firm productivity over the business cycle. Using Spanish plant-level survey data and unsupervised machine learning, we identify a “structured†management style positively correlated with performance before the 2008 financial crisis. Interestingly, this correlation turns negative during the crisis and positive again in the post-2013 recovery. Our evidence suggests structured firms focus on long-run profitability and innovation, prioritizing intangible investments. This strategy leads to higher short-run adjustment costs, evidenced by more fixed assets and lower employee turnover, making them less resilient during a severe downturn.
    Keywords: Culture; Productivity
    JEL: M12 D22 C38
    Date: 2025–10
    URL: https://d.repec.org/n?u=RePEc:cpr:ceprdp:20749
  13. By: Sinha, Rishabh
    Abstract: Using data from 116 countries spanning 2005-2023, this paper provides a global comparative assessment of the factors predicting medium-term job growth. I assemble 25 macroeconomic, demographic, health, and institutional variables and employ tuned machine-learning models to evaluate their predictive power. Gradient Boosting Regression delivers the strongest performance, and Shapley decompositions reveal three main results. First, GDP growth is not the dominant predictor of employment expansion, and its predictive strength has weakened over time, except during global recoveries. Second, past job growth is one of the most consistent predictors across income groups, indicating strong labor-market momentum, especially following global macroeconomic instability. Third, some structural forces, such as urbanization and quality of governance, contribute to job growth. But their effects shift with global macroeconomic conditions and a country's level of development. Overall, the evidence underscores the role of these contextual factors, suggesting that economic growth is often insufficient to sustain job creation.
    Keywords: Job growth; Economic growth; Medium-term forecasts of job creation; Machine learning; SHAP; Structural determinants of job growth
    JEL: J21 O15 O47
    Date: 2025–11
    URL: https://d.repec.org/n?u=RePEc:pra:mprapa:128251
  14. By: Hyung Joo Kim; Dong Hwan Oh
    Abstract: Despite documented heterogeneity in volatility dynamics across the option surface, standard implied volatility forecasting models apply homogeneous parameters throughout. We introduce a machine-learning framework that uses regression trees to partition the surface along both moneyness and maturity dimensions, identifying data-driven regions where distinct forecasting models perform best. Extending the Surface Heterogeneous Autoregressive (SHAR) framework of Dufays, Jacobs, and Rombouts (2025), we develop tree-based SHAR specifications that preserve interpretable structure while allowing model parameters to vary across the surface. Empirical analysis using S&P 500 options demonstrates that the boosted tree-based specification achieves the lowest out-of-sample forecast errors across all horizons, reducing one-month-ahead RMSE by 13 percent versus the benchmark SHAR model. The improvements are statistically significant and particularly pronounced during stress periods. The estimated tree presents economically interpretable segmentation: short-dated options exhibit higher daily persistence but lower monthly persistence than long-dated options, while deep out-of-the-money calls or puts display distinct dynamics from near-the-money contracts.
    Keywords: implied volatility forecasting; option surface; machine learning; regression trees; ensemble methods; heterogeneous autoregressive models
    JEL: C14 C22 C32 C51 C53 C58 G12
    Date: 2026–07–06
    URL: https://d.repec.org/n?u=RePEc:fip:fedgfe:103519
  15. By: Garau, Alessio
    Abstract: Can a language model improve how economists classify their own papers? Only 15% of four million IDEAS/RePEc records carry usable JEL codes, and similar papers often receive different ones. I use a large language model (LLM) to solve this problem and assign three-digit codes from titles and abstracts, evaluating it on 69, 503 coded articles published from 1991 to 2023 in the top 100 economics journals. Two tests do not assume that author codes provide the correct classification. Across semantic neighbors identified by a separate embedding model, model codes are 1.8 times as consistent as author codes. Holding codes per paper fixed, a blind check finds that 83% of model codes fit official American Economic Association (AEA) guidelines, compared with 67% of author codes. The classifier expands coverage, and the evaluation framework applies whenever human labels are incomplete or noisy.
    Keywords: JEL codes; field classification; generative artificial intelligence; large language models; text as data; semantic similarity
    JEL: A14 C45 C81
    Date: 2026–07
    URL: https://d.repec.org/n?u=RePEc:pra:mprapa:130163
  16. By: Xiangyu Ma; Mengmi Zhang; Shannon Ang; Minne Chen
    Abstract: Large-language models have proven to be remarkable if inconsistent parrots of public attitudes and opinions. The extent to which LLMs are able to produce reasonable approximations of cultural taste remains an open empirical question that becomes more urgent by the day, with market research companies already offering provisional `synthetic' survey panels and the contamination of standard survey data from LLM-generated responses. In this study, we build on past work on silicon sampling by extending considerations of its algorithmic fidelity and alignment to the domain of cultural consumption. We use large-language models from OpenAI, Anthropic, and DeepSeek to each produce 277, 470 (30x9249) silicon surrogates of survey respondents from the Survey of Public Participation in the Arts (SPPA). We find these silicon surrogates' tastes to be highly stylized facsimiles of human tastes. (1) Silicon samples have a systematic postive-bias for liking, resulting in inflated ecological estimates of tastes. The individual-level bias of silicon samples are not well-explained by the WEIRD-bias often discussed in the literature. (2) The complex relationality in real taste structures is completely lost among silicon samples. (3) Finally, very little of the known cultural alignment between tastes and social space are preserved. Silicon samples attenuate age-taste associations, resurrect anachronistic class-taste associations, caricaturize gender- and race-taste associations.
    Date: 2026–06
    URL: https://d.repec.org/n?u=RePEc:arx:papers:2606.30085
  17. By: Davis, Steven; Hansen, Stephen; Seminario-Amez, Cristhian
    Abstract: Macro shocks produce high dispersion in firm-level equity returns, sales growth, and other outcomes. We show that this dispersion reflects observable differences in business characteristics. To do so, we combine firm-level returns on stock market ``jump" days with text about business risks in prior 10-K filings to construct firm-specific shock exposures. Our exposure measures explain firm-level abnormal returns through interpretable variation in language. They also explain most of the increased dispersion in firm-level revenue growth after major shocks and much of the dispersion in employment growth, investment rates, and earnings surprises. Our evidence yields a novel interpretation for countercyclical dispersion, highlighting the key role of heterogeneous business characteristics in macro shock transmission.
    Keywords: Firm heterogeneity; Text data; Machine learning; Shock identification
    JEL: C55 E30 L20
    Date: 2025–06
    URL: https://d.repec.org/n?u=RePEc:cpr:ceprdp:20358

This nep-big issue is ©2026 by Tom Coupé. It is provided as is without any express or implied warranty. It may be freely redistributed in whole or in part for any purpose. If distributed in part, please include this notice.
General information on the NEP project can be found at https://nep.repec.org. For comments please write to the director of NEP, Marco Novarese at <director@nep.repec.org>. Put “NEP” in the subject, otherwise your mail may be rejected.
NEP’s infrastructure is sponsored by the Griffith Business School of Griffith University in Australia.