|
on Big Data |
| By: | Veni Arakelia; Guglielmo Maria Caporale; Mirto M. Gasparinatou; Menelaos Karanasos |
| Abstract: | This paper examines the forecasting of liquidity dynamics in European stock markets by means of traditional econometric models and machine learning techniques. It uses daily data for the DAX, CAC 40, FTSE 100, FTSE MIB, and IBEX 35 over 2010–2026, liquidity being measured by the logarithmic Amihud illiquidity indicator. The empirical framework compares ARIMA models and a dynamic panel specification with Random Forest, Extreme Gradient Boosting (XGBoost), and Support Vector Regression (SVR) within a common rolling one-step-ahead forecasting framework. The results show that liquidity is highly persistent and that the dynamic panel model achieves the lowest forecast errors, although Diebold–Mariano tests indicate no significant predictive advantage over the leading machine learning models. SHAP analysis reveals that trading activity, lagged liquidity, and market uncertainty are the main determinants of liquidity forecasts. The findings highlight the complementary role of explainable machine learning in empirical finance. |
| Keywords: | liquidity dynamics, european stock markets, forecasting, econometric models, machine learning (ML), artificial intelligence (AI) |
| JEL: | C22 C33 C53 G17 |
| Date: | 2026 |
| URL: | https://d.repec.org/n?u=RePEc:ces:ceswps:_12829 |
| By: | Yufei Wu; Daniel Schmierer; Dan Zylberglejd |
| Abstract: | In two-sided marketplaces with heterogeneous products, it is important to understand the causal relationship between additional supply and marketplace outcomes, such as the total quantity transacted or transaction value in the marketplace. This paper studies a causal machine learning approach to estimating this relationship across product segments. We use the Airbnb marketplace as an example, focusing on the impact of additional listing supply on total bookings, but the methodology applies to other two-sided marketplaces. Our approach combines double/debiased machine learning with a hierarchical Bayesian framework that leverages pre-existing knowledge as priors. We construct tractable and informative features for the model by leveraging measures of product segment similarity from the geospatial literature. We find that such a model provides plausible estimates of the marketplace returns to additional supply and strong out of sample performance. |
| Date: | 2026–06 |
| URL: | https://d.repec.org/n?u=RePEc:arx:papers:2606.30999 |
| By: | Aldasoro, Inaki; Hördahl, Peter; Schrimpf, Andreas; Zhu, Sonya |
| Abstract: | Using newly constructed market conditions indicators (MCIs) for three pivotal markets centered around the US dollar (Treasury, foreign exchange, and money markets), we demonstrate that tree-based machine learning (ML) models significantly outperform traditional time-series approaches in predicting the full distribution of future market stress. Through quantile regressions, we show that the random forest method achieves up to 27\% lower quantile loss than autoregressive benchmarks, particularly at longer horizons (up to 12 months). Shapley value analysis reveals that variables related to macro expectations and uncertainty — especially about the monetary policy stance — are important predictors of future tail realizations of market conditions. For individual market segments, the state of the global financial cycle, as well as liquidity conditions, also play important roles. These results highlight the value of ML in forecasting tail risks and identifying systemic vulnerabilities in real time, bridging the gap between high-frequency data and macroeconomic stability frameworks. |
| Keywords: | Shapley value |
| JEL: | G01 C53 G17 G12 G28 |
| Date: | 2025–07 |
| URL: | https://d.repec.org/n?u=RePEc:cpr:ceprdp:20439 |
| By: | Dina M Hamed |
| Abstract: | Making informed policy decisions is contingent upon the availability of reliable and timely data. The use of non-traditional data has been shown to be a powerful tool for enabling policymakers to conduct robust nowcasting—the practice of estimating the current period’s economic indicator(s), ahead of official releases, using a wide range of macroeconomic and high-frequency data. This paper showcases how different types of non-traditional data, such as indices extracted from satellite imagery, Google Trends, and flight tracking information, can be leveraged to complement official statistics and monitor economic activity, and how these timely signals can be incorporated into nowcasting models to provide early estimates of key macroeconomic variables in Morocco. The approach is applied to agricultural gross value added, tourism revenues, and the unemployment rate. The results demonstrate that non-traditional data substantially improves nowcasting models by enhancing predictive accuracy and enabling the rapid generation of nowcast estimates prior to the release of official data. |
| Keywords: | Nowcasting; Macroeconomic Forecasting; Non-traditional data; Satellite Imagery; Google Trends; Tourism Revenues; Agriculture GVA; Unemployment Rate; Machine learning; Morocco |
| Date: | 2026–06–05 |
| URL: | https://d.repec.org/n?u=RePEc:imf:imfwpa:2026/108 |
| By: | Chung, Wanyu; Dai, Duiyi; Elliott, Robert |
| Abstract: | In this paper, we investigate whether, and to what extent, UK newspapers exhibited image-based bias in their portrayal of politicians during the 2016 Brexit referendum, potentially shaping public perceptions. We use computer vision and machine learning techniques to first, identify the faces of politicians and assess the emotional content conveyed through their expressions and second, to measure the contextual sentiment of the overall image, including elements such as background, objects, and color. Our findings reveal that tabloid newspapers displayed significant partisan bias. Specifically, pro-leave tabloids were more likely to depict pro-leave politicians with positive facial expressions and in more favorable visual contexts, while pro-remain politicians were portrayed more negatively. This bias was especially pronounced in front-page images and those featuring key political figures, while no comparable patterns were found in broadsheets. However, the visual bias diminished immediately after the referendum vote. Our scalable framework offers a systematic way to detect visual bias in political imagery, with broader applicability to media coverage of other electoral or policy events. |
| Keywords: | Visual media bias; image analysis; Machine learning; Brexit referendum |
| JEL: | L82 D91 D83 |
| Date: | 2025–08 |
| URL: | https://d.repec.org/n?u=RePEc:cpr:ceprdp:20524 |
| By: | Aiken, Emily; Ashraf, Anik; Blumenstock, Joshua; Guiteras, Raymond; Mobarak, Ahmed |
| Abstract: | Innovations in big data and algorithms are enabling new approaches to target interventions at scale. We compare the accuracy of three different systems for identifying the poor to receive benefit transfers — proxy means-testing, nominations from community members, and an algorithmic approach using machine learning to predict poverty using mobile phone usage behavior— and study how their cost-effectiveness varies with the scale and scope of the program. We collect mobile phone records from all major telecom operators in Bangladesh and conduct community-based wealth rankings and detailed consumption surveys of 5, 000 households, to select 22, 000 poorest households for $300 transfers from 106, 000 listed households. While proxy-means testing is most accurate, algorithmic targeting becomes more cost-effective for national-scale programs where large numbers of households have to be screened. We explore the external validity of these insights using survey data and mobile phone records data from Togo, and cross-country information on benefit transfer programs from the World Bank. |
| Keywords: | Poverty; Development |
| JEL: | C55 I32 I38 |
| Date: | 2025–06 |
| URL: | https://d.repec.org/n?u=RePEc:cpr:ceprdp:20332 |
| By: | Juhász, Réka; Lane, Nathaniel; Oehlsen, Emily; Perez, Veronica |
| Abstract: | Since the 18th century, policymakers have debated the merits of industrial policy (IP). Yet, economists lack basic facts about its use due to measurement challenges. We propose a new approach to IP measurement based on information contained in policy text. We show how off-the-shelf supervised machine learning tools can be used to categorize industrial policies at scale. Using this approach, we validate longstanding concerns with earlier approaches to measurement which conflate IP with other types of policy. We apply our methodology to a global database of commercial policy descriptions, and provide a first look at IP use at the country, industry, and year levels (2010-2022). The new data on IP suggest that i) IP is on the rise; ii) modern IP tends to use subsidies and export promotion measures as opposed to tariffs; iii) rich countries heavily dominate IP use; iv) IP tends to target sectors with an established comparative advantage, particularly in high-income countries. |
| JEL: | O25 L52 C38 |
| Date: | 2025–06 |
| URL: | https://d.repec.org/n?u=RePEc:cpr:ceprdp:20333 |
| By: | Kevin Bauer; Andreas Grunewald; Florian Hett; Johanna Jagow; Maximilian Speicher |
| Abstract: | We study how behavioral economics and machine learning can jointly construct effective treatment-targeting rules. In a large field experiment at an online fashion retailer with approximately 500, 000 consumers, we test a loss-framed discount message. We elicit individual loss aversion in a nested incentivized behavioral measurement experiment (N=582) and use machine learning to impute it from digital footprints. Targeting based on scaled behavioral measurement yields statistically significant revenue gains and outperforms causal forests. The results show how scaling behavioral measurement can improve algorithmic treatment assignment relative to purely data-driven approaches, especially when pilot data are unavailable, noisy, or costly. |
| Keywords: | treatment targeting, behavioral measurement, machine learning |
| JEL: | C93 C55 D91 M31 L81 |
| Date: | 2026 |
| URL: | https://d.repec.org/n?u=RePEc:ces:ceswps:_12772 |
| By: | Duso, Tomaso; Harrington, Jr, Joseph E.; Kreuzberg, Carl; Sapi, Geza |
| Abstract: | Competition authorities increasingly rely on economic screening tools to identify markets where firms deviate from competitive norms. Traditional screening methods assume that collusion occurs through secret agreements. However, recent research highlights that firms can use public announcements to coordinate decisions, reducing competition while avoiding detection. We propose a novel approach to screening for collusion in public corporate statements. Using natural language processing, we analyze more than 300, 000 earnings call transcripts issued worldwide between 2004 and 2022. By identifying expressions commonly associated with collusion, our method provides competition authorities with a tool to detect potentially anticompetitive behavior in public communications. Our approach can extend beyond earnings calls to other sources, such as news articles, trade press, and industry reports. Our method informed the European Commission’s 2024 unannounced inspections in the car tire sector, prompted by concerns over price coordination through public communication. |
| Keywords: | Communication; Collusion; Screening |
| JEL: | C23 D22 L1 L4 L64 |
| Date: | 2025–07 |
| URL: | https://d.repec.org/n?u=RePEc:cpr:ceprdp:20490 |
| By: | Chi, Ta-Chung; Fan, Ting-Han; Ghigliazza, Raffaele; Giannone, Domenico; Wang, Zixuan (Kevin) |
| Abstract: | We forecast the full conditional distribution of macroeconomic outcomes by systematically integrating three key principles: using high-dimensional data with appropriate regularization, adopting rigorous out-of-sample validation procedures, and incorporating nonlinearities. By exploiting the rich information embedded in a large set of macroeconomic and financial predictors, we produce accurate predictions of the entire profile of macroeconomic risk in real time. Our findings show that regularization via shrinkage is essential to control model complexity, while introducing nonlinearities yields limited improvements in predictive accuracy. Out-of-sample validation plays a critical role in selecting model architecture and preventing overfitting. |
| Keywords: | Regularization |
| JEL: | C22 C52 C53 C55 |
| Date: | 2025–10 |
| URL: | https://d.repec.org/n?u=RePEc:cpr:ceprdp:20727 |
| By: | Menéndez, Luis; Montolio, Daniel; Mueller, Hannes; Slataper, Francesco |
| Abstract: | This article exploits data from a political conflict between language groups to show how political events can rapidly redefine how these groups interact on social media. Leveraging on a unique dataset of 26 million retweets by 120 000 Catalan- and Spanish-speaking Twitter users, we estimate individual exposure to tweets with a network-based model. We then compare two shocks in the same region and year: the Barcelona terror attack and the Catalan independence referendum of 2017. The referendum, and related police violence, triggered a sharp, symmetric jump in retweeting across language groups. The terror attack, by contrast, did not lead to a similar realignment. |
| Keywords: | Social Networks; Polarization |
| JEL: | D74 C55 C45 |
| Date: | 2025–08 |
| URL: | https://d.repec.org/n?u=RePEc:cpr:ceprdp:20559 |
| By: | Englmaier, Florian; Galdón Sánchez, José Enrique; Gil, Ricard; Kaiser, Michael; Strandt, Helene |
| Abstract: | This paper examines how management practices affect firm productivity over the business cycle. Using Spanish plant-level survey data and unsupervised machine learning, we identify a “structured†management style positively correlated with performance before the 2008 financial crisis. Interestingly, this correlation turns negative during the crisis and positive again in the post-2013 recovery. Our evidence suggests structured firms focus on long-run profitability and innovation, prioritizing intangible investments. This strategy leads to higher short-run adjustment costs, evidenced by more fixed assets and lower employee turnover, making them less resilient during a severe downturn. |
| Keywords: | Culture; Productivity |
| JEL: | M12 D22 C38 |
| Date: | 2025–10 |
| URL: | https://d.repec.org/n?u=RePEc:cpr:ceprdp:20749 |
| By: | Sinha, Rishabh |
| Abstract: | Using data from 116 countries spanning 2005-2023, this paper provides a global comparative assessment of the factors predicting medium-term job growth. I assemble 25 macroeconomic, demographic, health, and institutional variables and employ tuned machine-learning models to evaluate their predictive power. Gradient Boosting Regression delivers the strongest performance, and Shapley decompositions reveal three main results. First, GDP growth is not the dominant predictor of employment expansion, and its predictive strength has weakened over time, except during global recoveries. Second, past job growth is one of the most consistent predictors across income groups, indicating strong labor-market momentum, especially following global macroeconomic instability. Third, some structural forces, such as urbanization and quality of governance, contribute to job growth. But their effects shift with global macroeconomic conditions and a country's level of development. Overall, the evidence underscores the role of these contextual factors, suggesting that economic growth is often insufficient to sustain job creation. |
| Keywords: | Job growth; Economic growth; Medium-term forecasts of job creation; Machine learning; SHAP; Structural determinants of job growth |
| JEL: | J21 O15 O47 |
| Date: | 2025–11 |
| URL: | https://d.repec.org/n?u=RePEc:pra:mprapa:128251 |
| By: | Hyung Joo Kim; Dong Hwan Oh |
| Abstract: | Despite documented heterogeneity in volatility dynamics across the option surface, standard implied volatility forecasting models apply homogeneous parameters throughout. We introduce a machine-learning framework that uses regression trees to partition the surface along both moneyness and maturity dimensions, identifying data-driven regions where distinct forecasting models perform best. Extending the Surface Heterogeneous Autoregressive (SHAR) framework of Dufays, Jacobs, and Rombouts (2025), we develop tree-based SHAR specifications that preserve interpretable structure while allowing model parameters to vary across the surface. Empirical analysis using S&P 500 options demonstrates that the boosted tree-based specification achieves the lowest out-of-sample forecast errors across all horizons, reducing one-month-ahead RMSE by 13 percent versus the benchmark SHAR model. The improvements are statistically significant and particularly pronounced during stress periods. The estimated tree presents economically interpretable segmentation: short-dated options exhibit higher daily persistence but lower monthly persistence than long-dated options, while deep out-of-the-money calls or puts display distinct dynamics from near-the-money contracts. |
| Keywords: | implied volatility forecasting; option surface; machine learning; regression trees; ensemble methods; heterogeneous autoregressive models |
| JEL: | C14 C22 C32 C51 C53 C58 G12 |
| Date: | 2026–07–06 |
| URL: | https://d.repec.org/n?u=RePEc:fip:fedgfe:103519 |
| By: | Garau, Alessio |
| Abstract: | Can a language model improve how economists classify their own papers? Only 15% of four million IDEAS/RePEc records carry usable JEL codes, and similar papers often receive different ones. I use a large language model (LLM) to solve this problem and assign three-digit codes from titles and abstracts, evaluating it on 69, 503 coded articles published from 1991 to 2023 in the top 100 economics journals. Two tests do not assume that author codes provide the correct classification. Across semantic neighbors identified by a separate embedding model, model codes are 1.8 times as consistent as author codes. Holding codes per paper fixed, a blind check finds that 83% of model codes fit official American Economic Association (AEA) guidelines, compared with 67% of author codes. The classifier expands coverage, and the evaluation framework applies whenever human labels are incomplete or noisy. |
| Keywords: | JEL codes; field classification; generative artificial intelligence; large language models; text as data; semantic similarity |
| JEL: | A14 C45 C81 |
| Date: | 2026–07 |
| URL: | https://d.repec.org/n?u=RePEc:pra:mprapa:130163 |
| By: | Xiangyu Ma; Mengmi Zhang; Shannon Ang; Minne Chen |
| Abstract: | Large-language models have proven to be remarkable if inconsistent parrots of public attitudes and opinions. The extent to which LLMs are able to produce reasonable approximations of cultural taste remains an open empirical question that becomes more urgent by the day, with market research companies already offering provisional `synthetic' survey panels and the contamination of standard survey data from LLM-generated responses. In this study, we build on past work on silicon sampling by extending considerations of its algorithmic fidelity and alignment to the domain of cultural consumption. We use large-language models from OpenAI, Anthropic, and DeepSeek to each produce 277, 470 (30x9249) silicon surrogates of survey respondents from the Survey of Public Participation in the Arts (SPPA). We find these silicon surrogates' tastes to be highly stylized facsimiles of human tastes. (1) Silicon samples have a systematic postive-bias for liking, resulting in inflated ecological estimates of tastes. The individual-level bias of silicon samples are not well-explained by the WEIRD-bias often discussed in the literature. (2) The complex relationality in real taste structures is completely lost among silicon samples. (3) Finally, very little of the known cultural alignment between tastes and social space are preserved. Silicon samples attenuate age-taste associations, resurrect anachronistic class-taste associations, caricaturize gender- and race-taste associations. |
| Date: | 2026–06 |
| URL: | https://d.repec.org/n?u=RePEc:arx:papers:2606.30085 |
| By: | Davis, Steven; Hansen, Stephen; Seminario-Amez, Cristhian |
| Abstract: | Macro shocks produce high dispersion in firm-level equity returns, sales growth, and other outcomes. We show that this dispersion reflects observable differences in business characteristics. To do so, we combine firm-level returns on stock market ``jump" days with text about business risks in prior 10-K filings to construct firm-specific shock exposures. Our exposure measures explain firm-level abnormal returns through interpretable variation in language. They also explain most of the increased dispersion in firm-level revenue growth after major shocks and much of the dispersion in employment growth, investment rates, and earnings surprises. Our evidence yields a novel interpretation for countercyclical dispersion, highlighting the key role of heterogeneous business characteristics in macro shock transmission. |
| Keywords: | Firm heterogeneity; Text data; Machine learning; Shock identification |
| JEL: | C55 E30 L20 |
| Date: | 2025–06 |
| URL: | https://d.repec.org/n?u=RePEc:cpr:ceprdp:20358 |