|
on Big Data |
| By: | Liao, Yuan; Ma, Xinjie; Neuhierl, Andreas; Schilling, Linda |
| Abstract: | Machine learning in asset pricing typically predicts expected returns as point estimates, ignoring uncertainty. We develop new methods to construct forecast confidence intervals for expected returns obtained from neural networks. We show that neural network forecasts of expected returns share the same asymptotic distribution as classic nonparametric methods, enabling a closed-form expression for their standard errors. We also propose a computationally feasible bootstrap to obtain the asymptotic distribution. We incorporate these forecast confidence intervals into an uncertainty-averse investment framework. This provides an economic rationale for shrinkage implementations of portfolio selection. Empirically, our methods improve out-of-sample performance. |
| JEL: | G12 C45 C58 |
| Date: | 2025–03 |
| URL: | https://d.repec.org/n?u=RePEc:cpr:ceprdp:20080 |
| By: | Anamol Khadka; Milan Arjel; Ayush Lataula; Aayam Dhakal; Prajun Trital; Mingmar Sherpa; Biman Rimal |
| Abstract: | This study examines the dynamic relationship between the global oil prices and Nepal Stock Exchange (NEPSE) using an integrated approach which combines traditional econometric techniques with machine learning and explainable AI techniques. For this, Daily data of International Oil prices and NEPSE index is analyzed from approximately thirteen years (June 2013 to June 2026) using Granger causality, EGARCH(1, 1), and DCC-GARCH models to examine different properties like predictive relationships, asymmetric volatility behaviour, and time-varying correlations. To further supplement the econometric analysis, Machine Learning Models like Random Forest, LightGBM, and XGBoost algorithms were used to capture nonlinear relationships, along with explainable artificial intelligence techniques like SHAP values, Partial Dependence Plots, and Individual Conditional Expectation plots to further interpret the results of the model. The results from the econometric analysis showed a statistically significant unidirectional Granger causality from Brent crude oil to NEPSE with a four-day lag, high volatility persistence in both markets, and weak yet highly time-varying conditional correlations. Among the machine learning models, XGBoost achieves the best performance, and explainability analysis reveals that NEPSE own momentum and short-term volatility mainly influence its own behaviour and oil-related information serves as a minor, method-dependent contributor. The findings demonstrate that econometric and explainable machine learning approaches provide insights into the oil and equity market relationship in a way that each approach complements the result of one another. |
| Date: | 2026–07 |
| URL: | https://d.repec.org/n?u=RePEc:arx:papers:2607.11922 |
| By: | Gabaix, Xavier; Koijen, Ralph; Richmond, Robert; Yogo, Motohiro |
| Abstract: | Firm characteristics, based on accounting and financial market data, are commonly used to represent firms in economics and finance. However, investors collectively use a much richer information set beyond firm characteristics, including sources of information that are not readily available to researchers. We show theoretically that portfolio holdings contain all relevant information for asset pricing, which can be recovered under empirically realistic conditions. Such guarantees do not exist for other data sources, such as accounting or text data. We build on recent advances in artificial intelligence (AI) and machine learning (ML) that represent unstructured data (e.g., text, audio, and images) by high-dimensional latent vectors called embeddings. Just as word embeddings leverage the document structure to represent words, asset embeddings leverage portfolio holdings to represent firms. Thus, this paper is a bridge from recent advances in AI and ML to economics and finance. We explore various methods to estimate asset embeddings, including recommender systems, shallow neural network models such as Word2Vec, and transformer models such as BERT. We evaluate the performance of these models on three benchmarks that can be evaluated using a single quarter of data: predicting relative valuations, explaining the comovement of stock returns, and predicting institutional portfolio decisions. We also estimate investor embeddings (i.e., representations of investors and their strategies), which are useful for investor classification, performance evaluation, and detecting crowded trades. We discuss other applications of asset embeddings, including generative portfolios, risk management, and stress testing. Finally, we develop a framework to give an economic narrative to a group of similar firms, by applying large language models to firm-level text data. |
| Keywords: | Artificial intelligence; Asset pricing; Machine learning; Transformer models |
| JEL: | C53 G12 G23 |
| Date: | 2025–03 |
| URL: | https://d.repec.org/n?u=RePEc:cpr:ceprdp:20082 |
| By: | Hauzenberger, Niko; Huber, Florian; Klieber, Karin; Marcellino, Massimiliano |
| Abstract: | We propose a method to learn the nonlinear impulse responses to structural shocks using neural networks, and apply it to uncover the effects of US financial shocks. The results reveal substantial asymmetries with respect to the sign of the shock. Adverse financial shocks have powerful effects on the US economy, while benign shocks trigger much smaller reactions. Instead, with respect to the size of the shocks, we find no discernible asymmetries. |
| Keywords: | Bayesian neural networks |
| JEL: | C11 C30 C45 E3 E44 |
| Date: | 2025–02 |
| URL: | https://d.repec.org/n?u=RePEc:cpr:ceprdp:19964 |
| By: | Morgane Brenon; Géraldine Duthé |
| Abstract: | This paper contributes to the empirical literature on ne scale estimation of poverty using open-access satellite imagery in urban settings. Our main hypothesis is that urban housing morphology reects socioeconomic status, as households within similar physical environments share comparable characteristics. We rely on census data of Madagascar from 2018 that registered a total population of 1.27 million inhabitants living in 323 297 households distributed in 192 neighborhoods of the city of Antananarivo, to assess the performance of indicators produced from satellite imagery during the period 2014-2021 using spatial econometrics. We show that the building detection confidence metric provided by Google Open Buildings, combined with other readily available indicators of satellite imagery, can explain 74% of the variations of poverty at the neighborhood level. This approach demonstrates that open access satellite-derived indicators can serve as alternative source of data that bypass complex machine learning methods thereby signicantly reducing both financial cost and the technical expertise required for implementation. |
| Keywords: | Poverty estimation, Condence metric, Satellite imagery, Spatial Durbin Model, Antananarivo, Google Open Buildings, Madagascar, METHODOLOGIE / METHODOLOGY, MADAGASCAR / MADAGASCAR, CARTOGRAPHIE / CARTOGRAPHY, PAUVRETE / POVERTY |
| Date: | 2026 |
| URL: | https://d.repec.org/n?u=RePEc:idg:wpaper:8ptdf58bvmh4hxb48yt7 |
| By: | Beroud, Mohammed; Qiu, Feng; Wichmann, Bruno; Fan, Xiaoli |
| Abstract: | This paper develops a Machine Learning (ML) approach to quantify the impact of geospatial data quality on farmland valuation accuracy. The approach is useful when both the set of quality attributes and the predictor set are high-dimensional relative to the sample size, making standard non-market valuation methods difficult to implement. Conceptually, we draw on value of information (VoI) theory, which defines value as the difference in expected payoff under higherquality versus imperfect information. The VoI measure can be approximated as the difference between predicted farmland value under full-quality data and predicted value under incomplete data. We apply the approach to soil data and farmland values in the Canadian Prairies. Baseline and counterfactual values are predicted using a Random Forest model with hyperparameters tuned on spatially blocked folds and evaluated with blocked spatial cross-validation to limit spatial leakage. Counterfactual quality scenarios are implemented as random perturbations of soil features. Specifically, we simulate six scenarios: spatially localized bias, limited geographic coverage, coarse spatial resolution, measurement error, numeric rounding, and categorical misclassification. The results show that coarse spatial resolution generates the largest average valuation distortion (238.46 CAD/ha, 2006 CAD), followed by limited coverage (104.49 CAD/ha), while the remaining degradations have small effects. Quality–value curves traced over degradation intensities are nonlinear and concave, consistent with diminishing marginal returns to information improvements. The findings have policy implications for prioritizing investments in public geospatial data: budgets may yield higher returns by shifting from incremental precision upgrades toward improving spatial coverage and resolution. |
| Keywords: | Research Methods/ Statistical Methods |
| Date: | 2026 |
| URL: | https://d.repec.org/n?u=RePEc:ags:aaea26:404720 |
| By: | Catherine Chen; Chen Gao; Jonathon Hazell; Lihua Lei; Chen Lian |
| Abstract: | Does microeconomic heterogeneity help to forecast aggregate inflation in a non-stationary environment? We develop a scan test for whether one forecast outperforms another, over an interval with unknown starting point and duration. To exploit any occasional forecasting power that the scan test detects, we design an adaptive machine learning pipeline. We encode the distribution of price changes into a high-dimensional vector, which we combine with a gradient boosted trees algorithm. We then combine this micro forecast with other benchmark forecasts, using an adaptive algorithm that makes use of the micro forecast only when it performs well. We apply the pipeline to UK microdata, with four main results. First, the micro forecast outperforms a univariate benchmark, but only in the volatile period after 2020. Second, the scan test detects periods of micro outperformance, so the micro forecast enters the combined forecast. Third, the combined forecast performs comparably to the univariate benchmark before 2020 and better at every horizon after 2020. Fourth, the value of microdata for the combined forecast materializes after 2020. We conclude that microdata are valuable for forecasting aggregate inflation, but only after large shocks. |
| Date: | 2026–07 |
| URL: | https://d.repec.org/n?u=RePEc:arx:papers:2607.12345 |
| By: | Maria Saveria Mavillonio; Stefano Borgioli; Caterina Giannetti; Chiara Ongari; Giampiero M. Gallo |
| Abstract: | Measuring sentiment from financial news is a central task in economics and finance, yet most existing indicators rely on dictionary-based approaches that infer sentiment from word counts and only partially capture context, negation, and semantic structure. This paper proposes a framework for constructing daily news mood indices using transformer-based language models and evaluates whether they better represent sentiment than dictionary-based alternatives. Using 143, 755 financial news articles from Factiva, we classify sentiment at the sentence level with FinBERT and aggregate these predictions into article-level and daily sentiment measures through alternative normalization schemes. We compare the resulting indices with benchmark measures based on Shapiro et al., 2022 and Barbaglia et al., 2025. A central contribution is the validation of alternative sentiment measures against human judgments. We conducted an incentivized annotation exercise in which 444 participants evaluated a validation subsample of 588 financial news articles. Consensus ratings from independent human evaluations serve as an external benchmark for assessing the quality of automated sentiment measures. Across correlation, regression, and classification exercises, transformer-based measures show stronger agreement with human judgments than vocabulary-based alternatives and perform substantially better in distinguishing positive, neutral, and negative articles. Overall, the results suggest that incorporating contextual information through transformer-based language models produces sentiment measures that more closely reflect human assessments of financial news. |
| Date: | 2026–07 |
| URL: | https://d.repec.org/n?u=RePEc:arx:papers:2607.13968 |
| By: | Choudhury, Raj; Espinosa, Miguel; Khanna, Tarun; Makridis, Christos; Schirmann, Kyle |
| Abstract: | The rapid adoption of hybrid work by firms has led to a debate between managers and workers on the relative value of remote and in-person communication. Colocation between workers may be helpful for communication, aid with coordination, and affect the intensity of monitoring of workers by managers. Exploiting a hybrid work field experiment involving HR workers and using unique data related to the text of electronic communication between employees, this paper provides causal evidence of how colocation between employees affects internal communication within firms. A machine learning analysis of email content reveals that colocation is a substitute for horizontal, coordination-related communication, but---somewhat surprisingly---a complement to vertical, monitoring-related communication. |
| JEL: | J23 J24 O10 O33 |
| Date: | 2025–02 |
| URL: | https://d.repec.org/n?u=RePEc:cpr:ceprdp:19920 |
| By: | Hauzenberger, Niko; Marcellino, Massimiliano; Pfarrhofer, Michael; Stelzer, Anna |
| Abstract: | We develop Bayesian machine learning methods for mixed data sampling (MIDAS) regressions. This involves handling frequency mismatches and specifying functional relationships between many predictors and the dependent variable. We use Gaussian processes (GPs) and compress the input space with structured and unstructured MI-DAS variants. This yields several versions of GP-MIDAS with distinct properties and implications, which we evaluate in short-horizon now- and forecasting exercises with both simulated data and data on quarterly US output growth and inflation in the GDP deflator. Our proposed framework leverages macroeconomic Big Data in a computationally efficient way and offers gains in predictive accuracy along several dimensions. |
| JEL: | C11 C22 C53 E31 E37 |
| Date: | 2025–02 |
| URL: | https://d.repec.org/n?u=RePEc:cpr:ceprdp:19965 |
| By: | Ehmke, Mariah; Okrent, Abigail; Awad, Koroles; McCluskey, Jill; Restrepo, Brandon |
| Abstract: | Differentiating food products by level of industrial food processing is essential to measuring relationships among food processing, food demand, and public health. The Circana (formerly known as IRI) scanner data lacks identified measures of food product processing, but contains other food product descriptors that can be used to predict the processing of food products. This report specifies procedures to use Natural Language Processing with a Naïve Bayes model in Python’s scikit-learn to efficiently classify retail food purchases by level of food processing according to the NOVA food classification system. This method accurately predicts NOVA classification 94 percent using only the food item descriptions as predictors. Based on this classification we find 60 percent of retail food sales to be ultraprocessed in 2023, an increase from 55 percent in 2020. |
| Keywords: | Food Consumption/Nutrition/Food Safety |
| Date: | 2026 |
| URL: | https://d.repec.org/n?u=RePEc:ags:aaea26:404883 |
| By: | Cassidy, William; Kempf, Elisabeth |
| Abstract: | We construct a novel measure of partisan corporate speech using natural language processing techniques and use it to establish three stylized facts. First, the volume of partisan corporate speech has risen sharply between 2012 and 2022. Second, this increase has been disproportionately driven by companies adopting more Democratic-leaning language, a trend that is widespread across industries, geographies, and CEO political affiliations. Third, partisan corporate statements are followed by negative abnormal stock returns, with significant heterogeneity by shareholders' degree of alignment with the statement. Finally, we propose a theoretical framework and provide suggestive empirical evidence that these trends are at least in part driven by a shift in investors' nonpecuniary preferences with respect to partisan corporate speech. |
| Date: | 2025–05 |
| URL: | https://d.repec.org/n?u=RePEc:cpr:ceprdp:20242 |
| By: | Rossi-Hansberg, Esteban; Zhang, Jialing |
| Abstract: | We use high-resolution spatial data to build a novel global annual gridded GDP dataset at 1°, 0.5°, and 0.25° resolutions from 2012 onward. Our random forest model trained on local and national GDP achieves an R² above 0.92 for GDP levels and above 0.62 for annual changes in regions left out of the training sample. By incorporating diverse indicators beyond population and nighttime lights, our estimates offer more precise subnational GDP measurements for analyzing economic shocks, local policies, and regional disparities. We evaluate the precision of our estimates with a sample case of COVID-19’s impact on local GDP in China. |
| JEL: | E0 F0 R0 |
| Date: | 2025–03 |
| URL: | https://d.repec.org/n?u=RePEc:cpr:ceprdp:20023 |
| By: | Fr\'ed\'eric Godin |
| Abstract: | The deep hedging framework of Buehler et al. (2019) trains a neural network policy, via Monte Carlo simulation of price paths and stochastic gradient descent, to minimize a risk measure applied to the terminal hedging error. In a recent stream of papers, my coauthors and I have referred to this technique as reinforcement learning (RL), a characterization that referees on several submissions have challenged on two grounds, among others: first, that because feedback is generated only at the terminal date, with no intermediate reward signal, the method cannot constitute genuine RL; and second, that the absence of a value function, a Bellman equation, temporal-difference (TD) learning, and an explicit exploration mechanism disqualifies the method from the RL category altogether, so that it should instead be labeled a neural-network method for stochastic optimal control. I argue that both objections rest on an unduly narrow, TD-centric reading of what constitutes RL. Once RL is understood, as it is in the standard references of the field, to include Monte Carlo policy-gradient methods and direct (actor-only) policy search as first-class members, the deep hedging algorithm of Buehler et al. (2019) falls squarely within the RL umbrella. |
| Date: | 2026–07 |
| URL: | https://d.repec.org/n?u=RePEc:arx:papers:2607.13353 |
| By: | David Imhof; Thierry Madi\`es; Martin Huber |
| Abstract: | This paper analyzes the internal organization and economic effects of a bid-rigging cartel in the road construction sector of the Swiss canton of Ticino, active from 1999 to 2005. Using exceptionally rich documentary evidence, we reconstruct how cartel members coordinated bids and allocated contracts under a formal agreement known as the 'convention'. We show that, despite the absence of side payments, the cartel implemented a cost-based allocation mechanism that closely approximated the first-best collusive outcome. Regression and machine-learning analyses indicate that observable cost proxies systematically predict both winning bids and bid rankings. The evidence further suggests that cartel members strategically mimicked competitive bidding behavior, allowing them to evade standard econometric detection methods. Using double machine learning, we estimate average overcharges of at least 45\%, and potentially substantially higher, highlighting the significant financial harm caused by this sophisticated form of collusion. |
| Date: | 2026–06 |
| URL: | https://d.repec.org/n?u=RePEc:arx:papers:2606.30470 |
| By: | Sanggyu Sean Choi |
| Abstract: | Financial sentiment extraction has largely relied on news text and supervised extraction against return labels alone, leaving 10-K filings -- and volatility, the target risk disclosure is arguably best suited to informing -- comparatively unexplored. We extend a supervised lexicon-learning approach to 10-K filings and their Item 1A risk-factor sections, training sentiment scores against both return and volatility labels at three levels of aggregation: sector, portfolio, and individual firm. Across 1, 383 filings from 94 Nasdaq-100 technology constituents (2006--2023), we evaluate the resulting twelve sentiment metrics on classification accuracy, correlation with realised market outcomes, and qualitative lexical content. Full-filing text produces more accurate sentiment at the sector and portfolio level for both targets, but this reverses at the individual-firm level, where the narrower Item 1A section performs better -- an effect we attribute to the interaction between document volume and the amount of independent training signal available at each level of aggregation. A Loughran-McDonald dictionary baseline is consistently, strongly negatively correlated with price at every level tested, underscoring the value of a supervised approach for regulatory disclosure text. These findings, and the design choices they motivate, establish the sentiment-generation methodology underlying a subsequent, larger-scale, multi-source system. |
| Date: | 2026–07 |
| URL: | https://d.repec.org/n?u=RePEc:arx:papers:2607.14174 |
| By: | Ina Ganguli; Jeffrey Lin; Vitaly Meursault; Nicholas Reynolds |
| Abstract: | Over nearly two centuries, U.S. inventions have become increasingly dissimilar: not just fewer head-to-head collisions between inventors, but growing distance between neighboring inventions. We document this secular decline in similarity using validated neural language models applied to the full text of claims in over 11 million U.S. patents (1836–2023), corroborated by a 98 percent decline in patent interference rates, a measure of independent simultaneous invention. Measuring this correctly requires validation, since different representations of the same patent text can yield opposite conclusions about whether inventions are converging or spreading out. Our validation framework, the first systematic comparison for patent text, selects among these locations in idea space. The model explains spreading out and connects it to several independently documented patterns — rising R&D investment per inventor, increasing patent values, weakening knowledge spillovers, and declining research productivity. The mechanism is spatial; as inventors spread out to capture new territory, inventions become more valuable but also more costly for others to absorb. In doing so, the model turns spillover intensity, innovation step size, and research productivity from fixed primitives into outcomes of inventor positioning. A calibrated decomposition attributes roughly 40 percent of the long-run decline in U.S. research productivity to these spatial forces, alongside traditional explanations such as fishing out and the burden of knowledge. Where inventors stand relative to each other in idea space matters as much for growth as how many of them there are. |
| Keywords: | Idea Space; Knowledge Spillovers; Research Productivity; Endogenous Growth; Technological Distance; Patent Embeddings |
| JEL: | O31 O41 O47 C55 |
| Date: | 2026–08–05 |
| URL: | https://d.repec.org/n?u=RePEc:fip:fedpwp:103607 |
| By: | Anastasiou, Dimitris; Katsafados, Apostolos; Ongena, Steven; Tzomakas, Christos |
| Abstract: | Building on Gorodnichenko et al. (2023) we propose a novel measure that quantifies the voice sentiment of the Chair of the Federal Reserve press conference responses and examine its impact on the stock price crash risk of U.S. banks. We find that a more positive vocal sentiment, indicative of happiness, significantly reduces banks’ stock price crash risk, whereas negative emotions, such as sadness and anger, amplify it. These effects are economically meaningful and robust across various specifications, alternative crash risk proxies, and endogeneity checks, including an instrumental variables (IV) strategy and reverse causality tests. Additionally, the emotional sentiment has asymmetric effects on stock price crash risk, depending on bank size. Beyond the textual content of monetary policy statements, the emotional delivery of central bank communication plays a critical role in shaping financial stability outcomes, providing empirical evidence for the theoretical channels of uncertainty, systemic risk, and investor sentiment. |
| Keywords: | Financial stability |
| JEL: | G01 |
| Date: | 2025–05 |
| URL: | https://d.repec.org/n?u=RePEc:cpr:ceprdp:20308 |
| By: | Cole, Stephen J. (Department of Economics Marquette University); (Department of Economics Marquette University) |
| Abstract: | This paper uses an adaptive learning framework to study FOMC forecasts from the Summary of Economic Projections (SEP) dataset. FOMC expectations are modeled as the sum of two components: (1) an endogenous learning part and (2) a sentiment part capturing waves of optimism and/or pessimism. The results include key policy takeaways. FOMC forecasts are responsive to incoming macroeconomic information, consistent with adaptive learning, while sentiment is persistent, correlated across GDP growth and inflation forecasts, and becomes quantitatively more important during and around recessions. FOMC participants also rely more on their endogenous/learning model to form expectations, but sentiment plays a larger role during and around recessions. Finally, the model-implied sentiment measure is positively and significantly correlated with an external measure of FOMC sentiment and remains robust across alternative forecasting specifications. |
| Keywords: | summary of economic projections, FOMC, constant-gain learning, sentiment shocks, waves of optimism and pessimism, evolving beliefs, monetary policy |
| JEL: | C52 D84 E50 E52 E58 E60 E70 E71 |
| Date: | 2026–06 |
| URL: | https://d.repec.org/n?u=RePEc:mrq:wpaper:2026-03 |
| By: | Steffen, Sascha; Saunders, Anthony; Verhoff, Paulina |
| Abstract: | We introduce CovenantAI, an advanced AI-based approach to accurately identify loan covenant violations from SEC filings. CovenantAI outperforms traditional keyword-based and Dealscan approaches by precisely classifying complex renegotiation outcomes such as amendments, waivers, and technical defaults. Our analyses validate CovenantAI's higher accuracy and consistency against existing methods. By capturing nuanced creditor-borrower interactions, CovenantAI enables new research into the economic implications of covenant resolution outcomes and investor behaviors surrounding violations. This novel dataset thus significantly enhances empirical research capabilities in corporate finance by providing comprehensive, granular, and reliable data on covenant violations and resolutions. |
| Keywords: | Renegotiations |
| JEL: | G21 G32 G34 |
| Date: | 2025–04 |
| URL: | https://d.repec.org/n?u=RePEc:cpr:ceprdp:20117 |
| By: | Asef Y{\i}lk{\i} |
| Abstract: | This paper proposes a novel asset pricing framework that augments large language model (LLM) embeddings of annual report disclosures with supply chain knowledge graph (KG) propagation. Using FinBERT embeddings of 10-K MD&A sections for 255 S&P 500 firms over 2011-2025, two sets of return predictors are constructed: direct LLM embeddings and network-augmented embeddings, where firm-level signals propagate through inter-firm linkages. Fama-MacBeth cross-sectional regressions reveal that the network-augmented factor (net_pc_5) carries significant return predictability with a Newey-West t-statistic of -2.64, even after controlling for momentum, volatility, and firm size. A long-short portfolio sorted on net_pc_5 achieves an annualized Sharpe ratio of 0.86 and a Fama-French five-factor alpha of 7.27% per year (t = 2.30). The predictive power survives out-of-sample tests, placebo experiments, sector-neutralization, and subsample analysis. The findings suggest that inter-firm network structure contains pricing-relevant information beyond firm-level textual disclosures. |
| Date: | 2026–06 |
| URL: | https://d.repec.org/n?u=RePEc:arx:papers:2606.29290 |
| By: | Radoslaw Stefanski (University of St Andrews; University of Stavanger) |
| Abstract: | Long-run growth is driven by new ideas, yet the cultural environment shaping their production is difficult to measure over time. We use large language models to read 23, 000 books from the Western canon and score whether each endorses, rejects, or merely depicts six dimensions of culture. We accumulate the scores into inherited stocks and summarize them with an Innovation Wedge measuring cultural resistance to new ideas. Between 1000 and 1920 the wedge falls by 51 percent. Blinded expert readings and modern surveys validate the measure. An independent 5, 000-book archive reproduces the decline. In a calibrated semi-endogenous growth model, the falling wedge raises 1920 productivity to 1.78 times its counterfactual level, explains two-thirds of the first sustained acceleration in productivity growth between 1500 and 1700, and accounts for 38.7 percent of productivity growth in 1920. |
| Keywords: | culture and growth; ideas production; growth accounting; innovation barriers; text as data |
| JEL: | O41 O31 N13 Z10 |
| Date: | 2026–07–23 |
| URL: | https://d.repec.org/n?u=RePEc:san:econdp:2602 |