nep-big New Economics Papers
on Big Data
Issue of 2026–09–21
four papers chosen by
Tom Coupé, University of Canterbury


  1. Residential Price Modeling using Spatially Validated Machine Learning Methods: A Comparison Across Geographic Contexts By Usman Ahmed; Jason Hawkins
  2. Using machine learning metrics to provide deeper insights into the performance of choice models By Lorenzo Mu\~noz; Stephane Hess; Thomas O. Hancock; Georges Sfeir
  3. (How) Do We Teach Emotions? By Anjali Adukia; Matthew Bonci; Paula Dastres Gallardo; Emileigh Harrison; Jake Nicoll; Teodora Szasz
  4. One advisor for the whole world? Cross-country evidence on financial advice from large language models By Bäckman, Claes

  1. By: Usman Ahmed; Jason Hawkins
    Abstract: Land use and transportation infrastructure are tightly linked systems, and residential property prices play a central role in transportation planning. Transportation investments also influence property values, making accurate forecasting of both systems an interconnected research challenge. This study examines these interactions and evaluates machine learning methods for modeling residential real estate prices across contrasting urban contexts. We use XGBoost and Random Forest models to assess how land use and transportation infrastructure shape dwelling prices, comparing results between the Rawalpindi and Islamabad Metropolitan Area in Pakistan and the City of Toronto in Canada. We also examine the performance of machine learning models on spatial data by comparing nonspatial and spatial cross validation, and use SHAP values to interpret feature impacts. Despite differences in demographics and economic development, both cities show similar effects of transportation infrastructure and local amenities on prices. Proximity to major urban cores increases sale price, while high quality transit such as subway and bus rapid transit raises prices and conventional bus stop proximity lowers them. Nonspatial cross validation overestimates predictive accuracy for spatial datasets. XGBoost performs slightly better than Random Forest. We recommend careful application of machine learning methods when modeling spatially dependent land price data.
    Date: 2026–07
    URL: https://d.repec.org/n?u=RePEc:arx:papers:2609.00375
  2. By: Lorenzo Mu\~noz; Stephane Hess; Thomas O. Hancock; Georges Sfeir
    Abstract: Machine learning (ML) techniques are increasingly drawing interest in the choice modelling (CM) field. The focus has primarily been on comparing the performance of these contrasting approaches or on improving behavioural insights for ML techniques, rather than translating ideas from one field into the other. In the present paper, we specifically focus on knowledge transfer from ML into CM in the context of model performance evaluation. In CM, model performance is typically evaluated using log-likelihood and related indicators, which are aggregate fit metrics that focus on overall fit. Conversely, in ML, the focus is on alternative-level misclassifications and correct classifications, which provide a more nuanced view of the results. To bridge these approaches, we explore the use of a probabilistic version of the confusion matrix, which reports the average probability of the model predicting each alternative, conditional on which alternative was observed to be chosen, across all choice tasks. This enables the computation of probabilistic ML metrics for both classic choice models and ML algorithms. We analyse model performance jointly in terms of overall fit and alternative-level predictions. Our findings demonstrate that models with similar log-likelihood can exhibit substantially different confusion matrices, revealing different probability patterns that aggregate metrics cannot capture. This framework identifies where models systematically `confuse' alternatives, highlighting trade-offs between alternatives, and potentially guiding model specification. Furthermore, evaluating these matrices and metrics out-of-sample reveals alternative-level prediction shifts that significantly impact forecasting performance.
    Date: 2026–09
    URL: https://d.repec.org/n?u=RePEc:arx:papers:2609.20655
  3. By: Anjali Adukia; Matthew Bonci; Paula Dastres Gallardo; Emileigh Harrison; Jake Nicoll; Teodora Szasz
    Abstract: Emotional intelligence constitutes a key component of human capital, shaped partially by educational materials. Using machine learning and generative AI tools, we examine emotional content in public-school textbooks and children’s literature. A stark mismatch emerges: text exposes children to a broad emotional range, but images overwhelmingly depict happiness and calm, regardless of emotions described in text on the same page. Nearly half of pages show zero overlap between the emotions described in text and those shown in images. This pattern persists across time, contexts, and identities. Household purchases and library inventories suggest content may be endogenously shaped by consumer demand favoring "positive" cover imagery, implying market forces narrow the emotional landscape children encounter.
    Keywords: emotions, culture, content analysis, education policy, curriculum, artificial intelligence tools, computational social science, natural language processing, large language models, computer vision
    JEL: I20 I21 I24 L82 Z11 Z13
    Date: 2026
    URL: https://d.repec.org/n?u=RePEc:ces:ceswps:_12983
  4. By: Bäckman, Claes
    Abstract: Households around the world increasingly take portfolio advice from the same handful of large language models. Does that advice adjust to what local circumstances warrant? I pose the same portfolio problem to leading models across twenty-one countries, each in its own language. Advice is nearly uniform across countries, even though the explanations invoke local context. What moves the advice instead is the model a household consults, not the country it lives in. The models follow retail financial advice rather than what academic finance would prescribe, and fail to incorporate the household's balance sheet even when it is stated. While automated advice once held out the promise of tailoring guidance to everyone, large language models encode the same advice for everyone.
    Keywords: large language models, financial advice, household finance, portfolio choice, robo-advising, cross-country variation
    JEL: G11 G41 G51 D14 O33
    Date: 2026
    URL: https://d.repec.org/n?u=RePEc:zbw:safewp:343061

This nep-big issue is ©2026 by Tom Coupé. It is provided as is without any express or implied warranty. It may be freely redistributed in whole or in part for any purpose. If distributed in part, please include this notice.
General information on the NEP project can be found at https://nep.repec.org. For comments please write to the director of NEP, Marco Novarese at <director@nep.repec.org>. Put “NEP” in the subject, otherwise your mail may be rejected.
NEP’s infrastructure is sponsored by the Griffith Business School of Griffith University in Australia.