|
on Big Data |
| By: | Usman Ahmed; Jason Hawkins |
| Abstract: | Land use and transportation infrastructure are tightly linked systems, and residential property prices play a central role in transportation planning. Transportation investments also influence property values, making accurate forecasting of both systems an interconnected research challenge. This study examines these interactions and evaluates machine learning methods for modeling residential real estate prices across contrasting urban contexts. We use XGBoost and Random Forest models to assess how land use and transportation infrastructure shape dwelling prices, comparing results between the Rawalpindi and Islamabad Metropolitan Area in Pakistan and the City of Toronto in Canada. We also examine the performance of machine learning models on spatial data by comparing nonspatial and spatial cross validation, and use SHAP values to interpret feature impacts. Despite differences in demographics and economic development, both cities show similar effects of transportation infrastructure and local amenities on prices. Proximity to major urban cores increases sale price, while high quality transit such as subway and bus rapid transit raises prices and conventional bus stop proximity lowers them. Nonspatial cross validation overestimates predictive accuracy for spatial datasets. XGBoost performs slightly better than Random Forest. We recommend careful application of machine learning methods when modeling spatially dependent land price data. |
| Date: | 2026–07 |
| URL: | https://d.repec.org/n?u=RePEc:arx:papers:2609.00375 |
| By: | Lorenzo Mu\~noz; Stephane Hess; Thomas O. Hancock; Georges Sfeir |
| Abstract: | Machine learning (ML) techniques are increasingly drawing interest in the choice modelling (CM) field. The focus has primarily been on comparing the performance of these contrasting approaches or on improving behavioural insights for ML techniques, rather than translating ideas from one field into the other. In the present paper, we specifically focus on knowledge transfer from ML into CM in the context of model performance evaluation. In CM, model performance is typically evaluated using log-likelihood and related indicators, which are aggregate fit metrics that focus on overall fit. Conversely, in ML, the focus is on alternative-level misclassifications and correct classifications, which provide a more nuanced view of the results. To bridge these approaches, we explore the use of a probabilistic version of the confusion matrix, which reports the average probability of the model predicting each alternative, conditional on which alternative was observed to be chosen, across all choice tasks. This enables the computation of probabilistic ML metrics for both classic choice models and ML algorithms. We analyse model performance jointly in terms of overall fit and alternative-level predictions. Our findings demonstrate that models with similar log-likelihood can exhibit substantially different confusion matrices, revealing different probability patterns that aggregate metrics cannot capture. This framework identifies where models systematically `confuse' alternatives, highlighting trade-offs between alternatives, and potentially guiding model specification. Furthermore, evaluating these matrices and metrics out-of-sample reveals alternative-level prediction shifts that significantly impact forecasting performance. |
| Date: | 2026–09 |
| URL: | https://d.repec.org/n?u=RePEc:arx:papers:2609.20655 |
| By: | Anjali Adukia; Matthew Bonci; Paula Dastres Gallardo; Emileigh Harrison; Jake Nicoll; Teodora Szasz |
| Abstract: | Emotional intelligence constitutes a key component of human capital, shaped partially by educational materials. Using machine learning and generative AI tools, we examine emotional content in public-school textbooks and children’s literature. A stark mismatch emerges: text exposes children to a broad emotional range, but images overwhelmingly depict happiness and calm, regardless of emotions described in text on the same page. Nearly half of pages show zero overlap between the emotions described in text and those shown in images. This pattern persists across time, contexts, and identities. Household purchases and library inventories suggest content may be endogenously shaped by consumer demand favoring "positive" cover imagery, implying market forces narrow the emotional landscape children encounter. |
| Keywords: | emotions, culture, content analysis, education policy, curriculum, artificial intelligence tools, computational social science, natural language processing, large language models, computer vision |
| JEL: | I20 I21 I24 L82 Z11 Z13 |
| Date: | 2026 |
| URL: | https://d.repec.org/n?u=RePEc:ces:ceswps:_12983 |
| By: | Bäckman, Claes |
| Abstract: | Households around the world increasingly take portfolio advice from the same handful of large language models. Does that advice adjust to what local circumstances warrant? I pose the same portfolio problem to leading models across twenty-one countries, each in its own language. Advice is nearly uniform across countries, even though the explanations invoke local context. What moves the advice instead is the model a household consults, not the country it lives in. The models follow retail financial advice rather than what academic finance would prescribe, and fail to incorporate the household's balance sheet even when it is stated. While automated advice once held out the promise of tailoring guidance to everyone, large language models encode the same advice for everyone. |
| Keywords: | large language models, financial advice, household finance, portfolio choice, robo-advising, cross-country variation |
| JEL: | G11 G41 G51 D14 O33 |
| Date: | 2026 |
| URL: | https://d.repec.org/n?u=RePEc:zbw:safewp:343061 |