|
on Experimental Economics |
| By: | Guy Aridor; Winston Chou; Nathan Kallus; Antoine Scheid; Allen Tran; Kevin Zielnicki |
| Abstract: | We study an experiment with 8.5 million users on Netflix's recommender system to measure how improvements in recommendation technology affect the set of products that get consumed. Improvements increase total consumption and users' reliance on recommendations while diffusing recommendations and consumption away from the most popular products ("superstars") toward a larger number of moderately popular products ("middle-tail"), with minimal effects on the most niche products ("long-tail"). Our results challenge the notion that recommender systems polarize consumption — raising the consumption shares of the head and tail at the expense of the middle — and suggest that the returns to investing in middle-tail products grow as algorithms improve and platforms scale. |
| Keywords: | personalization, recommendation systems, field experiment |
| JEL: | D83 D82 C93 |
| Date: | 2026 |
| URL: | https://d.repec.org/n?u=RePEc:ces:ceswps:_12981 |
| By: | Takeshi Nishimura; Nobuyuki Hanaki |
| Abstract: | We provide experimental evidence on entry and bidding behavior in surplusextracting auctions built on the second-price format. In addition to the simple second-price auction, we consider full- and partial-surplus-extracting auctions. The latter leaves more surplus to bidders under full entry than the former, while preserving the same set of undominated strategies: value bidding and opting out. In contrast to the second-price auction, both auctions exhibit lower entry and overbidding among entrants. Our findings highlight the interplay between entry and bidding behavior, shedding light on the trade-off between surplus extraction and strategic simplicity in auction design with voluntary participation. |
| Date: | 2024–11 |
| URL: | https://d.repec.org/n?u=RePEc:dpr:wpaper:1266r |
| By: | Hilweg-Waldeck, Michael; Hild, Paul Ergün |
| Abstract: | Many donors leave tax benefits unclaimed even when doing so requires minimal effort and yields meaningful financial rewards. Findings from our representative survey point to confusion about how to deduct donations and to misperceived social norms about the moral appropriateness of doing so as the main drivers of this gap. We study how to tackle these two sources of the deduction gap by providing concise information on how to deduct donations and a one-sentence norm cue in an online experiment (n = 483), a door-to-door field experiment with address-level randomization (n = 6, 728), and a radio-based campaign spanning two Austrian federal states. We find that almost all donors deduct when donating through the anonymous online tool. By contrast, during face-to-face fundraising, where social-image concerns are salient, fewer than 1 in 100 donors choose to do so. Across settings, information on how to deduct donations alone leaves deduction behavior unchanged, whereas combining this information with the norm cue increases take-up in the door-to-door setting. Our findings show that financial incentives can falter when clashing with misperceived norms in social settings, unless paired with campaigns that reshape those norms. |
| Keywords: | social image, tax incentives, charitable giving |
| JEL: | C93 D64 D91 |
| Date: | 2026 |
| URL: | https://d.repec.org/n?u=RePEc:zbw:zewdip:343567 |
| By: | Hild, Paul Ergün; Hilweg-Waldeck, Michael |
| Abstract: | Career choice, earnings, and other key economic outcomes have been linked to gender differences in willingness to compete. This paper examines how gender stereotypes shape these differences. We conduct a meta-study of prior work and demonstrate that the wide variation in gender competition gaps can be explained by stereotypes: Men enter competitions more in traditionally male-stereotyped domains, whereas in female-stereotyped domains, the gap is smaller or even reversed. Importantly, these differences are not explained by gender gaps in performance. To explore mechanisms, we collect belief data in an elicitation experiment. We find that stereotyped beliefs about gender performance differences explain more than half of the variation in competition gaps in the literature. Next, we experimentally manipulate stereotypes through framing and informational cues about others beliefs. Although these interventions significantly shift beliefs, the effects do not translate into changes in competitive behavior. Our findings highlight the importance of stereotypes in shaping gender gaps in competitiveness while suggesting that shifting beliefs alone is unlikely to close these gaps without deeper or longer-term interventions. |
| Keywords: | Gender, Competitiveness, Stereotypes, Beliefs, Experiment |
| JEL: | D91 J16 C90 |
| Date: | 2026 |
| URL: | https://d.repec.org/n?u=RePEc:zbw:zewdip:343565 |
| By: | Carol Luengo (University of Waikato); Steven Tucker (University of Waikato); Yilong Xu (Utrecht School of Economics, Utrecht University); Frank Scrimgeour (University of Waikato) |
| Abstract: | This paper investigates the causal impact of artificial intelligence (AI) advice on price discovery in a controlled asset-market experiment. Using a call market setting in which the asset's value follows a stochastic geometric random walk, we compare a baseline "No AI" condition against two treatments: "Good AI" (advice aligned with long-term fundamentals) and "Bad AI" (myopic advice anchored to current buyout prices). Our results show that access to AI advice significantly reduces mispricing relative to the baseline. In particular, Good AI market prices started low, then converged and closely followed the fundamentals in the second half of the market. Interestingly, we find no meaningful difference in aggregate mispricing between the Good and Bad AI treatments. Analysis of trader behavior reveals that while over half of the participants follow AI recommendations, a significant portion actively filters out erroneous advice, particularly in the Bad AI treatment. This selective adherence explains why market prices remain resilient to poor advice. |
| Keywords: | artificial intelligence; financial advice; assest market experiment; mispricing; Human-AI interaction |
| JEL: | C90 C91 G12 |
| Date: | 2026–09–15 |
| URL: | https://d.repec.org/n?u=RePEc:wai:econwp:26/05 |
| By: | Gangadharan, Lata; Gsottbauer, Elisabeth; Leslie, Gordon; Pretto, Madeline |
| Abstract: | Do households comprehend the nature of tail-risks inherent to real-time electricity pricing (RTP) plans? We develop a randomized and incentivized experiment calibrated to real-world price distributions and find that (a) probabilistic risk disclosure, beyond that offered in standard marketing materials, elicits greater demand for real-time pricing products relative to a low-risk fixed-price alternative, (b) products with tail-risk protection might not be highly sought, and (c) the experience of a RTP bill shock drives choice away from RTP. Personal experience receiving a tail price plays a greater role in risk comprehension and moving subsequent choices toward less risky plans than receipt of an additional ex-ante probabilistic risk disclosure. We discuss the implications these findings may have for regulators with a consumer protection mandate. |
| Keywords: | consumer protection;dynamic pricing;experimental economics;information provision;retail electricity markets;risk perception;tail-risk |
| JEL: | C91 D18 D81 L94 Q41 |
| Date: | 2026–09–30 |
| URL: | https://d.repec.org/n?u=RePEc:ehl:lserod:140854 |
| By: | Samuel D. Higbee |
| Abstract: | We show how to optimally design experiments when the resulting data will be used to choose a welfare-maximizing policy subject to constraints. A decision maker seeks to maximize Bayes expected welfare by choosing a policy whose effects depend on an unknown finite-dimensional parameter. The decision maker has access to a first wave of experimental data with a fixed design but may choose the design of a second wave that will be collected before choosing the policy. The resulting experimental design--policy choice problem is a very high-dimensional dynamic program that is generally intractable in finite samples. We propose a tractable approximation based on the limit experiment and show it is asymptotically optimal using a new asymptotic representation theorem for adaptive experiments with continuous treatments. We apply the method to a conditional cash transfer experiment and demonstrate the potential for large gains from tailoring the experiment to the policy choice. |
| Date: | 2026–09 |
| URL: | https://d.repec.org/n?u=RePEc:arx:papers:2609.10971 |
| By: | Ali Moghaddasi Kelishomi (Loughborough University); Daniel Sgroi (University of Warwick); Andis Sofianos (Durham University) |
| Abstract: | How does rationality shape cooperation in strategic settings? We study this question in a laboratory experiment that links individual rationality, measured by consistency with the generalized axiom of revealed preference, to behaviour in an indefinitely repeated Prisoner’s Dilemma. Participants are grouped by pre-measured rationality before interacting repeatedly. We find that higher rationality substantially increases cooperation and payoffs. This effect operates through a novel mechanism: more rational individuals make fewer implementation errors when executing their intended strategies, thereby sustaining cooperative outcomes. By contrast, higher cognitive ability also promotes cooperation and higher payoffs, but through a distinct channel—reducing strategic errors in responding optimally to others’ actions. Our results provide the first experimental evidence linking rationality to cooperation via decision-making errors, and clarify the distinct roles of rationality and intelligence in shaping strategic behaviour. Together, the findings offer a unified account of how cognitive constraints affect cooperation in repeated games |
| Keywords: | Repeated Prisoners Dilemma, Cooperation, Rationality, Intelligence, Learning, Strategy Errors JEL codes: C73, C91, C92, D83 |
| Date: | 2026 |
| URL: | https://d.repec.org/n?u=RePEc:wrk:warwec:1630 |
| By: | Ashwini Deshpande (Ashoka University); Veronique Gille (Université Paris-Dauphine); Rajesh Ramachandran (Monash University Malaysia); Andis Sofianos (Durham University) |
| Abstract: | Affirmative action (AA) is highly controversial. This study asks whether opposition to AA reflects a genuine commitment to merit, or whether pro-merit discourse conceals deeper identity-based considerations. Specifically, does resistance vary with the social identity of those who benefit from the policy? Using an incentivized online experiment with university students in India, we compare perceptions and rewards toward participants selected through caste-based versus income-based affirmative action. Evaluators assess test-takers’ competence and allocate monetary rewards under different selection rules. We find that beneficiaries from India’s historically stigmatized caste groups (Scheduled Castes and Tribes, or SC-ST) are perceived as less competent than non-marginalized beneficiaries, even under income-based affirmative action. This suggests that responses to affirmative action vary systematically with the social identity of beneficiaries rather than reflecting a generalized aversion to preferential treatment alone. Yet these negative perceptions do not translate into corresponding material penalties: allocations toward caste-based beneficiaries are directionally compensatory, particularly among low-income beneficiaries, whereas allocations under income-based affirmative action align more closely with perceived competence. Overall, our findings reveal a duality in responses to affirmative action: identity-driven competence stigma can coexist with redistributive behavior that does not penalize marginalized beneficiaries. |
| Date: | 2026–05–23 |
| URL: | https://d.repec.org/n?u=RePEc:ash:wpaper:164 |
| By: | Yana Gallen; Melanie Wasserman |
| Abstract: | This paper provides the first causal evidence that gender affects the information an individual receives about careers. We conduct a large-scale field experiment in which real college students seek career information from 10, 000 working professionals. We randomize whether a professional receives a message from a male or a female student. When students ask broadly for information about a career, female students receive substantially more information on work/life balance than male students. This gender difference persists when students specifically ask about work/life balance. A survey of professionals suggests non-altruistic motives for discussing work/life balance with women. Combining findings from the field experiment and results from an information intervention, we conclude that gender gaps in information received about work/life balance are consequential for gender gaps in career intentions. |
| JEL: | C93 J16 J24 J71 |
| Date: | 2026–09 |
| URL: | https://d.repec.org/n?u=RePEc:nbr:nberwo:35750 |
| By: | Michael King (Department of Economics, Trinity College Dublin); Paolina Medina (University of Houston); Roland Umanan (Department of Economics, Trinity College Dublin, University of Dublin); Ray Charles Howard (University of Virginia) |
| Abstract: | Cash remains a dominant payment method in many markets despite the widespread availability of digital alternatives such as debit cards, creating costs for consumers, firms, and governments. This research examines which interventions most effectively increase debit‐card use, how effectiveness varies with incentive form and amount, and whether average effects conceal meaningful differences across customers. A large‐scale randomized field experiment with approximately 1.5 million customers of a Mexican bank tests cash and gift‐card incentives of 60 and 300 MXN, repeated benefit‐focused nudges, and combinations of nudges with cash incentives. Estimated treatment effects on debit‐card transaction frequency and spending range from approximately 0% to 5%, with larger incentives producing less‐than‐proportional gains in transaction frequency and incentive‐form effects varying across outcomes and incentive amounts. Causal‐forest estimates show substantial treatment‐effect heterogeneity: for example, the effect of the 300 MXN cash incentive is concentrated among customers with relatively high baseline debit‐card use, while some near‐zero average effects mask offsetting positive and negative subgroup responses. Nudges have small nonsignificant effects alone and no detectable incremental effect when combined with cash incentives. These findings show that effective payment‐activation strategies depend not only on intervention design but also on which customers receive them. |
| Keywords: | Financial Inclusion, Debit Cards, Cash and In-kind Incentives, Nudges, Spillover Effects |
| JEL: | C93 D12 D14 D18 G21 G51 |
| Date: | 2026–08 |
| URL: | https://d.repec.org/n?u=RePEc:tcd:tcduee:tep1826 |
| By: | Chizhe Cheng; Paan Jindapon; Ajalavat Viriyavipart |
| Abstract: | Informal risk sharing enables individuals to smooth consumption when access to formal insurance is limited. While previous studies have examined the roles of income correlation and initial income inequality, little is known about how differences in individual risk exposure affect voluntary risk sharing. We investigate this question using a laboratory experiment based on an indefinitely repeated risk-sharing game. Subjects are randomly assigned to either homogeneous income risk, in which both members of a pair face the same level of risk, or heterogeneous income risk, in which one subject faces high income risk and the other faces low income risk. The benchmark repeated-game model predicts that transfers should be highest under homogeneous high income risk and lower, but comparable, under homogeneous low income risk and heterogeneous income risk. Consistent with this prediction, transfers are higher when both subjects face high income risk than when both face low income risk. However, subjects facing heterogeneous income risk exhibit a higher incidence of positive transfers than those in the homogeneous-risk configurations. Although average transfers under heterogeneous income risk lie between those observed under homogeneous high and homogeneous low income risk, they are considerably closer to the former than the latter. These findings provide the first experimental evidence that heterogeneous income risk does not undermine voluntary risk sharing and may instead encourage greater participation in reciprocal transfer arrangements. |
| Keywords: | Risk sharing; Heterogeneous risk; Cooperation; Infinitely repeated games; Economic experiment |
| JEL: | D81 C91 C73 O17 |
| Date: | 2026–09 |
| URL: | https://d.repec.org/n?u=RePEc:pui:dpaper:262 |
| By: | Marcus Giamattei (Frankfurt School of Finance and Management); Pierfrancesco Alaimo di Loro (LUMSA University, Rome, Italy; HURfuture Research Center); Matteo Rizzolli (GEPLI Department, LUMSA University, Rome, Italy; HURfuture Research Center); Stefan Voigt (Hamburg University) |
| Abstract: | We test the replicability of ten stylized facts about voluntary contributions to public goods using a large crowd-sourced dataset of classroom experiments on the classEx platform. Our sample comprises 81, 391 contribution decisions by 10, 891 players across 377 sessions in 16 countries (2019-2025). Five facts replicate: initial contributions above Nash (37.6% of endowment), non-zero final-round contributions (38.7%), conditional cooperation as the modal behavioral type, imperfect matching, and cross-country heterogeneity. One is partially confirmed (a modest end-game effect) and one is weakly supported (a positive MPCR effect, sensitive to country fixed effects). Three do not replicate. Contributions follow a hump-shaped trajectory - rising through round 5, plateauing, then gently declining - rather than the canonical decline. Only 34.4% of players free-ride in the final round (vs. the laboratory benchmark of over 70%), while a substantial minority (22.1%) sustain full cooperation. Societal indicators do not predict country-level cooperation, partly because cross-country variation explains only 0.8% of individual contribution variance - an order of magnitude less than organization-level heterogeneity (6.8%). Taken together, the canonical behavioral building blocks of the public goods game reproduce in the classroom, whereas several of its dynamic and cross-country regularities - most notably the absence of the canonical decline - do not. (Includes Online Appendix with experimental interface and additional figures.) |
| Keywords: | Public goods game; Cooperation; Replication; Classroom experiments; ClassEx; Conditional cooperation |
| JEL: | C91 C92 D70 H41 |
| Date: | 2026–09 |
| URL: | https://d.repec.org/n?u=RePEc:lsa:wphurf:wphurf05 |
| By: | Munoz Boudet, Ana Maria; Kundu, Sayan; Moscoe, Ellen |
| Abstract: | This study reports results from a randomized controlled trial of a household-level behavioral intervention designed to improve complementary feeding practices among mothers of children aged 6–24 months in 2, 682 households in Bihar, India. Mothers in treatment households received a wall-mounted food-tracking journal intended to increase the salience and habitual practice of recommended feeding behaviors. The intervention increased mothers’ attention to and automaticity of key feeding practices and raised the probability that children achieved minimum dietary diversity. However, it had no detectable effects on the quantity of food consumed or children’s mid-upper arm circumference. Impacts varied by the frontline worker delivering the intervention, while limited household food availability constrained improvements in dietary diversity. Better-off households exhibited larger gains in knowledge, perceived norms, and habit formation, although treatment effects on dietary diversity did not differ by household economic status. These findings suggest that low-cost behavioral tools that make food choices more salient and easier to track can complement conventional infant and young child feeding programs, including home visits, by reinforcing reminders and supporting habit formation among caregivers. |
| Date: | 2026–09–21 |
| URL: | https://d.repec.org/n?u=RePEc:wbk:wbrwps:11461 |
| By: | Gerrit Bauch; Arthur Dolgopolov; Manuel Foerster |
| Abstract: | We investigate the strategic communication of narratives under model uncertainty. The sender has private information about the true data-generating process of publicly observable data. The receiver is uncertain about how to interpret the data, but aware of the sender's incentives to strategically provide interpretations (``narratives''). We theoretically show that the size of the conflict of interest between the sender and the receiver is a crucial determinant of equilibrium communication. In particular, the stronger the sender's bias, (i) the more senders exaggerate their information and (ii) the more receivers correct the sender's action recommendations. In a laboratory experiment, we find evidence in line with both predictions, suggesting that people in complex and uncertain environments take a narrator's strategic incentives into account. Additional analyses reveal that narrative likelihood does not drive receiver behavior and that narratives are, on average, slightly persuasive only when bias is low. |
| Date: | 2026–09 |
| URL: | https://d.repec.org/n?u=RePEc:arx:papers:2609.10074 |
| By: | Travis Baseler (University of Rochester); Lipeng Chen (Analysis Group); Thomas Ginn (Center for Global Development) |
| Abstract: | Survey data can be distorted through mistrust, discomfort, or demand effects on the part of respondents. We test whether familiar enumerators—those with whom a respondent has completed a prior survey—influence data quality in a panel survey with small business owners in Uganda. We randomly assign respondents to a familiar or a new enumerator and cross-cut a second randomization to an in-person or phone-based interview modality. Across a broad set of outcome types, we observe few impacts of either survey method on means, distributions, or attrition. However, respondents are significantly more likely to give socially desirable answers to new surveyors compared to familiar ones. The effect of familiarity does not interact significantly with a prior experiment conducted with the same sample, suggesting that familiarity influences estimates of levels but not of treatment impacts. These findings suggest that surveys by familiar enumerators can improve the measurement of sensitive beliefs. |
| Keywords: | survey design, enumerator effects, social desirability bias |
| JEL: | C42 C81 |
| Date: | 2026–09–01 |
| URL: | https://d.repec.org/n?u=RePEc:cgd:wpaper:754 |
| By: | Vinicius Ferraz; Leon Houf |
| Abstract: | People increasingly compete against AI agents rather than other human opponents. We distinguish two channels: an opponent effect and an information effect. These are different elements with different consequences: the opponent effect is specific to a given computational system, the information effect a property of the information environment that an organisation or policymaker can control. We separate them in a preregistered experiment (N = 1, 395) using a dynamic all-pay auction, a repeated contest in which escalation of commitment arises from the incentives. What participants are told about the opponent (human, an AI trained to imitate people, or an AI trained to compete well) is varied and crossed with who they actually face, in a deception-free design. What people are told influences escalation: the median price rises by 6.7 points when a human might be the opponent and falls by 8.8 when an optimising machine might be, a spread of about 15% of the prize value of the competition, produced by information alone. Competing against the AI agents lowers prices, yet reduces the chance that both sides finish with positive earnings, showing distinct effects of the opponent channel. The information effect is not explained by articulated strategy, or individual differences, and is consistent with a competitive response engaged when a human is a live possibility. This shows that describing an AI competitor is not behaviourally neutral. |
| Date: | 2026–09 |
| URL: | https://d.repec.org/n?u=RePEc:arx:papers:2609.21439 |
| By: | Zhongheng Qiao |
| Abstract: | People often face environments where multiple models compete to explain the same observations. This paper examines how people update beliefs in such settings and how preferences over payoff-relevant states shape model selection and belief updating. This paper first develops a framework where preference-driven bias distorts the perceived model, affecting Bayesian and best-fit updating differently. In a laboratory experiment, most participants are classified as Bayesian updaters, who average across models, while a substantial minority are classified as best-fit updaters, who select the model that best fits the observed signal. Within-participant comparisons between the symmetric payoff and asymmetric payoff conditions indicate that asymmetric payoffs shift reported beliefs toward the preferred state, particularly among participants classified as best-fit updaters. Relative to symmetric payoffs, asymmetric payoffs increase the reported belief of the preferred state by about 8 percentage points among best-fit updaters, while the estimated effect among Bayesian updaters is close to zero. These findings help us better understand model-based learning and have implications for domains such as political polarization and financial investment, where competing narratives and strong preferences often coexist. |
| Date: | 2026–09 |
| URL: | https://d.repec.org/n?u=RePEc:arx:papers:2609.22446 |
| By: | Lewin, Peter; , Danyelle; Zinn, Anna Kristina (University of Queensland); MacInnes, Sarah; Dolnicar, Sara (The University of Queensland) |
| Abstract: | Hotel sustainability research shows that pro-environmental behaviour‑change interventions are often effective in field experiments, but widespread adoption by hotels remains limited. One possible explanation is that hotels lack external recognition that would make adopting these measures visible to guests and rewarding to managers. This study investigates whether university-affiliated sustainability badges could serve as a recognition mechanism, testing their effectiveness across two studies drawing on Signalling Theory and Social Identity Theory. In an online experiment (N=975), prospective hotel guests viewed simulated hotel booking pages featuring one of the four badge options, or no badge, across budget, mid-range, and luxury hotel pricing segments. All badge variations increased perceived hotel sustainability and guests’ sense of belonging to a community of environmentally conscious travellers (group identification). Badges did not trigger negative emotions, but the university affiliated logo did not perform better than a generic environmental logo on almost all measures, suggesting these effects are based on the presence of any badge, rather than the university partnership specifically. Semi-structured interviews with ten hotel managers found conditional receptivity to university-affiliated badges, primarily through a sense of shared values and community orientation with the partnering institution (collective values) rather than the badge’s function as a credibility signal. These findings suggest that collective values and shared identities, rather than signalling, may be the more promising mechanism for engaging hotels and guests around sustainability initiatives, and offer an encouraging direction for potential university-hotel partnerships that may translate into the future implementation of pro-environmental behaviour change interventions. |
| Date: | 2026–09–18 |
| URL: | https://d.repec.org/n?u=RePEc:osf:socarx:2pv9b_v2 |
| By: | Coville, Aidan; Osman, Adam; Piza, Caio |
| Abstract: | Expanding market access via digital technologies is seen as a key pathway for growth, yet adoption remains low among small enterprises. This paper investigates barriers to entry through two randomized experiments in the country of Georgia. The findings show that a "supply-side" training intervention failed to increase digital participation, despite high initial interest. In contrast, a "demand-side" conditional purchase order increased market access by 24 percentage points, while a payment six times larger generated only a modest additional increase. The analysis finds no complementarity between training and demand incentives. The results highlight demand-side incentives as a cost-effective policy to kickstart adoption. Although the effects largely dissipate over time as control firms catch up, firms with higher baseline readiness for e-commerce remain more likely to engage in digital markets several years later, with suggestive evidence that the demand shock accelerated adoption among this group. The paper shows that the remaining barriers to growth are likely behavioral and organizational frictions rather than simple skill or capital deficits. |
| Date: | 2026–09–21 |
| URL: | https://d.repec.org/n?u=RePEc:wbk:wbrwps:11460 |
| By: | KINA, MEHMET FUAT; Ekmen, Helin Yaren |
| Abstract: | Large language models are increasingly used to simulate survey respondents, yet the inference pipeline that turns a persona into a response is rarely calibrated for the population and outcome it is applied to. This study makes population- and outcome-specific calibration the research task itself. Using a nationally representative survey in Turkey as the benchmark, we construct persona-conditioned synthetic respondents for all 2, 615 survey participants and search a structured space of persona content, sampling strategies, temperatures, model backbones, and prompt framings. Calibration is assessed against two held-out outcomes that carry the response structures on which social movement research typically rests, a rare binary measure of protest participation and a five-point attitude item, and the selected protocols differ between the two. We then run a within-persona synthetic experiment that places the same protest for Kurdish-language education in a political opportunity and a political threat context. The threat context lowers simulated movement legitimacy and willingness to attend, and the simulated contrasts differentiate by political orientation, ethnicity, and education in the directions theory anticipates. Calibration is a precondition for reading synthetic experiments, not a robustness check appended to them. |
| Date: | 2026–09–24 |
| URL: | https://d.repec.org/n?u=RePEc:osf:socarx:hp7gz_v1 |
| By: | Bremer, Björn (Central European University); Chwieroth, Jeffrey |
| Abstract: | Central banks relied for a decade on unconventional monetary policies (negative interest rates, large-scale asset purchases, and forward guidance) that carry significant distributive consequences and became intensely politicized. Yet little is known about the mass politics of central banking, or whether contestation over specific instruments threatens central banks' broader legitimacy. We provide evidence from two pre-registered survey experiments in Germany and the Netherlands. A conjoint experiment shows that citizens strongly oppose negative interest rates, the most penalized levels in the design, and are skeptical of unconditional asset purchases. A framing experiment shows support for negative rates responds to both egotropic and sociotropic arguments, with pocketbook concerns resonating most strongly. Exploiting a within-respondent measure of trust in the European Central Bank (ECB), we show that these same frames shift citizens' specific support for the policy without eroding their diffuse trust in the ECB, a highly insulated, technocratic, non-elected institution. |
| Date: | 2026–09–16 |
| URL: | https://d.repec.org/n?u=RePEc:osf:socarx:es7wf_v1 |
| By: | Dean Karlan; Natalia Rigol; Benjamin N. Roth |
| Abstract: | Evaluations of microenterprise credit typically measure effects only on borrowing firms. But what about their customers? If microenterprises sell relatively undifferentiated goods and services, as is often hypothesized, credit may simply reallocate sales across firms and create little consumer benefit. In a randomized controlled trial in Chile, large loans increased treated firms’ profits by USD 292 per month, a 13.4% increase. Customer survey data indicate even larger benefits for customers: a gain of USD 494 per month in consumer surplus. Furthermore, using a sample of more than 125, 000 non-treated firms operating in the same markets, we find little evidence of business stealing. The welfare gains from credit expansion thus extend well beyond the borrowers themselves. |
| JEL: | D53 L26 O12 |
| Date: | 2026–09 |
| URL: | https://d.repec.org/n?u=RePEc:nbr:nberwo:35729 |
| By: | Maxim Chupilkin |
| Abstract: | AI systems are becoming participants in the evaluation and use of scientific research. They encounter citation counts, download statistics and lists of popular articles developed around human readers, but the collective consequences of these signals for artificial readers remain uncertain. This paper adapts the Music Lab design to a market for academic attention. In the first experiment, 1, 000 AI agents choose papers from the titles and abstracts of all 114 regular research articles published in the American Economic Review in 2025. The experiment has five independent-choice communities and five social-influence communities, each with 100 sequential agents. Only agents in the social-influence condition observe earlier selections within their community. Agents may select any number of papers. Social-information communities select 17.2 percent fewer papers per agent, concentrate their choices more heavily, and collectively cover 73 papers, compared with 90 independently. Between-community variation is greater under social information. In a second experiment with 200 agents across twenty social communities, randomly assigning papers five initial selections raises their subsequent selection rate by 45.55 percentage points (95% CI: 41.20 to 49.90). Choices have modest correspondence with external citations and little correspondence with download counts. The results show how a simple information rule shapes the volume, breadth and distribution of scientific attention in an artificial population. |
| Date: | 2026–09 |
| URL: | https://d.repec.org/n?u=RePEc:arx:papers:2609.22408 |
| By: | Yu Liu; Wenwen Li; Yifan Dou; Guangnan Ye |
| Abstract: | In-context learning (ICL) enables large language model (LLM) agents to improve decisions using interaction history, yet it remains unclear whether such improvement reflects refined internal reasoning or mere extrapolation of statistical patterns. To disentangle these mechanisms, we study LLM agents in multi-agent incomplete-information games that require recursive belief reasoning. By constructing a public goods game and manipulating the statistical structure of historical feedback, we evaluate decision quality against a history-independent rational expectations equilibrium (REE) benchmark. Our experiments reveal that when historical statistical patterns are disrupted, the benefits of longer context largely vanish, degrading decision quality to the no-context baseline in a way sharply amplified by stronger strategic interdependence. These results suggest that, in such strategic environments, ICL behavior is more consistent with statistical extrapolation than with strategic reasoning. Our work extends the mechanistic study of ICL to strategic multi-agent settings, introduces REE as a diagnostic tool for distinguishing reasoning from extrapolation, and provides a reusable framework for probing the boundaries of LLM reasoning in recursive belief tasks. |
| Date: | 2026–09 |
| URL: | https://d.repec.org/n?u=RePEc:arx:papers:2609.18591 |
| By: | Karun Adusumilli; Jiaying Gu; Junfan Tao |
| Abstract: | We investigate Empirical Bayes (EB) methods in the context of compound adaptive experiments, where the arm distribution in each experiment follows a normal distribution with an unknown mean that we seek to estimate. There are two main EB strategies: $g$-modeling, which estimates the prior by maximizing the marginal likelihood, and $f$-modeling, which derives posterior means directly from the empirical distribution of the observations. We show that $g$-modeling continues to be a valid EB procedure even when it incorrectly assumes that data are collected exogenously; its validity does not depend on the particular sampling algorithm or on whether sample sizes are endogenous. In practice, one can apply standard $g$-modeling techniques by acting as though the data were exogenously sampled. We extend regret guarantees from exogenous sampling to adaptively generated data. By contrast, naively applying the Tweedie formula based on the marginal density of the observed data, as in standard $f$-modeling, can produce biased rules under adaptive sampling. We corroborate the robustness of $g$-modeling through simulations with widely used adaptive algorithms and demonstrate its applicability using a real-world dataset consisting of multiple sequential experiments. |
| Date: | 2026–09 |
| URL: | https://d.repec.org/n?u=RePEc:arx:papers:2609.17158 |
| By: | Anna Maria Boros (University of Warsaw, Faculty of Economic Sciences); Anna Małgorzata Bartczak (University of Warsaw, Faculty of Economic Sciences; University of Warsaw, Robert Zajonc Institute for Social Studies) |
| Abstract: | Perceived realism of the policy scenario presented in stated preference (SP) valuation depends on trust in the scientific soundness of the projected policy outcomes (epistemic credibility) and in policymakers' capability to deliver them (implementation credibility). Using a discrete choice experiment (DCE) with 1, 100 Polish urban residents, we examine how policy scenario credibility shapes preferences for air pollution mitigation reducing risks from chronic obstructive pulmonary disease (COPD). We contrast domain experts and artificial intelligence (AI) with no-source control as information sources on policy outcome projections. We find that epistemic and implementation credibility are strongly correlated; thus, we treat credibility as a single variable affecting preferences. Using a hybrid mixed logit (HMXL) model, we demonstrate that higher perceived credibility increases programme support, lowers cost sensitivity, and thus elevates willingness to pay (WTP) for the programme and morbidity risk reduction. We find no evidence that the information source influences preferences or scenario credibility. |
| Keywords: | air pollution, artificial intelligence, chronic obstructive pulmonary disease, credibility, discrete choice experiment, information source |
| JEL: | D91 Q51 Q53 |
| Date: | 2026 |
| URL: | https://d.repec.org/n?u=RePEc:war:wpaper:2026-33 |
| By: | Alex Smolin; Bryan Wilder |
| Abstract: | Inferring intelligence from observable behavior is a foundational challenge in artificial intelligence. We develop a theory of Bayesian intelligence for agents such as language models. Each prompt induces a possibly imperfect internal experiment; the agent updates a full-support prior by Bayes' rule and faithfully reports its posterior over the possible answers to the question. Repetitions draw fresh, independent outcomes from the same unobserved experiment at one fixed state. We show that the agent's behavior admits this explanation if and only if its reports are not fully contradictory, i.e., some state remains possible under every report across all prompts. Report frequencies and the sizes of positive probabilities impose no further restrictions. We further propose and characterize the behavioral implications of an intelligence order that makes the behavior of two agents consistent with one agent having access to a more informative experiment: there should exist a coupling of report distributions such that the more informative agent's report excludes every answer excluded by its counterpart. Finally, we show the difficulty of aggregating coarse reports from intelligent agents: unless the agent reports a belief about the complete state of the world, the optimal aggregation can assign arbitrary weights to states that have not been excluded. These results provide a basis for understanding when agents' behavior is intelligent and highlight the difficulty of rejecting Bayesian rationality. |
| Date: | 2026–09 |
| URL: | https://d.repec.org/n?u=RePEc:arx:papers:2609.14724 |
| By: | Johnston, Cliff Hurt; Green, Ted; Warren, Madeline; Meade, Valerie |
| Abstract: | Meta-analytic evidence links sustained officer fidelity to core correctional practices with substantially lower participant recidivism; caseloads supervised by untrained officers recidivate roughly 38% more often than those supervised by trained officers (Chadwick et al., 2015), and fidelity after training — not the training event itself — is the moderator that makes the difference (Labrecque et al., 2023). Our companion study validated the measurement layer: a frontier large language model can grade recorded supervision visits against a standards-based rubric about as reliably as expert human auditors (Green, Warren, & Meade, 2026). This paper tests the practice-change layer: whether an LLM can turn those transcripts and scores into officer-facing coaching that an independent expert judges to be as good as coaching written by expert human reviewers. On the same 30-visit corpus, we assembled 120 coaching outputs — four per visit, from two expert human sources (authentic prior grader coaching notes) and two LLM sources (Claude Sonnet 4.6 and GPT-5.5), randomized within visit — and had a third expert, Valerie Meade, score every output on eight pre-specified 1–5 dimensions, rank the four outputs within each visit, and make a product-display decision. The review was blinded: source identities and source types were withheld until scoring was complete, and the reviewer’s own historical coaching was excluded from the packet. Decoded against the internal source key, the strongest LLM source (Claude Sonnet 4.6) was rated equal to or better than both directly compared human sources on the composite quality index (vs. one human source +0.60, Holm-adjusted p = 0.006; vs. the other +0.28, not significantly different and formally non-inferior at a 0.4 point margin, one-sided p = 0.0002), earned the best mean within-visit rank (2.07 of 4), the most first-place outputs (10 of 30), and the most display-ready outputs (13 of 30). Pooled by type, LLM coaching outscored human coaching on the visit-level composite (3.44 vs. 3.14; paired difference +0.30, 95% CI +0.10 to +0.52, p = 0.019). We report the result with its boundaries: one expert reviewer, two directly compared human sources, non-standardized human comparison text, a 30-visit corpus, and no claim of field behavior change or recidivism impact. Together with the grader validation, the finding supports the working thesis that agencies can responsibly act on validated visit scores with display-filtered LLM-generated coaching — the mechanism the training literature ties to lower recidivism — while the outcome link itself awaits field study. |
| Date: | 2026–09–08 |
| URL: | https://d.repec.org/n?u=RePEc:osf:lawarc:m7d5g_v1 |
| By: | Pengfei Tian; Jizhou Liu; Lei Shi; Peng Ding |
| Abstract: | We study randomized experiments involving two interacting populations, such as buyers and sellers in a marketplace. In the two-sided experiments we consider, we randomize the two populations separately and independently. For a pair consisting of one member from each population, the two assignments jointly determine one of four exposure conditions. Under a local interference assumption, we consider a broad class of linear estimands, including total, interaction, and buyer- and seller-side spillover effects. Our first main result establishes that researchers can estimate these effects using ordinary least squares and conduct asymptotically valid design-based inference using the conventional two-way cluster-robust variance estimator, clustered at the buyers' and sellers' levels. Our second main result develops a sharper variance estimator for a single linear estimand that better preserves dependence within the buyer and seller dimensions and is asymptotically less conservative than the two-way clustered estimator and existing alternatives. Our third main result establishes the theory for covariate adjustment and recommends a two-way analysis-of-variance-type covariate representation to ensure efficiency gains. |
| Date: | 2026–09 |
| URL: | https://d.repec.org/n?u=RePEc:arx:papers:2609.22761 |
| By: | Small, Meg; Dancer, Sally; Wright, Jonathan; Colaianne, Blake Anthony; Sommerville, Gillian |
| Abstract: | Team effectiveness in high-stakes, time-limited work contexts remains an underexplored area of organizational and applied psychology research. This study examined whether daily process behaviors—specifically team communication and collaboration—predict end-of-program team effectiveness outcomes in the context of an intensive, eight-week technology entrepreneurship incubator program. Participants were members of five start-up teams enrolled in the start-up program at a large research university. As part of a larger mixed-methods study, participants completed daily diary measures of team communication and collaboration using a digital platform (Metricwire) for 15 days. At program end, participants completed a post-program survey assessing team effectiveness outcomes including creativity, psychological safety, team viability, collective efficacy, authentic relationships, opportunities for growth, and program satisfaction. Simple linear regressions revealed that daily collaboration significantly predicted creativity, psychological safety, team viability, and collective efficacy. Daily communication similarly predicted creativity and psychological safety. Authentic relationships were not significantly predicted by either behavioral measure. These findings contribute to the growing literature on team process and performance by demonstrating that observable daily team behaviors capture meaningful variance in team effectiveness outcomes, Implications for team-based intervention design and digital measurement methodology are discussed. |
| Date: | 2026–09–17 |
| URL: | https://d.repec.org/n?u=RePEc:osf:socarx:2ke9c_v1 |
| By: | Ying Jin; Naoki Egami |
| Abstract: | Policy learning aims to determine who should be treated based on individual characteristics. In high-stakes settings such as medicine and public policy where safety is a central concern, improving the average outcomes alone may not be sufficient: decision makers may also seek to protect individuals from harm, in line with the Hippocratic principle of ``do no harm.'' In this paper, we propose \textit{conformal policy learning} (CPL), a policy learning procedure with a new distribution-free safety guarantee that controls the probability of assigning treatment to an individual who would be harmed relative to control. CPL views each treatment decision as testing a hypothesis of counterfactual harm and assigns treatment by thresholding conformal p-values. These p-values use observable proxies and selective calibration to address the challenge that the potential outcomes under comparison are never simultaneously observed. For randomized experiments, under standard exchangeability conditions, CPL provides finite-sample safety guarantee at a user-specified level, without imposing any outcome modeling assumptions. Moreover, when the outcome model is consistently estimated, CPL achieves asymptotically optimal welfare subject to the safety constraint. In observational studies, CPL with learn-then-balance weights achieves doubly robust safety guarantees. We evaluate CPL through extensive simulations and apply it to an empirical study of AI-powered interventions designed to reduce conspiracy beliefs. |
| Date: | 2026–09 |
| URL: | https://d.repec.org/n?u=RePEc:arx:papers:2609.17296 |
| By: | Burak Agachan; Max van Duijn; Amirhossein Zohrehvand |
| Abstract: | Hierarchical orchestration, in which a Manager agent reviews worker output and can send it back for revision, is the default coordination pattern in production multi-agent LLM frameworks. Classical organizational theory predicts that the authority link speeds convergence on decisive output; work on sycophancy and Degeneration-of-Thought predicts that authoritative critique makes LLM output worse. Prior comparisons vary whole frameworks on tasks with checkable answers, leaving the authority link untested on open-ended work. We present a paired experiment that holds five LLM agents, their roles, prompts, tools, models, and data fixed and varies one link: whether the Manager may reject a worker's output and oblige a revision. Across 43 paired products and 86 runs of a business-intelligence reporting task, a five-model judge panel and a deterministic specification check score every report. The flat organization scores higher on Utility (d = 0.42, p = 0.009) and on Writing Clarity (d = 0.34, p = 0.030); the classical prediction fails. The reports are the same length, but hierarchical reports hedge 53% more, each revision loop is associated with a 0.14-point drop in Writing Clarity, and the hierarchical Writer's first draft is indistinguishable from the flat report: the gap opens inside the revision loop. Specification accuracy is at ceiling in both organizations, and the supervisory tier costs 51.5% more tokens for no quality gain. A supervisor pays for itself when it can verify and becomes a liability when it can only opine. |
| Date: | 2026–09 |
| URL: | https://d.repec.org/n?u=RePEc:arx:papers:2609.14767 |
| By: | Lakshya Katariya; Cindy Lopes Bento; Christoph Grimpe |
| Abstract: | Public money for research is distributed through peer-review but it has been criticized to disfavour risky science. Traditional review typically asks reviewers to compress multiple dimensions such as novelty, feasibility, impact, etc. into a single overall score, and this compression may penalise proposals whose potential is high but uncertain. The subjective expected utility (SEU) framework instead asks them to assess separately the utility and likelihood of a proposal's potential outcomes without issuing an overall compressed score. This paper examines whether structuring peer review in this way, systematically alters the ranking of grant proposals. Using a field experiment embedded in a real funding call, we have the same proposals evaluated under either traditional criteria or SEU criteria, with a treatment assignment that we argue to be as good as random. We find that the two methods produce very different orderings of the same applications: their percentile ranks are essentially uncorrelated, and individual applications shift considerably in relative position. SEU reorders applications along dimensions traditional review leaves implicit: 'High-Likelihood High-Utility' applications are ranked higher and 'Low-Likelihood Low-Utility' applications lower under SEU as compared to traditional review. Importantly, we do not observe a significant up- ranking of ‘Low-Likelihood High-Utility’ applications under SEU relative to the traditional review. On the contrary, 'High-Likelihood Low-Utility' application show a significant yet modest down-ranking under SEU relative to traditional review. These ranking differences translate into substantially different funding outcomes. Applying the funding cutoff that corresponds to the number of applications actually funded to each ranking, the 3 applications that would be funded under SEU were, without exception, those that referees had classified as high-likelihood (23 'High-Likelihood High-Utility' and 2 'High-Likelihood Low-Utility'). Traditional review, while also favouring high-likelihood submissions, would fund a mix across all four likelihood-utility categories, including the low-likelihood applications (roughly a third of the pool) that fail to clear the cutoff under SEU. On average, only 3.5 of the 25 applications funded under traditional review would also be funded under SEU, corresponding to a turnover rate of 86%. Our findings show that the structure of evaluation systematically shapes which applications rise to the top and, under a fixed funding cutoff, which would be funded, depending on whether referees issue a single overall judgment or score utility and likelihood separately. |
| Keywords: | Subjective expected utility, high-risk high-reward, feasibility, impact, grant proposal, peer-review, evaluation, funding |
| Date: | 2026–09–17 |
| URL: | https://d.repec.org/n?u=RePEc:ete:msiper:793640 |
| By: | Mira Frick; Ryota Iijima; Yuhta Ishii; Nicholas Wu |
| Abstract: | Consider an auction with buyers whose values depend on an underlying state (e.g., market fundamentals). How does the auction format shape the information that buyers' bids reveal about the state? We recast auctions as statistical experiments and compare different auction formats in terms of the (Lehmann) informativeness of the induced experiments. Our main finding is that among a large class of auctions (e.g., $k$th-price, all-pay), the first-price auction is the most informative. As a result, this auction guarantees the highest payoffs to a decision-maker who uses the information revealed by buyers' bids in a monotone decision problem (e.g., a prediction problem or the choice of a reserve price in a future auction). |
| Date: | 2026–09 |
| URL: | https://d.repec.org/n?u=RePEc:arx:papers:2609.21168 |
| By: | Janson, Joakim (Linnaeus University and); von Essen, Emma (Department of Sociology, Uppsala University) |
| Abstract: | We study how anonymity affects fake-news sharing on social media using a natural experiment on Flashback, a large Swedish online discussion forum. In 2014, journalists obtained information that allowed them to identify users registered before March 2007, sharply increasing these users’ risk of public exposure while leaving later registrants unaffected. Using a difference-in-differences design on 1.98 million posts, we find that exposed users reduce references to fake-news sites by about one-quarter of the mean. The effect is concentrated in immigration and domestic-politics discussions and does not extend to legitimate news sources, suggesting that accountability reduces misinformation without reducing news sharing more broadly. |
| Keywords: | Online; Anonymity; Text data; Fake news; Machine learning |
| JEL: | D10 D80 D83 D90 |
| Date: | 2026–09–16 |
| URL: | https://d.repec.org/n?u=RePEc:hhs:iuiwop:1566 |
| By: | Davood Wadi; Yu Ma |
| Abstract: | Large language models (LLMs) now serve as conversational shopping assistants on platforms that also sell advertising. These AI agents face a conflict of duty. They advise consumers who rely on their judgment, yet are deployed by platforms that benefit when sponsored listings are chosen. Sponsorship disclosures, designed to allow consumers to penalize paid placements, now reach the AI agent rather than the consumer, and the agent's evaluation of them is hidden from the consumer. Drawing on the fiduciary concept of conflict of duty, we argue that an agent's evaluation of a sponsored listing should not depend on which party deployed it. In controlled choice experiments, we manipulate assigned roles in the system prompt to name either a traveler or a booking platform as the agent's principal. Platform delegation significantly attenuates the penalty that agents apply to sponsored listings and weakens the skepticism that disclosure triggers in their reasoning traces. We replicate out findings across LLMs and reasoning depths. A second study decomposes the disclosure label and shows that the divergence between the two delegates widens significantly when the paid placement is attributed to the platform. Stricter terminology ("Sponsored" instead of "Promoted") lowers choice of paid listings but does not close this gap when the platform is named. The findings show that disclosure mandates designed for human consumers cannot by themselves protect consumers in AI-mediated commerce. |
| Date: | 2026–09 |
| URL: | https://d.repec.org/n?u=RePEc:arx:papers:2609.17989 |
| By: | Pedro Cadahia |
| Abstract: | This paper presents a causal decision-making framework for estimating price elasticity in retail channels, a process typically confounded by promotions, competitor movements, and market frictions. Rather than forcing a calculation when data is ambiguous, the system introduces decision abstention (\textsc{wait}) as an active diagnostic tool rather than an estimation failure. Combining Double Machine Learning and conformal prediction, the tool evaluates whether reliable conditions exist to adjust prices or if pausing the decision is preferable. When the system abstains, it exhaustively classifies the reason for the pause, identifying which products require designed pricing experiments or whether aggregating data to the brand level restores usable estimates. Tested on controlled synthetic data, the model shows that this operational discipline drastically reduces estimation error (lowering RMSE from 0.571 to 0.159) and offers a practical, secure alternative to blind estimation in thin-data retail environments. |
| Date: | 2026–09 |
| URL: | https://d.repec.org/n?u=RePEc:arx:papers:2609.10615 |