Equity appeals are increasingly used by digital fundraising platforms, nonprofits, and public institutions to direct attention and resources toward disadvantaged communities. However, it remains unclear whether and when equity appeals actually increase giving. We examine this question in the context of educational crowdfunding, where platforms explicitly focus on reducing funding disparities across schools, particularly for students from racial or ethnic minority and low-income backgrounds. Leveraging large-scale data from DonorsChoose, one of the largest educational crowdfunding platforms in the United States, and exploiting arbitrary cutoffs in the platform’s deployment of equity appeals based on the student composition of benefitting schools, we show that equity appeals increase fundraising when they highlight student disadvantage in terms of poverty while providing little to no measurable benefit when they highlight student disadvantage in terms of race. These differential effects reflect how donors interpret disadvantage. Many donors appear to view poverty as a legitimate and actionable barrier to learning, making poverty-based appeals effective. In contrast, perceptions of race as a structural barrier to educational opportunity are more heterogeneous and politically sensitive, limiting the impact of race-based appeals. For platform designers and policymakers seeking to reduce educational fundraising disparities, our findings highlight the importance of how equity appeals are framed. More broadly, our results contribute to understanding under what conditions behavioral nudges can meaningfully reduce inequality versus when alternative approaches may be necessary to achieve equitable outcomes.
Large language models (LLMs) often produce homogeneous outputs, raising concerns that AI coding assistants may lead to convergence in the software artifacts that developers create. Whether this occurs in practice is unclear because developers interactively prompt, evaluate, modify, and reject model outputs, and because outputs vary with prompt and repository context. I examine code homogenization using Kaggle contest submissions from 2019 to mid-2026. I first document widespread convergence toward the random seed value 42, consistent with LLMs reinforcing a longstanding convention in programming culture. I then study homogenization more broadly, at two levels of aggregation and abstraction. At the submission level, I measure the average pairwise similarity of submissions within contests. At the contest level, I measure the conceptual span of submitted code, motivating distinct measures for each: TF-IDF representations, which capture surface syntax, and Voyage 3 code embeddings, which capture code intent and semantics. The results demonstrate substantial syntactic homogenization at both the individual and collective levels: individual submissions have become more alike in literal syntax and code structure, while the latent dimensionality of syntactic variation has narrowed. In contrast, I find little evidence of semantic homogenization, individually and collectively. Average semantic distance remains essentially flat, and the contest-level latent dimensional span of semantic approaches remains stable, with evidence suggesting it has even expanded modestly. These findings suggest that AI coding assistants are certainly standardizing implementation details, yet they have not yet produced evidence of homogenization in the approaches and problem-solving strategies coders employ.
Telehealth technologies are transforming the U.S. healthcare system by enabling more efficient and costeffective access to quality care. However, this shift may also create new challenges, including increased opportunities for doctor shopping, where patients acquire controlled substances from multiple sources simultaneously. We examine the impact of telehealth expansion on doctor shopping by studying policy changes that increased telehealth adoption, including states' entry into the Interstate Medical Licensure Compact (IMLC) and temporary federal licensure waivers during the COVID-19 pandemic. Using private insurance claims data from 2012 to 2022, we find that telehealth expansion increases doctor shopping and drug overdoses. The effect is concentrated among telehealth users without a prior history of doctor shopping, with null effects for non-telehealth users and for patients who engaged in doctor shopping prior to telehealth expansion. This indicates that the effects operate primarily at the extensive margin of doctor shopping and that telehealth adoption is a necessary condition for the effect to emerge. Furthermore, the effect size is approximately three times larger for rural patients (residing outside an MSA) than for urban patients, suggesting a cost-reduction mechanism. Finally, we examine the moderating role of Prescription Drug Monitoring Programs (PDMPs) in the telehealth era. Our findings document an unintended consequence of telehealth expansion and offer practical insights for policymakers and healthcare system designers.
Despite increasing popularity in empirical studies, the integration of machine learning generated variables into regression models for statistical inference suffers from the measurement error problem, which can bias estimation and threaten the validity of inferences. In this paper, we develop a novel approach to alleviate associated estimation biases. Our proposed approach, EnsembleIV, creates valid and strong instrumental variables from weak learners in an ensemble model, and uses them to obtain consistent estimates that are robust against the measurement error problem. Our empirical evaluations, using both synthetic and real-world datasets, show that EnsembleIV can effectively reduce estimation biases across several common regression specifications, and can be combined with modern deep learning techniques when dealing with unstructured data.
Crypto donations now represent a significant fraction of charitable giving worldwide. Nonfungible token (NFT) charity fundraisers, which involve the sale of NFTs of artistic works with the proceeds donated to philanthropic causes, have emerged as a novel development in this space. A unique aspect of NFT charity fundraisers is the significant potential for donors to reap financial gains from the rising value of purchased NFTs. Questions may arise about donors' motivations in these charity fundraisers, potentially resulting in a negative social image. NFT charity fundraisers, thus, offer a unique opportunity to understand the economic consequences of a donor's social image. We investigate these effects in the context of a large NFT charity fundraiser, leveraging random variation in transaction processing times on the blockchain to identify the causal effect of purchasing a charity NFT on a donor's later market outcomes. Further, we demonstrate a clear pattern of heterogeneity based on an individual's decision to relist (versus hold) the purchased charity NFT (a sign of perceived strategic generosity) and based on an individual's social exposure within the NFT marketplace. We show that charity-NFT relisters experience significant market penalties with an estimated 15.3% decrease in the prices they can command for other NFTs in their portfolio. This negative effect is particularly pronounced among those who are more socially exposed. Two controlled online experiments (one incentivecompatible and one scenario-based) corroborate our findings, demonstrating that the relisting of a charity NFT for sale at a profit leads onlookers to perceive the initial donation as strategic generosity and reduces their willingness to purchase NFTs from the donor. Our study underscores the growing importance of digital visibility and traceability, features that characterize crypto-philanthropy and online philanthropy more broadly.
Problem definition: We examine moviegoers’ choices between consuming content via legal theatrical channels and illegal piracy channels. We focus on how two important factors affect this choice: the picture quality of piracy sources (a function of studio security investments) and the costs associated with legal channels, including consumer transportation costs (a function of screening volumes) and ticket prices. Methodology/results: We formulate a structural model and conduct counterfactual simulations. Our findings indicate that consumers’ choice between legal and illegal channels is significantly influenced by the quality of pirated sources and the costs of legal consumption. The emergence of high-quality piracy sources in the first week of a movie’s theatrical release leads to a 7.9% reduction in theatrical revenue during the first eight weeks of release, as compared with a scenario where only low-quality pirated sources are available. We also find that high-quality piracy sources pose a greater threat to the sales of smaller movies than of blockbusters. Managerial implications: Our work provides a rare insight into the operational decisions faced by movie studios around the theatrical delivery of movies. Our counterfactual simulations explore the potential for movie studios to manipulate the supply, cost, and quality associated with legal content to mitigate piracy and increase profit. Our results indicate that the potential for cost reductions is limited, as changes in ticket prices or screen volume have minimal impact on legal consumption. However, a moderate improvement in the value or quality of the theatrical experience (e.g., via one-time investment by theaters in upgrading equipment or technology) can effectively offset the impact of high-quality pirated content. Supplemental Material: The online appendix is available at https://doi.org/10.1287/msom.2023.0407 .
Matching platforms seek to facilitate market clearing, but congestion arises with imbalances in supply and demand. In online labor markets, when workers can apply to jobs without restriction, they may apply to an excessive volume of positions to maximize their likelihood of securing work, overwhelming employers and making screening difficult if not altogether unmanageable. Prior research has theoretically argued that imposing application costs on workers can mitigate this issue. However, the efficacy of such a solution has yet to be evaluated empirically in matching platforms where employers incur significant screening costs. We address that gap here, considering a prominent online labor market that imposed application costs on workers. We report evidence that application costs successfully improved matching outcomes via at least two channels. First, the application costs reduced bid volumes, lowering employer screening costs in turn. Second, the application costs led workers to become more selective in their applications, focusing on the employers and jobs that they were best suited for and putting greater effort into differentiating themselves, including actively reaching out to employers via direct messages. These worker-side, secondary responses enhance the first-order benefit of the application costs (i.e., constraining application volumes). Finally, the workers' increased selectivity comes paired with risk aversion among workers, who become more likely to apply for jobs requiring familiar skills and less likely to apply for jobs posted by employers in different time zones or speaking a different language. We discuss the implications of our findings for workers' longer-term career trajectories and market sustainability generally.
This study examines the impact of a query recommender system on user search behavior, sales volume, and consumption diversity.
Recent studies suggest large language models (LLMs) can generate human-like responses, aligning with human behavior in economic experiments, surveys, and political discourse. This has led many to propose that LLMs can be used as surrogates or simulations for humans in social science research. However, LLMs differ fundamentally from humans, relying on probabilistic patterns, absent the embodied experiences or survival objectives that shape human cognition. We assess the reasoning depth of LLMs using the 11-20 money request game. Nearly all advanced approaches fail to replicate human behavior distributions across many models. The causes of failure are diverse and unpredictable, relating to input language, roles, safeguarding, and more. These results warrant caution in using LLMs as surrogates or for simulating human behavior in research.
Many digital platforms offer advertisers experimentation tools like Meta's Lift and A/B tests to optimize their ad campaigns. Lift tests compare outcomes between users eligible to see ads versus users in a no-ad control group. In contrast, A/B tests compare users exposed to alternative ad configurations, absent any control group. The latter setup raises the prospect of divergent delivery: ad delivery algorithms may target different ad variants to different audience segments. This complicates causal interpretation because results may reflect both ad content effectiveness and changes to audience composition. We offer three key contributions. First, we make clear that divergent delivery is specific to A/B tests and intentional, informing advertisers about ad performance in practice. Second, we measure divergent delivery at scale, considering 3,204 Lift tests and 181,890 A/B tests. Lift tests show no meaningful audience imbalance, confirming their causal validity, while A/B tests show clear imbalance, as expected. Third, we demonstrate that campaign configuration choices can reduce divergent delivery in A/B tests, lessening algorithmic influence on results. While no configuration guarantees eliminating divergent delivery entirely, we offer evidence-based guidance for those seeking more generalizable insights about ad content in A/B tests.
Crowd-voting mechanisms are commonly used to implement scalable evaluations of crowdsourced creative submissions. Unfortunately, the use of crowd-voting also raises the potential for gaming and manipulation. Manipulation is problematic because (1) submitters' motivation depends on their belief that the system is meritocratic, and (2) manipulated feedback may undermine learning, as submitters seek to learn from received evaluations and those of peers. In this work, we consider a design approach to addressing the issue, focusing on the notion of strategic opacity, i.e., purposefully obfuscating evaluation procedures. On the one hand, opacity may reduce the incentive and thus the prevalence of vote manipulation, and submitters may instead dedicate that time and effort to improving their submission quantity or quality. On the other hand, because opacity makes it difficult for submitters to discern the returns to legitimate effort, submitters may also reduce their submission effort or simply exit the market. We explored this tension via a multimethod study employing field experiments at 99designs and a controlled experiment on Amazon Mechanical Turk. We observed consistent results across all experiments: opacity leads to reductions in gaming in these crowdsourcing contests and significant increases in the allocation of effort toward legitimate vs. illegitimate activities, with no discernible influence on contest participation. We discuss boundary conditions and the implications for contest organizers and contest platform operators.
Background:Generative artificial intelligence (GenAI) has rapidly emerged as a promising tool in health care. Despite its growing adoption, how physicians make use of it in medical practice has not been qualitatively studied. Existing literature has largely focused on theoretical applications or experimental validations, with limited insight into real-world physician engagement with GenAI technologies. Objective:The aim of this study was to leverage a fine-grained dataset at the query level to quantitatively examine how physicians incorporate GenAI into their clinical and research workflows. The primary objective was to analyze usage patterns over time and across physician demographics. A secondary goal was to assess potential risks to patient privacy arising from physicians' interactions with GenAI platforms. Methods:This study collected 106,942 query-and-answer pairs by 989 physicians between August 29, 2023, and April 16, 2024. We performed topic classification to identify the most prevalent use cases, examining how these use cases evolved over time and across demographics. We also developed sensitivity classifiers to detect personally identifiable information in physicians' queries to explore the potential privacy breach risks around physicians' use of GenAI. Results:Approximately 40% (396/989) of the enrolled physicians were female, 45.9% (454/989) were younger than 25 years, and 54.1% (535/989) were between 25 and 56 years of age. The majority of them worked in clinical departments (680/989, 68.8%) or medical technology departments (127/989, 12.8%). Our classification-based quantitative analyses suggest the following. First, physicians use GenAI predominantly for medical research (64,379/106,942, 60.2%) rather than clinical practice (13,100/106,942, 12.25%). Second, physicians focus more on health care-related questions (rising from 64,165/106,942, 60% to 83,415/106,942, 78%) within the first 15% (16,041/106,942) of their query sequence. Third, the use of GenAI differed across physician demographics and features. Specifically, female physicians asked a larger proportion of clinical questions (female: 0.154 vs male: 0.108; P<.001) and administration questions (female: 0.027 vs male: 0.018; P<.001) than male physicians; younger physicians posed more clinical questions (age ≤25: 0.146 vs age ∈ (25, 40]: 0.115 vs age >40: 0.103; P<.001) but fewer research questions (age ≤25: 0.580 vs age ∈ (25, 40]: 0.607 vs age >40: 0.664; P<.001) than senior physicians; and physicians accessing GenAI via computers asked more research questions (computer: 0.637 vs mobile: 0.296; P<.001), whereas physicians using mobile devices asked more clinical questions (computer: 0.107 vs mobile: 0.264; P<.001). Fourth, only 2.68% (2866/106,942) of physician queries contained sensitive information, the majority of which were primarily derived from writing and editing. Conclusions:Physicians are actively integrating GenAI into their professional routines, primarily leveraging it for research but also increasingly for clinical support. Usage patterns vary significantly across demographic lines, including gender, age, and device preference. Despite the presence of sensitive information in some queries, the risk of privacy breaches appears to be low.
Businesses have long sought to shape their online reputations, not only by soliciting positive reviews but also by suppressing negative ones through legal threats. Such practices distort consumer decision-making and undermine trust in digital marketplaces. To address this problem, the U.S. Congress enacted the Consumer Review Fairness Act (CRFA) in 2016, prohibiting businesses from using contractual clauses or intimidation to silence consumer feedback. Leveraging millions of hotel reviews from TripAdvisor, this study provides the first systematic evaluation of the CRFA’s effectiveness. We find that after the law’s enactment, U.S. hotel reviews became longer, more candid, and systematically more negative. Effects were especially strong for less reputable hotels, highly competitive markets, and experienced reviewers, pointing to where suppression had been most prevalent. Supplemental analyses confirm similar patterns using Google Places data and state-level repeals of criminal defamation laws. These findings demonstrate that regulation can effectively restore transparency and authenticity in online review platforms. For policymakers, the results provide concrete evidence that consumer-protection laws can counteract harmful business practices. For managers and platforms, they underscore the value of fostering open feedback ecosystems that promote trust and informed consumer choice.
Objectives. To evaluate the impact of state-level abortion restrictions enacted between 2018 and 2023 on infant mortality in the United States, comparing mortality trends across restricting and nonrestricting states. Methods. Using a difference-in-differences approach, we drew infant mortality data from the Centers for Disease Control and Prevention's WONDER (Wide-ranging Online Data for Epidemiologic Research) database and categorized deaths by age and cause, incorporating information on abortion restrictions and legal exceptions from the Center for Reproductive Rights and the Kaiser Family Foundation. We calculated estimates at the state-year level. Results. Infant mortality increased by 7.2% in restricting states. The increase is predominantly attributable to early (aged < 1 day) and late (aged 1 month to 1 year) infant deaths. Effects were largest for perinatal and noncongenital causes of death. Health exceptions did not significantly moderate the effects. Conclusions. Curtailing abortion access increases infant deaths. Fetal and maternal health exceptions do not moderate this effect. Excess deaths are not exclusively attributable to congenital abnormalities. Further work is needed to understand how restrictions contribute to late deaths and the long-term effects of infant deaths on families and communities. (Am J Public Health. 2025;115(11):1895-1902. https://doi.org/10.2105/AJPH.2025.308228).