Large language models (LLMs) hold great promise for generating social science data, potentially expanding the methodological toolkit of quantitative social research. Prior studies have primarily focused on individual-level predictability or behavioral plausibility of LLM-generated data. We propose a framework for assessing the validity of LLM-generated data by returning to the foundational principles of survey research in the social sciences. Just as surveys based on representative samples yield statistics that approximate the corresponding statistical moments of the target population, assessment should center on the ability of LLM-generated data to reproduce real-world, population-level statistical patterns. We introduce SSDataBench, a systematic benchmark designed to evaluate population-level statistical realism in LLM-generated social science data. The benchmark assesses five types of statistical patterns central to social research: univariate distributions, bivariate associations, multivariate outcome predictions, life event sequence distributions, and associations between life event sequences and covariates. We illustrate SSDataBench using four longitudinal datasets and three cross-sectional datasets spanning six major social domains: demographics, socioeconomic status, marriage, health, abilities, and attitudes. Our study reveals representational limitations in current LLMs under sparse conditioning settings, manifested in a pronounced tendency to compress real-world heterogeneity into simplified typological structures. Finally, we outline a roadmap toward improved statistical realism and report preliminary results indicating that domain-specific training can enhance population-level realism.
With the growing prevalence of generative artificial intelligence (AI), an increasing amount of content is no longer exclusively generated by humans but by generative AI models with human guidance. This shift presents notable challenges for the delineation of originality due to the varying degrees of human contribution in AI-assisted works. This study raises the research question of measuring human contribution in AI-assisted content generation and introduces a framework to address this question that is grounded in information theory. By calculating mutual information between human input and AI-assisted output relative to self-information of AI-assisted output, we quantify the proportional information contribution of humans in content generation. Our experimental results demonstrate that the proposed measure effectively discriminates between varying degrees of human contribution across multiple creative domains. We hope that this work lays a foundation for measuring human contributions in AI-assisted content generation in the era of generative AI.
Generative AI has been found, and will likely be found increasingly, useful in education. However, existing AI-for-education studies provide inconsistent evidence on its average effects. More broadly, research on prior educational technologies shows that average effects often mask substantial heterogeneity across student populations. Motivated by this evidence, this study examines heterogeneity in students' learning behavior with AI, which students benefit from AI assistance, and how learner profiles and learning behavior shape these patterns. To this end, we recruited 318 university students to participate in structured learning experiments lasting up to 125 minutes. Our findings indicate that students' learning behavior is strongly associated with learning outcomes, with behaviors characterized by proactive and critical engagement, rather than limited engagement, associated with significantly better performance. These behavioral differences are related to learner profiles, with students from higher-ranking universities and those with greater prior knowledge tending to benefit more, consistent with their greater likelihood of adopting proactive interaction strategies. Accounting for learning behavior substantially weakens or eliminates the associations between learner profiles and learning outcomes, suggesting that how students use AI is a key pathway through which background differences are linked to learning gains. Overall, this work provides a deeper understanding of AI assistance in education by showing how differences in learner profiles and learning behavior shape who benefits from AI-supported learning. These insights can help educators and students better navigate and integrate AI into educational practices.
Research on social networks now sits at a productive intersection between network science and computational social science [...]
This paper evaluates the propensity score as a tool for summarizing treatment effect heterogeneity and facilitating extrapolation from experimental settings to broader populations. We argue that propensity score methods capture the most consequential treatment effect heterogeneity for population-level inference, achieving high accuracy at the aggregate level and enabling extrapolation. Using benchmark data from the National Supported Work (NSW) Demonstration and a comparison group derived from the Current Population Survey (CPS) and the Panel Study of Income Dynamics (PSID), we assess the utility of propensity score methods for recovering heterogeneous effects and population-level estimates. The results indicate that a simple propensity score approach substantially reduces confounding and produces average treatment effect estimates close to those obtained experimentally, while also enabling extrapolation to target populations.
As large language models (LLMs) gradually demonstrate their potential to boost productivity and become integral tools for problem-solving in daily life worldwide, understanding the linguistic inequalities they introduce is becoming increasingly important. Prior research has primarily focused on static analyses of disparities in existing knowledge and capabilities of LLMs across languages. However, LLMs are continuously evolving, acquiring new knowledge to provide current, relevant responses and deliver precise, expert-level answers in specific domains. Investigating linguistic inequalities within this dynamic learning process is, therefore, also essential. In this paper, we explore inequalities in new knowledge learning by LLMs across different languages and four key dimensions: effectiveness, transferability, prioritization, and robustness. Through extensive experiments in both in-context learning and fine-tuning settings, with proprietary and open-source models, we reveal four key findings: 1) LLMs face greater challenges in efficiently and accurately learning new knowledge in lowerresource languages; 2) knowledge learned by LLMs tends to be more easily transferred to higher-resource languages than to lower-resource ones; 3) new knowledge in higherresource languages is more likely to be retained and prioritized; and 4) LLMs are more robust against incorrect or misleading information in higher-resource languages. We further analyze the underlying causes of these inequalities from linguistic perspectives, pretraining characteristics, and tokenizer design, and propose a preliminary mitigation strategy through the lens of linguistic neurons. This work highlights the urgent need to recognize and address emerging linguistic inequalities in the development of LLMs.
We present **Lean4PHYS**, a comprehensive reasoning framework for college-level physics problems in Lean4. **Lean4PHYS** includes *LeanPhysBench*, a college-level benchmark for formal physics reasoning in Lean4, which contains 200 hand-crafted and peer-reviewed statements derived from university textbooks and physics competition problems. To establish a solid foundation for formal reasoning in physics, we also introduce *PhysLib*, a community-driven repository containing fundamental unit systems and theorems essential for formal physics reasoning. Based on the benchmark and Lean4 repository we composed in **Lean4PHYS**, we report baseline results using major expert Math Lean4 provers and state-of-the-art closed-source models, with the best performance of DeepSeek-Prover-V2-7B achieving only 16
In this paper, we present findings from four separate studies using different data sources and methods to examine Chinese attitudes toward the United States amid the COVID-19 pandemic. The empirical results consistently indicate a marked and significant decline in Chinese attitudes toward the US between late 2019 and the end of 2022. Using a quasi-experimental design and granular survey data that exploit daily variations in public opinion, we offer additional evidence that the decline in Chinese attitudes toward the United States followed a distinct pattern not true for Chinese attitudes toward other countries. Specifically, the rise in Chinese unfavorability toward the United States closely corresponded to the heightened Chinese attention to the pandemic’s progression in the United States. These results collectively suggest a causal effect of COVID-19, shedding light on how public health crises, international relations, and media jointly shape the increasing enmity between the two great powers.
Past research has studied social determinants of attitudes toward foreign countries. Confounded by potential endogeneity biases due to unobserved factors or reverse causality, the causal impact of these factors on public opinion is usually difficult to establish. Using social media data, we leverage the suddenness of the COVID-19 pandemic to examine whether a major global event has causally changed American views of another country. We collate a database of more than 297 million posts on the social media platform Twitter about China or COVID-19 up to June 2020, and we treat tweeting about COVID-19 as a proxy for individual awareness of COVID-19. Using regression discontinuity and difference-in-difference estimation, we find that awareness of COVID-19 causes a sharp rise in anti-China attitudes. Our work has implications for understanding how self-interest affects policy preference and how Americans view migrant communities.
The use of pooled data from different repeated survey series to study long-term trends is handicapped by a measurement difficulty: different survey series often use different scales to measure the same attitude and thus generate scale-incomparable data. In this article, the authors propose the latent attitude method (LAM) to address this scale-incomparability problem, on the basis of the assumption that attitudes measured by ordinal categories reflect a latent attitude with cut points. The method extends the latent variable method in the case of a single survey series to the case of multiple survey series and leverages overlapping years for identification. The authors first assess the validity of the method with simulated data. The results show that the method yields accurate estimates of mean attitudes and cut point values. The authors then apply the method to an empirical study of Americans’ attitudes toward China from 1974 to 2019.
The declaration of COVID-19 as a pandemic has largely amplified the spread of related information on social platforms, such as Twitter, Facebook and WeChat. In this work, we investigate how the disease and information co-evolve in the population. We focus on COVID-19 and its information during the period when the disease was widely spread in China, i.e., from January 25th to March 24th, 2020. The co-evolution between disease and information is explored via the spatial analysis of the two spreading processes. We visualize the geo-location of both disease and information at the province level and find that disease is more geo-localized compared to information. High correlation between disease and information data is observed, and also people care about the spread of disease only when it comes to their neighborhood. Regard to the content of the information, we obtain that positive messages are more negatively correlated with the disease compared to negative and neutral messages. Additionally, two machine learning algorithms, i.e., linear regression and random forest, are introduced to further predict the number of infected using characteristics, such as disease spatial related and information-related features. We obtain that both the disease spatial related characteristics of nearby cities and information-related characteristics can help to improve the prediction accuracy. The methodology proposed in this paper may shed light on new clues of emerging infections prediction.
The American public’s perception of China, an important aspect of the relationship between the world’s two largest economies, has become unfavorable in recent years, with a sudden decline in 2020 after the outbreak of the COVID-19 pandemic. Although the attitude decline was concomitant with the pandemic, it is not easy to establish a causal relationship between them, given other parallel events in this time period, such as the US–China trade war. Identifying the causal effect of the COVID-19 pandemic on the American public’s perception of China will help understand how Americans evaluate a foreign country. In this study, we examine the mediating role of social media in shaping the American public’s opinion in the context of the COVID-19 pandemic in 2020. We empirically analyze a large-scale dataset of 2.2 million posts on the social networking platform Twitter in a 12-month window and identify that the outbreak of COVID-19 triggered American social media users to become more negative toward China, when posts related to COVID-19 begin spreading in their “cyber neighborhood”. This analysis is performed with a deep neural network model, BERT, to quantify American attitudes toward China, and a before-and-after event analysis and a difference-in-difference analysis to show the causal relationship between COVID-19 and declining favorability toward China. This finding confirms the mediating role of social media in shaping American public opinion on China.
The US global leadership in science and technology has greatly benefitted from immigrants from other countries, most notably from China in the recent decades. However, feeling the pressure of potential federal investigations since the 2018 launch of the China Initiative, scientists of Chinese descent in the United States now face higher incentives to leave the United States and lower incentives to apply for federal grants. Analyzing data pertaining to institutional affiliations of more than 200 million scientific papers, we find a steady increase in the return migration of scientists of Chinese descent from the United States to China. We also conducted a survey of scientists of Chinese descent employed by US universities in tenured or tenure-track positions (n = 1,304), with results revealing general feelings of fear and anxiety that lead them to consider leaving the United States and/or stop applying for federal grants. If the situation is not corrected, American science will likely suffer the loss of scientific talent to China and other countries.
Do mass media influence people’s opinions of other countries? Using BERT, a deep neural network-based natural language processing model, this study analyzes a large corpus of 267,907 China-related articles published by The New York Times since 1970. The output from The New York Times is then compared to a longitudinal data set constructed from 101 cross-sectional surveys of the American public’s views on China, revealing that the reporting of The New York Times on China in one year explains 54% of the variance in American public opinion on China in the next. This result confirms hypothesized links between media and public opinion and helps shed light on how mass media can influence the public opinion of foreign countries.
Millions of people are surveyed every year regarding their attitudes toward various topics. Together these surveys have produced a large corps of data that document how people think collectively toward various aspects of contemporary social life.The wealth of the attitude surveys has promoted scholars to move beyond the single-survey analysis. However, the use of survey data for studying trends in attitudes is handicapped by a measurement difficulty: different surveys have used different survey instruments to measure the same attitude and thus have generated data that strictly non-comparable. We propose the Latent Attitude Method (LAM) to address this issue. Our method borrows strength from two research traditions: (1) the latent variable method in attitude research and (2) the comparable distribution condition in survey design and evaluation. The core of this method is that, when two or more surveys overlap in a given year, we assume that the same latent attitude is measured as if two measurement scales are randomly given to two independent samples drawn from the same population. Thus, we can assume the same statistical properties for the latent attitude. In so doing, we are able to reduce the number of unknowns to be less than the number of established equations and estimate the best-fit parameters with maximum likelihood method. We demonstrate the utility of the method with simulated data, and apply the method to an empirical example of estimating America’s attitude toward China from 1974 to 2019.
Modern science is dominated by scientific productions from teams. A recent finding shows that teams of both large and small sizes are essential in research, prompting us to analyze the extent to which a country's scientific work is carried out by big or small teams. Here, using over 26 million publications from Web of Science, we find that China's research output is more dominated by big teams than the rest of the world, which is particularly the case in fields of natural science. Despite the global trend that more papers are written by big teams, China's drop in small team output is much steeper. As teams in China shift from small to large size, the team diversity that is essential for innovative work does not increase as much as that in other countries. Using the national average as the baseline, we find that the National Natural Science Foundation of China (NSFC) supports fewer small teams than the National Science Foundation (NSF) of the United States does, implying that big teams are preferred by grant agencies in China. Our finding provides new insights into the concern of originality and innovation in China, which indicates a need to balance small and big teams.
There are many conflicting theories about the relationship between media reports and public opinion, but few of them are supported by empirical work on large data sets. We use sentiment results on over 260,000 China-related articles in The New York Times to show that events in international relations affect media sentiment which, in turn, affects public opinion. We find that sudden shifts in US-China relations are accompanied by changes in how The New York Times covers China and that the news reporting on China leads public opinion on China by 1 year. Our work illustrates how The New York Times, a prestigious mass media institution, propagates international relation signals to shape American views of the Chinese state and the Chinese people.
There is extensive, yet fragmented, evidence of gender differences in academia suggesting that women are under-represented in most scientific disciplines, publish fewer articles throughout a career, and their work acquires fewer citations. Here, we offer a comprehensive picture of longitudinal gender discrepancies in performance through a bibliometric analysis of academic careers by reconstructing the complete publication history of over 1.5 million gender-identified authors whose publishing career ended between 1955 and 2010, covering 83 countries and 13 disciplines. We find that, paradoxically, the increase of participation of women in science over the past 60 years was accompanied by an increase of gender differences in both productivity and impact. Most surprisingly though, we uncover two gender invariants, finding that men and women publish at a comparable annual rate and have equivalent career-wise impact for the same size body of work. Finally, we demonstrate that differences in dropout rates and career length explain a large portion of the reported career-wise differences in productivity and impact. This comprehensive picture of gender inequality in academia can help rephrase the conversation around the sustainability of women's careers in academia, with important consequences for institutions and policy makers.