The rapid proliferation of AI models has underscored the importance of thorough documentation, which enables users to understand, trust and effectively use these models in various applications. Although developers are encouraged to produce model cards, it’s not clear how much or what information these cards contain. In this study we conduct a comprehensive analysis of 32,111 AI model documentations on Hugging Face, a leading platform for distributing and deploying AI models. Our investigation sheds light on the prevailing model card documentation practices. Most AI models with a substantial number of downloads provide model cards, although with uneven informativeness. We find that sections addressing environmental impact, limitations and evaluation exhibit the lowest filled-out rates, whereas the training section is the one most consistently filled-out. We analyse the content of each section to characterize practitioners’ priorities. Interestingly, there are considerable discussions of data, sometimes with equal or even greater emphasis than the model itself. Our study provides a systematic assessment of community norms and practices surroinding model documentation through large-scale data science and linguistic analysis. As the number of AI models has rapidly grown, there is an increased focus on improving the documentation through model cards. Liang et al. explore questions around adoption practices and the type of information provided in model cards through a large-scale analysis of 32,111 model card documentation from 74,970 models.
The rapid proliferation of AI models has underscored the importance of thorough documentation, as it enables users to understand, trust, and effectively utilize these models in various applications. Although developers are encouraged to produce model cards, it's not clear how much information or what information these cards contain. In this study, we conduct a comprehensive analysis of 32,111 AI model documentations on Hugging Face, a leading platform for distributing and deploying AI models. Our investigation sheds light on the prevailing model card documentation practices. Most of the AI models with substantial downloads provide model cards, though the cards have uneven informativeness. We find that sections addressing environmental impact, limitations, and evaluation exhibit the lowest filled-out rates, while the training section is the most consistently filled-out. We analyze the content of each section to characterize practitioners' priorities. Interestingly, there are substantial discussions of data, sometimes with equal or even greater emphasis than the model itself. To evaluate the impact of model cards, we conducted an intervention study by adding detailed model cards to 42 popular models which had no or sparse model cards previously. We find that adding model cards is moderately correlated with an increase weekly download rates. Our study opens up a new perspective for analyzing community norms and practices for model documentation through large-scale data science and linguistics analysis.
A promising approach to estimate the causal effects of peer review policies is to analyze data from publication venues that shift policies from single-blind to double-blind from one year to the next. However, in these settings the content of the manuscript is a confounding variable—each year has a different distribution of scientific content which may naturally affect the distribution of reviewer scores. To address this textual confounding, we extend variable ratio nearest neighbor matching to incorporate text embeddings. We compare this matching method to a widely-used causal method of stratified propensity score matching and a baseline of randomly selected matches. For our case study of the ICLR conference shifting from single- to double-blind review from 2017 to 2018, we find human judges prefer manuscript matches from our method in 70% of cases. While the unadjusted estimate of the average causal effect of reviewers’ scores is -0.25, our method shifts the estimate to -0.17, a slightly smaller difference between the outcomes of single- and double-blind policies. We hope this case study enables exploration of additional text-based causal estimation methods and domains in the future.
Expert feedback lays the foundation of rigorous research. However, the rapid growth of scholarly production and intricate knowledge specialization challenge the conventional scientific feedback mechanisms. High-quality peer reviews are increasingly difficult to obtain. Researchers who are more junior or from under-resourced settings have especially hard times getting timely feedback. With the breakthrough of large language models (LLM) such as GPT-4, there is growing interest in using LLMs to generate scientific feedback on research manuscripts. However, the utility of LLM-generated feedback has not been systematically studied. To address this gap, we created an automated pipeline using GPT-4 to provide comments on the full PDFs of scientific papers. We evaluated the quality of GPT-4's feedback through two large-scale studies. We first quantitatively compared GPT-4's generated feedback with human peer reviewer feedback in 15 Nature family journals (3,096 papers in total) and the ICLR machine learning conference (1,709 papers). The overlap in the points raised by GPT-4 and by human reviewers (average overlap 30.85% for Nature journals, 39.23% for ICLR) is comparable to the overlap between two human reviewers (average overlap 28.58% for Nature journals, 35.25% for ICLR). The overlap between GPT-4 and human reviewers is larger for the weaker papers. We then conducted a prospective user study with 308 researchers from 110 US institutions in the field of AI and computational biology to understand how researchers perceive feedback generated by our GPT-4 system on their own papers. Overall, more than half (57.4%) of the users found GPT-4 generated feedback helpful/very helpful and 82.4% found it more beneficial than feedback from at least some human reviewers. While our findings show that LLM-generated feedback can help researchers, we also identify several limitations.
What conditions enable novel intellectual contributions to diffuse and become integrated into later scientific work? Prior work tends to focus on whole cultural products, such as patents and articles, and emphasizes external social factors as important. This article focuses on concepts as reflections of ideas, and we identify the combined influence that social factors and internal intellectual structures have on ideational diffusion. To develop this perspective, we use computational techniques to identify nearly 60,000 new ideas introduced over two decades (1993 to 2016) in the Web of Science and follow their diffusion across 38 million later publications. We find new ideas diffuse more widely when they socially and intellectually resonate. New ideas become core concepts of science when they reach expansive networks of unrelated authors, achieve consistent intellectual usage, are associated with other prominent ideas, and fit with extant research traditions. These ecological conditions play an increasingly decisive role later in an idea's career, after their relations with the environment are established. This work advances the systematic study of scientific ideas by moving beyond products to focus on the content of ideas themselves and applies a relational perspective that takes seriously the contingency of their success.
Until the 19th century, the UK state stayed out of education. Only in 1833 would Parliament first pass an act that subsidized education for the poor. By 1914, 160 education acts had been passed, consolidating into the state schooling system we recognize today. This paper seeks to explain this remarkable progression. I argue that the emergence of social-knowledge institutions across the West was a powerful force of cultural construction. What I term social scientization, this process was multidimensional and translocal, entailing the elaboration, reification, and diffusion of functionalist theories of the nation-state that centered national education as means to greater cultural rationalization. Longitudinal analyses on comprehensive population data comprising over 10,100 UK parliamentary acts support the core historical insight of this piece: increasingly routine and aggressive forms of state intervention in education were the progressive instantiation of the 19th-century nation-state model, which was fundamentally epistemic in character and inextricably linked to the expansive cultural content of the ascendant social sciences.
The advance of science rests on a robust peer review process. However whether or not a paper is accepted can depend on random external factors--e.g. the timing of the submission, the matching of editors and reviewers--that are beyond the quality of the work. This article systematically investigates the impact of these random factors independent of the paper’s quality on peer review outcomes in a major biomedical journal, eLife . We analyzed all of the submissions to eLife between 2016 to 2018, with 23,190 total submissions. We examined the effects of random factors at each decision point in the review process, from the gatekeeping senior editors who may desk-reject papers to review editors and reviewers who recommend the final outcome. Our results suggest that the peer-review process in eLife is robust overall and that random external factors have relatively little quantifiable bias.
AbstractEarly in the nineteenth century, members in the UK Parliament (MPs) hardly ever debated education. When they did, it was nearly always in the context of aid for the religious instruction of the poor. Indeed, even by 1850, nearly two decades after the first Great Reform Act (1832), the Prime Minister Lord John Russell made the case that a system of compulsory state schooling would be immoral and un-British. Yet, by the ‘80s, MPs debating in Westminster routinely drew connections between schooling and the most critical social issues of the day: social-class mobility and equity, child welfare, national development, emigration, and the civil service, among others. What explains the expanding, and expansive, political uses that elite policymakers put to schooling? How did schooling and education take on such an aggrandized role in society for British statesmen? To address these questions, this paper combines natural language processing techniques, semantic network, discourse, and regression analyses to read and interpret the ∼1.1 million political speeches given in the UK Houses of Parliament during the long nineteenth century (1804–1913). In contrast to explanations emphasizing the direct role that economic, social, and political development as well as conflict played in the UK state’s historic expansion, this piece demonstrates how social scientization, the sweeping international epistemic movement that institutionalized and diffused functionalist social theory, created the context that made it possible for political elites to see and promote schooling as an effective policy instrument of greater cultural rationalization supporting the development of capitalist industrial society.
Women and men often contribute differently to research knowledge. Do differences in these contributions partially explain disparities in academic career outcomes? We explore this by looking at how gender is embodied in research language, and then ascertain whether the adoption of more gendered research language affects career outcomes beyond the researcher's attributes. We identify different forms of gendered knowledge—gender referents (explicit references to sex and gender) and gender-associated terms (words that are implicitly associated with women or men researchers)—by applying natural language processing techniques to nearly one million doctoral dissertations published in the United States between 1980 and 2010. We then determine whether employing gender referents and gender-associated terms affects the course of PhDs' ensuing careers. We find women researchers have lower chances of securing academic positions than men in every field; explicit references to women as research subjects are modestly rewarded in comparison to references to men; and more career opportunities are afforded to research knowledge associated with men. These results suggest that academia is slowly correcting the traditional and explicit bias of studying men at the exclusion of women. Still, there remains a stronger implicit bias against knowledge associated with women scholars. We discuss relative differences between humanities and social sciences versus natural sciences, technology, engineering, and math, as well as potential treatments for offsetting bias in those fields.
Abstract Traditional accounts of state expansion and of the rise of state schooling in the nineteenth century emphasize economic, political, and social development as well as conflict and domination. These accounts explain the introduction of new state structures, like ministries of education, rules of compulsion, and the general elaboration of bureaucracies. This article contributes to the historical sociological study of state expansion with specific regard to schooling by refocusing on the role that macrocultural processes of social scientization played in shaping the discursive construction and expansion of the state. Designed to analyze the 1.3 million speeches given in the UK parliament during the nineteenth century, the research reported here supports the argument that the development, professionalization, and institutionalization of the social sciences—social scientization—was a powerful force of cultural construction across the West and was positively associated with expanded notions of the state, as evidenced with the case of the United Kingdom. This article therefore not only provides an important alternative view to those who emphasize economic and social transformation but it also advances the empirical study of the powerful role that social science, as generative institution of cultural construction, played in shaping official discourses of the state—in this instance, the schooling state.
To understand the relationship between social background and sex in schooling, we use Bourdieu's theory of social reproduction and a feminist perspective of gender as practice. We pose two questions: (1) What is the relationship between economic and cultural capital and achievement for 4th-grade females versus males studying in Germany? (2) Is the relationship between school composition and student achievement different for 4th-grade females versus males? We report no differences between females and males in the relationships between social background and achievement (p > 0.05). However, the relationship between class-aggregated social background and achievement is halved in female-majority mathematics classrooms (beta = -12.6, p < 0.05).