The issue of gender disparity in scientific publications has been a topic of ongoing debate. One aspect of this debate concerns whether women receive equal credit for their contributions compared to men. While prior qualitative and quantitative studies have suggested that women are more likely to be acknowledged than listed as co-authors, large-scale empirical evidence across multiple disciplines remains limited. In this study, we analyze data from over 20,000 authors and 60,000 acknowledged individuals across nine disciplines in open-access journals. Our results confirm persistent gender disparities: women are more frequently acknowledged than credited as co-authors, especially in roles involving investigation and analysis. To account for status and disciplinary effects, we examined collaboration pairs composed of highly cited and less cited scholars. In collaborations, highly cited scholars are more likely to be listed as an author regardless of gender. Notably, highly cited women in such pairs are even more likely to be co-authors than their men counterparts. Our findings suggest that power dynamics and perceived success heavily influence the distribution of credit in scientific publishing. These results underscore the role of status dynamics in shaping authorship and call for a more nuanced understanding of how gender, power, and recognition interact in scientific publishing. Our findings offer valuable insights for scholars, editors, and funding bodies committed to advancing equity in science.
Most automated peer review systems rely on textual manuscript content alone, leaving visual elements such as figures and external scholarly signals underutilized. We introduce REM-CTX, a reinforcement-learning system that incorporates auxiliary context into the review generation process via correspondence-aware reward functions. REM-CTX trains an 8B-parameter language model with Group Relative Policy Optimization (GRPO) and combines a multi-aspect quality reward with two correspondence rewards that explicitly encourage alignment with auxiliary context. Experiments on manuscripts across Computer, Biological, and Physical Sciences show that REM-CTX achieves the highest overall review quality among six baselines, outperforming other systems with substantially larger commercial models, and surpassing the next-best RL baseline across both quality and contextual grounding metrics. Ablation studies confirm that the two correspondence rewards are complementary: each selectively improves its targeted correspondence reward while preserving all quality dimensions, and the full model outperforms all partial variants. Analysis of training dynamics reveals that the criticism aspect is negatively correlated with other metrics during training, suggesting that future studies should group multi-dimension rewards for review generation.
Open scholarly meta-databases are becoming an indispensable component of quantitative studies of science. In particular, reference information from these databases is essential for citation analysis, including impact evaluation and disruption measurement. Recent studies have examined the reference coverage of different scholarly databases; nevertheless, the underlying factors driving the disparity in reference quality across publications remain largely unexplored. Here, we examine the reference quality of OpenAlex, one of the largest open scholarly meta-databases to date, and find pervasive reference missingness. Importantly, we observe that this effect is more pronounced in articles associated with less prestigious institutions, lower-impact journals, and authors affiliated with countries of lower economic standing. A striking 57% of articles below the median prestige and journal impact that belong to the Global South are completely missing references. A regression analysis revealed that these disparities are mainly associated with the prestige of the journal and the institution. We discuss how we should address this issue more systematically and exercise caution when using metadata to assess impact.
The promises and perils about scientific team diversity are still debated in the scholarly literature, partly because the importance of underrepresented groups is not fully recognized or valued. In this paper, we summarize two perspectives on team diversity in science: horizontal differences and vertical disparity. Horizontal differences refer to variations across individuals on equal levels, such as differences in gender, nationality, and occupation. Vertical disparity reflects broader social and structural inequalities that limit historically marginalized groups. We introduce team hierarchy, defined as the distribution of power and influence among team members, as a moderating mechanism that helps explain the mixed findings surrounding team diversity and performance. Understanding how hierarchical structures shape diversity is not only important to maximize the benefits of diverse teams but also for enhancing the representation and impact of underrepresented voices. By analyzing 64,038 papers from PloS One and 75,260,139 teams from Microsoft Academic Graph (MAG), our comprehensive study underscores the critical role of team hierarchy in significantly affecting team performance. We also investigate how team hierarchy interacts with team diversity along three dimensions: authors' gender, sector, and country. Interestingly, we find that flat team structures are more positively associated with performance in diverse teams than in homogeneous teams, where members share similar identities. This effect is particularly strong in science compared to social science & arts disciplines. Drawing from social identity and social dominance theories, we propose that flat team structures foster conditions for diverse teams to flourish, enabling minority groups to assume significant roles and wield influential power. Our study contributes valuable insights into team diversity within the scientific community, emphasizing the significance of meaningful inclusion beyond mere numerical representation.
Research has shown that most resources shared in articles (e.g., URLs to code or data) are not kept up to date and mostly disappear from the web after some years (Zeng et al., 2019). Little is known about the factors that differentiate and predict the longevity of these resources. This article explores a range of explanatory features related to the publication venue, authors, references, and where the resource is shared. We analyze an extensive repository of publications and, through web archival services, reconstruct how they looked at different time points. We discover that the most important factors are related to where and how the resource is shared, and surprisingly little is explained by the author's reputation or prestige of the journal. By examining the places where long-lasting resources are shared, we suggest that it is critical to disseminate and create standards with modern technologies. Finally, we discuss implications for reproducibility and recognizing scientific datasets as first-class citizens.
The rapid advancement of Natural Language Processing (NLP) has greatly improved text-generation tools like ChatGPT and Claude, offering significant utility but also posing risks to media credibility through paraphrased plagiarism—a subtle yet widespread form of content misuse. Despite progress in automated paraphrase detection, inconsistencies in training datasets often limit their effectiveness. This study examines traditional and modern approaches to paraphrase identification, revealing how the under-representation of certain paraphrase types in widely-used datasets, including those for training Large Language Models (LLMs), undermines plagiarism detection accuracy. To address these issues, we introduce and validate ReParaphrased, a refined paraphrase typology, and extend the Extended Typology Paraphrase Corpus (ETPC) with meticulous manual annotations to enhance reliability. Using the augmented ETPC, we fine-tune the LLama3.1-7B-instruct model, uncovering significant disparities in paraphrase type distribution across existing datasets. A detailed analysis of the MRPC benchmark dataset further highlights critical distributional issues and their implications. We propose four key solutions to address dataset limitations, providing both theoretical and practical guidance for improving dataset quality. These contributions aim to establish a more robust foundation for NLP model training and evaluation. Finally, we outline future research directions and suggest improvements for dataset development to advance AI-driven paraphrase detection.
Questionable journals threaten global research integrity, yet manual vetting can be slow and inflexible. Here, we explore the potential of artificial intelligence (AI) to systematically identify such venues by analyzing website design, content, and publication metadata. Evaluated against extensive human-annotated datasets, our method achieves practical accuracy and uncovers previously overlooked indicators of journal legitimacy. By adjusting the decision threshold, our method can prioritize either comprehensive screening or precise, low-noise identification. At a balanced threshold, we flag over 1000 suspect journals, which collectively publish hundreds of thousands of articles, receive millions of citations, acknowledge funding from major agencies, and attract authors from developing countries. Error analysis reveals challenges involving discontinued titles, book series misclassified as journals, and small society outlets with limited online presence, which are issues addressable with improved data quality. Our findings demonstrate AI's potential for scalable integrity checks, while also highlighting the need to pair automated triage with expert review.
Artificial intelligence (AI) has seen tremendous development in industry and academia. However, striking recent advances by industry have stunned the world, inviting a fresh perspective on the role of academic research in this field. Here, we characterize the impact and type of AI produced by both environments over the last 25 years and establish several patterns. We find that articles published by teams consisting exclusively of industry researchers tend to get greater attention, with a higher chance of being highly cited and citation-disruptive, and several times more likely to produce state-of-the-art models. In contrast, we find that exclusively academic teams publish the bulk of AI research and tend to produce higher novelty work, with single papers having several times higher likelihood of being unconventional and atypical. The respective impact-novelty advantages of industry and academia are robust to controls for subfield, team size, seniority, and prestige. We find that academic-industry collaborations struggle to replicate the novelty of academic teams and tend to look similar to industry teams. Together, our findings identify the unique and nearly irreplaceable contributions that both academia and industry make toward the healthy progress of AI.
Download This Paper Open PDF in Browser Add Paper to My Library Share: Permalink Using these links will ensure access to this page indefinitely Copy URL Copy DOI
Leadership is evolving dynamically from an individual endeavor to shared efforts. This paper aims to advance our understanding of shared leadership in scientific teams. We define three kinds of leaders, junior (10-15), mid (15-20), and senior (20+) based on career age. By considering the combinations of any two leaders, we distinguish shared leadership as heterogeneous when leaders are in different age cohorts and homogeneous when leaders are in the same age cohort. Drawing on 1,845,351 CS, 254,039 Sociology, and 193,338 Business teams with two leaders in the OpenAlex dataset, we identify that heterogeneous shared leadership brings higher citation impact for teams than homogeneous shared leadership. Specifically, when junior leaders are paired with senior leaders, it significantly increases team citation ranking by 1-2 in comparison with two leaders of similar age. We explore the patterns between homogeneous leaders and heterogeneous leaders from team scale, expertise composition, and knowledge recency perspectives. Compared with homogeneous leaders, heterogeneous leaders are more adaptive in large teams, have more diverse expertise, and trace both the newest and oldest references.
Datasets are critical for scientific research, playing an important role in replication, reproducibility, and efficiency. Researchers have recently shown that datasets are becoming more important for science to function properly, even serving as artifacts of study themselves. However, citing datasets is not a common or standard practice in spite of recent efforts by data repositories and funding agencies. This greatly affects our ability to track their usage and importance. A potential solution to this problem is to automatically extract dataset mentions from scientific articles. In this work, we propose to achieve such extraction by using a neural network based on a Bi-LSTM-CRF architecture. Our method achieves F1 = 0.885 in social science articles released as part of the Rich Context Dataset. We discuss the limitations of the current datasets and propose modifications to the model to be done in the future.
Correctness is a key aspiration of the scientific process, yet recent studies suggest that many high-profile findings may be difficult to replicate or require considerable evidence for verification. Proposals to fix these issues typically ask for tighter statistical controls (e.g., stricter p-value thresholds or higher statistical power). However, these approaches often overlook the importance of contemplating research outcomes' potential costs and benefits. Here, we develop a framework grounded in Bayesian decision theory that seamlessly integrates cost-benefit analysis into evaluating research programs with potentially uncertain results. We derive minimally acceptable prestudy odds and positive predictive values for cost and benefit levels. We show that tolerance to inaccurate results changes dramatically due to uncertainties posed by research. We also show that reducing uncertainties (e.g., by recruiting more subjects) may have limited effects on the expected benefit of continuing specific research programs. We apply our framework to several types of cancer research and their funding. Our analysis shows that highly exploratory research designs are easily justifiable due to their potential benefits, even when probabilistic models suggest otherwise. We discuss how the cost and benefit of research could and should always be part of the toolkit used by scientists, institutions, or funding agencies.
The more science advances, the more questions are asked. This compounding growth can make it difficult to keep up with current research directions. Furthermore, this difficulty is exacerbated for junior researchers who enter fields with already large bases of potentially fruitful research avenues. In this paper, we propose a novel task and a recommender system for research directions, RecSOI, that draws from statements of ignorance (SOIs) found in the research literature. By building researchers’ profiles based on textual elements, RecSOI generates personalized recommendations of potential research directions tailored to their interests. In addition, RecSOI provides context for the recommended SOIs, so that users can quickly evaluate how relevant the research direction is for them. In this paper, we provide an overview of RecSOI’s functioning, implementation, and evaluation, demonstrating its effectiveness in guiding researchers through the vast landscape of potential research directions.
Same-race mentorship preference refers to mentors or mentees forming connections significantly influenced by a shared race. Although racial diversity in science has been well-studied and linked to favorable outcomes, the extent and effects of same-race mentorship preferences remain largely underexplored. Here, we analyze 465,355 mentor-mentee pairs from more than 60 research areas over the last 70 years to investigate the effect of same-race mentorship preferences on mentees' academic performance and survival. We use causal inference and statistical matching to measure same-race mentorship preferences while accounting for racial demographic variations across institutions, time periods, and research fields. Our findings reveal a pervasive same-race mentorship propensity across races, fields, and universities of varying research intensity. We observe an increase in same-race mentorship propensity over the years, further reinforced inter-generationally within a mentorship lineage. This propensity is more pronounced for minorities (Asians, Blacks, and Hispanics). Our results reveal that mentees under the supervision of mentors with high same-race propensity experience significantly lower productivity, impact, and collaboration reach during and after training, ultimately leading to a 27.6% reduced likelihood of remaining in academia. In contrast, a mentorship approach devoid of racial propensity appears to offer the best prospects for academic performance and persistence. These findings underscore the importance of mentorship diversity for academic success and shed light on factors contributing to minority underrepresentation in science.
Figures are an essential part of scientific communication. Yet little is understood about how accessible (e.g., color-blind safe), readable (e.g., good contrast), and explainable (e.g., contain captions and legends) they are. We develop computational techniques to measure these features and analyze a large sample of them from open access publications. Our method combines computer and human vision research principles, achieving high accuracy in detecting problems. In our sample, we estimated that around 20.6% of publications contain either accessibility, readability, or explainability issues (around 2% of all figures contain accessibility issues, 3% of diagnostic figures contain readability issues, and 23% of line charts contain explainability issues). We release our analysis as a dataset and methods for further examination by the scientific community.
Citation data, along with other bibliographic datasets, have long been adopted by the knowledge and data discovery community as an important direction for presenting the validity and effectiveness of proposed algorithms and strategies. Many top computer scientists are also excellent researchers in the science of science. The purpose of this workshop is to bridge the two communities (i.e., the knowledge discovery community and the science of science community) together as the scholarly activities become salient web and social activities that start to generate a ripple effect on broader knowledge discovery communities. This workshop will showcase the current data-driven science of science research by highlighting several studies and constructing a community of researchers to explore questions critical to the future of data-driven science of science, especially a community of data-driven science of science in Data Science so as to facilitate collaboration and inspire innovation. Through discussion on emerging and critical topics in the science of science, this workshop aims to help generate effective solutions for addressing environmental, societal, and technological problems in the scientific community.
Mentorship in science is crucial for topic choice, career decisions, and the success of mentees and mentors. Typically, researchers who study mentorship use article co-authorship and doctoral dissertation datasets. However, available datasets of this type focus on narrow selections of fields and miss out on early career and non-publication-related interactions. Here, we describe Mentorship, a crowdsourced dataset of 743176 mentorship relationships among 738989 scientists primarily in biosciences that avoids these shortcomings. Our dataset enriches the Academic Family Tree project by adding publication data from the Microsoft Academic Graph and “semantic” representations of research using deep learning content analysis. Because gender and race have become critical dimensions when analyzing mentorship and disparities in science, we also provide estimations of these factors. We perform extensive validations of the profile–publication matching, semantic content, and demographic inferences, which mostly cover neuroscience and biomedical sciences. We anticipate this dataset will spur the study of mentorship in science and deepen our understanding of its role in scientists’ career outcomes.