
Understanding how scientific contributions evolve from their inception to long-term influence is central to the study of knowledge dynamics and research evaluation. This study investigates the evolution of actors' cognitive perceptions of scientific contributions in academic papers by harnessing large language models (LLMs) to conduct a combined analysis of full texts, peer review comments, and citation contexts of over 29,000 articles from PLOS ONE. Drawing on the social construction theory of scientific knowledge, we propose a dynamic construction lifecycle comprising three stages-generation, external evaluation, and impact diffusion-shaped respectively by authors, reviewers, and citers. By classifying contribution types from several textual sources, we examined their consistency and evolution in terms of both type and strength across lifecycle stages. The results reveal that authors highlight experimental and application-oriented contributions; reviewers focus on conceptual dimensions and apply stricter criteria; citers emphasize experimental and methodological aspects. Furthermore, a cognitive gap emerges between reviewers and other actors, indicating that reviewers often rely on perceived significance rather than contribution typologies. In contrast, alignment between authors and citers suggests that citation-based indicators, despite limitations, capture meaningful dimensions of scholarly recognition. Analysis of contribution evolution patterns shows both divergent and stable trajectories with non-linear fluctuations in perceived strength.
Abstract Information behavior is a longstanding area of scholarly interest in library and information science. This scoping review focused specifically on information behavior in early childhood, defined as birth to 8 years of age. Following a systematic search of five key databases, 36 publications that met the inclusion criteria were identified. Data related to the study participants, research methods and settings, conceptual and theoretical frameworks, and key study findings were extracted from the publications, then collated and summarized. Reviewed publications spanned three decades (1992–2022) and focused predominantly on young children's information seeking and information searching. Children under 5 years of age and/or in preschool or kindergarten were the least studied groups in this corpus. Formal educational settings were the most common research setting. Interviews and task‐centered activities were the most frequently used methods of data collection. Half of the reviewed studies were explicitly informed by information behavior theories and models, while less than a third applied theories of child development and/or learning. This analysis reveals existing trends and gaps in the body of research concerned with young children's information behavior, which have implications for the provision of relevant and appropriate information resources and services for this population.
Abstract Research of health information on Question and Answer (Q&A) platforms has expanded considerably in recent years. To provide a comprehensive understanding of this growing field, this study employed a systematic literature review methodology to examine research of health‐related information on Q&A platforms. The review examined study characteristics, methodological patterns, theoretical frameworks, and thematic trends across literature. The analysis of study characteristics identified the Q&A platforms used across studies and developed coding schemes for user groups and health topics based on prior research. The methodological analysis established a two‐layer coding scheme to classify research methodologies. The analysis of theoretical frameworks created a disciplinary coding scheme for theoretical approaches. The thematic synthesis identified major research themes and described the current state of the field. Together, these findings offer a comprehensive overview of the current state of research in this domain. They highlight key research gaps, show major research trends, provide methodological and ethical insights, and present theoretical and practical recommendations for researchers, institutions, platform management, public health education, and policymaking.
A cornerstone of digital humanities (DH) is its engagement with diverse data practices across disciplines, which create academic, educational, cultural, historical, social, technological, and economic value. However, data value theories derived from business and management domains are not entirely applicable to the DH domain. Humanities data practices remain insufficiently theorized, with limited insights into the implementation of these value-centric practices. To address these gaps, this study employs a multiple-case study and content analysis to examine data practices in 17 DH projects. The analysis identifies four stages of the humanities data value chain (DVC): data collection, data processing, data exposition, and data sharing. Humanities data practices are clustered into three archetypes: data transformation by expanding data scale, data enrichment by enriching data context, and data revitalization by facilitating data reuse. Synthesizing these findings, the study proposes a conceptual model illustrating how humanities data practices create multi-dimensional value across the DVC stages. Furthermore, four guidelines are proposed to support the implementation of these value-centric data practices. As a pivotal effort to theorize value-centric data practices in DH, this study extends data value theory in the DH domain, thereby fostering interdisciplinary collaboration and advancing the development of robust humanities data infrastructure.
Abstract Cross‐disciplinary research is a priority for many academic institutions, with a growing body of scholarship dedicated to studying the central practice of cross‐disciplinarity: integration, or the synthesis of knowledge, information, and data across disciplines and domains. Discussions on this topic are regularly published in JASIST and ARIST and are of interest to ARIST readership because of the centrality of information to definitions of integration. Emerging challenges in this area are that (1) cross‐disciplinary integration practices are not only discussed by information scholars, but in fragmented, rarely overlapping discourses across domains, and (2) despite being linked to information, cross‐disciplinary integration is rarely explicitly positioned as an information practice. In this review, to respond to these challenges, I adopt critical interpretive synthesis methods to review the literature on cross‐disciplinary integration. I consolidate definitions of integration across 12 areas of scholarship and highlight the roles of information in these definitions. I find that certain areas of work are more likely to describe integration as the homogenization of disparate datasets, while others describe integration as a synthesis of knowledge or perspectives, and that the relationship between these integrative activities—for example, how combination of multi‐source datasets can support interdisciplinary conclusions—is rarely explored in depth. I identify concepts from the information sciences that offer ways forward to address this knowledge gap, and suggest roles for information practitioners in supporting integrative work.
Abstract This review analyzed 241 scholarly articles published between 2010 and 2025 in information science venues to examine how affect shapes refugees' information behavior during forced migration and to identify additional contextual factors. It identifies seven affective dimensions: anxiety, shame and stigma, grief and loss, frustration, (mis)trust, (loss of) control, and (not) belonging. It outlines related information practices, including calibration, collective or selective practices, personal information management, language‐ and design‐adaptive navigation, mapping, and digital wandering. Refugees' information behavior is shaped by fractured information landscapes, situational information triggers (e.g., policy changes, deadlines, and health crises), information mediators (technologies, people, and institutions), and the forms information takes, from tangible records to intangible memories. The review shows that affect is not incidental but a central force that guides, amplifies, or inhibits how information is sought, used, or avoided in refugees' lives. It also contributes a framework that positions affect at the heart of refugee information practices in contexts of forced displacement and highlights implications for information system design, services, and policies that are linguistically accessible, trauma‐informed, and responsive to urgent, trigger‐driven needs.
Abstract The integrity of scholarly communication depends critically on the accuracy and verifiability of cited references. Citations enable readers to trace prior work, assess evidence, and situate new contributions within the existing literature. However, concerns have emerged regarding the presence of references in published papers that cannot be resolved to any identifiable source, including citations that appear structurally complete yet correspond to non‐existent publications. While such problems have long been recognized in the form of citation errors and fabricated references, recent developments in automated text generation using large language models have renewed attention on the reliability of bibliographic data in academic publishing. This study addresses two research questions: Can we develop an automated, auditable tool to detect non‐existent or unverifiable references in scholarly manuscripts? What is the prevalence and distribution of such unverifiable references in peer‐reviewed papers at the intersection of artificial intelligence and education published since 2023? To answer these questions, we develop a reference verification pipeline that combines reference extraction with DOI/URL checks and cross‐database matching, and apply it to a corpus of 3201 peer‐reviewed papers (published between 2023 and 2025). We flag 69 papers (2.16%) that contain at least one reference that remains unverifiable under our protocol; such unverifiable references can mislead readers and propagate through citation chains. This suggests that greater care is needed when using, reviewing, and reusing bibliographies.
Data literacy has gained significant momentum as an essential lifelong learning competency to address challenges arising from growth in the availability and accessibility of data. Much work has been done to address the data literacy gap, including both conceptual and empirical pieces from a wide range of disciplines. This paper critically reviews the literature on data literacy to provide a deeper understanding of how it is conceptualized, framed, and assessed across different contexts. Through a systematic search of the literature from various academic fields published from 2000 to 2025, relevant works in this area were identified and evaluated. By adopting a critical review approach, we conducted a conceptual analysis of data literacy by tracing its evolution, examining current interpretations, and exploring the theoretical frameworks that underpin it, including competency‐based models, critical theory approaches, and learning‐centered perspectives. The review also discusses emerging perspectives and gaps in the conceptual understanding of data literacy, highlighting areas for future research and scholarly inquiry.
Abstract The theme of this ARIST chapter is the digital rights of children and youth. It is grounded on a belief in the dignity of the child and an understanding that young people have protection, provision, and participation rights in the technology and online environments in which they live, learn, and play. The chapter is framed by the United Nations Convention on the Rights of the Child (U.N., 1989) and its recent addendum, General Comment No. 25 on Children's Rights in Relation to the Digital Environment (U.N., 2021). Within this structure, the chapter surveys selected topics, such as privacy, the emergence of AI and data surveillance, age verification, access to accurate information, the right to play, minoritized children and inequities of inclusion. Throughout, the chapter connects with young people's voices and practices, as evidenced by the research literature and concludes with a discussion about young people's inclusion in the design of their digital worlds, in terms of participatory design. While the chapter references some policy and regulations, it is not intended to be a comprehensive policy review since the policy environment in relation to children's digital rights is changing rapidly.
Content-based citation analysis seeks to capture the meaning and functions of citations but continues to face unresolved methodological challenges. This study analyzes a stratified sample of Library and Information Science publications to examine how citance segmentation and annotator expertise influence the consistency of classification. Using two annotators with different professional backgrounds, the findings show that agreement is high when citances are defined identically, but reliability decreases sharply once text boundaries diverge. Citance length, rather than subject category or citation density, emerges as the strongest predictor of disagreement. These results identify segmentation as a methodological rather than a purely technical issue, shaping both human and automated tagging outcomes. By highlighting the interplay between expertise effects and boundary definitions, the study underscores the need for clearer operational frameworks in citation analysis. The contribution lies in demonstrating that methodological refinements in citance identification are essential for improving reproducibility, enhancing hybrid human-machine approaches, and strengthening the validity of citation-based indicators in research evaluation.
This study examines the digital curation of the Taipei City Council's political special collections by integrating Morville's User Experience (UX) Honeycomb Model with Bardzell's Feminist Human-Computer Interaction framework. A mixed-methods design was employed, combining a survey of 190 participants with in-depth interviews and participant observation involving 20 respondents with feminist backgrounds. The quantitative analysis evaluated usability, findability, accessibility, credibility, desirability, usefulness, and value, while the qualitative component explored feminist design principles including pluralism, advocacy, participation, embodiment, ecology, and self-disclosure. Findings indicate that the curated platform was strongly recognized for its educational value, clarity of information, and cross-device accessibility, but interactivity and visual appeal were identified as weaker dimensions. The interviews further revealed that the exhibition successfully highlighted women's diverse political trajectories, fostered emotional resonance, and raised awareness of gender equality, though opportunities for participatory engagement and cross-issue connections remained limited. By demonstrating the complementarity of UX evaluation and feminist interpretive analysis, this study underscores how libraries can reposition political archives from static repositories to dynamic platforms for civic engagement and social inclusion.
This study proposes a framework for understanding international collaboration in problem-oriented research that addresses geographically specific issues through distinguishing between investigating and investigated countries. We employ Merton's insider-outsider theory to categorize authors from the countries under study as insiders and those from outside the studied countries as outsiders. Based on the combinations of their shared perspectives, we develop a typology of five collaboration patterns (CPs)-Internal Perspective (CP1), Combined Perspective (CP2), Expanded Perspective (CP3), Partially Overlapping Perspective (CP4), and External Perspective (CP5). An empirical analysis of research related to "Sustainable Development Goal 1: No Poverty" reveals that CP1 is the most prevalent perspective around this topic. Whereas CP5 has seen a gradual decline, CP2 has risen over the years. A case study on international collaboration in poverty research in African countries reveals significant benefits from outsider involvement, including substantial funding from developed countries, enhanced research productivity on specific topics, as well as higher and broader research impact. However, we suggest being attentive to the potential shaping effect of outsiders on the perspectives and research agendas of insiders, which may complicate internal efforts to develop research topics rooted in the local context and the pursuit of domestic development priorities.
Collaboration has become important at all stages of research careers. In data-intensive research fields such as wind energy, many PhD fellows are socialized to such collaboration in networks that train a cohort of PhD fellows. Based on interviews with 23 PhD fellows in four wind-energy training networks, we investigate their expectations and early experiences regarding scientific collaboration. We find that expectations for collaboration are high, but also that the PhD fellows' expectations and early experiences differ. Their experiences are influenced by conditions that make collaboration a balancing act at the organizational, content, process, and personal levels. At the organizational level, learning scientific collaboration involves negotiating interdependencies that are created at the network-proposal stage and presume collaborations about data among the PhD fellows in the network. At the content and process levels, the PhD fellows are dependent on collaboration to collect, analyze, share, validate, and compare data. At the personal level, devising a doable dissertation involves that the individual PhD fellow succeeds in realizing these collaborations. While the studied wind-energy research networks nurture a collaborative environment, the collaboration they presume and the rewards it promises come with the risk of imposing imbalanced collaborations on some PhD fellows.
Consumer-grade sleep-tracking technologies (CSTs) have brought sleep into everyday data practices, reframing it from a clinical concern into a site of personal optimization and reflection. Yet existing taxonomies of sleep-tracking often medicalize users and overlook the complexity of sleep-tracking technologies. This paper presents SleepTax, a naturalistic, multifaceted taxonomy of sleep-tracking technologies based on metadata from 350 consumer devices. Using faceted classification, it identifies five dimensions-user purpose, technological form, functionality, contextual mobility, and temporal mode-and 120 concepts to characterize sleep tracking "in the wild," including 17 forms, 83 functionalities, and multiple engagement styles. We further operationalize SleepTax through a design map and demonstrate how this framework supports scenario-building and speculative inquiry into the sociotechnical consequences that emerge across different personal informatics (PI) infrastructure configurations. Together, SleepTax and the design map form a bridge between classification and design, supporting both systematic description and speculative inquiry in information science, PI, and human-computer interaction (HCI).
Academic discussions on Twitter have increased over the past decade. This study analyzes ca. 1.3 million scholarly conversations on Twitter, examining their temporal evolution, topical distribution, and associated publications. The results show a steady growth in the number of scholarly conversations, accompanied by rising proportions of longer and multi-topic conversations. The visibility of different disciplines and topics varies significantly, with popular themes often closely related to societal events. By comparing topic co-occurrence networks in Twitter conversations with co-citation networks in scholarly literature, the study reveals both commonalities and differences in topic structures across academic and non-academic communication contexts. Moreover, publications that appear in multiple conversations are associated with higher citation counts, suggesting a potential link between conversational engagement and scholarly impact. These findings highlight that the feature of conversations on social media has value not only for refining indicators of scholarly attention and influence beyond simple mention counts, but also to capture more broadly how research is disseminated online.
Large language models (LLMs) generate texts that increasingly circulate as documents in knowledge infrastructures, yet their documentary status remains theoretically underdetermined. Unlike traditional documents, LLM outputs lack identifiable authorship, stable provenance, or testimonial grounding. This challenges foundational assumptions in document theory about authority, accountability, and evidentiary value. This study investigates how AI-generated text acquires documentary force through situated use. Through reflexive case study analysis of an extended ChatGPT dialogue, I apply systematic thematic coding guided by the Model of Documentation Activity (MoDA) to trace documentation activity across physical, mental, and social dimensions. The analysis advances three contributions to document theory and AI governance. First, I demonstrate the analytic utility of the concept of artificially blended testimony (ABT) for examining how LLM outputs provisionally stabilize as documentary artifacts despite lacking testimonial grounding. Second, I show how AI outputs acquire documentary status through human practices of framing boundaries, establishing form, and attributing responsibility. Third, I reconceptualize human oversight mandates in AI governance as documentation thresholds, identifying the documentation practices required for AI outputs to function as preservable, citable, and accountable documents within information infrastructures.
Autocomplete is a search feature that algorithmically generates information cues for any keywords entered in the search bar. While this feature makes the search process more efficient, it also frequently produces biased, misleading, offensive, or otherwise inappropriate suggestions. To address this problem, commercial search systems like Google Search now moderate Autocomplete information cues. However, we know relatively little about how users perceive this moderation process. Conducting interviews with 20 users of web search systems, I examine user attitudes toward the ethical tradeoffs in enacting moderation, reliance on search systems for making regulation decisions, and users' own role in moderating autocomplete information. My findings show that users desire greater visibility into the autocomplete moderation process. They also see flags and personal moderation mechanisms as promising avenues for themselves to exert greater agency within contemporary information infrastructures. My analysis bridges the fields of content moderation and search engine critiques, and lays the groundwork for enacting fair, accountable, and transparent search moderation.
Open Access has changed how research is published and discovered. Studies generally report that OA articles are cited and mentioned more often than non-OA, describing an open access advantage (OAA). The mechanisms causing the OAA are under-investigated: this research analyzes citation and altmetrics post-publication, reporting on the development of OAAs over time. A set of journal articles (N = 44.4 M) is analyzed using seven metrics (citations from journal articles and six altmetrics), using a longitudinal methodology. These data are presented as a ratio of OA:non-OA. OAAs are confirmed for many metrics: early advantages are strong for citation and policy metrics, moderate for news and blogs, and mostly absent for patent citations. Many OAAs are sustained in the years post-publication, although citation OAAs, which typically drop in the year after publication, grow strongly in the following years. OAAs do not generally disappear over time, with the exception of policy citations. Most OAAs appear shortly after publication, and although some develop over time, they are generally sustained in subsequent years. The evidence favors an explanation associated with early activity; other explanations are required to understand subsequent changes in OAAs. There are significant implications for funders, researchers, and publishers.
The promises and perils about scientific team diversity are still debated in the scholarly literature, partly because the importance of underrepresented groups is not fully recognized or valued. In this paper, we summarize two perspectives on team diversity in science: horizontal differences and vertical disparity. Horizontal differences refer to variations across individuals on equal levels, such as differences in gender, nationality, and occupation. Vertical disparity reflects broader social and structural inequalities that limit historically marginalized groups. We introduce team hierarchy, defined as the distribution of power and influence among team members, as a moderating mechanism that helps explain the mixed findings surrounding team diversity and performance. Understanding how hierarchical structures shape diversity is not only important to maximize the benefits of diverse teams but also for enhancing the representation and impact of underrepresented voices. By analyzing 64,038 papers from PloS One and 75,260,139 teams from Microsoft Academic Graph (MAG), our comprehensive study underscores the critical role of team hierarchy in significantly affecting team performance. We also investigate how team hierarchy interacts with team diversity along three dimensions: authors' gender, sector, and country. Interestingly, we find that flat team structures are more positively associated with performance in diverse teams than in homogeneous teams, where members share similar identities. This effect is particularly strong in science compared to social science & arts disciplines. Drawing from social identity and social dominance theories, we propose that flat team structures foster conditions for diverse teams to flourish, enabling minority groups to assume significant roles and wield influential power. Our study contributes valuable insights into team diversity within the scientific community, emphasizing the significance of meaningful inclusion beyond mere numerical representation.