
Abstract Purpose Most minimally invasive surgery (MIS) literature analyses focus on minimally invasive diagnosis and treatment of specific diseases or single procedures, lacking a systematic depiction of the field’s overall structure and evolution, while traditional methods suffer from strong subjectivity and limited data scale. This study aims to employ a topic modeling approach to comprehensively map the knowledge graph and evolutionary pathway of the MIS field. Design/methodology/approach A combination of latent dirichlet allocation (LDA) and dynamic topic model (DTM) was used to analyze all MIS articles published between 1985 and 2024 in 37 high-quality, authoritative PubMed journals (including both specialized surgical and comprehensive journals), examining development maturity, chronological distribution, country distribution, author distribution, hot topic identification, and evolution. Findings The number of MIS publications is projected to grow continuously, with the United States leading and China ranking second; a cohort of highly productive scholars has formed. The field can be grouped into five major topic clusters, among which “minimally invasive diagnosis and treatment of digestive system (gastrointestinal) diseases” currently has the highest influence and attention. Core subtopics running through the hot topics include minimally invasive diagnosis and treatment of colorectal diseases, hernia, and bariatric/metabolic conditions. The research focus of colorectal minimally invasive diagnosis and treatment has shifted from general laparoscopic techniques to refined operations for specific diseases, deeply integrating with precision oncology. Hernia minimally invasive diagnosis and treatment has extended from safety evaluation of basic surgical procedures to instrument innovation, simulation training, and complex hernia management. Bariatric/metabolic minimally invasive diagnosis and treatment has expanded from simple weight-loss surgery to multidisciplinary intersecting areas such as management of metabolic comorbidities, prevention and control of postoperative complications, and health economics. Overall, MIS research is continuously evolving toward greater precision, more systematic integration, and deeper interdisciplinarity. Research limitations The limited journal scope may affect the generalizability of the findings. Practical implications This study reveals the hot topics and evolutionary patterns of MIS research, helping researchers quickly grasp frontier dynamics and providing empirical evidence for hospitals to optimize resource allocation and formulate disciplinary development strategies. Originality/value It pioneers the introduction of topic modeling to systematically mine the vast MIS literature, expanding its application in healthcare services and offering an objective and dynamic cognitive perspective.
Abstract Purpose The discovery, interpretation, and transformation of knowledge into practical insight are being revolutionized by generative artificial intelligence. In light of these developments, this conceptual paper critically reexamines human-guided KDD and knowledge discovery in databases. As a revised framework for understanding knowledge discovery in human AI environments, it suggests Generative Knowledge Discovery in Databases, or Gen-KDD. Design/methodology/approach Using a conceptual approach based on envisioning, the paper seeks to advance theory. In order to create a layered framework that incorporates process, agency, governance, and epistemological reflection, it synthesizes literature on KDD, knowledge creation, human AI collaboration, responsible AI, and generative AI. Findings Knowledge discovery is rethought as a generative and reflexive human AI ecology by the suggested Gen-KDD framework. Generative contextualization, autonomous data curation, generative feature engineering, generative discovery, and co-evolutionary sensemaking are its five stages. These stages are arranged according to human-dominant, AI-dominant, and hybrid layers, making it clear where AI can take the lead, where human judgment is still crucial, and where shared sensemaking is necessary. Research limitations Rather than providing actual evidence, the paper provides a conceptual framework. Future studies should operationalize Gen-KDD across domains and investigate how it affects organizational learning, accountability, transparency, fairness, and decision quality. Practical implications Gen-KDD’s design guidelines can be used by organizations that want to ethically incorporate generative AI into analytics and decision-making. It demonstrates how human oversight, ethics, and governance can be integrated into discovery processes rather than being added as external controls. Originality/value This paper advances a theory-oriented framework that integrates KDD, knowledge-creation theory, human-AI collaboration, responsible generative AI, and epistemological reflection into a single-layered ecology. Gen-KDD offers a richer foundation for knowledge discovery than linear process models by explicitly addressing the changing roles and limits of both human and machine cognition.
Abstract Purpose This study aims to improve fine-grained understanding of citation contexts in scientific literature by jointly modeling two complementary aspects of citation behavior: citation intent and citation evaluation. Design/methodology/approach We propose a joint multi-task learning (Joint-MTL) framework that simultaneously models citation intent classification and citation evaluation classification through a shared encoder with task-specific prediction heads. To support the evaluation task, we construct a new manually annotated dataset, CiteEva. The model is trained using an alternating optimization strategy across the two tasks, enabling the shared encoder to learn representations that capture both functional and evaluative characteristics of citation contexts. Findings Experiments conducted on the public SciCite dataset and the newly constructed CiteEva dataset show that the joint learning framework consistently outperforms strong single-task baselines and several recent multi-task or feature-based approaches. Additional analysis reveals a moderate representation overlap between the two tasks, indicating that they share meaningful semantic signals while maintaining task-specific distinctions. These findings empirically support the effectiveness of jointly modeling citation intent and citation evaluation. Research limitations The current study focuses on two citation-related tasks and evaluates the framework on a limited set of datasets and citation categories. Future work could extend the approach to additional citation analysis tasks and larger-scale corpora. Practical implications The proposed approach can support applications such as scientific impact assessment, literature review assistance, and scholarly knowledge mining by enabling more nuanced interpretation of citation roles and evaluative stances. Originality/value This study provides empirical evidence for the benefit of jointly modeling citation intent and citation evaluation and introduces CiteEva, a new high-quality dataset for citation evaluation research, contributing to more comprehensive citation context analysis.
Abstract Purpose Although citation-based indicators are widely used, they are not useful for recently published research, directly reflect only one of the three common dimensions of research quality, and have little value in some social sciences, arts and humanities. Large Language Models (LLMs) may address some of these weaknesses. Design/methodology/approach This article reports a science-wide assessment of the research quality scoring capability of ChatGPT-4o mini, ChatGPT-4o, and ChatGPT-5 mini. It correlates ChatGPT scores, averaged over 5 repetitions, with departmental average quality scores for 107,212 UK-based journal articles. Findings ChatGPT-4o is marginally better than ChatGPT-4o mini in most of the 34 field-based Units of Assessment (UoAs) tested. ChatGPT-4o scores have a positive correlation with research quality in 33 of the 34 UoAs, with the results being statistically significant in 31. ChatGPT-4o scores had a higher correlation with research quality than long term citation rates in 21 out of 34 UoAs and a higher correlation than short term citation rates in 26 out of 34 UoAs. The most substantial exception is Physics, for which citations are more useful. ChatGPT-5 mini has even stronger correlations overall and for departmental averages, it correlates more strongly with quality scores than do citations in 31 out of 34 UoAs, with correlations reaching 0.905. Research limitations All articles assessed are from the UK. The practical value of LLM scores for decision making is not assessed. Only the Normalised Log-transformed Citation Score (NLCS) was tested against ChatGPT rather than other citation rate indicators. Practical implications ChatGPT can be considered as a research quality indicator to support expert judgement in almost all academic fields. Originality/value The results give science-wide evidence that ChatGPT-4o mini, ChatGPT-4o, and ChatGPT-5 mini are competitive with citations as new research quality indicator sources, and technically better in most fields.
Purpose This article assumes that embracing the values of open science is central to the ongoing reform of research assessment. Two main objectives are pursued: (1) to analyze whether responsible research assessment (RRA) and open science principles are being implemented in the Severo Ochoa (SO) call for proposals for centers of excellence in Spain; (2) to study the engagement of a selection of centers holding the excellence seal with open science practices.Methodology First, a longitudinal study of the content of the SO calls from 2011 to 2025 is conducted to identify changes reflecting adoption of the RRA and open science principles. Second, the open science practices of the selected centers are analyzed across two operational dimensions: i) open scientific knowledge and infrastructure and ii) openness to non-academic actors.Findings The call has evolved over the years, adapting its evaluation criteria to the principles of RRA: drastic decrease of the role of publication-based indicators, wider range of contributions considered and use of narrative CVs. Open science practices are increasingly promoted, particularly in centers' strategic plans. Centers with the seal of excellence show higher levels of open publications and societal impact than the average in their fields. These centers show quite a strong commitment to the values of open science, although there is still room for improvement.Research limitations The analysis is limited by the lack of solid, reliable sources and indicators for some open science practices. Enhanced standardization and interoperability of information sources are needed to generate reliable indicators for monitoring these practices.Practical implications The RRA principles are being integrated into the SO call, but some open science practices still require greater promotion. This is relevant, because centers of excellence can set the benchmarks for the country's entire scientific community.Originality/value Studying how RRA and open science principles are implemented by funding agencies is crucial, as funders have a clear influence on researchers' behavior.
Purpose Organizations are major actors in multiple dimensions of the research landscape. Subsequently, accurate organization identification is a prerequisite for precise studies about, for example, collaboration in publication practice or project participation. Global reference databases are essential instruments in this context, but recent experience in Flanders showed that ROR (the worldwide Research Organization Registry) does not contain a sufficient set of records to code the full spectrum of organizations involved in research and innovation. This raises an important question: how can we structurally map all relevant science and technology actors in analysis and studies of different aspects of the research and innovation landscape? We elaborate how this question is being addressed in Flanders through the development of the Flemish Organization Registry (FOR).Design/methodology/approach FOR, as a local expansion of ROR, provides unique identifiers for all additional organizations that appear in author affiliation datasets or project participant files. We describe the source data, the consolidation, the addition of metadata and the relation to the international reference database.Findings We find, through a case study, that the use of FOR delivers structured, fine-grained information about a full set of organizations in multiple dimensions of science in a Flemish context.Research limitations The paper introduces FOR as a proof of concept and discusses one case study only.Practical implications FOR and similar regionally anchored infrastructures can be deployed as tools to enrich data when analyzing a national research context organization-wise.Originality The article presents a national extension of ROR as an example of a way to map all actors in a national research system.
Abstract Purpose This study examines how data-driven research (DDR) has diffused across U.S. higher education institutions and investigates its relationship with scientific impact as measured by citation counts. It aims to clarify how institutional context, disciplinary affiliation, and publication characteristics shape the visibility of data-intensive research. Design/methodology/approach Using publications indexed in the Web of Science from 2013 to 2023, DDR is identified through a transparent text-mining approach based on abstract-level keywords related to artificial intelligence, machine learning, big data, and data science. Disciplinary, institutional, and demographic patterns in DDR output are analyzed. Citation counts are modeled using Zero-Inflated Negative Binomial regression, Random Forest, eXtreme Gradient Boosting, and Support Vector Regression. Feature variables capture DDR status, institutional classification, population served, and publication-level characteristics. Findings Results indicate that DDR output expanded across all research areas, with the strongest growth in computer science and engineering and disproportionately large contributions from R1 institutions. Across modeling approaches, publication age, number of authors, and DDR status consistently emerge as the most influential predictors of citation impact. Research limitations DDR identification relies on keyword-based text classification, which may omit relevant studies or include unrelated work. The Web of Science database emphasizes English-language and high-impact journals, potentially limiting coverage. In addition, the observational design supports association rather than causal inference. Practical implications The findings inform research evaluators, policymakers, and institutional leaders about patterns of methodological diffusion and citation visibility. The proposed framework can support evidence-based assessment of data-driven scholarship and guide capacity-building efforts, particularly in less-resourced institutions. Originality/value This study integrates bibliometric analysis, text mining, and machine-learning methods to examine data-driven research at scale. It provides a reproducible approach for identifying DDR and offers new empirical evidence on how institutional and disciplinary contexts shape research visibility, contributing to science-of-science and research evaluation literature.
Purpose This study aims to investigate the interdisciplinary trend of human-computer interaction (HCI) through scientometric methods, which has been widely discussed in academia but still lacks quantitative evidence.Design/methodology/approach In this study, combining scientometric measures of disciplinary diversity with network coherence, we examined the evolution of interdisciplinarity of HCI over the past 20 years from the perspective of knowledge integration.Findings Our findings indicate that the disciplines contributing knowledge to HCI have become increasingly diverse. While tending towards high disciplinary heterogeneity, this knowledge has also led to a more even distribution of disciplinary sources within HCI. In terms of network coherence, the structure of knowledge sources of HCI has tended towards decentralization. The phenomenon of the 'rich-club' has shown a trend of 'obvious presence-fluctuation-absence-obvious presence'.Research limitations The study has only analyzed the evolution of the interdisciplinarity of HCI from the perspective of knowledge integration, and have not yet analyzed the subsequent knowledge diffusion of HCI research. Meanwhile, this study focuses on the development trend of human-computer interaction as an interdisciplinary field, and does not further analyze the "rich-club" nodes in the annual knowledge integration network, whose attributes and changes may be related to the evolution of HCI.Practical implications Practitioners and policymakers in HCI should foster collaborations across distant disciplines through targeted seminars, funding initiatives, and open research environments to reduce knowledge barriers and unlock innovative integration. Additionally, leveraging the cyclical "rich-club" dynamics in knowledge networks can aid in balancing core knowledge reliance with interdisciplinary inclusivity.Originality/value The revealed trends in the evolution of interdisciplinarity can provide insights for interdisciplinary research and the construction of knowledge system in HCI.
Purpose This paper presents a unified framework for evaluating synthetic data across utility, fidelity, and privacy, with the goal of improving trust and reliability in scientific applications where realism alone is not sufficient.Design/methodology/approach The article defines synthetic data generation as a constraint-aware process guided by domain knowledge and introduces a structured pipeline that includes data auditing, controlled generation, and multi-criteria evaluation. The framework is validated through a Harmful Algal Bloom case study comparing statistical, deep, and quantum generative models under consistent training and evaluation settings.Findings Results show that no single model dominates across all criteria. Statistical models best preserve correlation structure and achieve high predictive performance, deep models increase variability but reduce fidelity, and quantum models improve privacy by increasing separation from real data at the cost of accuracy. Combining real and synthetic data improves segmentation results, with the best model achieving mIoU of 0.553 and Dice of 0.668. Overall, synthetic data quality depends on trade-offs rather than a single metric.Research limitations The evaluation focuses on one environmental dataset and a limited set of generative models. Results depend on data quality and chosen configurations, and fairness is not fully explored.Practical implications Model selection should match the application goal: high-risk scientific tasks require strong fidelity, while privacy-sensitive settings benefit from models that reduce reconstruction risk. The framework supports structured, repeatable deployment of synthetic data pipelines from analysis, synthetic data is reliable and performance is as good as observation data.Originality/value This paper introduces a multi-objective evaluation framework that integrates utility, fidelity, and privacy into a single decision process, shifting synthetic data from a realism-focused task to a governed, context-dependent system.
Purpose This study aims to systematically compare patent-to-patent and patent-to-paper citations, and to examine how their differences reflect distinct modes of knowledge flow between technological development and scientific research.Design/methodology/approach Using United States patent data from PATSTAT, combined with the Reliance on Science dataset and the Microsoft Academic Graph, we conduct a multi-dimensional analysis across four measures: citation frequency, citation time lag, semantic similarity, and science intensity. Patents are further classified by technological domain, innovation type, and citation source to capture heterogeneity in citation patterns.Findings Patent-to-patent citations are more frequent and exhibit higher semantic similarity, whereas patent-to-paper citations tend to occur with shorter time lags. Patents with higher impact are more likely to be associated with stronger linkages to scientific knowledge, particularly within a moderate range of influence. Emerging technological fields show a greater tendency to integrate recent scientific knowledge, while exploratory innovations draw on more technologically similar and interdisciplinary sources and exhibit shorter citation lags. Applicant citations are slightly more frequent, whereas examiner citations are more temporally proximate and semantically aligned with the citing patents.Research limitations The analysis is based on citation data, which capture observable linkage patterns but do not directly identify causal mechanisms. In addition, classification of technological domains and innovation types may not fully account for within-field heterogeneity.Practical implications The findings provide empirical insights for policymakers, research institutions, and corporate R&D practitioners seeking to better understand and manage the interaction between scientific research and technological innovation, and to design more effective innovation and patent strategies.Originality/value This study offers a comprehensive and systematic comparison of patent-to-patent and patent-to-paper citations across multiple dimensions and contexts, contributing to a more nuanced understanding of science-technology linkages and the structure of knowledge flows in innovation systems.
Purpose To address the limitations of traditional patent metrics in capturing technical substance and the high cost of expert review, this study proposes a hybrid evaluation framework integrating Large Language Models (LLMs) with machine learning to achieve automated, highly accurate identification of high-value patents.Design/methodology/approach Adopting a "Virtual Assessor" paradigm, we constructed a dataset based on the China Patent Gold Awards. The study integrated semantic scores from three diverse LLMs (DeepSeek, Qwen, GLM) under zero-shot and few-shot prompt strategies into a Stacking ensemble learning model (combining XGBoost, Random Forest, and SVM) to predict patent value across nine comparative experimental setups.Findings Direct LLM evaluation revealed a "Knowledge Injection Paradox," where explicit expert prior knowledge caused negative transfer and reduced accuracy due to over-conditioning. However, the Stacking model successfully rectified these biases, transforming subjective LLM evaluations into robust predictive features. The hybrid model achieved over 97 % accuracy in identifying high-value patents, demonstrating strong robustness even in high-noise environments.Research limitations The study relies on a binary classification of extreme samples (Gold Award vs. non-awarded), potentially oversimplifying the continuous distribution of patent value. Furthermore, the interpretability of the "black box" feature fusion mechanism requires further exploration.Practical implications The proposed framework offers IP managers and policymakers a scalable, cost-effective tool for automated patent screening, effectively bridging the gap between qualitative expert intuition and quantitative data precision.Originality/value This research introduces a "Semantic Enhancement + Algorithmic Rectification" paradigm. It empirically demonstrates how machine learning can correct LLM hallucinations and biases, marking a significant shift from data-driven perception to AI-driven cognitive decision-making in patent valuation.
Purpose Academic documents require expert time to evaluate, and Large Language Models (LLMs) might support this through score or decision predictions. For confidential structured academic texts, such as grants and Impact Case Studies (ICSs), medium-sized LLMs can be run offline without expensive computing infrastructures, enhancing security.Design/methodology/approach This study evaluates for the first time how well medium-sized LLMs can score structured academic documents using the UK Research Excellence Framework (REF) 2021 ICSs, and whether LLMs can guess scores from individual sections. We obtained score estimates from five recent popular LLMs (DeepSeek R1 32B, Qwen 3 32B, Magistral Small 24B, Gemma 3 27B, and Llama 4 Scout 27B) across 6,010 REF 2021 ICSs, correlating the scores with a proxy quality rating (departmental average score).Findings Scoring the full texts was only moderately effective (in terms of correlations with the proxy quality rating) and Llama 4 failed to score most of the longest. Surprisingly, all LLMs except Magistral were able to make statistically significantly above random guesses at ICS scores from each of the individual component sections (summary, underpinning research, references, details of the impacts, and sources to support the impact). A logical two-stage approach mimicking the human reviewer instructions did not outperform focusing on impact alone. The best strategy was to score the summary and the details of the impact sections combined (five times, averaged) with Gemma 3. This gave the highest Spearman correlation (0.37) with departmental average proxy quality scores (0.55 for department-level correlations).Practical implications Medium sized LLMs can be used to score structured academic documents to support research assessments.Research limitations This uses a single large case study with a public, albeit obscured, gold standard.Originality/value This improves on the state of the art despite the additional restrictions and with a much cheaper and potentially private open weights LLM approach.
Purpose This study investigates the diversity of national scholarly journal publishing ecosystems in seven countries across Europe and Latin America: Argentina, Brazil, Colombia, Finland, Mexico, Poland, and T & uuml;rkiye. It challenges the common perception that global scholarly publishing is dominated by international commercial publishers by examining national publishing structures beyond English speaking contexts.Design/methodology/approach Using ISSN Centre data and national sources, we analyse journal-level publishing structures rather than article- or citation-level outputs. Publishers were categorized according to their institutional and organizational characteristics. The analysis focuses on active journals, defined as those with a recorded start year and no identified termination date. Journal coverage in Web of Science, Scopus, and OpenAlex was examined to assess how national publishing landscapes are represented in major bibliometric databases.Findings Educational institutions emerge as the primary publishers in most countries, representing more than 75 % of journals in Colombia and Brazil and more than 50 % in Mexico, Argentina, and Poland. Finland stands out, with scientific and professional associations leading journal publication at 62 %. Commercial publishers hold comparatively small shares, reaching their highest levels in T & uuml;rkiye at 12.1 % and Poland at 8.2 %. In terms of database representation, OpenAlex indexes over half of the journals in most countries, whereas Web of Science (WoS) and Scopus cover only a small portion.Research limitations The study relies on ISSN and national datasets, which differ in completeness and standardization. Variations in national reporting practices and database indexing policies may influence coverage comparisons. As the analysis is based on currently active journals, historical trends reflect surviving journals only.Practical implications The results provide evidence for policymakers, database providers, and research evaluators to recognize the diversity of national publishing systems. They highlight the importance of improving data sources and analytical approaches to ensure more accurate assessment of scholarly communication outside heavily commercialized environments.Originality/value The study offers a comparative analysis of national journal publishing ecosystems across seven countries, revealing structural differences that challenge the assumption of a globally uniform publishing model. It underscores the need for bibliometric research frameworks that include and accurately represent national and regional publishing structures.
Purpose Retraction count is a widely used metric for research integrity, but it overlooks the heterogeneous impact of retracted papers. Erroneous information from retracted papers tends to spread to subsequent research, compromising downstream reliability. The citation impact and lifespan of retracted papers represent a critical yet underexplored dimension of research integrity, warranting inclusion in evaluation frameworks.Design/methodology/approach This study proposes a framework for analyzing the citation impact and lifespan of retracted papers, applied to over 50,000 retracted articles indexed in the Web of Science (WoS).Findings Retracted papers accumulate substantial citations, yet their distribution is highly skewed, with 20 % of articles capturing 75 % of all citations. While the majority exhibit short citation lifespans, a small subset sustains influence for decades. Significant disparities exist across disciplines: papers in Clinical & Life Science tend to attract more citations and maintain them over longer periods, whereas papers in Computer Science generally receive fewer citations that decline rapidly. Network-level analysis further reveals that retractions are not randomly scattered but tend to cluster within citation networks, forming patterns resembling propagation along citation pathways. Articles citing retracted works are associated with substantially higher subsequent retraction rates, and this association appears to strengthen with the number of retracted papers cited.Research limitations The analysis was confined to WoS data, which may underestimate citation lifespan. Citation counts do not distinguish between positive, negative, and neutral citations. As an observational study, we identify associations rather than establish causation.Practical implications These findings provide a nuanced understanding of how retracted papers influence subsequent research across disciplines. The framework supports targeted integrity governance.Originality/value This study proposes integrating the citation impact and lifespan of retracted papers into scientific integrity metrics, thereby expanding the dimensional framework for evaluating research integrity beyond simple retraction counts.
Purpose The Open Researcher and Contributor Identifier (ORCID) is becoming the de facto standard for researcher identification in scholarly communication, providing a persistent unique identifier and a registry that functions as a digital CV. The purpose of this study is to analyse the ORCID profiles of a selected group of leading researchers to analyse creation, completion, and updating of their records, with particular attention to the Works section.Design/methodology/approach We focus on the 357 grants awarded by the European Research Council (ERC) to researchers with a Spanish host institution between 2014 and 2020, for whom a high degree of ORCID adoption has been reported. Data included in their ORCID records were downloaded and the completion and dynamics of Personal Information and Activities sections are studied. Differences by domain and researcher career stage are explored.Findings All ERC researchers have an ORCID iD, and in most cases, their records are publicly available. ORCID profile completion is quite high in the Activities sections, particularly Works and Employment, and lower in the Personal Information sections. Most of the profiles were created by the users themselves and had been updated recently (75 % in the last three months). Sections that allow for automatic input tend to show higher completion rates and are more recently updated. Although ORCID accepts all kind of research outputs, journal articles form the majority (84 %) in our study. Works are added to profiles mainly by commercial entities (especially Elsevier and Clarivate), with non-profit organisations (e.g., Crossref) a distant second. Only 10 % of works are included by researchers themselves.Research limitations As the study examines a particular group of elite researchers, their practices regarding profile completion and updating cannot be generalized to other researcher populations.Practical implications ERC-funded researchers show quite high engagement with the ORCID system, but there is uneven completion of record sections and data quality issues that need to be addressed to improve the system and consolidate identifier use.Originality/value ORCID record completion, work collection and updating behaviour have been insufficiently explored to date. The study of a population of leading researchers helps reveal ORCID record completion practices, as well as data quality issues and controversial approaches to work collection.
Purpose This study seeks to understand how the map of science has evolved and to identify the forces driving that evolution. By examining changes in these maps over time, we track the development of science through the growing interdisciplinarity of subject categories and the way they cluster together. Design/methodology/approach We integrate multiple classification schemes from Web of Science products to build a multilevel framework that connects journals, categories, groups, and broad domains. Using Journal Citation Reports (JCR) data from 2011, 2016, and 2024, we construct two types of maps of science: one based on citation relationships and another based on the sharing of journals across categories. We then examine how these maps evolve, identify factors influencing their development, and analyze how knowledge percolates through a multilevel structure. Findings The map of science has evolved from a bipolar structure into a more interconnected, rounded triangular configuration. In 2016, Arts & Humanities and the Social Sciences comprised a single cluster; by 2024, they had separated, while Biological Sciences and Medical Sciences, once distinct, had merged into a unified cluster. At the same time, categories such as education, special education, and applied psychology shifted toward hearing and speech pathology, forming a new special education cluster at the intersection of the social sciences, biomedicine, and technology. The expansion of Arts & Humanities categories, along with the addition of new categories in the Journal Citation Reports (JCR), and the resulting growth in interdisciplinarity across categories and journals, has played a key role in reshaping the overall map of science. Although categories within a cluster may disperse across multiple groups and broader domains through knowledge percolation in a multilevel system, they nonetheless remain relatively concentrated. Research limitations As with any empirical investigation, this study has some limitations. Most notably, our analysis is based on only three points in time and relies on a single data source, namely the Web of Science (WoS). Yet, we described our results in far more detail than is usually done. Practical implications Maps of science serve as tools for navigating the research landscape, helping to inform strategic investments and shape future research directions. Examining how these maps evolve over time and how such changes influence research trajectories is a central concern of the science of science. Originality/value Mapping clusters to their respective groups and broad categories reveals a hierarchical classification system in which clusters extend beyond disciplinary boundaries. Overlaps among categories, groups and broad categories indicate that scientific knowledge percolates through traditional field divisions.
Purpose While the imperative of enterprise digital transformation (EDT) has been widely acknowledged, a systematic understanding of its intricate network of antecedents and consequences remains fragmented. This study proposes a novel knowledge representation framework that leverages large language models (LLMs) to construct a variable relational network (VRN), offering a panoramic, micro-level perspective on EDT. Design/methodology/approach We extract five types of variable relationships from a vast corpus of academic publications on EDT to generate the VRN. Subsequently, we apply network topology analysis to uncover the temporal and regional characteristics of the VRN. Its hierarchical structure is then analyzed through K-shell decomposition. Findings Our results show that, over the past two decades, the scale of the VRN has experienced rapid growth, driven collectively by multi-layered external factors such as the rapid advancement of digital technologies, and its internal connections have become increasingly tighter. Regional comparisons of the VRN reveal that different economies, shaped by institutional theories, exhibit distinct transformation paradigms while striving toward common goals. K-shell analysis uncovers a clear hierarchical structure, distinguishing peripheral, intermediate, and core variables, with these layers corresponding to varying degrees of strategic significance and transformation maturity. Research limitations The study's limitations primarily concern the accuracy of the VRN, which depends on the LLM's extraction performance and its potential for hallucinations, which may introduce noise into the network topology. Practical implications The VRN and its network topology structure serve as a diagnostic tool for strategic decision-making, enterprises and policymakers can also use these insights to design targeted support programs. Originality/value This study contributes a data-driven, LLM-assisted framework for mapping the evolving and multidimensional landscape of enterprise digital transformation, thereby validating and extending the theoretical boundaries of EDT.
Purpose This study analyzes 25 years of retraction trends in Sub-Saharan Africa, focusing on underlying reasons, publisher involvement, disciplinary patterns, authorship characteristics, and country-level distributions. Design/methodology/approach The whole database of Retraction Watch was downloaded into a spreadsheet file. Retraction data for all 46 countries in the Sub-Saharan African region was extracted Publication data of all 46 countries in the Sub-Saharan African region was retrieved from Web of Science to calculate retractions rate (i.e. retractions per 1,000 articles). Findings Approximately 75 % of retracted papers from Sub-Saharan Africa involved international collaboration based on Retraction Watch data. Ethiopia leads in total retractions, accounting for about half of all retractions from the region. India, Saudi Arabia, and China emerged as the most prolific external contributor to retractions from Sub-Saharan Africa. Compromised peer review, unethical generative AI use and paper mills emerged as major concerns by 2023. Research limitations The analysis relies on data from Retraction Watch and Web of Science, which may not capture all retractions, especially those from local Sub-Saharan African journals that are not indexed in international databases. Secondly, retraction records are often delayed relative to publication dates, potentially skewing 2024 trends. Practical implications In practice, emerging patterns indicate a growing crisis in research integrity, marked by the rise of new retraction categories over the past three years-such as paper mills, peer review compromise, and unethical use of generative AI-and a sharp increase in longstanding issues like plagiarism. Originality/value This study offers a comprehensive analysis of retractions in Sub-Saharan Africa, addressing a significant gap in the literature. Besides, covering a 25-year period and examining key themes-including retraction reasons, publisher involvement, disciplinary patterns, authorship characteristics, and country-level distributions-it presents original insights that contribute meaningfully to the global discourse on research integrity and retraction trends.
Purpose This empirical case study employs institutional theory to analyze the strategic publication behavior of individuals conducting practice-based research (PBR) at Flemish universities of applied sciences (UASs). It examines how competing institutional logics and pressures within the higher education field shape dissemination practices. Design/methodology/approach To supplement bibliometric studies, a dataset of articles (2014-2024) from selected Flemish UASs was compiled from institutional repositories and supplementary sources. The study traces how article types, language, and authorship align with different institutional demands. Findings Analyses reveal a dual dissemination pattern shaped by isomorphic pressures. State reporting systems' coercive pressures and mimetic pulls toward academic legitimacy drive substantial output in English-language, peer-reviewed journals. Simultaneously, professional fields' normative pressures sustain a parallel, diverse circuit of Dutch-language, practice-oriented outputs (e.g., magazines, blogs). UASs and researchers navigate these logics strategically, demonstrating institutional entrepreneurship through self-publishing platforms and cross-channel dissemination. Research limitations The study's focus on articles underrepresents the full spectrum of practice-based research dissemination. Furthermore, given institutional repositories' under-registration, the dataset may not capture all relevant publications, reflecting the very registration challenges the study examines. As a case study of Flemish UASs, generalizability to other contexts requires caution. Practical implications Findings highlight a systemic misalignment between PBR dissemination and academic metadata standards. They underscore the need for research information systems and evaluation frameworks that recognize plural institutional logics to adequately valorize the impact of practice-based research. Originality/value This study provides a novel institutionalist perspective on PBR dissemination practices, framing them as strategic responses to a complex organizational field. It advances understanding of how competing logics manifest in knowledge dissemination and identifies institutional work undertaken to sustain and legitimize a practice-based mission.
Purpose The study examines how local topics in the Flemish Academic Bibliographic Database for the Social Sciences and Humanities (VABB-SHW) are positioned within a disciplinary framework. It explores their size, language profile, and disciplinary profiles compared to the broader topic landscape.Design/methodology/approach Topics were extracted using the clustering strategy of (Guns, R. 2024. "A Bibliometric Map of Local Research in the Social Sciences and Humanities." In Research Evaluatuion in Social Sciences and Humanities 2024. Galway, Ireland) with BERTopic, combining multilingual embeddings, UMAP dimensionality reduction, and HDBSCAN. Descriptions were generated with GPT-4o-mini, labelled with Gemini-2.5-Flash, and classified with a content-based model trained on Web of Science data and applied to VABB-SHW (Arhiliuc, C., R. Guns, and T. C. E. Engels. 2025b. "Text-Based Classification of all Social Sciences and Humanities Publications Indexed in the Flemish VABB Database." In Proceedings of the 20th International Conference on Scientometrics & Informetrics (ISSI, 2025)).Findings Out of 517 topics, 76 (17.2 % of publications) were identified as local. They contain more non-English publications, and cluster mainly in "History", "Law", "Literature", "Political science", and "Art". Contrasts emerge in their profiles: "Law" topics are internally consistent, "History" topics diffuse across disciplines, and "Literature" is consistently classified when modal but tends to be overattributed otherwise.Research limitations/implications The results reflect the scope of VABB-SHW and the narrow definition of "local". Topic descriptions and disciplinary expectations may introduce uncertainty. The findings are not directly generalizable, but the approach can be replicated in other national databases and with broader definitions to test robustness.Practical implications The approach illustrates how national bibliographic databases can be systematically analysed to identify and profile locally anchored research, offering a basis for comparative studies across regions.Originality/value This is the first study to systematically analyse local topics in VABB-SHW, combining topic modelling and content-based classification to highlight how SSH research engages with nationally specific issues.