
The idea of a Hierarchy of the Sciences (HoS), first proposed by Auguste Comte in the 19th Century, argues that scientific disciplines can be ordered in a hierarchy of complexity, with the inorganic sciences (Astronomy, Physics) being the least complex and the organic sciences (Biology, Sociology) being the most complex. Nearly 200 years later, this order still reflects widely held attitudes towards the relationships between and within academic (sub-) disciplines. HoS-related hypotheses have been supported by a variety of studies using bibliographic, linguistic and scientometric data. The present study investigates whether there is further linguistic, more precisely, grammatical evidence to support or qualify the HoS. We collected a corpus of 39,809 research article abstracts (7,630,993 words) from the Web of Science (WoS) database, performed a multidimensional analysis to identify typical patterns of co-occurring grammatical features and then compared the results of the sub-corpora of the disciplines commonly included in research on the HoS (Arts Humanities, Sociology, Psychology, Biology, Chemistry, Physics, Astronomy Astrophysics and Mathematics). Although some of our findings replicated those of previous studies, notably that the grammatically most similar disciplines are those adjacent in traditional arrangements on the HoS, we did not find evidence of a simple one-dimensional hierarchy. Instead, our findings indicate that typical grammatical patterns in the disciplines at the extremes of the traditional hierarchy also resemble one another. This suggests that, as far as grammatical patterns are concerned, representing disciplines in a two-dimensional space of the sciences is superior to a one-dimensional hierarchy.
Scholars exhibit gender differences across various aspects, including research productivity, research influence, collaboration patterns, funding, and self-citation. To examine these gender differences from the perspective of knowledge profiles, we employed three indicators to analyze the gender differences in the scholars’ knowledge profiles in the field of Library and Information Science (LIS), including knowledge breadth (KB), knowledge depth (KD), and knowledge structuration degree (KS). Our results indicate significant gender differences in the knowledge profiles of top LIS scholars with the highest h-index. Specifically, although male scholars show significant advantages in overall KB, KD and KS, female scholars exhibit stronger performance in KS when authorship order weighting is considered, reflecting more interconnected and systematically organized knowledge points within female scholars’ own knowledge profiles. This study offers a new research perspective on gender differences in LIS academia and contributes to the theoretical discussion of the relationship between knowledge profiles and scholarly performance.
As representatives of Chinese prominent scientists, the academic directions of Chinese Academy of Sciences academicians and Yangtze Scholars are exemplary examples of successful research strategies. Employing techniques including the BERTopic model, Dynamic Time Warping (DTW), and trajectory clustering analysis, this study analyzes 505,639 publications from 1,513 top scientists. It formulates and calculates their R^2 , e , and JS at various times, constructing time series to represent the changes in the breadth, depth, and degree of shift in their research directions. The study identifies several categories and aspects of research direction evolution patterns by analyzing the morphological properties of the time series of numerous indicators for diverse researchers. The results indicate that most prominent scientists’ research directions initially diverge and later converge, showing an overall trend of increasing depth in their studies. Their research directions often undergo significant adjustments in the early stages, while scientists in the humanities tend to make substantial changes to their research directions in the later stages. Furthermore, Chinese selection mechanism for prominent scientists has shifted from an early emphasis on the breadth and divergence of their research directions to a focus on the depth and concentration of those directions. Typically, there is a requirement for their research directions to be flexible and adjustable in the early stages while remaining stable and concentrated in the mid to later stages. Additionally, scientists who delve deeply into a specific research area tend to grow at a faster rate.
Bibliodiversity has emerged as a central concept in debates on open science, scholarly communication, and research evaluation, highlighting the need to sustain linguistic, geographic, epistemic, and infrastructural diversity in knowledge production and dissemination. While its normative foundations are well established, operationalising bibliodiversity at scale remains methodologically challenging and strongly dependent on the availability of inclusive, interdata infrastructures. This paper presents a methodological demonstration of how bibliodiversity can be systematically analysed using the OpenAIRE Graph, a large-scale, open Scholarly Knowledge Graph that aggregates metadata from publishers, repositories, and research infrastructures worldwide. Drawing on the multidimensional framework proposed in the literature, we illustrate how key dimensions of bibliodiversity—linguistic, geographic, disciplinary, publisher, business model, and infrastructure diversity—can be operationalised using the OpenAIRE data model and queried reproducibly in Google BigQuery. Through a series of exploratory analyses, we show how the Graph captures the “long tail” of scholarly communication, including non-English products, diverse research formats such as datasets and software, and contributions from institutional and thematic repositories, particularly beyond the Global North. Rather than identifying new empirical trends, this work demonstrates the analytical potential of the OpenAIRE Graph as an open, community-governed infrastructure for bibliodiversity research. By making diversity measurable across multiple, interconnected dimensions, the OpenAIRE Graph provides a robust empirical basis for future studies and supports evidence-informed policies to foster a more inclusive and equitable scholarly communication ecosystem.
Accurately quantifying scientific innovation is a central issue in research evaluation. This paper proposes and validates a unified three-dimensional framework for measuring scientific innovation at the paper level. To address limitations of existing single-axis indicators, we conceptualize innovation as an integrated construct composed of textual novelty, scholarly impact, and citation-based disruptiveness. Methodologically, we leverage the deep semantic reasoning and comparative capabilities of Large Language Models (LLMs), augmented by a Retrieval-Augmented Generation (RAG) that retrieves relevant abstracts, to quantify textual novelty; we operationalize impact via normalized citation counts and capture disruptiveness using the CD index; these heterogeneous signals are integrated using an entropy-based weighting strategy to generate an interpretable 3D innovation metric. We empirically validate the framework on a corpus of 10,875 documents from S2ORC, evaluating convergent, discriminant, and incremental validity. Findings indicate that the proposed 3D measure outperforms single-dimension metrics in distinguishing breakthrough, disruptive and nascent research, and supports an interpretable eight-quadrant typology for innovation classification.The framework offers a practical and valuable tool to research assessment and prioritization, and provides a methodological foundation for integrating deep semantic representations with citation-topology analysis in future studies. The source code for this study is publicly available at: https://github.com/sdx5256/Scientific-Innovation_3d .
Large language models (LLMs) are increasingly employed in scientometric studies for scientific text analysis and knowledge discovery. A key factor influencing their in-context learning (ICL) performance is the organization of demonstrations, which guides how LLMs interpret complex scientific texts. However, existing organization strategies often involve high computational costs and lack interpretability, limiting their effectiveness and transferability across models. We propose In-Context Curriculum Learning (ICCL), a lightweight and explainable framework that organizes demonstrations in ascending order of complexity to enhance LLMs’ deep semantic understanding of scientific text. ICCL comprises three stages: demonstration retrieval, label recall, and curriculum-guided ordering based on a novel Instruction Alignment Score (IAS) to quantify contextual difficulty. We evaluate ICCL on three representative scientific text mining benchmarks, SciCite, SciNLI, and SciERC, covering citation intent classification, scientific language inference, and entity extraction tasks. Relative to few-shot prompting, ICCL yields consistent average F1-score gains of 4.65 https://github.com/61peng/sci-iccl .
In recent years, the proliferation of open data initiatives has increased the use of large-scale data platforms in scientific research. While these platforms are anticipated to transform research practices, the mechanisms through which open data influences the evolution of scientific fields remain poorly understood. This study proposes an analytical framework for investigating shifts in researchers’ topical orientations before and after adopting data platforms. This framework is applied through a case study leveraging data from the Global Biodiversity Information Facility (GBIF), demonstrating how engagement with open data can trigger systematic changes in research agendas. The findings revealed a tendency among researchers to transition from themes like “Climate Change and Forest Ecology” and “Forest Ecosystem Structure” toward areas such as “Plant Biodiversity and Soil Ecology” and “Conservation and Habitat Management”. These findings suggest that GBIF directs scholarly attention toward specific research fields and influences topic selection. Research aligned with European policy priorities, such as pest management and risk assessment, has also expanded following GBIF adoption. Comparative analyses revealed distinct patterns of topic transitions among GBIF users compared with non-GBIF users. Notably, transitions toward themes such as “Plant Biodiversity and Soil Ecology” and “Plant-Host Interactions” were prominent, indicating that these themes have been further explored with the use of GBIF. Conversely, diminished attention was observed in topics such as “Taxonomic Studies” and “Forest Modelling”. These results underscore the role of open data platforms—exemplified by GBIF—as foundational infrastructures shaping the evolution of research trajectories.
Academic mobility is a central concern in science policy, yet research has focused largely on international mobility or migration, leaving national mobility in middle-income countries underexplored. This study examines how national geographic mobility between cities and municipalities is associated with the career performance of Colombian researchers classified under the national evaluation system of the Ministry of Science, Technology, and Innovation (MinCiencias). Drawing on open administrative data from six assessment calls conducted between 2013 and 2021, the analysis covers 40,485 researcher-call observations. An ordinal logistic regression model estimates the association between researcher rank (Junior, Associate, Senior) and a set of lagged individual, institutional, and contextual covariates, including a binary indicator of inter-city relocation. Because the proportional odds assumption is rejected, an unrestricted specification is adopted to capture heterogeneous effects across rank thresholds. Inter-city mobility is uncommon, affecting 2.7
Women tend to be underrepresented in almost all scientific fields in Germany. This study analyses publications of scientists in a German pharmaceutical non-peer-reviewed journal over a period of 50 years (1972–2021). The journal Pharmakon (previously named Pharmazie in unserer Zeit) is the member journal of the German Pharmaceutical Society (DPhG). We performed a gender analysis of the journal and analysed 1577 articles by 2509 authors. Overall, the percentage of women is 26
The interdisciplinarity of scientific publications is often reflected in their citation patterns, and prior studies suggest that interdisciplinary knowledge integration facilitates novel and disruptive outcomes. However, citation distributions vary across paper sections, reflecting differences in section-based interdisciplinarity, which has long been overlooked. This study elucidates the correlation between interdisciplinarity and innovation at the section level. Using bioinformatics articles, we firstly calculate interdisciplinarity in the Introduction (I), Methods (M), Results (R), and Discussion (D) sections, respectively. Secondly, the two measures of innovation, the content-based ex-ante novelty and citation-based ex-post disruption are calculated by semantic similarity and the D index. Finally, the explainable machine learning model with SHAP is employed to investigate the importance and nonlinear patterns of the section-based interdisciplinarity. The results show that the XGBoost model outperforms baseline models, with classification performance for ex-ante novelty being better than for ex-post disruption. This suggests that section-based interdisciplinarity is more closely related to novelty, whereas disruption, as a lagged indicator, is subject to multiple external factors, such as the calculation mechanism of the D index. Moreover, ex-ante novelty is primarily associated with disparity in the Introduction, but the relationship follows an inverted-U pattern, with moderate disparity corresponding to higher novelty. By contrast, variety is the most important feature associated with ex-post disruption, though the association weakens as variety increases. The balance and RS (Rao-Stirling) in the Introduction and Results sections show a higher likelihood of a paper being classified as disruptive, which means that disruption involves more balanced and comprehensive knowledge integration. Overall, these findings provide a micro-level perspective on how the distribution patterns of section-level interdisciplinarity relate to ex-ante novelty and ex-post disruption, offering valuable implications for scientific evaluation and the early detection of innovative contributions.
Citation function analysis (CFA) is used to explain why scholars cite particular works by identifying the functional roles that citations play in academic writing. Conventional approaches treat CFA as a single-label classification task, but a cited work may serve multiple functions within the citation context. To overcome the limitations of subjectivity, data imbalance, and restricted interpretability, this study introduces CiteFuncRanker, a pairwise ranking framework that reformulates CFA as an ordered comparison problem. The framework models the relative importance among multiple potential citation functions using systematic label pairing and aggregation, enabling large language models to perform structured comparative reasoning under minimal supervision. Experimental results on the ACL-ARC dataset demonstrate the effectiveness of this paradigm, with an average MR of 2.04, MRR of 0.68, and NDCG@3 of 0.66, alongside 86.67
This paper presents a methodology for predicting the economic value of technology using a multimodal-based joint deep learning model trained on patent and firm-level data. Existing data-based technology valuation studies have largely relied on structured patent indicators or single-modality inputs for patent evaluation. In contrast, this study incorporates key factors relevant to technology valuation, including patent structured data, patent unstructured text data, and firm structured financial data, and conducts joint representation learning across heterogeneous data types. The proposed multimodal-based joint deep learning model demonstrates improved prediction performance compared to baseline models, achieving a MAPE of 0.2377. These results suggest that jointly learning technological characteristics and firm-level commercialization capacity contributes to more stable prediction of technology value. Importantly, unlike prior studies that predict proxy measures of patent value, this study directly predicts monetary technology valuation outcomes based on institutional appraisal data. In this sense, the study provides an empirically grounded data-driven approach to technology valuation that reflects both patent characteristics and firm-level factors.
This study examined different facets of scientific mobility, including local, regional and global mobility, as well as international cooperation through an analysis of co-authorship among researchers publishing in the Scopus database. The study focuses on three selected South American countries, Peru, Colombia and Venezuela, for the period between 2000 and 2022. The analysis is based on the use of metadata available in the Scopus database for the period 2000 to 2022, the affiliation of the authors of each identified paper was used to analyze scientific mobility flows and also international collaboration, in order to determine the dynamics of internal and international mobility and its impact on peripheral South American countries. This process was analyzed for Peru, Colombia and Venezuela, the main destinations of each country were geolocated and recognizable patterns were established for each of them in the period studied. The analysis indicates that concepts such as brain drain or brain gain, depending on the specific case and historical moment, are more appropriate than brain circulation (scientific mobility), which is applied to the central countries.
This study introduces a capability-based network framework for analyzing actor heterogeneity in the Korean national R D system. Using project-level data from the National Science and Technology Information System (NTIS)—comprising 63,305 R D projects conducted by 593 innovation actors across 222 science and technology fields in 2021—we construct an R D Actor Space in which universities, firms, and government research institutes (GRIs) are connected by similarity in their research portfolios. The framework complements existing Triple Helix and innovation system perspectives by focusing on capability heterogeneity within actor categories. The resulting network is positively assortative: high-degree actors tend to connect with other highly connected actors, with within-type rich-club patterns among firms and GRIs but a more degree-egalitarian connectivity pattern among universities. Hierarchical clustering reveals a stratification that is type-coherent yet sectorally organized: universities anchor a densely connected central region, firms form multiple capability-coherent industry sub-clusters, and a heterogeneous mixed region interleaves all three types in sectoral groupings that cut across institutional categories. The findings highlight a capability-composition mismatch between universities and firms and show that cross-type structure is organized more by sectoral portfolio similarity than by institutional category alone. We discuss implications for Korean R D policy and outline comparative and longitudinal extensions.
Topic modeling remains a key approach for uncovering latent thematic structures in large text corpora. Classical Bag-of-Words (BoW) models such as LDA and NMF now coexist with embedding-based methods like Top2Vec and BERTopic, prompting renewed assessment of their respective strengths and limitations. This study advances this comparison in two ways. First, building on a previously proposed BoW-based method combining Growing Neural Gas clustering and feature maximization (CFMf), we introduce a reclassification variant (CFMf-R) to refine ill-formed clusters, an embedding-based version (CFMf-e-R) that integrates contextual text representations and a dimensionality-reduced variant (CFMf-e-U-R) that applies UMAP to the contextual embeddings. Second, using a corpus of 16,917 philosophy of science research articles, we develop a pipeline to systematically evaluate eight topic modeling approaches—LDA, NMF, CFMf-R, CFMf-e-R, CFMf-e-U-R, Top2Vec, BERTopic, and HD-BERTopic—across coherence, diversity, and recall metrics, while also examining document distributions and topic interpretability. Results reveal distinct trade-offs: Top2Vec achieves high coherence and diversity but low recall and interpretability; BERTopic slightly outperforms LDA in coherence but not in recall. Document distributions appear more regular with LDA, CFMf-e-R, and CFMf-e-U-R; CFMf-e-R and BERTopic yield higher interpretability. Overall, CFMf-e-R tends to offer a better balance across dimensions, though no single model dominates. Notably, topic content and size vary substantially across methods, revealing distinct corpus perspectives. Reaffirming the continuing relevance of BoW-based models and emphasizing the modularity of topic modeling pipelines, these studies suggest that ensemble approaches combining multiple models may yield more robust corpus representations.
Despite the growing body of research on green technology innovation efficiency, existing studies predominantly focus on single-country analyses or static frameworks, neglecting systematic cross-country heterogeneity and dynamic inter-stage synergy within the macro-innovation ecosystem. To accurately diagnose the structural bottlenecks of global green transition, this study conducts a cross-country comparative research driven by a methodological innovation. Specifically, building on the innovation value chain theory, this study constructs a novel two-stage dynamic stochastic nonparametric envelopment of data (StoNED) model. Unlike traditional deterministic models, this framework uniquely accommodates both statistical noise and nonparametric frontier flexibility, enabling a more robust inter-temporal evaluation of the R D and production stages. Furthermore, a coupling coordination degree (CCD) model is integrated to evaluate the systemic synergy between stages. Incorporating the R D time lag, this study utilizes 2007–2019 panel data from 39 countries to measure the two-stage efficiency for the 2007–2018 period. The results reveal: (1) Significant cross-country heterogeneity exists in efficiency distribution. Countries are classified into four distinct profiles (Dual-high synergy, R D-driven, Double-low, and Production-driven). Notably, massive R D capabilities in major economies often coexist with lagging production efficiency, highlighting the “technology-institutional complex locking” effect as a key bottleneck in the commercialization process. (2) Dynamically, while R D efficiency rises steadily, production efficiency exhibits an “inverted U-shaped” trajectory that highly synchronizes with global energy price fluctuations. (3) The overall systemic coordination reveals a tiered global landscape, demonstrating that sustainable green innovation relies on balanced inter-stage coherence rather than single-stage superiority. By advancing the dynamic StoNED-CCD framework, this study enriches the methodological literature and provides critical insights for designing differentiated, dynamically adjustable green technology policies globally.
Adapted scientific literature (ASL) is increasingly used in educational contexts to make authentic science texts accessible to students. One hallmark of scientific writing is the use of hedging - linguistic means to express uncertainty or tentativeness. This study investigates whether ASL maintains this characteristic feature, and to what extent its usage of hedging aligns with original scientific articles. Using a machine learning approach, we trained classifiers on the BioScope corpus to automatically detect hedging in English and German texts. We then analyzed corpora consisting of original research articles, their adapted counterparts, and textbook excerpts. Our results show that ASL retains the typical distributional pattern of hedging - particularly its increased frequency in the introduction and discussion sections - but the overall proportion of hedging expressions is reduced compared to the original texts. In contrast, textbook texts contain very little hedging and present scientific knowledge as more certain. These findings suggest that ASL offers a linguistically authentic alternative to original scientific literature and may help learners engage with the epistemic practices of science more realistically. While some reduction in hedging occurs during the adaptation process, ASL preserves essential rhetorical features of scientific discourse and can thus serve as a valuable resource in science education.
Bibliometric analyses rely on accurate citation counts, yet bibliographic databases routinely contain variant representations of the same cited reference, differing in journal abbreviation style, author name format, punctuation, or metadata completeness, that fragment citation links and distort standard indicators such as the h-index and journal impact metrics. We propose an unsupervised, multi-phase reference matching algorithm designed to consolidate these variants without requiring training data or external authority files beyond the ISO 4 List of Title Word Abbreviations (LTWA). The pipeline operates in seven phases: (i) format detection and string normalisation, which parses heterogeneous reference styles and standardises author names, titles, and pagination; (ii) ISO 4 journal-name normalisation, which maps both abbreviated and full journal names to a canonical short form using the LTWA; (iii) exact matching on DOI identifiers and normalised reference strings; (iv) blocking by first-author surname and publication year; (v) within-block fuzzy matching that combines Jaro-Winkler similarity with agglomerative hierarchical clustering to group near-duplicate references; (vi) post-processing metadata reconciliation, which merges complementary fields across matched records; and (vii) canonical representative selection, which elects the most informative variant as the group representative. Evaluation on a synthetic benchmark of 1 064 source articles under 17 controlled perturbation scenarios, yields precision, recall, and F_1 scores above 0.95 in 15 of 17 scenarios, with the two lowest-scoring scenarios still achieving F_1 ≥ 0.78 . Validation on two real-world Scopus datasets demonstrates that the algorithm reduces unique cited-reference counts by 4.6-−11.1