In the future of intelligent education, enhancing the cognitive ability of students is a promising avenue worth investigating. Although most educational services address students' explicit knowledge needs, exploring beyond cognitive boundaries to upgrade cognition are rare. In this study, we employ a hybrid approach combining quantitative questionnaire and qualitative epistemic network analysis to investigate the 'cognitive upgrading' experiences of 235 university students. Findings of the inquiry indicate: (1) 'Cognitive upgrading' represents a prevalent and objectively discernible necessity for students, indicating the significance for facilitating its promotion; (2)Current information resources cannot directionally guide students towards cognitive upgrading, necessitating the intervention of a third-party to render sporadic events manageable; (3)The existing paradigms of cognitive upgrading can be categorized into eight quadrants. There still lacks a data-driven approach to be fully explored for more effectively facilitating the happening occurrence; (4)Students with disparity in genders, educational backgrounds, and disciplines exhibit significant differences in cognitive upgrading, suggesting that training programs tailored to students' characteristics can provide targeted support for enhancing cognitive upgrading efficacy. The study provides valuable insights for educators to implement cognitive upgrading strategies and paves a new path towards cognitive upgrading in intelligent education.
Papers, patents, and clinical trials are essential scientific resources in biomedicine, crucial for knowledge sharing and dissemination. However, these documents are often stored in disparate databases with varying management standards and data formats, making it challenging to form systematic and fine-grained connections among them. To address this issue, we construct PKG 2.0, a comprehensive knowledge graph dataset encompassing over 36 million papers, 1.3 million patents, and 0.48 million clinical trials in the biomedical field. PKG 2.0 integrates these dispersed resources through 482 million biomedical entity linkages, 19 million citation linkages, and 7 million project linkages. The construction of PKG 2.0 wove together fine-grained biomedical entity extraction, high-performance author name disambiguation, multi-source citation integration, and high-quality project data from the NIH Exporter. Data validation demonstrates that PKG 2.0 excels in key tasks such as author disambiguation and biomedical entity recognition. This dataset provides valuable resources for biomedical researchers, bibliometric scholars, and those engaged in literature mining.
In today’s pluralistic society, cross-disciplinary career transitions are critical for adapting to ever-changing job demands. Important influencing factors have been identified to explain these transition actions. However, studies on the prevalence of career transitions in different disciplinary areas are rare, especially those investigating the patterns of disciplinary changes accompanied by career transitions. This study applies text mining techniques, presenting data to demonstrate and compare how career transition exists widely across various disciplines. Taking library science, linguistics, electronic engineering, and radiology as representative areas and extracting resume data from Indeed, the empirical study yields noteworthy findings: Overall, the prevalence of career transition exists, with patterns varying across disciplines, illuminating the necessity for developing more inclusive curriculums or career planning frameworks to accommodate diverse needs of the job market. From the dimension of major discipline, the permeation pattern of disciplines varies. Humanities and natural sciences typically extend to other target domains, social sciences exhibit significant permeability, while engineering sciences generally revert to the initial domain. From the sub-discipline dimension, disciplines closer to the initial discipline do not necessarily have a higher employment rate. Rankings of disciplinary interaction strength and employment probabilities have a notable rank-turbulence divergence spanning from 0.4 to 0.7; from the occupation dimension, occupations spanning across different disciplines are abundant in the job market; from the skill dimension, the representativeness of occupational skills varies both within and outside the initial disciplines. Discipline-specific skills in electronic engineering and linguistics exhibit high representativeness both within and outside the initial discipline, while soft skills rank high outside the initial discipline in library science and radiology. These findings indicate how individuals develop professionally throughout transitions between various disciplines, providing strategic guidance for policymakers to synchronize career transition policies with specific patterns of discipline change, thereby promoting interdisciplinary talent cultivation objectives.
Scientific novelty is the essential driving force for research breakthroughs and innovation. However, little is known about how early-career scientists pursue novel research paths, and the gender disparities in this process. To address this research gap, this study investigates a comprehensive dataset of 277,288 doctoral theses in the biomedical sciences authored by US Ph. D. graduates. Spanning from 1980 to 2016, the data originates from the ProQuest Dissertations & Theses Database. This study aims to shed light on Ph.D. students' pursuit of scientific novelty in their doctoral theses and assess gender-related differences in this process. Using a combinatorial approach and a pre-trained Bio-BERT model, we quantify the scientific novelty of doctoral theses based on bio-entities. Applying fractional logistic and quantile regression models, this study reveals a decreasing trend in scientific novelty over time and heterogeneous gender disparities in doctoral theses. Specifically, female students consistently exhibit lower scientific novelty levels than their male peers. Under the supervision of female advisors, students tend to produce doctoral theses that exhibit lower levels of novelty compared to those supervised by male advisors. The significant interaction effect of female students and female advisors suggests that female advisors may amplify gender disparities in scientific novelty. Moreover, heterogeneous gender disparities in scientific novelty are identified, with non-top-tier universities displaying more pronounced disparities, while the gender differences at higher percentile ranges of scientific novelty scores were comparatively more minor. These findings indicate a potential underrepresentation of earlycareer female scientists pursuing novel research. Notably, the outcomes of this study hold significant policy implications for advancing the careers of female scientists.
Purpose Proper topic selection is an essential prerequisite for the success of research. To study this, this article proposes an important concerned factor of topic selection-topic popularity, to examine the relationship between topic selection and team performance. Design/methodology/approach The authors adopt extracted entities on the type of gene/protein, which are used as proxies as topics, to keep track of the development of topic popularity. The decision tree model is used to classify the ascending phase and descending phase of entity popularity based on the temporal trend of entity occurrence frequency. Through comparing various dimensions of team performance – academic performance, research funding, relationship between performance and funding and corresponding author's influence at different phases of topic popularity – the relationship between the selected phase of topic popularity and academic performance of research teams can be explored. Findings First, topic popularity can impact team performance in the academic productivity and their research work's academic influence. Second, topic popularity can affect the quantity and amount of research funding received by teams. Third, topic popularity can impact the promotion effect of funding on team performance. Fourth, topic popularity can impact the influence of the corresponding author on team performance. Originality/value This is a new attempt to conduct team-oriented analysis on the relationship between topic selection and academic performance. Through understanding relationships amongst topic popularity, team performance and research funding, the study would be valuable for researchers and policy makers to conduct reasonable decision making on topic selection.
Measuring diversity in research works is critical to the promotion of scientific evaluation. Existing approaches usually focus on quantifying the diverse interdisciplinary states of research fields, lacking fine-grained detection and comparative analysis between different diversity calculation schemas at the level of individual scholars. To clearly recognize essential characters of research content and reflect important research diversity characteristics of scholars, we propose to use entitymetrics analysis for the scholar-level diversity measurement. Compared to disciplines or topics, minimal stable knowledge units embedded in scientific papers have advantages in mining knowledge usage and capturing the microscopic differences among contents. In this study, we compare six diversity calculation schemas based on the three primary properties of entity diversity: variety, balance, and disparity. The comparison demonstrates the usefulness of entities in portraying manifold subject categories, detecting content disparity, and characterizing content distribution patterns of research, which reflects research diversity of scholars in a more granular way. It also points out the respective features, scopes, and application scenarios of each schema, which provides further guidance for selections and appropriate usage of schemas, ultimately fostering the accurate scientific assessment of scholar characteristics (e.g., the cross-disciplinary collaboration potentiality, academic aptitude, multi-subject problem-solving skills).
[目的]面对网络舆情事件,从情感分歧角度出发,为舆情分析提供新的分析角度.[方法]引入情感分歧度概念,创建多层次情感分歧度算法,构建网络舆情事件多层次情感分歧度分析模型,对网络舆情事件层、评论对象层、用户层进行情感值及情感分歧度的计算,并将三个层次进行关联分析.[结果]实验结果表明,引入情感分歧度可以弥补原有情感分析研究中对网民意见分歧角度的缺失,本模型可以实现舆情事件关键节点及争议较大的评论对象的识别、判断舆论引导效果,并对舆情争议产生的原因实现精准定位.[局限]仅选取微博作为数据源,未从豆瓣、知乎等其他平台获取数据.[结论]本文模型可应用于监控舆情事件关键节点、根据争议原因选择不同的舆论引导方式以及判断舆论引导效果.
[目的/意义]将认知升级理论融入图书馆智慧推荐服务中,以实现知以藏往、见贤思齐的智慧化推荐服务.[方法/过程]首先,从兴趣热度、内容质量评价和专指度3个指标入手,构建图书馆智慧推荐系统的指标体系:其次,基于认知升级理论,将用户分为"前辈"和"后辈",通过改进协同过滤推荐算法计算用户相似度,将"前辈"的成功学习路径推荐给相似的后辈:最后,利用精准率、召回率、AUC、MMR、F1值等指标对离线实验和在线实验结果进行检验.[结果/结论]实验结果表明,改进后的智慧推荐算法相比传统协同过滤算法的实现效果有明显提高;对比离线实验和在线实验结果发现,在线实验的推荐效果显著提升,意味着若将基于认知升级理论的智慧推荐服务加以推广,将会对高校学生的专业素质培养和认知层次升级产生积极影响.
For the accurate scientific evaluation and advancement of science, it is crucial to understand the development state of a given disciplinary domain. Existing comprehension methodologies concentrate on quantitatively analyzing broad subject trends without considering the underlying complex status attributes of the words that support and enrich these surface-level trends. Through the perspective of the word role, this study deepens the examination of domain development in a more granular way. We use Word2vec to identify the representative semantic neighbors of a domain-specific feature word from literature. Then, a word status observation model is provided that classifies the role of these representative words as tree structures, including young leaves, dead leaves, roots, and trunk-branches, based on changes in their similarity ranking divergence and information entropy. The static word role provides a new insight into detecting the internal detailed organizational composition and maturing status, while the dynamic role shift helps track historical development changes in the domain. Taking attention in psychology as a case by extracting psychology articles from Microsoft Academic Graph, the empirical study illustrates the results of the observation model and yields intriguing findings. First, the static word role positioning shows that a word's status can be obviously different. Different status indicates areas currently keeping the states of thriving, obsolete, becoming firm supportive pillars (precipitating), and irregularly developing within the domain, respectively. Second, during different development periods, most of the words are maintaining the role of irregularly changing or tending towards precipitation. A certain proportion of words are tending towards prosperity, while only a small set of words are deviating from the prosperity state. The growth of attention-related studies as a whole is well-grown and gradually approaching maturity. Our study further supports researchers’ understanding of the domain development status from a more granular perspective of word roles.
[目的/意义]人工智能技术的更迭应用驱动着数据集在科学计量研究领域发挥着日渐重要的作用.从传统面向信息供应的数据资料集合到如今辅助知识发现和关系网络构建的知识资源,数据集各功能取得了快速发展,进而为拓展科学计量研究的深度和广度提供支持.[研究设计/方法]以Scientometrics期刊2016-2020年收录的论文为数据源,分析数据集的整体使用情况,探究数据集的使用热度与文献数量之间的关系,针对典型数据集进行特征分析,并探讨人工智能技术对于数据集工作的影响,展望数据集的未来建设方向.[结论/发现]数据集的被使用频次与其收录论文数量之间存在一定正相关关系,同一科学计量研究倾向于同时使用多种数据集,且基于科研文献数据集的科学计量研究与人工智能技术之间的关系日益紧密.[创新/价值]旨在通过分析科学计量相关论文所使用数据集的特征,总结归纳近年来数据集的建设发展规律,并为开展科学计量研究选用数据集提供参考.
Comparing the aging of scholars in different regions is critical to have an integrative understanding of its causes and effects on academic performance. Using descriptive statistics and comparative analysis methods, we categorize aging trends in four types and find correlations between aging and academic performance by regression analysis. Findings show that: (1) Aging phenomenon is widespread in different regions, but their aging trends are obviously different at the regional level. Aging types include: Tending towards youth, tending towards maturity, maintaining maturity, and tending towards senility; (2) the type of aging largely depends on the variation in the proportion of scholars in different age groups; (3) aging can be further categorized into positive and negative ones based on the academic performance across of a region. The research is of great significance for understanding aging mechanism, solving problems it brings and informing decision-making for policy-makers.
[目的]面向网络招聘广告提出一个完整、系统的岗位人才需求分析的框架,并基于框架对我国互联网行业人才需求进行分析.[方法]采集互联网行业招聘广告,构建LDA模型以实现岗位需求的主题挖掘与分类,利用Word2Vec模型与依存句法分析得到主题词-程度词词表并构建主题本体.[结果]实证分析发现互联网行业岗位主要分布于我国的东南沿海与一线城市,计算机技术和个人素质能力是互联网行业最为看重的两项主题能力,不同类别的岗位对人才的能力需求差异较大;并基于框架构建了对不同岗位需求的量化评价.[局限]校园招聘的数据样本较少,导致分析结果与实际情况存在偏差;构建LDA模型时分词不够完善,某些主题代表性不强.[结论]实证分析表明岗位人才需求分析框架对人才市场需求和岗位能力要求的分析是有效的,并依据分析结果提出了制定职业规划、提高培养计划灵活性等建议.
The journal Scientometrics published a paper presenting an entitymetric analysis on COVID-19 (the COVID-19 paper, Yu et al., 2021), and the key part of the paper is an entity-entity co-occurrence network based on the bio-entities extracted from titles and abstracts of COVID-19 publications.According to the original paper proposing entitymetrics (the entitymetrics paper, Ding et al., 2013), there are two aspects of entitymetrics: (1) the co-occurrence of entities as mentioned the COVID-19 paper, and (2) citation relationship between entities.The COVID-19 paper extracted the entity co-occurrence relations instead of the citation relations.However, we believe there are some issues that need to be discussed about the citation relationship aspect of entitymetrics.The entitymetrics paper (Ding et al., 2013) assumes that if one paper cites another paper, then an entity in the citing paper will be considered to cite an entity in the cited paper, and entity citation network is built based on the assumption (see Fig. 1).The first issue about this assumption is that where should the cited entities be extracted.Both the COVID-19 paper (Yu et al., 2021) and the entitymetrics paper (Ding et al., 2013) extracted the entities from the abstracts and titles of the PubMed Central (PMC) papers.However, instead of the titles and abstracts, we might consider extracting the entities from the full text, specifically the citation context.An abstract is a summary of a research article, and the title is the name for the work.Both are provided by the authors, whereas the citation contexts are the description of the cited paper from the citing authors' point of view.Even though from the authors' perspective, it seems that the most important entities should be in the titles and abstracts, however, the citing authors' perception could be different with the cited authors.The entities extracted from the titles and abstracts might not be in the citation contexts, and vice versa.One paper contains a lot of concepts, entities, and ideas, but only a few of them are used or cited by the citing authors.What entities are used is directly revealed or
With the increasing pressure on the National Institutes of Health (NIH) budget nowadays, it is such a major challenge to cut waste and improve efficiency in the research funding allocation. To meet this challenge, this paper explores research hotspots and disciplinary trends of the biomedical area, and discusses the relationship between these factors and the government funding, thereby uncovering biomedical hotspots of interest to academia and the evolution law of the U.S. federal government funding through an entitymetrics analysis. Considering that the rapid proliferation of biomedical literature provides large amounts of information resources for knowledge discovery, entities extracted from articles in PubMed and NIH-funded projects during 1988–2017 are taken as experimental data. They are divided into four categories: species, diseases, genes, and drugs. Subsequently, a comparative analysis of entity trajectories in the four domains is performed, which includes occurrence frequency calculations of disease entities to explore frequency variation trends in high-frequency entities and the situation of the distribution of research funds. Finally, we conduct an evolutionary analysis of two sides, respectively: the relationship between research popularity and the amount of funding; the relationship between research popularity and the number of funded projects. The results suggest that research on gene and disease entities is at the stage of rapid development. Diseases with high prevalence rate and mortality and diseases associated with genetic factors will be the emphasis of research trends in the future. The distribution of NIH grant appears obvious long tail effect and can influence overall trends in the heat of research topics.. We also find that there is a strong linear correlation between the research popularity of bio-entities, and the amount and number of funding grants, respectively. However, the impact of the amount and number of grant funds on the entity research popularity is decreasing. The above results indicate the extensive applicability of entitymetrics in funding research.
Scientific novelty drives the efforts to invent new vaccines and solutions during the pandemic. First-time collaboration and international collaboration are two pivotal channels to expand teams' search activities for a broader scope of resources required to address the global challenge, which might facilitate the generation of novel ideas. Our analysis of 98,981 coronavirus papers suggests that scientific novelty measured by the BioBERT model that is pre-trained on 29 million PubMed articles, and first-time collaboration increased after the outbreak of COVID-19, and international collaboration witnessed a sudden decrease. During COVID-19, papers with more first-time collaboration were found to be more novel and international collaboration did not hamper novelty as it had done in the normal periods. The findings suggest the necessity of reaching out for distant resources and the importance of maintaining a collaborative scientific community beyond nationalism during a pandemic.
COVID-19 cases have surpassed the 109 + million markers, with deaths tallying up to 2.4 million. Tens of thousands of papers regarding COVID-19 have been published along with countless bibliometric analyses done on COVID-19 literature. Despite this, none of the analyses have focused on domain entities occurring in scientific publications. However, analysis of these bio-entities and the relations among them, a strategy called entity metrics, could offer more insights into knowledge usage and diffusion in specific cases. Thus, this paper presents an entitymetric analysis on COVID-19 literature. We construct an entity-entity co-occurrence network and employ network indicators to analyze the extracted entities. We find that ACE-2 and C-reactive protein are two very important genes and that lopinavir and ritonavir are two very important chemicals, regardless of the results from either ranking.
[目的/意义]人工智能、大数据等领域的快速发展使得商业发展与隐私保护之间的矛盾愈发尖锐.通过对不同类型的网络隐私争议事件微博评论进行情感及话题对比分析,以探究不同情境下网络用户的隐私态度的异同点与背后机理.[方法/过程]采集2012年至2019年网络隐私争议事件的相关微博评论,对其进行预处理,作为实验数据;基于情感词典计算各评论的情感强度值,并将隐私争议事件分为隐私收集类、隐私曝光类及隐私协议类,对比分析不同情境下的用户评论情感趋势;构建用户隐私讨论对象-情感表达二分网络,并通过二分网络投影构建单顶点网络,结合节点中心性等指标进行二分网络及投影分析.[结果/结论]结果表明,用户整体隐私关注呈现上升趋势;不同类型隐私争议事件的用户负面情感强度水平不同;不同隐私争议情境下用户的关注热点差异较大,情感表达各有特点.以上结果表明不同情境中的用户隐私关注及情感表现具有明显差异.