
There is a sharp contrast between the high number of patent applications and the low commercialization rate of universities in China.Through an empirical study of patents successfully commercialized by 25 universities,this paper finds that,first,inventor team sizes for commercialized patents have an inverted U-shaped distribution,and four-person teams perform best.Second,university-enterprise collaborative teams of multi-source inventors are not conducive to patent transformation.Third,there is no positive correlation between team members'invention experience and patent commer-cialization.Finally,the multiple knowledge structures of the team are not conducive to patent conversion,but a single cen-tralized knowledge structure is more advantageous.This paper uses the theory of team conflict,sticky information,and en-dowment effect to make attributions to the research results and puts forward corresponding suggestions for university re-search management.
Discovering potential collaborators automatically from massive data is a hot topic in scientific collaboration pre-diction.Considering that research interests change over time and people with social relationships are more likely to collab-orate,a potential scientific collaborator recommendation model"SimTrustRec"is proposed,which integrates dynamic re-search interest and academic social trust.First,the Latent Dirichlet Allocation model is used to learn the topic distribution of published papers,and dynamic research interests of scholars are mined to calculate the similarity of research interests be-tween two scholars.Second,an academic social network is constructed based on the co-occurrence relationship of scholars and units in papers.Direct academic social trust values are calculated and indirect academic social trust values are then cal-culated based on the transitivity of social trust.Finally,the possibility of potential collaboration between two scholars is calculated by combining research interest similarity and academic social trust value,and a list of potential collaborators is generated.Experimental results using ArnetMiner datasets demonstrate that the proposed method achieves better perfor-mance in terms of recall,hit rate,and mean reciprocal rank compared to existing methods.
It has been shown that Altmetrics-based"sleeping beauties"(A-SBs)are of great value.After developing meth-ods to identify A-SBs,we must analyze the features,sleep causes,and awakening laws of A-SBs to realize their early pre-diction.This research considered the articles with the top 1%of attention as its experimental data and ASB index as the identification method.To explore the literature features and awakening laws of A-SBs,this research took the 10 articles with the highest ASB values as representative A-SBs and analyzed their features and attention accumulation processes in detail.The content of A-SBs is of high quality and can be divided into two categories,innovation and classic,and their top-ics frequently appear in daily life.Non-open-access or non-funded articles are more likely to become A-SBs.With complex content that is beyond the knowledge level of the general public,an article has more possibilities to fall asleep.The reasons for A-SBs being awakened include(1)improvement of the public's knowledge,(2)encountering similar problems,(3)re-lated public events,and(4)influential account recommendations.For the early prediction of A-SBs,it is helpful to pay more attention to innovative or classic articles,articles that gained attention on Patent or Wikipedia,non-open access arti-cles,or non-funded articles.Scholars who use social media play an important role in realizing the early prediction of A-SBs.
This study explored the relationship between the research dynamics of top scholars and the development of a discipline,to understand their influence,explore the regularity in iteration of scientific knowledge,and obtain insights for optimizing talent evaluation.We developed a quantitative model that focused on top scholars in a discipline,utilizing three dimensions(time,topic,and influence relationships)for the basic indicators.We combined this with the development life cycle of topic popularity,defined the types of influence that scholars have on the development of the discipline,construct-ed scholars'impact functions,and analyzed the degree of impact of top scholars'research on the development of the disci-pline.Based on empirical research in the field of gene editing,the most significant impact of top scholars on the develop-ment of the discipline was seen in the promotion of the third-generation gene editing technology CRISPR(clustered regu-larly interspaced short palindromic repeats)to become the most mainstream research topic,thereby achieving technologi-cal innovation.Compared to all scholars,top scholars played a prominent role in pioneering innovation and supporting oth-er scholars and maintained advantages as the field continued to develop.We also analyzed the research dynamics of two typical scholars'impact on the development of their discipline and identified two distinct impact patterns:the promoting pattern and the pattern that is a combination of pioneering,guiding,and supporting.The computed results are consistent with reality,which verifies the reliability of the constructed model.
As a powerful tool for judging scientific and technological innovation capabilities and identifying market trans-formation trends,patent analysis is an important basis for a new round of technological revolution and industrial transfor-mation in China.The establishment of a reasonable and efficient patent retrieval strategy is a prerequisite for patent analy-sis.This study entailed the development of a set of retrieval strategies based on deep learning algorithms,which addresses the shortcomings of insufficient dynamics and intelligence in existing research.The proposed model consists of two main parts:construction of retrieval strategy and revision of retrieval results.As regards the construction of retrieval strategies,this study aimed to systematically analyze the principle of technology composition,integrate the deep learning algorithm,and train the model from two dimensions of scientific and technological corpus and domain corpus so as to preliminarily screen and expand the search terms.
Identifying technology opportunities in specific fields is an important part of the R&D organizations'innova-tion management.This study proposes a new method to identify technical opportunities in specific fields.Based on the combination attributes of knowledge elements(i.e.,combination breadth,combination intensity,and combination dis-tance),transitivity,and homogeneity,exponential random graph models are constructed.The probability of combination be-tween knowledge elements is calculated based on the model parameter,and subsequently technical opportunities in the field are discovered.To validate the effectiveness of the proposed method,patent data in the field of the Internet of Things from 2016 to 2021 is used.To check the robustness of the proposed method,patent data in the field of industrial robots is used.
The evolution of scientific research institution names caused by the development and changes of institutions has seriously affected the quality and effect of knowledge services based on institution names,such as in information retrieval and research evaluation.This study developed a method for institution name normalization based on the author and re-search theme so as to eliminate heterogeneity between the names of scientific research institutions and optimize informa-tion retrieval and knowledge discovery services based on institution names.By conducting a performance analysis of the name evolution of scientific research institutions in the signature of academic papers,this study constructed a recognition model of the name evolution relationship of scientific research institutions based on the author and research theme and identified the renaming,splitting,merging,and restructuring relationships between the names of scientific research institu-tions.The model was then verified using small-scale academic paper data.The experimental results indicate that the pro-posed method achieves a better accuracy and recall rate in identifying the name evolution relationship of primary and sec-ondary institutions and can also identify name evolution relationships between unpopular institutions.
In addition to online customer reviews,management responses have become an essential dual-channel commu-nication for customers to obtain related information online.Customer reviews and management responses can influence the engagement behaviors of prospective customers;however,this internal mechanism remains unclear.Therefore,this study aimed to combine each customer review with its corresponding management response to explore whether impacts subsequent customer engagement behaviors.Based on cognitive consistency theory,this study devised a research model for revealing the effect of both the topic and length consistency of each combination of reviews and responses on customer engagement,including review and response likes.An empirical study was conducted utilizing data collected from Huawei Mall with 80,329 combinations of reviews and responses.The topic and length consistency were measured using the Bidi-rectional Encoder Representations from Transformers(BERT)technique and cosine similarity,and text mining,respective-ly.A zero-inflated negative binomial model was applied for regression analysis.The results indicate that both consistencies of the customer reviews and management responses positively impact customer engagement behaviors.Moreover,review sentiment positively moderates the impact of the length consistency on response likes.The results of this study contribute to identifying matches of reviews and responses to predict future customer engagement behaviors.This assists practitioners in formulating management response strategies using the matching effect between reviews and responses to attract poten-tial customers.
The increasing scale of social media data and data heterogeneity pose new challenges.In contrast,complex structural features in rumor propagation networks are difficult to explore;however,the interpretability of a deep neural net-work-based rumor detection model must be further investigated.In this study,we design and implement an interpretable graph neural network model for a rumor detection task.Specifically,we train a graph neural network interpreter based on mask learning using a residual graph neural network model.This framework incorporates the structural features of rumor propagation and provides an automatic interpretation of the graph neural network model,considering both network struc-ture and node features.The experiments are conducted on two online rumor datasets sourced from Sina Weibo(Chinese)and Twitter(English).Interpretive analyses are performed at both the global and case levels.The experimental results show that the proposed graph neural network model effectively exploits the communication structure features and outperforms a series of baseline models in the rumor detection task.Using the trained graph neural network interpreter,we discovered that long propagation chains are the key network topology for rumors in larger-scale rumor propagation trees and text fea-tures are the key node attributes in smaller-scale rumor propagation trees.
Facing increasing information overload and the rise of cross-cutting research,studies on the filtering ability of the information retrieval system to provide effective search term recommendation services are becoming increasingly im-portant.This study proposes a search term recommendation method that integrates dependent syntax and language network theories using the PageRank algorithm.By constructing a search term set and dependent syntax network,and sorting the search terms using the PageRank algorithm,the search term recommendation is realized.The method is validated using the Web of Science platform 124,516 literature abstracts in the field of information science&library science as an example.We also invited ten MLIS graduate students to participate in the user study,which combined the comparison results with similar methods and systems.The results show that the accuracy of the recommended method is 80%,average Cosine simi-larity in the recommendation list is 0.53,and average Jaccard similarity in the table is 0.39.Compared with other methods and systems, the diversity of our approach reacts better with a higher degree of surprise. Overall, the results show that our method increased the coverage of the search terms based on the user's information requirements. Our method is expected to provide references on methodological perspectives on the representation of information retrieval. It can be directly applied for terminological organization in the back end of information retrieval, as well as indirectly for knowledge discovery and inter-disciplinary study.
The evolution of research topics is critical for clarifying scientific development and predicting frontiers.The cal-culation of the similarity of topics in adjacent time periods and identification of their evolutionary paths is the core step in the topic evolution analysis.In this study,the matrix similarity algorithm is innovatively proposed to identify the topic evo-lution path.Based on the local network structure of the research topic in the co-word network,the similarity of the research topic in terms of words and relations is considered in the calculation of topic similarity.Subsequently,an analytical frame-work for research topic evolution based on matrix similarity is constructed.Considering piecewise linear representation,the framework divides the data into time periods to build the temporal co-word networks.After identifying the topic com-munities in the co-word network at each time period using the community discovery algorithm,multi-dimensional feature indexes such as novelty,popularity,core,and maturity are calculated to represent the types of research topics.The evolu-tionary paths of the research topics are then identified using matrix similarity calculations.Finally,the evolution process of the research topics is visualized by a Sankey diagram and a multi-dimensional strategic coordinate plot.Specifically,this study uses the field of library and information science as an example of empirical analysis.The results show that the pro-posed method can effectively support the evolution analysis of research topics in a research area and provide methodologi-cal support for research decision-making.
公共数据是指国家机关、事业单位、经依法授权具有公共事务管理与公共服务职能的组织在履行职责过程中收集和产生的各类数据,对其进行收集、处理、共享和开放等多途径治理是发挥数据要素作用的重要保障.本文通过分析我国公共数据治理政策发展变化的特点、内在规律和基本趋势等为公共数据治理政策的完善和创新提供参考依据.在全面搜集我国公共数据治理政策文本基础上,运用统计分析和文本挖掘等方法对政策文本形式与内容特征进行抽取和分析,重点对数据治理政策的高频词组与主题内容变化、关键线索词分布等进行了分析.研究结果发现,数据治理政策客体对象在发展演变中不断变化,其覆盖的主题内容范围明显扩大;在数据治理政策中开始关注有关主体的权利或权益保护;数据治理政策更加关注数据开发利用和数据对经济社会发展的作用.
丰富的互联网文献数据库是科研人员了解领域发展和前沿的重要资源,从全局视角对领域的海量科研成果进行高效信息挖掘,可以在知识洪流中为科研人员提供更加明确的方向.本研究基于经典生物医学文献数据库PubMed收录的发表于2010-2021年的13万篇文章,挖掘科研人员的历史行为信息,构建同时包含作者、论文、关键词的异质信息网络,利用异质信息网络表示学习算法metapath2vec将该网络嵌入成为异质向量空间,并通过计算异质向量空间中向量的相似度指标,同时实现科研合作者推荐与科研兴趣关键词推荐.与已有研究相比,本研究的方法更加重视多任务协同,不仅在新增的科研兴趣关键词的任务中获得了有意义的推荐结果,还显著提高了科研合作者推荐的准确度.同时,本研究在作者空间与关键词空间进行了深入挖掘,并证明其在科研兴趣的语义理解方面具有指导意义.本研究在科研兴趣的研究、挖掘与推荐方面提供了新的研究视角.
专利技术互补性作为各类组织进行技术创新的重要参考,近年来受到国内外学者的广泛关注.本文回顾了技术互补性概念的发展沿革,从产业/行业分类、专利分类、专利引用关系以及专利内容特征关联四个角度归纳其测度方法,最后综述专利技术互补性的多种应用.基于此,总结形成专利技术互补性的概念内涵,发现相关研究主要利用专利分类或专利引用网络来形成技术互补测度指标和方法,并主要应用于创新绩效因素判定、企业并购决策制定以及潜在合作伙伴发现等.未来,建议继续细化和具体化技术互补性概念,综合利用专利文本、图表、市场信息等多模异构数据,设计细粒度定量测度指标,引入深度学习等方法,提升专利技术互补测度的准确性,进一步拓宽专利技术互补性的应用范围,提升应用效果.
在中国式现代化建设要求的背景下,情报事业进入了新的发展阶段.情报赋能理念的提出,不仅为情报学术探索注入了新的动力,也为信息行为研究开辟了新的空间.本文运用情报赋能理念,审视信息行为研究中积累的成果,分析通过信息行为研究实现情报赋能的可行性,探讨利用信息行为研究的成果增强情报工作中的线索发现、分析研究和决策保障的效能,进而提升情报用户的决策能力以达到赋能目标;同时,在情报赋能的关切下,拓展信息行为研究成果的应用范围,为信息行为研究的未来探索新的途径.
在移动互联网时代,移动阅读、碎片化阅读已经成为人们阅读的主流方式.在用户阅读过程中,提供摘要内容以提高阅读效率是解决信息过载问题的重要途径之一.科技研究论文文本长、内容广且包含领域知识,其摘要生成任务相比于新闻等普通文本更具有挑战性.本文提出了一种科技论文结构化摘要方法.首先,将科技论文划分为不同的语步;其次,分别对不同语步文本进行抽取式摘要,将文本多特征按权重融入TextRank算法的迭代计算过程中,引入MMR(maximal marginal relevance)算法对预选摘要集进行冗余处理;最后,使用依存句法分析对文本进行语义分析,进一步精简摘要,并组合成结构化摘要.研究结果表明,相比于基准模型,该方法在不同语步的相关性、多样性和可读性指标提升上具有一定差异;结合人工评价发现,该方法在显著提升摘要多样性的同时,一定程度上提升了摘要的相关性和可读性.
本质上,Altmetrics与被引频次均是对于学术成果影响力的计量,那么在引文中出现的"睡美人"现象,在Altmetrics中也同样存在.睡美人文献具有重要科研价值,实现基于Altmetrics的睡美人文献的早期识别,可以提升文献利用率,提升公众智慧,反映公众对科学的关注.设计识别方法用于识别现有的睡美人文献,是实现睡美人文献早期识别的第一步.基于Altmetrics的睡美人文献最重要的两个特点是较长的睡眠时间和关注的突增,以这两个特点为核心,本文参考四分位数和Bcp指数的思路,设计了一种基于Altmetrics的睡美人文献识别方法——ASB指数(alt-metrics sleeping beauty index).以关注度排名前1%的文章作为高关注度文献,形成实验集对ASB指数的识别效果进行检验,比较ASB值最高、中位、末位各10篇文章的关注累积曲线和指标特征,观察ASB指数的识别效果.检验结果表明,ASB指数对基于Altmetrics的睡美人文献识别效果良好.
知识单元的特征信息是其学科归属判定的基础,挖掘关键特征有助于提升学科判定方法的性能,从而更好地服务于知识内容层面的跨学科规律研究.本文借助16种知识单元学科归属判定方法,通过对比分析这些方法,判别不同词频、不同学科覆盖度词汇的学科归属,评价方法所蕴含的学科重要度、学科相关度和学科区分度3种特征和特征组合效果,以挖掘效果最好的特征子集.本文以"计算医学"这一交叉领域数据为基础构建测试数据集,研究分析表明,综合使用3种特征的方法在各组数据上均取得了较好的性能,同时学科重要度的性能优势表明其在3种特征中最为重要;高频词的学科归属判定需要注重学科区分度,而低频词需要重点考虑学科重要性;对多学科覆盖度的知识单元,需要在学科重要度基础上补充对学科区分度的考虑.本文的发现能够为知识单元学科归属判定方法优化提供理论指导和实践建议.
团队创新是"大科技时代"的显著特征.过去关于科学团队创新过程的研究较多,但对技术团队创新的关注不足,关于技术团队知识多样性及其对创新绩效的影响仍不清晰.本文考量技术团队的特征,提出一种新方法用于识别团队,进而构建团队知识网络、测度知识多样性,从创新数量、创新质量、创新广度3个维度探究其对团队创新绩效的影响.研究结果发现,知识多样性可以在上述3个维度显著提升技术团队的创新绩效.当以创新数量为目标时,大团队更能发挥知识多样性的正向效应;但当侧重创新成果的新颖性和创造性时,小团队中知识多样性对创新绩效的提升更明显.本文的研究结论,一方面,可以与科学团队相关研究形成对话,帮助理解科学团队与技术团队在创新过程中的不同特征与机制;另一方面,可以辅助技术研发主体与管理者参照创新目标合理配置资源,以实现理想的创新绩效.
在新一轮科技革命背景下,技术机会发现相关研究已受到国内外学界的广泛关注.技术机会发现旨在发现新的技术动向,推测该领域可能出现的技术形态或技术发展点,对于技术创新、产业发展具有重要意义.本文系统梳理了专利挖掘方法在技术机会发现中的应用研究现状,总结了5种底层共性分析方法的代表性研究,厘清分析方法与研究内容的适配关系,为该领域开展后续相关研究和实践的技术选型提供参考依据.研究结果表明,技术机会发现所采用的方法手段虽然一直跟进深度学习的发展,但是尚未针对性地进行方法应用上的创新.最后,本文从数据、方法应用和评价体系3个角度提出了改进思路.