Synthesis programs, designed to address critical scientific challenges through interdisciplinary solutions, have garnered substantial attention from funding agencies. This study quantitatively evaluates the detailed impact of synthesis programs on promoting these interdisciplinary solutions, using data drawn from the Major Research Plan (MRP) instituted by the National Natural Science Foundation of China (NSFC). We measure and compare the interdisciplinarity in knowledge absorption and integration of articles funded by the MRP with that of articles published in the same journal and year without synthesis intervention, as well as those supported by the NSFC’s General Program. Key dimensions of interdisciplinarity, encompassing variety, balance, and disparity, along with their aggregation, are measured using article references and, more importantly, the main content of these articles. Findings indicate that synthesis programs have fostered interdisciplinary research, but their effects on bolstering knowledge absorption and integration differ. These initiatives motivate researchers to absorb more disparate knowledge from a wider range of disciplines and pave the way for a more balanced integration of dissimilar knowledge. Our findings support the implementation of synthesis programs as a catalyst to accelerate the integration of knowledge across disciplines and domains, and offer insights for funding agencies, expert groups, and researchers engaged in the design, execution, and assessment of such programs.
Introduction Drug synergism may occur when two or more drugs are used in combination. Synergistic drug pairs can enhance efficacy and reduce drug dosage and side effects. Therefore, employing computational methodologies to identify specific synergistic drug combinations for clinical application is of significant importance.Methods We proposed a multiple kernel-based non-negative matrix factorization, MK-NMF, specifically for mining specific synergistic drug pairs in cell lines. In this method, we treated the features of drug pair space and cell line space in the form of two kernel matrices. We incorporated feature kernel matrices into the matrix factorization process.Results MK-NMF achieved an area under the curve (AUC) of 0.884 and an area under the precision versus recall curve (AUPR) of 0.537 on the NCI ALMANAC dataset. Both measures were more than a 5% improvement over the previous matrix factorization model. MK-NMF had good robustness with the missing input data. Its performance was stable when the amount of matrix data input was at least 40%. Literature and experimental verification confirmed some of our predictions.Discussion The increase in data volume and the introduction of more high-quality features will further enhance the performance of MK-NMF. Single-drug response data will help address the challenge of predicting synergistic combinations of new drugs.Conclusion MK-NMF could assist medical professionals in rapidly screening synergistic drug combinations against specific cancer cell lines. The source code of MK-NMF is freely available at https://github.com/XDRFDH/MK-NMF.
Topic evolution is essential for exploring a field; however, the journal’s contribution has not been explored in topic evolution research. In this work, we interpret a journal’s contribution as a journal preference and investigate the concept based on topic focus, as shown in the journal-topic distribution. To analyse the topic focus, we first processed the data into documents consisting of only fine-grained topic words. Document vectors were generated using Sci-BERT and clustered using the k-means algorithm after dimensionality reduction. By matching journals with topic clusters, we calculated the journal preference score based on topic focus and then added a time factor to represent the evolution of journal preference. Simultaneously, we used the Zipfian distribution to classify fine-grained topic words into core and rare topic words, which were then used to establish topic relations in the evolutionary analysis and calculate the novelty scores of journal topic words. We use the technology innovation management (TIM) field to conduct a case study. There were 8 typical and 16 derivative topics, totalling 24 different topics. We focused on four important topics: R D activity, technology management, innovation activity, and climate change, and found that they all have a relatively innovative evolution in a given year. The study indicates that within a given topic, while the composition and ranking of top journal preferences fluctuate over time, a subset of journals consistently exhibits dominance, appearing in the top ranks across most years. Although no clear relationship exists between journal preferences and ratings, A- and B-rated journals often dominate preferences for specific topics. Additionally, A- and B-rated journals with high or long preferences showed limited novelty. Most journals that preferred to interact with novel issues were C-rated.
In the context of considering the double-edged sword of corporate innovation, this study investigates whether and how corporate innovation affects the cost of equity. Using a patent-based innovation dataset of China's A-share publicly traded companies from 2009 to 2020, the study reveals that corporate innovation is associated with the cost of equity. The results show corporate innovation and the cost of equity have an inverted U-shaped relationship. In the initial stage of innovation, corporate innovation leads to an increase in the cost of equity, exacerbating the problems it brought to the company; but in the middle and late stage of innovation, as the level of innovation increases, the cost of equity experiences a sharp decrease, suggesting that continuing participating in innovative activities can mitigate the problems and even reverse problems into benefits. This study also explores the role of government subsidies as mediator on the association between corporate innovation and the cost of equity. The findings show that subsidies mediate the relationship between corporate innovation and the cost of equity. Innovation can attract more government subsidies and meanwhile, there is an inverted U-shaped relationship between subsidies and the cost of equity. The results provide empirical evidence to encourage managers to invest in innovative activities and also suggestions to policy makers to inject more funds to initial-stage innovative companies to foster innovation development.
Female scientific and technological talents are vital to scientific advancement worldwide. Governments and research organizations around the world have applied more inclusive and flexible time restrictions for female scientists. The National Natural Science Foundation of China extended the application age limit of the Young Scientists Fund (YSF) for female scientists from 35 to 40. Based on the unique dataset of 416 female applicants to the YSF and the model of difference-in-differences, we found that the policy significantly improved the quantity and quality of the female applicants' scientific publications during a 15-year period. The policy's impact is more pronounced among more competitive institutions and research fields with higher proportions of women. Furthermore, the impact varies by the age at which female scientists are exposed to the policy. The policy impact can be explained by the increased collaboration opportunities. The study may provide instructive and feasible guidance for empowering women in science.
Patent citation data is widely used in the study of technology evolution, but existing research has overlooked an issue that there may be potential differences between examiner citations and applicant citations, which may introduce biases from examiner citations. Yet, there is still a lack of systematic comparative study on the differences between applicant citations and examiner citations for technology evolution. To address this, we conducted a comprehensive comparison using USPTO patent data across four dimensions: technology profiling, technology relevance, technology diversity, and technology evolution pathways. For our case study, we selected the promising research area of photovoltaic cells. After comparing nine sub-technologies in this area, we have drawn some conclusions: (1) Applicants tend to provide more citations than examiners, and examiners tend to cite more recent patents than applicants; (2) There is no apparent inclination for applicants to avoid citing particularly relevant patents. On average, examiner citations are slightly closer in technological proximity to their invention than those cited by applicants; (3) The degree of diversity for applicant citations, examiner citations, and applicant examiner citations at a single patent level lacks consistency. However, their average trend by year or by sub-technology is similar after adding examiner citations; (4) Merging family members strongly impacts main pathways through added examiner citations, which is quite contrary in the citation network with only USPTO-granted patents without merging patent members; (5) In sub-technologies at the growth stage, applicants and examiners both cite more recent patents and tend to integrate border technologies from other fields, which can be used as an indicator for evaluating the potential to become emerging. The findings remind us to pay extra attention to the context in which citation data is used to measure technology evolution, and can serve as signals for technology assessment as well.
The system of scientific innovation can be characterized as a complex, multi-layered network of actors, their products and knowledge elements. Despite the progress that has been made, a more comprehensive understanding of the interactions and dynamics of this multi-layered network remains a significant challenge. This paper constructs a multilayer longitudinal network to abstract institutions, products and ideas of the scientific system, then identifies patterns and elucidates the mechanism through which actor collaboration and their knowledge transmission influence the innovation performance and network dynamics. Aside from fostering a collaborative network of institutions via co-authorship, fine-grained knowledge elements are extracted using KeyBERT from academic papers to build knowledge network layer. Empirical studies demonstrate that actor collaboration and their unique and diverse ideas have a positive impact on the performance of the research products. This paper also presents empirical evidence that the embeddedness of the actors, their ideas and features of their research products influence the network dynamics. This study gains a deeper understanding of the driving factors that impact the interactions and dynamics of the multi-layered scientific networks.
Understanding the evolution of knowledge has been and will continue to be the key task of science, technology, and innovation management. Existing research on evolutionary path identification relies primarily on traditional co-occurrence analysis and bag-of-words (BOW)–based models for topic extraction. However, these approaches have limitations in effectively capturing the underlying semantics and linkages of the topics. In this article, we propose a novel embedding-based methodology for scientific evolution analysis, in which word embedding, document embedding, clustering, and network analysis are applied to extract topics, measure topical semantic similarities, and quantitatively distinguish topics’ evolutionary states. We first perform benchmark experiments to demonstrate that doc2vec generally outperforms the BOW-based models in topic extraction before evolution analysis. We then consider topic consistency in vector spaces to identify evolutionary states including newborn, convergence, inheritance, and extinction. Scientific evolutionary paths are finally unraveled based on topic similarity matrixes and evolutionary states. We conduct a case study on object detection research to validate the effectiveness of our methodology. The empirical results, validated by domain experts, demonstrate that the proposed methodology is capable of effectively revealing patterns of knowledge inheritance and integration. Consequently, this methodology can be used to improve decision-making processes in future innovation management.
Identifying potentially disruptive technologies is challenging but important for innovators. The existing research based on the tech mining approach pays much attention to technological change but less attention to characterizing its disruptive process and effects. In this article, we suggest a novel perspective to help understand and identify potentially disruptive technology that displaces the mainstream technology (termed "alternative disruption") by modeling it as a process in which alternative technologies compete against incumbent technology. Accordingly, we propose a systematic framework to identify this type of technological disruptors by quantitatively characterizing its disruptive process and effects. To illustrate the alternative features, uniquely, the framework is solutions focused that answer the same technological problem as the mainstream one in which subject-action-object semantic analysis and community detection algorithm are used to mine and cluster the found solutions into groups as candidate technologies, including mainstream technologies and potentially alternative ones. Incorporating disruptive characteristics of technological advance, technology applicability, and market niche, the alternative one with a highest comprehensively competitive position against mainstream technologies remains as the most disruptive potential. Finally, the case of cancer treatments verifies the feasibility and effectiveness of this framework. Also, this proposed framework can provide quantitative information for decision making in promising technologies deployment and resource allocation.
Science, technology, and innovation are becoming increasingly collaborative, prompting concerted efforts to understand and measure the factors influencing these collaborations. This study aims to explore the driving factors and underlying mechanisms of collaboration dynamics based on patent data. Multilayer longitudinal networks are constructed to scrutinize interactions among organizations as well as the embedding of their knowledge elements in the network fabric. We then analyze the structures and characteristics of collaboration and knowledge networks from global and local perspectives, in which process topological indicators and graphlets are used to feature each organization's collaborative patterns and knowledge stock. Knowledge elements are extracted to present the core concepts of patents, overcoming the limitations of predefined categorizations, such as IPC, when representing technological content and context. By performing a longitudinal analysis using a stochastic actor-oriented model, we integrate network structures, node characteristics, and different dimensions of proximity to model collaboration dynamics and reveal the driving factors behind them. An empirical study in the field of lithography finds that organizations with a larger number of partners or a higher number of annular graphlets in their collaboration networks are less likely to collaborate with others. If an assignee has a more extensive range of knowledge elements and demonstrates a higher capability for knowledge combination, or if its local knowledge network exhibits weaker connectivity, its propensity to seek new collaborators increases. Both cognitive and organizational proximity play important roles in fostering collaboration.
With advances in medicine and biotechnology, the variety of available drugs has become more and more abundant. However, along with these innovations come complex adverse drug reactions (ADRs) as well. Extensive clinical trials are one of the best ways to reduce the incidence of drug reactions, but as the number of potential drug interactions grows, trials are becoming enormously time-consuming and costly. Hence, we set out to develop an alternative that could widely identify potential ADRs. Our solution is a "drug-ADR" network built from semantic subject-action-object structures, combined with complex network analysis and link prediction methods to reveal likely adverse reactions. Some similarity calculating methods also be used to improve our prediction accuracy. Evaluations of the results against the medical literature show that the predictions produced can be used as a weather vane for clinical trials, helping to save R&D time and capital costs. In addition, the framework can be used to provide useful guidance for discovering new drug indications or to inform the development of new drugs.
Early career funding is usually the first prestigious funding young scientists receive, allowing them to make their debut on a nationally recognised foundation. In this study, we examined the impact of an early debut on young scientists' research productivity. First-movers and late-comers are distinguished based on the years between the first application to the final award of early career funding. We then explored the variations between 3353 first-movers and 4650 late-comers of the Young Scientists Fund sponsored by the National Natural Science Foundation of China. We find that an early debut has a strong positive short-term effect on research productivity in terms of both quantity and quality, and the positive effect amplifies with the increasing time span of the final award between first-movers and late-comers. However, the strong positive effect on long-term productivity presents only in the three- and four-year early debuts. These results suggest that the productivity gains of young scientists with an early debut tend to decrease over time. The significant gap between first-movers and three-, and four-year late-comers in the long term demonstrates a time threshold which distinguishes scientists' long-term research productivity. In addition, we find that the research productivity gap can be explained by the expanding research network and increasing funding opportunities.
This article presents an improved method of measuring technology similarity by introducing a subject-action-object (SAO) analysis that uses the feature weights of semantic structure and professional vocabulary to measure technology similarity in the medical field. First, the SAO semantic structures are extracted and cleaned; then the structures related to technology are identified using a semantic network of the unified medical language system (UMLS). Second, the similarity between the SAO semantic structures is evaluated using semantic information from the Metathesaurus of the UMLS. Third, the feature weights of the SAO semantic structure are introduced to represent the importance of the patentees’ technology features. Finally, using the SAO and weight information, each patentee's vector is constructed to measure the technology complementarity between different patentees. This study conducts empirical research on Alzheimer's disease. The results indicate that the propose method for measuring technology similarity enables finer distinctions with more reliable outcomes than the traditional methods that are based on keywords and international patent classification.
Science and technology are becoming increasingly collaborative. This paper aims to explore the factors and mechanisms that impact the dynamic changes of collaborative innovation networks. We consider both collaborative interactions of organizations and their knowledge element exchanges to reveal how social and knowledge network embeddedness affects the collaboration dynamics. Knowledge elements are extracted to present the core concepts of scientific and technical information, overcoming the limitations of using predefined categorizations such as IPC when representing the content. Based on multiple collaboration and knowledge networks, we then conduct a longitudinal analysis and apply a stochastic actor-oriented model (SAOM) to model network dynamics over different periods. The influence of network features and structures, individual node characteristics, and various dimensions of proximity on collaboration dynamics is tested and analyzed.
As the scale of data shows rapid growth in various fields, big data’s vast amount of information can facilitate scientific discovery or decision-making. Deep neural network prevails in modeling big data such as images and text in computer vision and natural language processing. However, there is currently no widespread deep neural network for high-dimensional tabular data (HTD), as HTD could increase the model’s complexity and make estimating the parameters more difficult. Therefore, this paper proposes CLDNSR, a contrastive learning-enhanced deep neural network with serial regularization. This method combines relaxed Bernoulli distribution-based L0 regularization and adaptive L2 regularization for important feature selection and adaptive redundancy control to effectively handle high-dimensional input features. In addition, a tabular contrastive pre-training method is proposed to stabilize the supervised training process through better parameter initialization. Experiments on eleven real-world high-dimensional tabular datasets demonstrate that CLDNSR outperforms the baseline models designed for high-dimensional data.
概念内涵及其特征是科学研究的重要基础,现有文献对突破性创新和颠覆性创新的研究存在概念内涵交叉重叠的现象,对各自特征的分析和演化过程研究也尚未达成共识.本文采用系统性综述的方法,通过对国内外主要文献的梳理和分析,明确了突破性创新和颠覆性创新的概念内涵,探讨了突破性创新和颠覆性创新演化的底层逻辑,厘清了颠覆性创新的源起、发展及其与突破性创新之间的复杂关系,为颠覆性创新的系统化研究提供了基础.本文认为,突破性创新与颠覆性创新具有显著差异,从突破性创新到颠覆性创新是一个梯度演化的过程.突破性创新属于技术创新层面的概念,是基于组织战术层面的创新活动,而颠覆性创新是基于组织战略层面的创新活动,是以"颠覆性打破"为导向的复杂创新,注重技术和组织资源的深度融合,涵盖了创新的所有范畴.本文提出,在数字经济时代,进一步探讨突破性创新和颠覆性创新的内在演化逻辑、构建颠覆性创新理论分析框架、寻求颠覆性创新开放与控制的战略平衡治理方式、发现颠覆性技术的早期识别方法以及剖析中国情境下颠覆性创新的演化机制与实现路径等都是未来研究的重要内容.
Technology opportunity analysis (TOA) is of great help to technology innovation and R&D strategy of enterprises. Most previous studies of TOA only focused on scientific research or technology development phases, seldom linking technology opportunities to business applications. Using both patent and trademark data, we focused on the linkage of technology and business areas and discovered technology opportunities based on firms’ existing technology base. First, we extracted common product keywords from two data sources to construct document-keyword matrices. Then, we matched patents and trademarks with common keywords to build the relevancy network between technology and the business areas. Finally, we discovered technology opportunities from potential undeveloped business areas, taking into account technology-business relevancy and firms’ technology base. We conducted a case study on 3D printing to test our method.
Cancer prediction based on microarray data can facilitate the molecular exploration of cancers, thus building more accurate cancer prediction models is essential. This study focuses on a deep learning-based cancer prediction model. However, using a deep neural network to predict cancer is a difficult task due to the complexity of the underlying biological patterns and high dimension low sample size (HDLSS) of microarray data, which could bring about over-fitting and large training gradient variance. Therefore, a tree-enhanced deep adaptive network (TEDAN) is proposed to address these issues. Firstly, we employ the idea of the ensemble tree as a feature transformation method to alleviate the over-fitting problem, which generates a feature with a lower dimension and a more discriminative pattern. Secondly, a deep adaptive network (DAN) based on a self-attention mechanism is proposed to model the underlying biological interaction between different genes. Thirdly, a low sample size training (LSST) method is proposed to further reduce the large training gradient variance. Experiment results on six public cancer prediction datasets demonstrate that the TEDAN outperforms other strong baseline models.
[目的/意义]突破性创新对科技发展具有关键作用.大数据环境下,科学技术发展本身所具有的复杂、多维、不断进化等特征越发凸显.以动态视角进行突破性创新主题识别,对于为国家、企业及高校详析突破性创新领域、合理配置创新资源以及提供创新升级解决方案具有重要意义.[方法/过程]综合运用主题模型、词嵌入算法以及复杂网络分析等方法构建动态主题网络,全面考量主题在时间窗口内的结构特性以及时间窗口间的演化状态,并以其为基础结合突破性创新的新颖性、突变性、影响力和学科交叉性特征识别突破性创新主题.[结果/结论]面向区块链领域展开实证研究,识别出神经网络(Neural Network)和边缘计算(Edge Computing)两个主题的突破性创新特征最为显著.结合区块链现有研究及美国国家科学技术委员会发布的关键和新兴技术清单,验证了本文方法的可行性和有效性.但有关结果的定量验证,以及融合多源数据的突破性创新主题识别有待进一步研究.
Since the first Global Tech Mining (GTM) conference was held in Atlanta in 2011, the GTM conference has created a platform to connect tech mining researchers, exchange ideas and research progress, and promote collaborations. When it came to its 10th anniversary in 2020, COVID-19 forced the GTM conference into an online format. In tumultuous times for ST&I research activity, the GTM conference sought to focus on several issues: How to better collect and combine multiple "large data" sources? How to analyze these data effectively? And how to utilize these results more powerfully in ST&I management? In this collection, 15 papers are selected after evaluating by the science advisory committee, the guest editor team, and our peer review experts to address the following aspects regarding "tech mining": (1) DATA: Maximizing the potential of traditional and novel data; (2) METHODS: Advancing and integrating methods; (3) APPLICATIONS: Innovative analyses translating to usefulintelligence.