In big data era, multi-source heterogeneous data become the biggest obstacle to data sharing due to its high dimension and inconsistent structure. Using text classification to solve the ontology construction and mapping problem of multi-source heterogeneous data can not only reduce manual operation, but also improve the accuracy and efficiency. This paper proposes an ontology construction and mapping scheme based on hybrid neural network and autoencoder. Firstly, the proposed text classification method uses the multi-core convolutional neural network to capture local features and uses the improved Bidirectional Long Short-Term Memory network to compensate for the shortcomings of the convolutional neural network that cannot obtain context-related information. Secondly, a similarity matching method is used for ontology mapping, which integrate autoencoder to improve anti-interference ability. We have carried out several sets of experiments to test the validity of the proposed ontology construction and mapping scheme.
In recent years, the related research of entity alignment has mainly focused on entity alignment via knowledge embeddings and graph neural networks; however, these proposed models usually suffer from structural heterogeneity and the large-scale problem of knowledge graph. A novel entity alignment model based on graph isomorphic network and compressed sensing is proposed. First, for the problem of structural heterogeneity, graph isomorphic network encoder is applied in knowledge graph to capture structural similarity of entity relation. Second, for the problem of large scale, key node and community are integrated for priority entity alignment to improve execution speed. However, the exiting node importance ranking algorithm cannot accurately identify key node in knowledge graph. So the compressed sensing is adopted in node importance ranking to improve the accuracy of identifying key node. The authors have carried out several experiments to test the effect and efficiency of the proposed entity alignment model.
At present, the relation completion mainly study the influence of single-path or first-order information, but ignoring the more complex relation information widely existing between entities. Meanwhile, recommendation systems based on collaborative filtering algorithms are susceptible to data sparsity issues, leading to a cold start of the system. Therefore, a novel embedding learning framework for relation completion and recommendation based on graph neural network and multi-task learning is proposed. The graph neural network relation completion model GNNRC predicts the relation between entities by embedding learning, which fuses the high-order semantic features of the two target entities’ subgraph based on graph neural network. A comparison with TransD model on embeding learning for relation completion is reported. On this basic, the multi-task learning relation recommendation model MLRR add the bridging unit into deep end-to-end recommendation model, which realizes the alternating learning of knowledge graph embedding and recommendation algorithm. Experimental results show that the performance of the proposed recommendation model is significantly better than other baselines. Moreover, strong recommendation performance can be maintained in cold start scenarios where data are sparse.
With the increasing complexity of scientific research, it has gradually turned to a collaborative approach, which can promote knowledge sharing, resource sharing and improve the efficiency of scientific research achievements. Therefore, It is of great significance to study the internal organizational structure and evolution mechanism of scientific research collaboration, which plays a crucial role in the management of scientific research work and the formulation of scientific and technological policies. This paper focuses on three aspects: core node evaluation, community detection and visual layout algorithm of scientific research collaboration network, which is constructed based on the network embedding of the scientific research achievements' attributes. Considering network topology and node heterogeneity, a core node evaluation method is proposed, and a community detection algorithm and a visual layout algorithm is improved to display the community structure of scientific research collaboration network from many aspects. The experimental results show that the proposed method can more clearly show the internal structure of scientific research collaboration community. (c) 2021 Elsevier B.V. All rights reserved.
Data reuse strategy is an effective method to save storage space and improve data utilization in data management. In view of the successful application of deep learning in the field of text mining, a data reuse strategy based on deep learning is proposed for high dimensional data’s pattern and instance similarity. With traditional feature analysis and deep learning model of convolutional neural network, the pattern similarity of data dimension is analyzed so as to optimize the similar dimension pairs among high dimensional data sets. Combining inner-attention mechanism, a semantic similarity model IA-LSTM is designed for instance similarity, which can build the association mapping among data entities by the calculation of the similarity of short text. Based on the pattern and instance similarity in the proposed strategy, reusable data entities are discovered, and column storage is designed to improve data reuse efficiency.
The spreading of information in network is different from epidemics in the population; meanwhile, the node is heterogeneous, and the structure is going in the direction of double or even multi-layer. It is of great practical significance to study the anti-risk capability of coupled network. Based on the subjective heterogeneity and memory effect heterogeneity, a two-layer SIR information propagation model is constructed and an important node selection method for the coupled network based on technique for order preference by similarity to an ideal solution (TOPSIS) is proposed. The effectiveness of the constructed model and the proposed method is verified by simulation experiment which selects the important nodes as the immune nodes of TOPSIS immunization strategy and adopts random immunization strategy, partial nodes immune layer strategy and TOPSIS immunization strategy on BA_BA, WS_WS and BA_WS coupled network. The experimental results show that subjective heterogeneity can hinder the dissemination of information, while the memory effect heterogeneity can facilitate the dissemination of information. In addition, different immune strategies have different effects on different coupled networks, for example, the TOPSIS immune strategy has the best effect in BA_BA network.
Currently, large data sets are deployed on large-scale clusters, which require a large amount of physical resources. However, current network architecture does not have flexible deployment, making it difficult to adjust physical resources after deployment. Based on software definition network, this article proposes a framework for building virtual data domain, which establishes a multi-attribute decision model by network nodes for optimizing the deployment of control layer, so as to realize large-scale deployment. By analyzing the actual usage and virtual distribution of the underlying data resources, a mapping algorithm of network resource overhead–allocation ratio is proposed for adjusting the mapping space of network flow based on the mapping result, so as to meet more virtual domain applications. At the same time, the reasonable utilization of resources is helpful to reduce the communication delay in domain. Simulation results show that compared with the shortest path mapping and greedy resource mapping algorithms, the virtual domain established by the network resource overhead–allocation algorithm can improve the resource utilization rate by 10% and reduce the intra-domain communication delay by 30%. Therefore, under the background of the expanding scale of data domain, this framework can solve the problem that the current backward network architecture cannot adapt to the development trend of the information field.
SummaryThe emergence of big data makes more and more enterprise change data management strategy, from simple data storage to OLAP query analysis; meanwhile, NoSQL‐based data warehouse receive more increasing attention than traditional SQL‐based database. By improving the JFSS model for ETL, this paper proposes the uniform distribution code (UDC), model identification code (MIC), standard dimension code (SDC), and attribute dimensional code (ADC); defines the data storage format of ; and identifies the extraction, transformation, and loading strategies of data warehouse. Several experiments are carried out to analyze single record and range record queries as typical OLAP based on Hadoop database (HBase). The results show the proposed scheme can provide lower overhead than the traditional SQL‐based database while facilitating the scope and flexibility of data warehouse services.
With the development of science and technology, the interactions among scientific research teams become more and more frequent, and their relationship and behavior become more and more complex. Many researches mainly adopt complex network to analyze, but these researches only consider some aspects of scientific research factors, so lack of comprehensive consideration. From the aspect of ability, resource, activity, and familiarity, scientific research factors are quantified based on multi-source data of scientific and technological big data, and some factors of text information are similarly quantified. Based on paper citation and project cooperation, a complex network which takes scientific research team as node is constructed and is weighted by quantification of scientific research factor. The experiment of influence spread is carried out by the comparison of unweighted network and weighted network, the comparison of single node and multiple nodes, and the comparison of influence spread and other index. The results show that the scientific research factor is closely related to the influence spread; the proposed scientific research factor quantification improves the analysis of scientific research team relationship. The relationship between influence spread and the number of related communities is greater than the number of adjacent nodes. In addition, the influence spread can effectively reflect the importance of scientific research team.
Science research has general rules of development, is like any other social activity. With the improvement of science and technology, scientific problems have become more complex and systematic, individual approach has been replaced by teamwork in scientific research. This paper takes scientific research team cooperative network as research object, analyzes the influence of scientific research teams in the cooperative network from the aspect of node heterogeneity and node similarity of content and structure, and puts forward the influence evaluation method of scientific research team. A scientific research team cooperation network is constructed as the unweighted and undirected graph by the cooperation relationship data of scientific research teams, including co-author, citation, project cooperation and son on. In this network, the scientific research teams are take nodes, and the cooperative relationships between scientific research teams are take as edges. The major factors of scientific research team influence are analyzed, including node heterogeneity and relationship strength between nodes, then a weight and attributed graph is constructed by the research direction of scientific research team and is weighted based on the similarity of nodes’ content and structure by the SimRank model and the Jaccard similarity method. An influence evaluation method was proposed based on the impact of node subjective heterogeneity and node domain heterogeneity, and An influence spread model based on SIR model was given for verifying the proposed influence evaluation method.
In this study, a path prediction method based on trusted central nodes is proposed for information flow transmission among multi-layer of social network. With the complex, sensitive and the burn-in of information protection strategies, the regulation and control of information flow transmission is becoming difficult in social network. By exacting the trusted central nodes from the community in social network, the feedback mechanism is used to realize the time-varying selecting of trusted central nodes. Then, an information spread link model for multi-layer of social network is obtained through the trusted central nodes. Finally, the shortest transmission path among layers of social network is calculated. The experimental results show that time-varying selection strategy of trusted central nodes restrains the rumor transmission which increases the reliability of information in social network. The information spread link algorithm for multi-layer of social network can reduce the path length and transmission time and improve the transmission efficiency.
The development and improvement of Internet technology has made network information richer and more attractive. Users can access networks more easily and enjoy greater network services. While the Internet provides rich and convenient service, Internet networks also provide greater conditions now for the breeding and spreading of harmful information. The user is subject to information dissemination on the network, and the user's behaviour in information flow has a tremendous impact on information dissemination. In this paper, based on the heterogeneity of the nodes in a network, an information flow model is established, wherein the factors influencing a node's information flow behaviour are researched and categorized as factors internal and external to the node. The internal factor entails the autonomy of a node, which contains the degree of interest and the subjective judgment of the node. The external factors comprise the network structure, the location of the node, and the relationships among the nodes. Considering the issue of the spread of harmful information, a complex network model is built based on the information flow behaviour of the nodes, and a control policy to address harmful information is established.
Reasonable selection of node deployment controller for software defined network will effectively improve the entire network performance.This paper introduced the reliability of the node betw eenness centrality and node as a parameter ,establishing nodes and the parameters of the matrix ,the matrix parameters normalize ,on the basis of ordered w eighted operator to sort the parameters ,and found the optimal node deployment controller .At the end of paper ,the control information transmission time is compared with the centrality method based on net‐work topology .The deployment node of multi‐parameter selection controller will effectively reduce the control path propagation delay and improve the netw ork performance of SDN .
Based on the content attributes of Weibo and the characteristics of the information dissemination law of social network ,the paper combines the Weibo text with the user's follo‐wer relationship as the standard of user interest classification ,so that the user's interest is more accurate and effective.Using the established user interest classification model to solve the problem of user interest classification ,the paper chooses Sina Weibo as the research ob‐ject ,in which the main topic extraction algorithm is LDA ,classification algorithm is LIBSVM . The experimental results show that the method can be used to classify the user information comprehensively and has higher classification accuracy than other methods .
The data analysis is closely related to data attribute dimension. The traditional extraction and partition of data attribute dimension is so manual and inefficiency as to not meet the needs of analysing big data. This paper proposed an attribute dimension partition scheme based on SVM classifying and MapReduce for analysing big data. This scheme improve traditional SVM classifying method by combining Euclidean distance theory for overcoming its disadvantages, and adopts punish coefficient to reduce the unbalance of data distribution. With the improved SVM classifying method, the implementation of attribute dimension partition take MapReduce model of Hadoop as process engine, use TF-IDF vector to save the extracted attribute dimension, and use k-means clustering algorithm to clustering partition. The experiment result shows that the execution efficiency of the proposed method is enhanced, and while the rationality of partition is guaranteed, the increasing of data attributes does not significantly increase the execution time.
Hebei Province science and technology innovation big data public platform is based on massive data resources ,the construction is based on data warehouse and data mining tech‐nology ,oriented management departments to carry out the decision‐making service ,netw ork information platform for the public to provide information service .How ever ,during the con‐struction of data warehouse ,there are all kinds of data quality problems ,resulting in various error analysis results ,so ,before the data get into the data w arehouse ,data cleaning should be done ,so as to ensure the quality of the data into data w arehouse.According to the scientific and technological big data standardization processing and application system of science and technology project in Hebei Province ,put forw ard the innovation of science and technology of data cleaning framework ,on the basis of the framework ,the definition of data cleaning rules , improved data cleaning algorithm ,experiments were carried out on the technological innova‐tion of large data system on real data sets ,solving the problems of data quality in data ware‐house ,so as to ensure consistency and the correctness of the datain the data warehouse ,provi‐ding a solid foundation for data analysis and processing of the late .
Text label is a kind of text keywords,can simplify extraction of effective information from science and technology policy.For science and technology policy,this paper divides text label into four kinds,such as science and technology investment,intellectual property rights, rural science and technology,tax.Aimed at the shortcoming of the traditional SVM algorithm's label data unbalance,this paper provides a text label classification method of sci-ence and technology policy,w hich combines the Euclidean distance algorithm and ESVM algo-rithm with penalty factor.Finally,with comparing SVM and ESVM,the validity of the pro-pose method on science and technology policy text label is verified.
Currently, more and more scholars believe that most networks are not isolated, but interact with each other. Internet-based social network (online network), and physical contact networks (offline network) are a coupled network that interact with each other. Person have online virtual identity and offline social identity. In this paper, the double SIR information spreading model with unilateral effect is constructed according to the characteristics of online and offline information dissemination. And two typical types of coupled networks, BA_BA network and WS_WS network, are used to simulate. The experimental results show that the online information propagation can inhibit the scope of offline. This inhibition is slightly weaker for the BA_BA network that the inter degree-degree correlation (IDDC) is positive. At the same time, it is found that the increase of the interlayer influence rate can enhance the synchronization of information transmission on online and offline so that effectively promote the information propagation of online/offline.
Based on mobility support of IPv6 and node discovery mechanism, this paper proposes a method using node discovery mechanism for mobile node auxiliary connection, in which we detect the real-time received signal strength indicator, by computing and analyzing the change of average value to predict variation trend. In addition, we drew up a new generation of Internet protocol implementation scheme in 6LoWPAN, which is the access of the combination of the mobile node and IoT gateway, etc. In the local IoT, things connect with IPv6, which satisfy the demand of mobility, the whole IoT architecture model of IPv6 connection will be implemented. The analysis and evaluation of network performance was completed in Cooja simulator based on Contiki system after the hardware test. Results show that this method is suitable for the stub network of Internet of things in the common dynamic scenes and more efficient and less packet loss probability when compare to the original 6LoWPAN.
In cloud services, users’ data and applications are stored in remote cloud servers, their safety and security can be guaranteed by their cloud servers, but it is difficult for guarantee their security in interaction among them. As software defined network arise, its architecture realizes the separation between control plane and data plane, provide a promising way of dealing with privacy leakage of cloud data. This paper study a data security mechanism among cloud service based on software define network architecture, which maintains security policy on SDN controller cluster, which controls forwarding based on the mapping between physical network and logic network. This mechanism contains data service provider, data service user and privacy service provider. Data service provider customizes data attribute restriction based on data security protection requirements. Privacy service provider which is realized based on SDN controller, is responsible for the security of data interacting between data service user and data service provider, such as identity authentication, source data partition and data block recovering in accordance with data attribute restriction. Data service provider stores source data on cloud server. Through experiments and analysis, this data security mechanism is effective and feasible.