Knowledge Graphs (KG) and associated Knowledge Graph Embedding (KGE) models have recently begun to be explored in the context of drug discovery and have the potential to assist in key challenges such as target identification. In the drug discovery domain, KGs can be employed as part of a process which can result in lab-based experiments being performed, or impact on other decisions, incurring significant time and financial costs and most importantly, ultimately influencing patient healthcare. For KGE models to have impact in this domain, a better understanding of not only of performance, but also the various factors which determine it, is required. In this study we investigate, over the course of many thousands of experiments, the predictive performance of five KGE models on two public drug discovery-oriented KGs. Our goal is not to focus on the best overall model or configuration, instead we take a deeper look at how performance can be affected by changes in the training setup, choice of hyperparameters, model parameter initialisation seed and different splits of the datasets. Our results highlight that these factors have significant impact on performance and can even affect the ranking of models. Indeed these factors should be reported along with model architectures to ensure complete reproducibility and fair comparisons of future work, and we argue this is critical for the acceptance of use, and impact of KGEs in a biomedical setting.
The drug discovery and development process is a long and expensive one, costing over 1 billion USD on average per drug and taking 10-15 years. To reduce the high levels of attrition throughout the process, there has been a growing interest in applying machine learning methodologies to various stages of drug discovery and development in the recent decade, especially at the earliest stage - identification of druggable disease genes. In this paper, we have developed a new tensor factorisation model to predict potential drug targets (genes or proteins) for treating diseases. We created a three-dimensional data tensor consisting of 1,048 gene targets, 860 diseases and 230,011 evidence attributes and clinical outcomes connecting them, using data extracted from the Open Targets and PharmaProjects databases. We enriched the data with gene target representations learned from a drug discovery-oriented knowledge graph and applied our proposed method to predict the clinical outcomes for unseen gene target and disease pairs. We designed three evaluation strategies to measure the prediction performance and benchmarked several commonly used machine learning classifiers together with Bayesian matrix and tensor factorisation methods. The result shows that incorporating knowledge graph embeddings significantly improves the prediction accuracy and that training tensor factorisation alongside a dense neural network outperforms all other baselines. In summary, our framework combines two actively studied machine learning approaches to disease target identification, namely tensor factorisation and knowledge graph representation learning, which could be a promising avenue for further exploration in data-driven drug discovery.
Drug discovery and development is a complex and costly process. Machine learning approaches are being investigated to help improve the effectiveness and speed of multiple stages of the drug discovery pipeline. Of these, those that use Knowledge Graphs (KG) have promise in many tasks, including drug repurposing, drug toxicity prediction and target gene-disease prioritisation. In a drug discovery KG, crucial elements including genes, diseases and drugs are represented as entities, whilst relationships between them indicate an interaction. However, to construct high-quality KGs, suitable data is required. In this review, we detail publicly available sources suitable for use in constructing drug discovery focused KGs. We aim to help guide machine learning and KG practitioners who are interested in applying new techniques to the drug discovery field, but who may be unfamiliar with the relevant data sources. The datasets are selected via strict criteria, categorised according to the primary type of information contained within and are considered based upon what information could be extracted to build a KG. We then present a comparative analysis of existing public drug discovery KGs and a evaluation of selected motivating case studies from the literature. Additionally, we raise numerous and unique challenges and issues associated with the domain and its datasets, whilst also highlighting key future research directions. We hope this review will motivate KGs use in solving key and emerging questions in the drug discovery domain.
Online markets for mental health care (OMMH) allow clients to connect remotely with counselors to receive psychological therapy. Rooted in signaling theory and in the specific context of an OMMH, we theorize relative credibility of signals as the boundary condition that determines whether the demonstration signal of responsiveness to client questions substitutes or complements the two description signals of professional qualifications and counseling style in predicting market demand for counselors from new clients in an OMMH. Based on a panel dataset of 823 observations from 309 counselors participating on YiXinLi, a leading OMMH in China, we tested our hypotheses using linguistic and sentiment analysis methods and zero-inflated negative binomial models. We found broad support for nine out of ten hypotheses. Findings are robust with respect to different measures of variables, potential endogeneity from the simultaneity of responsiveness and counselor demand, and potential selection bias from both observable and endogenous covariates. Our study extends the literature on signaling in online markets in the unique context of OMMH by showing that: (1) relative credibility of signals is the boundary condition that determines when a demonstration signal will complement and when it will substitute for a description signal, (2) previous clients’ feedback on counselors’ empathy and warmth was deemed not credible by new clients in the context of online counseling, and (3) responsiveness to client questions is the most influential predictor of market demand from new clients in an OMMH.
The drug discovery and development process is a long and expensive one, costing over 1 billion USD on average per drug and taking 10-15 years. To reduce the high levels of attrition throughout the process, there has been a growing interest in applying machine learning methodologies to various stages of drug discovery process in the recent decade, including at the earliest stage - identification of druggable disease genes. In this paper, we have developed a new tensor factorisation model to predict potential drug targets (i.e.,genes or proteins) for diseases. We created a three dimensional tensor which consists of 1,048 targets, 860 diseases and 230,011 evidence attributes and clinical outcomes connecting them, using data extracted from the Open Targets and PharmaProjects databases. We enriched the data with gene representations learned from a drug discovery-oriented knowledge graph and applied our proposed method to predict the clinical outcomes for unseen target and dis-ease pairs. We designed three evaluation strategies to measure the prediction performance and benchmarked several commonly used machine learning classifiers together with matrix and tensor factorisation methods. The result shows that incorporating knowledge graph embeddings significantly improves the prediction accuracy and that training tensor factorisation alongside a dense neural network outperforms other methods. In summary, our framework combines two actively studied machine learning approaches to disease target identification, tensor factorisation and knowledge graph representation learning, which could be a promising avenue for further exploration in data-driven drug discovery.
Background Genome-wide ligation-based assays such as Hi-C provide us with an unprecedented opportunity to investigate the spatial organization of the genome. Results of a typical Hi-C experiment are often summarized in a chromosomal contact map, a matrix whose elements reflect the co-location frequencies of genomic loci. To elucidate the complex structural and functional interactions between those genomic loci, networks offer a natural and powerful framework. Results We propose a novel graph-theoretical framework, the Corrected Gene Proximity (CGP) map to study the effect of the 3D spatial organization of genes in transcriptional regulation. The starting point of the CGP map is a weighted network, the gene proximity map, whose weights are based on the contact frequencies between genes extracted from genome-wide Hi-C data. We derive a null model for the network based on the signal contributed by the 1D genomic distance and use it to “correct” the gene proximity for cell type 3D specific arrangements. The CGP map, therefore, provides a network framework for the 3D structure of the genome on a global scale. On human cell lines, we show that the CGP map can detect and quantify gene co-regulation and co-localization more effectively than the map obtained by raw contact frequencies. Analyzing the expression pattern of metabolic pathways of two hematopoietic cell lines, we find that the relative positioning of the genes, as captured and quantified by the CGP, is highly correlated with their expression change. We further show that the CGP map can be used to form an inter-chromosomal proximity map that allows large-scale abnormalities, such as chromosomal translocations, to be identified. Conclusions The Corrected Gene Proximity map is a map of the 3D structure of the genome on a global scale. It allows the simultaneous analysis of intra- and inter- chromosomal interactions and of gene co-regulation and co-localization more effectively than the map obtained by raw contact frequencies, thus revealing hidden associations between global spatial positioning and gene expression. The flexible graph-based formalism of the CGP map can be easily generalized to study any existing Hi-C datasets.
Background: Normal users' voluntary behaviors (e.g., knowledge sharing) in virtual communities (VCs) has been well investigated; however, research on health professionals' voluntary behaviors in online health communities (OHCs) is limited. Objective: This paper focuses on OHCs for mental health and aims to explore how intrinsic and extrinsic motivations influence mental health service providers' voluntary behaviors. Methods: Based on motivation theory and prior studies, we incorporated technical competence as intrinsic motivation and online reputation and economic rewards as extrinsic motivations, and proposed five hypotheses. We crawled objective data from YiXinLi, a Chinese OHC for mental health, and tested the hypotheses based on the Poisson regression model. All hypotheses are supported. Results: 1) Technical competence, online reputation, and economic rewards positively influence mental health service providers' voluntary behaviors; 2) the interaction effect between technical competence and online reputation negatively influences mental health service providers' voluntary behaviors; 3) the interaction effect between technical competence and economic rewards negatively influences mental health service providers' voluntary behaviors. Conclusions: Both intrinsic motivations and extrinsic motivations positively influence mental health service providers' voluntary behaviors, and their interaction effects negatively influence mental health service providers' voluntary behaviors. This study first contributes to the literature on health professionals' voluntary behaviors in OHCs by verifying the positive effect of economic rewards. It then contributes to motivation theory by incorporating a situation where intrinsic motivations and extrinsic motivations could negatively interact.
Cheng Ye, ∗ César H. Comin, † Thomas K. DM. Peron, ‡ Filipi N. Silva, § Francisco A. Rodrigues, ¶ Luciano da F. Costa, ∗∗ Andrea Torsello, †† and Edwin R. Hancock ‡‡ Department of Computer Science, University of York, York, YO10 5GH, UK. Institute of Physics at São Carlos, University of São Paulo, PO Box 369, São Carlos, São Paulo, 13560-970, Brazil. Institute of Mathematical and Computer Sciences, University of São Paulo, PO Box 668, São Carlos, São Paulo, 13560-970, Brazil. Department of Environmental Sciences, Informatics and Statistics, Ca’ Foscari University of Venice, Dorsoduro 3246 30123 Venezia, Italy. (Dated: September 2015)
With the rapid growth of social networking in the health industry, the online health community (OHC) has become an important channel for people to conduct mental health care activities. People can seek psychological knowledge to selfeducate, communicate with other patients like them to look for support, or interact with psychological counselors for mental health care services. Online health services belong to credence goods.[1] In the credence market, experts not only provide professional services but also act as the experts who determine how much treatment is necessary, due largely to information
Structural complexity measures have found widespread use in network analysis. For instance, entropy can be used to distinguish between different structures. Recently, we have reported an approximate network von Neumann entropy measure, which can be conveniently expressed in terms of the degree configurations associated with the vertices that define the edges in both undirected and directed graphs. However, this analysis was posed at the global level, and did not consider in detail how the entropy is distributed across edges. The aim in this paper is to use our previous analysis to define a new characterization of network structure, which captures the distribution of entropy across the edges of a network. Since our entropy is defined in terms of vertex degree values defining an edge, we can histogram the edge entropy using a multi-dimensional array for both undirected and directed networks. Each edge in a network increments the contents of the appropriate bin in the histogram, indexed according to the degree pair in an undirected graph or the in/out-degree quadruple for a directed graph. We normalize the resulting histograms and vectorize them to give network feature vectors reflecting the distribution of entropy across the edges of the network. By performing principal component analysis (PCA) on the feature vectors for samples, we embed populations of graphs into a low-dimensional space. We explore a number of variants of this method, including using both fixed and adaptive binning over edge vertex degree combinations, using both entropy weighted and raw bin-contents, and using multi-linear PCA, aimed at extracting the tensorial structure of high-dimensional data, as an alternative to classical PCA for component analysis. We apply the resulting methods to the problem of graph classification, and compare the results obtained to those obtained using some alternative state-of-the-art methods on real-world data.
In this paper, an extended car-following model with consideration of the combination effect in heterogeneous urban traffic flow with slow and fast vehicles is proposed. The combination effect means the order of vehicles in a heterogeneous flow matters, since a fast vehicle would degenerate to a middle state between a fast vehicle and a slow vehicle when it follows a slow vehicle. Specifically, we consider three types of vehicle combinations in the urban traffic flow, namely fast vehicle following fast vehicle (FF), fast vehicle following slow vehicle (FS), and slow vehicle following fast/slow vehicle (SX). Linear stability analysis is conducted to obtain stability criterion, from which we can conclude that the higher penetration of the FS combination can increase the stability of traffic flow.
In this thesis, we address problems encountered in complex network analysis using graph theoretic methods. The thesis specifically centers on the challenge of how to characterize the structural properties and time evolution of graphs. We commence by providing a brief roadmap for our research in Chapter 1, followed by a review of the relevant research literature in Chapter 2. The remainder of the thesis is structured as follows. In Chapter 3, we focus on the graph entropic characterizations and explore whether the von Neumann entropy recently defined only on undirected graphs, can be extended to the domain of directed graphs. The substantial contribution involves a simplified form of the entropy which can be expressed in terms of simple graph statistics, such as graph size and vertex in-degree and out-degree. Chapter 4 further investigates the uses and applications of the von Neumann entropy in order to solve a number of network analysis and machine learning problems. The contribution in this chapter includes an entropic edge assortativity measure and an entropic graph embedding method, which are developed for both undirected and directed graphs. The next part of the thesis analyzes the time-evolving complex networks using physical and information theoretic approaches. In particular, Chapter 5 provides a thermodynamic framework for handling dynamic graphs using ideas from algebraic graph theory and statistical mechanics. This allows us to derive expressions for a number of thermodynamic functions, including energy, entropy and temperature, which are shown to be efficient in identifying abrupt structural changes and phase transitions in real-world dynamical systems. Chapter 6 develops a novel method for constructing a generative model to analyze the structure of labeled data, which provides a number of novel directions to the study of graph time-series. Finally, in Chapter 7, we provide concluding remarks and discuss the limitations of our methodologies, and point out possible future research directions.
In this paper, we present a new method for modeling time-evolving correlation networks, using a Mean Reversion Autoregressive Model, and apply this to stock market data. The work is motivated by the assumption that the price and return of a stock eventually regresses back towards their mean or average. This allows us to model the stock correlation time-series as an autoregressive process with a mean reversion term. Traditionally, the mean is computed as the arithmetic average of the stock correlations. However, this approach does not generalize the data well. In our analysis we utilize a recently developed generative probabilistic model for network structure to summarize the underlying structure of the time-varying networks. In this way we obtain a more meaningful mean reversion term. We show experimentally that the dynamic network model can be used to recover detailed statistical properties of the original network data. More importantly, it also suggests that the model is effective in analyzing the predictability of stock correlation networks.
In this paper, we present a novel method for constructing a generative model to analyze the structure of labeled data. Given a time-series of sample graphs, we aim to learn a so-called “supergraph” that best describes the underlying average connectivity structure presenting in the data. In this time-series the vertex set is fixed and labeled and the set of possible connections between vertices change with time. The supergraph represents these changes with a Gaussian probability distribution for the connection weights on each individual edge. This structure is fitted to the time-series data by minimizing a description length criterion, with the von Neumann entropy controlling the complexity of the fitted model structure and the Gaussian log-likelihood controlling the mean edge weights and variances. We further show this fitting process can be optimized by using a new fixed-point iteration scheme which locates the elements of the optimal weighted adjacency matrix of the supergraph. We show the iteration process is in fact governed by the partial derivative of the von Neumann entropy. In the experiments, the resulting generative model is shown to be an effective tool for analyzing the underlying connectivity structure of time-evolving networks in the financial domain, and in particular locating critical events and distinct time epochs in their evolution.
Recently, kernel methods have been widely employed to solve machine learning problems such as classification and clustering. Although there are many existing graph kernel methods for comparing patterns represented by undirected graphs, the corresponding methods for directed structures are less developed. In this paper, to fill this gap in the literature we exploit the graph kernels and graph complexity measures, and present an information theoretic kernel method for assessing the similarity between a pair of directed graphs. In particular, we show how the Jensen-Shannon divergence, which is a mutual information measure that gauges the difference between probability distributions, together with the recently developed directed graph von Neumann entropy, can be used to compute the graph kernel. In the experiments, we show that our kernel method provides an efficient tool for classifying directed graphs with different structures and analyzing real-world complex data.
Quantification of symmetries in complex networks is typically done globally in terms of automorphisms. Extending previous methods to locally assess the symmetry of nodes is not straightforward. Here we present a new framework to quantify the symmetries around nodes, which we call connectivity patterns. We develop two topological transformations that allow a concise characterization of the different types of symmetry appearing on networks and apply these concepts to six network models, namely the Erdos-Renyi, Barabasi-Albert, random geometric graph, Waxman, Voronoi and rewired Voronoi. Real-world networks, namely the scientific areas of Wikipedia, the world-wide airport network and the street networks of Oldenburg and San Joaquin, are also analyzed in terms of the proposed symmetry measurements. Several interesting results emerge from this analysis, including the high symmetry exhibited by the Erdos-Renyi model. Additionally, we found that the proposed measurements present low correlation with other traditional metrics, such as node degree and betweenness centrality. Principal component analysis is used to combine all the results, revealing that the concepts presented here have substantial potential to also characterize networks at a global scale. We also provide a real-world application to the financial market network. (C) 2015 Elsevier Inc. All rights reserved.
In this paper, we present a method for characterizing the evolution of time-varying complex networks by adopting a thermodynamic representation of network structure computed from a polynomial (or algebraic) characterization of graph structure. Commencing from a representation of graph structure based on a characteristic polynomial computed from the normalized Laplacian matrix, we show how the polynomial is linked to the Boltzmann partition function of a network. This allows us to compute a number of thermodynamic quantities for the network, including the average energy and entropy. Assuming that the system does not change volume, we can also compute the temperature, defined as the rate of change of entropy with energy. All three thermodynamic variables can be approximated using low-order Taylor series that can be computed using the traces of powers of the Laplacian matrix, avoiding explicit computation of the normalized Laplacian spectrum. These polynomial approximations allow a smoothed representation of the evolution of networks to be constructed in the thermodynamic space spanned by entropy, energy, and temperature. We show how these thermodynamic variables can be computed in terms of simple network characteristics, e.g., the total number of nodes and node degree statistics for nodes connected by edges. We apply the resulting thermodynamic characterization to real-world time-varying networks representing complex systems in the financial and biological domains. The study demonstrates that the method provides an efficient tool for detecting abrupt changes and characterizing different stages in network evolution.
The financial market is a complex dynamical system composed of a large variety of intricate relationships between several entities, such as banks, corporations and institutions. At the heart of the system lies the stock exchange mechanism, which establishes a time-evolving network of trades among companies and individuals. Such network can be inferred through correlations between time series of companies stock prices, allowing the overall system to be characterized by techniques borrowed from network science. Here we study the presence of communities in the inferred stock market network, and show that the knowledge about the communities alone can provide a nearly complete representation of the system topology. This is done by defining a simple null model, a randomized version of the studied network sharing only the sizes and interconnectivity between communities observed. We show that many topological characteristics of the inferred networks are carried over the networks generated by the null model. In particular, we find that in periods of instability, such as during a financial crisis, the network strays away from a state of well-defined community structure to a much more uniform topological organization. We show that the framework presented here provides a good null model representation of topological variations taking place in the market during crises. Also, the general approach used in this work can be extended to other systems.
In this paper, we present a novel and effective method for better understanding the evolution of time-varying complex networks by adopting a thermodynamic representation of network structure. We commence from the spectrum of the normalized Laplacian of a network. We show that by defining the normalized Laplacian eigenvalues as the microstate occupation probabilities of a complex system, the recently developed von Neumann entropy can be interpreted as the thermodynamic entropy of the network. Then, we give an expression for the internal energy of a network and derive a formula for the network temperature as the ratio of change of entropy and change in energy. We show how these thermodynamic variables can be computed in terms of node degree statistics for nodes connected by edges. We apply the thermodynamic characterization to real-world time-varying networks representing complex systems in the financial and biological domains. The study demonstrates that the method provides an efficient tool for detecting abrupt changes and characterizing different stages in evolving network evolution.