BackgroundThe identification of gene-phenotype relationships is important in medical genetics as it serves as a basis for precision medicine. However, most of the gene-phenotype relationship data are buried in the biomedical literature in textual form.ObjectiveWe propose RelCurator, a curation system that extracts sentences including both gene and phenotype entities related to specific disease categories from PubMed articles, provides rich additional information such as entity taggings, and predictions of gene-phenotype relationships.MethodsWe targeted neurodegenerative disorders and developed a deep learning model using Bidirectional Gated Recurrent Unit (BiGRU) networks and BioWordVec word embeddings for predicting gene-phenotype relationships from biomedical texts. The prediction model is trained with more than 130,000 labeled PubMed sentences including gene and phenotype entities, which are related to or unrelated to neurodegenerative disorders.ResultsWe compared the performance of our deep learning model with those of Bidirectional Encoder Representations from Transformers (BERT), Support Vector Machine (SVM), and simple Recurrent Neural Network (simple RNN) models. Our model performed better with an F1-score of 0.96. Furthermore, the evaluation done using a few curation cases in the real scenario showed the effectiveness of our work. Therefore, we conclude that RelCurator can identify not only new causative genes, but also new genes associated with neurodegenerative disorders' phenotype.ConclusionRelCurator is a user-friendly method for accessing deep learning-based supporting information and a concise web interface to assist curators while browsing the PubMed articles. Our curation process represents an important and broadly applicable improvement to the state of the art for the curation of gene-phenotype relationships.
RNA decay is an important regulatory mechanism for gene expression at the posttranscriptional level. Although the main pathways and major enzymes that facilitate this process are well defined, global analysis of RNA turnover remains under-investigated. Recent advances in the application of next-generation sequencing technology enable its use in order to examine various RNA decay patterns at the genome-wide scale. In this study, we investigated human RNA decay patterns using parallel analysis of RNA end-sequencing (PARE-seq) data from XRN1-knockdown HeLa cell lines, followed by a comparison of steady state and degraded mRNA levels from RNA-seq and PARE-seq data, respectively. The results revealed 1103 and 1347 transcripts classified as stable and unstable candidates, respectively. Of the unstable candidates, we found that a subset of the replication-dependent histone transcripts was polyadenylated and rapidly degraded. Additionally, we identified 380 endonucleolytically cleaved candidates by analyzing the most abundant PARE sequence on a transcript. Of these, 41.4% of genes were classified as unstable genes, which implied that their endonucleolytic cleavage might affect their mRNA stability. Furthermore, we identified 1877 decapped candidates, including HSP90B1 and SWI5, having the most abundant PARE sequences at the 5′-end positions of the transcripts. These results provide a useful resource for further analysis of RNA decay patterns in human cells.
The advent of high-throughput Next-Generation Sequencing (NGS) technology has brought in a new genomics era. It is widely utilised not only for genomic but also for transcriptome sequencing. NGS-based transcriptomic methods, including RNA-seq, are widely used to study gene expression regulation. Recently, post-transcriptional gene regulation has been studied by several RNA degradome sequencing methods that differ from RNA-seq and directly profile degraded RNAs. Parallel Analysis of RNA Ends (PARE) sequencing is one of these methods. It is widely used for identifying microRNA targets and nonsense-mediated mRNA decay events as well as for global analysis of RNA degradation. We have developed WebPORD, a user-friendly web-based pipeline for analysing the PARE sequencing data. This pipeline trims and applies several filters to PARE sequencing. In addition, this pipeline provides a degradome table and degradome-plots for users to further analyse RNA degradation at a genomic scale. This is the first web-based pipeline for global analysis of RNA degradation using PARE data from animals and plants, which provides mechanistic insights into post-transcriptional gene regulation.
The Earth Mover's Distance (EMD) is one of the most-widely used distance functions to measure the similarity between two multimedia objects. While providing good search results, the EMD is too much time consuming to be used in large multimedia databases. To solve the problem, we propose an approximate k-nearest neighbor (k-NN) search method based on the EMD. In the proposed method, the overhead for both disk accesses and EMD computations is reduced significantly, thanks to the approximation. First, the proposed method builds an index using the M-tree, a distance-based multi-dimensional index structure, to reduce the disk access overhead. When building the index, we reduce the number of features in the multimedia objects through dimensionalityreduction. When performing the k-NN search on the M-tree, we find a small set of candidates from the disk using the index and then perform the post-processing on them. Second, the proposed method uses the approximate EMD for index retrieval and post-processing to reduce the computational overhead of the EMD. To compensate the errors due to the approximation, the method provides a way of accuracy improvement of the approximate EMD. We performed extensive experiments to show the efficiency of the proposed method. As a result, the method achieves significant improvement in performance with only small errors: the proposed method outperforms the previous method by up to 67.3% with only 3.5% error.
MicroRNAs (miRNA) of approximately 22 nucleotides are fundamental molecules in cellular biology that act by modulating target gene expression. As a post-transcriptional/translational regulator, the role of miRNAs in human disease is highly expected. Recently, a rising number of miRNAs that are associated with a variety of human diseases have been reported in the literature. Aiming to provide a comprehensive of resource of miRNAs that are related to the human disease, we built the Disease miRNA Database (Disease-miRNAdb). MiRNAs from three groups of diseases including kidney disease, cancer, and metabolic disease were manually curated in the current version of database located on a user-friendly online website ( http://disease-miRNAdb.hallym.ac.kr ). In addition, this database also includes information regarding single nucleotide polymorphisms (SNPs) located in pre-miRNAs, miRNA flanking regions, and 3′-UTRs of miRNA target genes. The generation of genetic variation in these regions may result in a loss or gain of miRNA-target gene interactions, which may influence the development of disease. Thus, information regarding miRNA-related SNPs will be useful in the discovery of disease-associated miRNAs in human populations by employing SNP association analysis for the disease of interest. Disease-miRNAdb will be continually maintained with up-to-date information to provide a current and valuable resource for investigating the roles of miRNAs in human disease.
1Research Center of Information and Electronic Engineering, Hallym University, Korea, Hallymdaehak-gil, Chuncheon, Gangwon-do 200-702, Korea E-mail: jiwon@hallym.ac.kr 2Division of Information Engineering and Telecommunications, Hallym University, Korea E-mail: {backhyun@sysgate.co.kr, jhyoon@hallym.ac.kr} 3Department of Computer Science, Yonsei University, Korea. E-mail:sanghyun@cs.yonsei.ac.kr 4College of Information and Communications, Hanyang University, Korea 17 Haengdang-dong, Seongdong-gu, Seoul 133-791, Korea E-mail: wook@hanyang.ac.kr
Over the past several decades, biologists have conducted numerous studies examining both general and specific functions of proteins. Generally, if similarities in either the structure or sequence of amino acids exist for two proteins, then a common biological function is expected. Protein function is determined primarily based on the structure rather than the sequence of amino acids. The algorithm for protein structure alignment is an essential tool for the research. The quality of the algorithm depends on the quality of the similarity measure that is used, and the similarity measure is an objective function used to determine the best alignment. However, none of existing similarity measures became golden standard because of their individual strength and weakness. They require excessive filtering to find a single alignment. In this paper, we introduce a new strategy that finds not a single alignment, but multiple alignments with different lengths. This method has obvious benefits of high quality alignment. However, this novel method leads to a new problem that the running time for this method is considerably longer than that for methods that find only a single alignment. To address this problem, we propose algorithms that can locate a common region (CORE) of multiple alignment candidates, and can then extend the CORE into multiple alignments. Because the CORE can be defined from a final alignment, we introduce CORE* that is similar to CORE and propose an algorithm to identify the CORE*. By adopting CORE* and dynamic programming, our proposed method produces multiple alignments of various lengths with higher accuracy than previous methods. In the experiments, the alignments identified by our algorithm are longer than those obtained by TM-align by 17 % and 15.48 %, on average, when the comparison is conducted at the level of super-family and fold, respectively.
Generating a shader is an important factor for portraying realistic objects on the screen using 3D computer graphics. However, since generating a shader requires significant amounts of time and effort, the user must be able to search for a desired shader from a large shader database. This paper proposes a novel hierarchical clustering algorithm that structuralizes a shader database by clustering similar shaders in order to facilitate the shader search process. Since conventional hierarchical clustering methods do not take into account the cluster diameter and the number of shaders in a cluster, the resulting databases structure is highly likely to be a skewed tree. The proposed method represents a shader database into a graph and recursively uses a graph partitioning algorithm to perform hierarchical clustering. In this process, since the cluster diameter and the number of shaders in a cluster are taken into account preventing a single cluster from becoming excessively large, we can construct a hierarchical database structure close to a balanced state. Various experiments were conducted to verify the merits of the proposed method. The experimental results indicate that the proposed method provides a high level of accuracy as well as a structure close to a balanced state.
Embryo is a very early stage of the development of multicellular organism such as animals and plants. It is an important research target for studying ontogeny because the fundamental body system of multicellular organism is determined during an embryo state. Researchers in the developmental biology have a large volume of embryo image databases for studying embryos and they frequently search for an embryo image efficiently from those databases. Thus, it is crucial to organize databases for their efficient search. Hierarchical clustering methods have been widely used for database organization. However, most of previous algorithms tend to produce a highly skewed tree as a result of clustering because they do not simultaneously consider both the size of a cluster and the number of objects within the cluster. The skewed tree requires much time to be traversed in users' search process. In this paper, we propose a method that effectively organizes a large volume of embryo image data in a balanced tree structure. We first represent embryo image data as a similarity-based graph. Next, we identify clusters by performing a graph partitioning algorithm repeatedly. We check constantly the size of a cluster and the number of objects, and partition clusters whose size is too large or whose number of objects is too high, which prevents clusters from growing too large or having too many objects. We show the superiority of the proposed method by extensive experiments. Moreover, we implement the visualization tool to help users quickly and easily navigate the embryo image database.
Alternative splicing is a main source of generating a highly dynamic human proteome. It enables exons of genes to recombine in different ways when a segment of DNA is transcribed into messenger RNAs (mRNAs). There are seven patterns of alternative splicing which are known as typical ways in human genes. Alternative splicing has been found to be associated with many human diseases. This work proposes a novel method to identify alternative splicing patterns and analyze the distribution variations in a given gene expression process from mRNA read sequence data generated by next-generation sequencing (NGS) technology. The proposed method is based on parallel mapping of mRNA read sequence data to both genomic and transcriptomic reference sequences, and analyzes the mRNA read sequence data coverage and junctions spanning two exons. The preliminary results conducted with simulated data showed the prominent efficiency in identifying and quantifying events of alternative splicing.
A spatial join is a query that searches for a set of object pairs satisfying a given spatial relationship from a database. It is one of the most costly queries, and thus requires an efficient processing algorithm that fully exploits the features of the underlying spatial indexes. In our earlier work, we devised a fairly effective algorithm for processing spatial joins with double transformation (DOT) indexing, which is one of several spatial indexing schemes. However, the algorithm is restricted to only the one-dimensional cases. In this paper, we extend the algorithm for the two-dimensional cases, which are general in Geographic Information Systems (GIS) applications. We first extend DOT to two-dimensional original space. Next, we propose an efficient algorithm for processing range queries using extended DOT. This algorithm employs the quarter division technique and the tri-quarter division technique devised by analyzing the regularity of the space-filling curve used in DOT. This greatly reduces the number of space transformation operations. We then propose a novel spatial join algorithm based on this range query processing algorithm. In processing a spatial join, we determine the access order of disk pages so that we can minimize the number of disk accesses. We show the superiority of the proposed method by extensive experiments using data sets of various distributions and sizes. The experimental results reveal that the proposed method improves the performance of spatial join processing up to three times in comparison with the widely-used R-tree-based spatial join method.
This paper proposes a new trajectory clustering scheme for objects moving on road networks. A trajectory on road networks can be defined as a sequence of road segments a moving object has passed by. We first propose a similarity measurement scheme that judges the degree of similarity by considering the total length of matched road segments. Then, we propose a new clustering algorithm based on such similarity measurement criteria by modifying and adjusting the FastMap and hierarchical clustering schemes. To evaluate the performance of the proposed clustering scheme, we also develop a trajectory generator considering the fact that most objects tend to move from the starting point to the destination point along their shortest path. The performance result shows that our scheme has the accuracy of over 95%.
This paper addresses an approach that recommends investment types to stock investors by discovering useful rules from past changing patterns of stock prices in databases. First, we define a new rule model for recommending stock investment types. For a frequent pattern of stock prices, if its subsequent stock prices are matched to a condition of an investor, the model recommends a corresponding investment type for this stock. The frequent pattern is regarded as a rule head, and the subsequent part a rule body. We observed that the conditions on rule bodies are quite different depending on dispositions of investors while rule heads are independent of characteristics of investors in most cases. With this observation, we propose a new method that discovers and stores only the rule heads rather than the whole rules in a rule discovery process. This allows investors to impose various conditions on rule bodies flexibly, and also improves the performance of a rule discovery process by reducing the number of rules to be discovered. For efficient discovery and matching of rules, we propose methods for discovering frequent patterns, constructing a frequent pattern base, and its indexing. We also suggest a method that finds the rules matched to a query from a frequent pattern base, and a method that recommends an investment type by using the rules. Finally, we verify the effectiveness and the efficiency of our approach through extensive experiments with real-life stock data.
This paper addresses an approach that recommends investment types for stock investors by discovering useful rules from past changing patterns of stock prices in databases. First, we define a new rule model for recommending stock investment types. For a frequent pattern of stock prices, if its subsequent stock prices are matched to a condition of an investor, the model recommends a corresponding investment type for this stock. The frequent pattern is regarded as a rule head, and the subsequent part a rule body. We observed that the conditions on rule bodies are quite different depending on dispositions of investors while rule heads are independent of characteristics of investors in most cases. With this observation, we propose a new method that discovers and stores only the rule heads rather than the whole rules in a rule discovery process. This allows investors to define various conditions on rule bodies flexibly, and also improves the performance of a rule discovery process by reducing the number of rules. For efficient discovery and matching of rules, we propose methods for discovering frequent patterns, constructing a frequent pattern base, and indexing them. We also suggest a method that finds the rules matched to a query issued by an investor from a frequent pattern base, and a method that recommends an investment type using the rules. Finally, we verify the superiority of our approach via various experiments using real-life stock data.
The choice of an effective indexing method is crucial to guarantee the performance of the spatial join operator which is heavily used in geographical information systems. The -tree based method is renowned as one of the most representative indexing methods. In this paper, we propose an efficient spatial join technique based on the DOT(Double Transformation) index, and compare it with the spatial Join technique based on the -tree index. The DOT index transforms the MBR of an spatial object into a single numeric value using a space filling curve, and builds the -tree from a set of numeric values transformed as such. The DOT index is possible to be employed as a primary index for spatial objects. The proposed spatial join technique exploits the regularities in the moving patterns of space filling curves to divide a query region into a set of maximal sub-regions within which space filling curves traverse without interruption. Such division reduces the number of spatial transformations required to perform the spatial join and thus improves the performance of join processing. The experiments with the data sets of various distributions and sizes revealed that the proposed join technique is up to three times faster than the spatial join method based on the -tree index.