The number of annual scientific publications is growing year by year, which has led to the accumulation and formation of large databases. This increases the complexity of the search for relevant articles. Modern search engines leverage keyphrases to improve the performance of search results. Keyphrases (or keywords) are a set of single or multi-word expressions which provide a very compact summary of contents and describe the overall topic of a document. Implementation of keyphrase extraction can be varied for specific-domain databases. This motivated us to evaluate these methods on a database of influenza scientific literature. In this work, we considered 9 well known methods for extracting keyphrases. Our preliminary results show txhat graph-based methods which employ topics outperform others using our database. We also determined that influenza-related papers can be grouped into three general topics: public health & medical care; molecular biology & immunology; and phylogenetic & epidemiological studies.
Assessment of antigenic similarity between strains of the influenza virus is a crucial factor when planning vaccine compositions. To perform this, a gold-standard laboratory procedure, hemagglutination inhibition assay, is conventionally used. Despite its theoretical importance and accuracy, this procedure suffers from several shortcomings, including high time consumption. Therefore, various computer-aided and mathematical methods have been developed to acquire earlier knowledge on the antigenic characteristics of currently circulating viruses. In this paper, we introduce a state-of-the-art ensemble artificial neural network model based on features derived from multi-representation of antigenicity. Generally, each feature is generated from an optimized convolutional neural network whose input describes the genetic difference between viruses in a specific numerical space. The space is determined based on embedding of the genetic sequence by a reduced amino acid alphabet and the Word2Vec framework. Our experiments indicated that the proposed model outperformed approaches from the literature by achieving an accuracy level of 0.933 for the HINI subtype. This implies possible application of our model as a promising exploratory tool in practical tasks of virus control.
Analysis of viral evolution is a key element of epidemiological surveillance and control. One of the fundamental tools which is widely used to illustrate evolutionary history is the phylogenetic tree. Recently, we have proposed an alternative visualization for the phylogenetic tree using the evolutionary trajectory of its taxa. An evolutionary trajectory is a path starting from a taxon and ending at the root of the tree. In this paper, we propose an embedding of tree nodes by encoding their genetic sequence using a reduced amino acid alphabet and employing the Word2Vec framework. The suggested visualization maintains the phylogenetic relationship between nodes, while their proximity in 3D space depends on three factors: the type of reduced amino acid alphabet; fixed-length genetic patterns used in Word2Vec; and the neighbor effect of adjacent signatures. The results of our experiments showed that the majority of evolutionary history can be described in the embedded space. Moreover, they suggest potential application of our approach as an explanatory tool in studying various aspects: evolutionary dynamics; evolutionary deviation of viral variants; and phylogenetic characteristics, such as formation of new clades. Besides the usual local analysis of point mutations, the developed framework enables studying these aspects based on a more comprehensive global context, including neighboring effects, genetic signatures.