Background and Objective:differential expression analysis is one of the most popular activities in transcriptomic studies based on next-generation sequencing technologies. In fact, differentially expressed genes (DEGs) between two conditions represent ideal prognostic and diagnostic candidate biomarkers for many pathologies. As a result, several algorithms, such as DESeq2 and edgeR, have been developed to identify DEGs. Despite their widespread use, there is no consensus on which model performs best for different types of data, and many existing methods suffer from high False Discovery Rates (FDR).Methods:we present a new algorithm, DeClUt, based on the intuition that the expression profile of differentially expressed genes should form two reasonably compact and well-separated clusters. This, in turn, implies that the bipartition induced by the two conditions being compared should overlap with the clustering. The clustering algorithm underlying DeClUt was designed to be robust to outliers typical of RNA-seq data. In particular, we used the average silhouette function to enforce membership assignment of samples to the most appropriate condition.Results:DeClUt was tested on real RNA-seq datasets and benchmarked against four of the most widely used methods (edgeR, DESeq2, NOISeq, and SAMseq). Experiments showed a higher self-consistency of results than the competitors as well as a significantly lower False Positive Rate (FPR). Moreover, tested on a real prostate cancer RNA-seq dataset, DeClUt has highlighted 8 DE genes, linked to neoplastic process according to DisGeNET database, that none of the other methods had identified.Conclusions:our work presents a novel algorithm that builds upon basic concepts of data clustering and exhibits greater consistency and significantly lower False Positive Rate than state-of-the-art methods. Additionally, DeClUt is able to highlight relevant differentially expressed genes not otherwise identified by other tools contributing to improve efficacy of differential expression analyses in various biological applications.
In this work, we introduce a similarity-network-based approach to explore the role of interacting single-cell histone modification signals in haematopoiesis—the process of differentiation of blood cells. Histones are proteins that provide structural support to chromosomes. They are subject to chemical modifications—acetylation or methylation—that affect the degree of accessibility of genes and, in turn, the formation of different phenotypes. The concentration of histone modifications can be modelled as a continuous signal, which can be used to build single-cell profiles. In the present work, the profiles of cell types involved in haematopoiesis are built based on all the major histone modifications (i.e., H3K27ac, H3K27me3, H3K36me3, H3K4me1, H3K4me3, H3K9me3) by counting the number of peaks in the modification signals; then, the profiles are used to compute modification-specific similarity networks among the considered phenotypes. As histone modifications come as interacting signals, we applied a similarity network fusion technique to integrate these networks in a unique graph, with the aim of studying the simultaneous effect of all the modifications for the determination of different phenotypes. The networks permit defining of a graph-cut-based separation score for evaluating the homogeneity of subgroups of cell types corresponding to the myeloid and lymphoid phenotypes in the classical representation of the haematopoietic tree. Resulting scores show that separation into myeloid and lymphoid phenotypes reflects the actual process of haematopoiesis.
MicroRNAs (miRNAs) are short non-coding RNAs engaged in cellular regulation by suppressing genes at their post-transcriptional stage. Evidence of their involvement in breast cancer and the possibility of quantifying the their concentration in the blood has sparked the hope of using them as reliable, inexpensive and non-invasive biomarkers.
Hypergraphs and simplical complexes both capture the higher-order interactions of complex systems, ranging from higher-order collaboration networks to brain networks. One open problem in the field is what should drive the choice of the adopted mathematical framework to describe higher-order networks starting from data of higher-order interactions. Unweighted simplicial complexes typically involve a loss of information of the data, though having the benefit to capture the higher-order topology of the data. In this work we show that weighted simplicial complexes allow one to circumvent all the limitations of unweighted simplicial complexes to represent higher-order interactions. In particular, weighted simplicial complexes can represent higher-order networks without loss of information, allowing one at the same time to capture the weighted topology of the data. The higher-order topology is probed by studying the spectral properties of suitably defined weighted Hodge Laplacians displaying a normalized spectrum. The higher-order spectrum of (weighted) normalized Hodge Laplacians is studied combining cohomology theory with information theory. In the proposed framework we quantify and compare the information content of higher-order spectra of different dimension using higher-order spectral entropies and spectral relative entropies. The proposed methodology is tested on real higher-order collaboration networks and on the weighted version of the simplicial complex model "Network Geometry with Flavor."
Objectives: Dilated cardiomyopathy (DCM) is characterized by a specific transcriptome. Since the DCM molecular network is largely unknown, the aim was to identify specific disease-related molecular targets combining an original machine learning (ML) approach with protein-protein interaction network. Methods: The transcriptomic profiles of human myocardial tissues were investigated integrating an original computational approach, based on the Custom Decision Tree algorithm, in a differential expression bioinformatic framework. Validation was performed by quantitative real-time PCR. Results: Our preliminary study, using samples from transplanted tissues, allowed the discovery of specific DCM-related genes, including MYH6, NPPA, MT-RNR1 and NEAT1, already known to be involved in cardiomyopathies Interestingly, a combination of these expression profiles with clinical characteristics showed a significant association between NEAT1 and left ventricular end-diastolic diameter (LVEDD) (Rho = 0.73, p = 0.05), according to severity classification (NYHA-class III). Conclusions: The use of the ML approach was useful to discover preliminary specific genes that could lead to a rapid selection of molecular targets correlated with DCM clinical parameters. For the first time, NEAT1 under-expression was significantly associated with LVEDD in the human heart.
MicroRNAs (miRNAs) are short endogenous molecules of RNA that influence cell regulation by suppressing genes. Their ubiquity throughout all branches of the tree of life has suggested their central role in many cellular functions. Nowadays, several personalized medicine applications rely on miRNAs as biomarkers for diagnoses, prognoses, and prediction of drug response. The increasing ease of sequencing miRNAs contrasts with the difficulty of accurately quantifying their concentration. The use of general purpose aligners is only a partial solution as they have limited possibilities to accurately solve ambiguous mapping due to the short length of these sequences. We developed E Z c o u n t, an all-in-one software that, with a single command, performs the entire quantification process: from raw fastq files to read counts. Experiments show that E Z c o u n t is more sensitive and accurate than methods based on sequence alignment, independently of the library preparation protocol and sequencing machine. The parallel architecture of E Z c o u n t makes it fast enough to process a sample in minutes using a standard workstation. E Z c o u n t runs on all of the most common operating systems (Linux, Windows and MacOS) and is freely available for download at https://gitlab.com/BioAlgo/miR-pipe . A detailed description of the datasets, the raw experimental results, and all the scripts used for testing are available as supplementary material. Display Omitted • EZcount searches miRNAs in the reads, instead of aligning reads to the reference. • EZcount resolves ambiguous alignments using quality scores. • EZcount outperforms the existing tools in terms of properly matched reads. • EZcount self-tune without the need to know the adapter sequence or other library parameters.
We review the current applications of artificial intelligence (AI) in functional genomics. The recent explosion of AI follows the remarkable achievements made possible by "deep learning", along with a burst of "big data" that can meet its hunger. Biology is about to overthrow astronomy as the paradigmatic representative of big data producer. This has been made possible by huge advancements in the field of high throughput technologies, applied to determine how the individual components of a biological system work together to accomplish different processes. The disciplines contributing to this bulk of data are collectively known as functional genomics. They consist in studies of: i) the information contained in the DNA (genomics); ii) the modifications that DNA can reversibly undergo (epigenomics); iii) the RNA transcripts originated by a genome (transcriptomics); iv) the ensemble of chemical modifications decorating different types of RNA transcripts (epitranscriptomics); v) the products of protein-coding transcripts (proteomics); and vi) the small molecules produced from cell metabolism (metabolomics) present in an organism or system at a given time, in physiological or pathological conditions. After reviewing main applications of AI in functional genomics, we discuss important accompanying issues, including ethical, legal and economic issues and the importance of explainability.
MicroRNAs (miRNAs) are short endogenous molecules of RNA that influence cell regulation by suppressing genes. Their ubiquity throughout all branches of the tree of life has suggested their central role in many cellular functions. Nowadays, several personalized medicine applications rely on miRNAs as biomarkers for diagnoses, prognoses, and prediction of drug response. The increasing ease of sequencing miRNAs contrasts with the difficulty of accurately quantifying their concentration. The use of general purpose aligners is only a partial solution as they have limited possibilities to accurately solve ambiguous mapping due to the short length of these sequences. We developed EZcount, an all-in-one software that, with a single command, performs the entire quantification process: from raw fastq files to read counts. Experiments show that EZcount is more sensitive and accurate than methods based on sequence alignment, independently of the library preparation protocol and sequencing machine. The parallel architecture of EZcount makes it fast enough to process a sample in minutes using a standard workstation. EZcount runs on all of the most common operating systems (Linux, Windows and MacOS) and is freely available for download at https://gitlab.com/BioAlgo/miR-pipe. A detailed description of the datasets, the raw experimental results, and all the scripts used for testing are available as supplementary material.
Frontiers is more than just an open-access publisher of scholarly articles: it is a pioneering approach to the world of academia, radically improving the way scholarly research is managed.The grand vision of Frontiers is a world where all people have an equal opportunity to seek, share and generate knowledge.Frontiers provides immediate and permanent online open access to all its publications, but this alone is not enough to realize our grand goals. Frontiers Journal SeriesThe Frontiers Journal Series is a multi-tier and interdisciplinary set of open-access, online journals, promising a paradigm shift from the current review, selection and dissemination processes in academic publishing.All Frontiers journals are driven by researchers for researchers; therefore, they constitute a service to the scholarly community.At the same time, the Frontiers Journal Series operates on a revolutionary invention, the tiered publishing system, initially addressing specific communities of scholars, and gradually climbing up to broader public understanding, thus serving the interests of the lay society, too. Dedication to QualityEach Frontiers article is a landmark of the highest quality, thanks to genuinely collaborative interactions between authors and review editors, who include some of the world's best academicians.Research must be certified by peers before entering a stream of knowledge that may eventually reach the public -and shape society; therefore, Frontiers only applies the most rigorous and unbiased reviews.Frontiers revolutionizes research publishing by freely delivering the most outstanding research, evaluated with no bias from both the academic and social point of view.By applying the most advanced information technologies, Frontiers is catapulting scholarly publishing into a new generation.
In this paper, we show that quantifying histone modifications by counting the number of high– resolution peaks in each gene allows to build profiles of these epigenetic marks, associating them to a phenotype. The significance of this approach is verified by applying graph–cut techniques for assessing the differentiation between myeloid and lymphoid cells in haematopoiesis, i.e. the process through which all the different types of blood cells originate starting from a unique cell type. The experiments are conducted on a population of samples from 24 cell types involved in haematopoiesis. Six profiles are constructed for each cell type, based on a different histone modification signal. Following the experimentally verified idea that the peak number distribution per gene behaves similarly to gene expression, the profile computation employs standard differential analysis tools to find genes whose epigenetic modifications are related to a given phenotype. Next, six similarity networks of cell types are constructed, based on each histone modification, and then combined into a unique one through similarity network fusion. Finally, the similarity networks are transformed into dissimilarity graphs, to which two different cuts are applied and compared to evaluate the classic differentiation between myeloid and lymphoid cells. The results show that all histone modifications contribute almost equally to the myeloid/lymphoid differentiation, and this is also confirmed by the analysis of the fused network. However, they also suggest that histone modifications may not be the only mechanism for regulating the differentiation of hematopoietic cells.
MOTIVATION:Large-scale sequencing projects have confirmed the hypothesis that eukaryotic DNA is rich in repetitions whose functional role needs to be elucidated. In particular, tandem repeats (TRs) (i.e. short, almost identical sequences that lie adjacent to each other) have been associated to many cellular processes and, indeed, are also involved in several genetic disorders. The need of comprehensive lists of TRs for association studies and the absence of a computational model able to capture their variability have revived research on discovery algorithms. RESULTS:Building upon the idea that sequence similarities can be easily displayed using graphical methods, we formalized the structure that TRs induce in dot-plot matrices where a sequence is compared with itself. Leveraging on the observation that a compact representation of these matrices can be built and searched in linear time, we developed Dot2dot: an accurate algorithm fast enough to be suitable for whole-genome discovery of TRs. Experiments on five manually curated collections of TRs have shown that Dot2dot is more accurate than other established methods, and completes the analysis of the biggest known reference genome in about one day on a standard PC. AVAILABILITY AND IMPLEMENTATION:Source code and datasets are freely available upon paper acceptance at the URL: https://github.com/Gege7177/Dot2dot. SUPPLEMENTARY INFORMATION:Supplementary data are available at Bioinformatics online.
Highly accurate genotyping is essential for genomic projects aimed at understanding the etiology of diseases as well as for routinary screening of patients. For this reason, genotyping software packages are subject to a strict validation process that requires a large amount of sequencing data endowed with accurate genotype information. In-vitro assessment of genotyping is a long, complex and expensive activity that also depends on the specific variation and locus, and thus it cannot really be used for validation of in-silico genotyping algorithms. In this scenario, sequencing simulation has emerged as a practical alternative. Simulators must be able to keep up with the continuous improvement of different sequencing technologies producing datasets as much indistinguishable from real ones as possible. Moreover, they must be able to mimic as many types of genomic variant as possible. In this paper we describe OmniSim: a simulator whose ultimate goal is that of being suitable in all the possible applicative scenarios. In order to fulfill this goal, OmniSim uses an abstract model where variations are read from a.vcf file and mapped into edit operations (insertion, deletion, substitution) on the reference genome. Technological parameters (e.g. error distributions, read length and per-base quality) are learned from real data. As a result of the combination of our abstract model and parameter learning module, OmniSim is able to output data in all aspects similar to that produced in a real sequencing experiment. The source code of OmniSim is freely available at the URL: https://gitlab.com/geraci/omnisim
miRandola (http://mirandola. iit. cnr. it/) is a database of extracellular non-coding RNAs (ncRNAs) that was initially published in 2012, foreseeing the relevance of ncRNAs as non-invasive biomarkers. An increasing amount of experimental evidence shows that ncRNAs are frequently dysregulated in diseases. Further, ncRNAs have been discovered in different extracellular forms, such as exosomes, which circulate in human body fluids. Thus, miRandola 2017 is an effort to update and collect the accumulating information on extracellular ncRNAs that is spread across scientific publications and different databases. Data are manually curated from 314 articles that describe miRNAs, long non-coding RNAs and circular RNAs. Fourteen organisms are now included in the database, and associations of ncRNAs with 25 drugs, 47 sample types and 197 diseases. miRandola also classifies extracellular RNAs based on their extracellular form: Argonaute2 protein, exosome, microvesicle, microparticle, membrane vesicle, high density lipoprotein and circulating. We also implemented a new web interface to improve the user experience.
Manuela Montangero合作论文数Dipartimento di Ingegneria dell'Informazione6
Fabrizio Sebastiani合作论文数Networked Multimedia Information Access Laboratory, Institute for the Science and Technologies of Information, Italian National Council of Research3
Mauro Leoncini合作论文数Dipartimento di Ingegneria dell'Informazione;Universit?? di Modena e Reggio Emilia3