Delineating the natural variations in the Chinese population is a crucial initial step towards the establishment of precision medi-cine.This comprehensive characterization will not only elucidate the population's unique genetic diversity but also facilitate the identification of variants unique to the population.
Nanopore sequence technology has demonstrated a longer read length and enabled to potentially address the limitations of short-read sequencing including long-range haplotype phasing and accurate variant calling. However, there is still room for improvement in terms of the performance of single nucleotide variant (SNV) identification and computing resource usage for the state-of-the-art approaches. In this work, we introduce miniSNV, a lightweight SNV calling algorithm that simultaneously achieves high performance and yield. miniSNV utilizes known common variants in populations as variation backgrounds and leverages read pileup, read-based phasing, and consensus generation to identify and genotype SNVs for Oxford Nanopore Technologies (ONT) long reads. Benchmarks on real and simulated ONT data under various error profiles demonstrate that miniSNV has superior sensitivity and comparable accuracy on SNV detection and runs faster with outstanding scalability and lower memory than most state-of-the-art variant callers. miniSNV is available from https://github.com/CuiMiao-HIT/miniSNV.
Non-coding RNAs (ncRNAs) participate in multiple biological processes associated with cancer as tumor suppressors or oncogenic drivers. Due to their high stability in plasma, urine, and many other fluids, ncRNAs have the potential to serve as key biomarkers for early diagnosis and screening of cancers. During cancer progression, tumor heterogeneity plays a crucial role, and it is particularly important to understand the gene expression patterns of individual cells. With the development of single-cell RNA sequencing (scRNA-seq) technologies, uncovering gene expression in different cell types for human cancers has become feasible by profiling transcriptomes at the cellular level. However, a well-organized and comprehensive online resource that provides access to the expression of genes corresponding to ncRNA biomarkers in different cell types at the single cell level is not available yet. Therefore, we developed the SCancerRNA database to summarize experimentally supported data on long ncRNA, microRNA, PIWI-interacting RNA, small nucleolar RNA, and circular RNA biomarkers, as well as data on their differential expression at the cellular level. Furthermore, we collected biological functions and clinical applications of biomarkers to facilitate the application of ncRNA biomarkers to cancer diagnosis, as well as monitoring of progression and targeted therapies. SCancerRNA also allows users to explore interaction networks of different types of ncRNAs, and build computational models in the future. SCancerRNA is freely accessible at http://www.scancerrna.com/BioMarker.
Understanding how genetic variants influence molecular phenotypes in different cellular contexts is crucial for elucidating the molecular and cellular mechanisms behind complex traits, which in turn has spurred significant advances in research into molecular quantitative trait locus (xQTL) at the cellular level. With the rapid proliferation of data, there is a critical need for a comprehensive and accessible platform to integrate this information. To meet this need, we developed xQTLatlas (http://www.hitxqtl.org.cn/), a database that provides a multi-omics genetic regulatory landscape at cellular resolution. xQTLatlas compiles xQTL summary statistics from 151 cell types and 339 cell states across 55 human tissues. It organizes these data into 20 xQTL types, based on four distinct discovery strategies, and spans 13 molecular phenotypes. Each entry in xQTLatlas is meticulously annotated with comprehensive metadata, including the origin of the tissue, cell type, cell state and the QTL discovery strategies utilized. Additionally, xQTLatlas features multiscale data exploration tools and a suite of interactive visualizations, facilitating in-depth analysis of cell-level xQTL. xQTLatlas provides a valuable resource for deepening our understanding of the impact of functional variants on molecular phenotypes in different cellular environments, thereby facilitating extensive research efforts. Graphical Abstract
Long-read RNA sequencing (RNA-seq) has revolutionized our ability to comprehensively study transcriptomes, enabling the detection of full-length transcripts. The accurate alignment of long-read RNA-seq is a critical step in downstream analysis. However, with the development of alignment tools and each tool declaring the ability to handle long-read RNA-seq alignment on their own, researchers face the challenge of distinguishing and selecting the most suitable tool for their tasks. But to the best of our knowledge, there is still a lack of comprehensive evaluations on the aligners designed for long-read sequencing data. Here, we conducted a benchmark on five state-of-the-art tools on simulated long-read RNA-seq data from three levels to comprehensively assess the performance of each aligner, and provide valuable insights into the strengths and limitations of each aligner, aiding researchers in selecting the most appropriate tool for their long-read RNA-seq studies. We expected our study to advance the field of transcriptomics and promote the accurate analysis of long-read RNA-seq data in diverse biological contexts.
Individuals have different symptoms of COVID-19 infection, which is not only caused by environmental and socioeconomic factors. Many studies have pointed out that it is closely related to genetic factors. We read with interest the publication by Li et al.1 reporting the relationship between Interferon-induced transmembrane protein 3 gene (IFIM3) polymorphisms and susceptibility and severity of COVID-19. Through a genotype study of 1989 samples, they found that the rs12252 of IFIM3 polymorphism was associated with COVID-19 susceptibility and severity, while rs34481144 was not.
Motivation: Single-cell RNA sequencing (scRNA-seq) technologies allow us to interrogate the state of an individual cell within its microenvironment. However, prior to sequencing, cells should be dissociated first, making it difficult to obtain their spatial information. Since the spatial distribution of cells is critical in a few cir-cumstances such as cancer immunotherapy, we present MLSpatial, a novel computational method to learn the relationship between gene expression patterns and spatial locations of cells, and then predict cell-to-cell distance distribution based on scRNA-seq data alone.Results: We collected the drosophila embryo dataset, which contains both the fluorescence in situ hybridization (FISH) data and single cell RNA-seq (scRNA-seq) data of drosophila embryo. The FISH data provided the spatial position of 3039 cells and the expression of 84 genes for each cell. The scRNA-seq data contains the expressions of 8924 genes in 1297 high-quality cells with cell location unknown. For a comparison, we also collected the MERFISH data of 645 osteosarcoma cells with cell location and the expression status of 10,050 genes known. For each data, the cells were randomly divided into a training set and a test set, in the ratio of 7:3. The cell-to-cell distances our model extracted had a higher correspondence (i.e., correlation coefficient 0.99) with those of the real situation than those of existing methods in the FISH data of drosophila embryo. However, in the osteosarcoma data, our model captured the spatial relationship between cells, with a correlation of 0.514 to that of the real situation. We also applied the model trained using the FISH data of drosophila embryo into the single cell data of drosophila embryo, for which the real location of cells are unknown. The reconstructed pseudo drosophila embryo and the real embryo (as shown by the FISH data) had a high similarity in the spatial distribution of gene expression.Conclusion: MLSpatial can accurately restore the relative position of cells from scRNA-seq data; however, the performance depends on the type of cells. The trained model might be useful in reconstructing the spatial distributions of single cells with only scRNA-seq data, provided that the scRNA-seq data and the FISH data are under similar background (i.e., the same tissue with similar disease background).
Cancer development and progression are significantly influenced by cancer driver genes. Understanding cancer driver genes and their mechanisms of action is essential for developing effective cancer treatments. As a result, identifying driver genes is important for drug development, cancer diagnosis, and treatment. Here, we present an algorithm to discover driver genes based on the two-stage random walk with restart (RWR), and the modified method for calculating the transition probability matrix in random walk algorithm. First, we performed the first stage of RWR on the whole gene interaction network, in which we employ a new method for calculating the transition probability matrix and extracted the subnetwork based on nodes that had a high correlation with the seed nodes. The subnetwork was then applied to the second stage of RWR and the nodes were re-ranked in the subnetwork. Our approach outperformed existing methods in identifying driver genes. The outcome of the effect of three gene interaction networks, two rounds of random walk, and the seed nodes' sensitivity were all compared at the same time. In addition, we identified several potential driver genes, some of which are involved in driving cancer development. Overall, our method is efficient in various cancer types, significantly outperforms existing methods, and can identify possible driver genes.
基因组注释是识别出基因组序列中功能组件的过程,其可以直接对序列赋予生物学意义,由此方便研究者探究和分析基因组功能.基因组注释可以帮助研究从三个层次上理解基因组,一种是在核苷酸水平的注释,主要确定DNA序列中基因、RNA、重复序列等组件的物理位置,包括转录起始,翻译起始,外显子边界等具体位置信息.同时可以注释得到变异在不同人群中的变异频率差异,这是解读不同人群表型差异的图谱基础.第二种是蛋白水平的注释,主要解读基因或变异的可能功能异常,评估变异所在基因位置、变异类型等对蛋白质改变的影响.第三种是生物学功能/过程注释,主要解读不同基因相互作用的对生物学过程和通路的影响,可以从系统生物学角度解释基因或调控元件对生命生化过程或功能的影响.自从人类基因组计划完成之后,各国陆续启动了基因组测序计划,完成绘制了人类基因详尽的基因多态性谱图,记录了不同表型群体的变异分布和频率差异情况等注释信息.我们结合已有的注释数据库知识,开发了具有高准确性和高效的面向大规模人群的基因组注释系统,实现对大规模的人群变异数据进行全自动化的功能性注释分析计算,进一步助力未来人群遗传变异分布等方面的研究.
BACKGROUND:The technologies advances of single-cell Assay for Transposase Accessible Chromatin using sequencing (scATAC-seq) allowed to generate thousands of single cells in a relatively easy and economic manner and it is rapidly advancing the understanding of the cellular composition of complex organisms and tissues. The data structure and feature in scRNA-seq is similar to that in scATAC-seq, therefore, it's encouraged to identify and classify the cell types in scATAC-seq through traditional supervised machine learning methods, which are proved reliable in scRNA-seq datasets. RESULTS:In this study, we evaluated the classification performance of 6 well-known machine learning methods on scATAC-seq. A total of 4 public scATAC-seq datasets vary in tissues, sizes and technologies were applied to the evaluation of the performance of the methods. We assessed these methods using a 5-folds cross validation experiment, called intra-dataset experiment, based on recall, precision and the percentage of correctly predicted cells. The results show that these methods performed well in some specific types of the cell in a specific scATAC-seq dataset, while the overall performance is not as well as that in scRNA-seq analysis. In addition, we evaluated the classification performance of these methods by training and predicting in different datasets generated from same sample, called inter-datasets experiments, which may help us to assess the performance of these methods in more realistic scenarios. CONCLUSIONS:Both in intra-dataset and in inter-dataset experiment, SVM and NMC are overall outperformed others across all 4 datasets. Thus, we recommend researchers to use SVM and NMC as the underlying classifier when developing an automatic cell-type classification method for scATAC-seq.
Abstract Background With the rapid development of long-read sequencing technologies, it is possible to reveal the full spectrum of genetic structural variation (SV). However, the expensive cost, finite read length and high sequencing error for long-read data greatly limit the widespread adoption of SV calling. Therefore, it is urgent to establish guidance concerning sequencing coverage, read length, and error rate to maintain high SV yields and to achieve the lowest cost simultaneously. Results In this study, we generated a full range of simulated error-prone long-read datasets containing various sequencing settings and comprehensively evaluated the performance of SV calling with state-of-the-art long-read SV detection methods. The benchmark results demonstrate that almost all SV callers perform better when the long-read data reach 20× coverage, 20 kbp average read length, and approximately 10–7.5% or below 1% error rates. Furthermore, high sequencing coverage is the most influential factor in promoting SV calling, while it also directly determines the expensive costs. Conclusions Based on the comprehensive evaluation results, we provide important guidelines for selecting long-read sequencing settings for efficient SV calling. We believe these recommended settings of long-read sequencing will have extraordinary guiding significance in cutting-edge genomic studies and clinical practices.
The de Bruijn graph, a fundamental data structure to represent and organize genome sequence, plays important roles in various kinds of sequence analysis tasks. With the rapid development of HTS data and ever-increasing number of assembled genomes, there is a high demand to construct the very large de Bruijn graph for sequences up to Tera-base-pair level. Current approaches may have unaffordable memory footprints to handle such a large de Bruijn graph. We propose a lightweight parallel de Bruijn graph construction approach: de Bruijn Graph Constructor in Scalable Memory (deGSM). The main idea of deGSM is to efficiently construct the Burrows-Wheeler Transformation (BWT) of the unipaths of the de Bruijn graph in constant RAM space and transform the BWT into the original unitigs. The experimental results demonstrate that, just with a commonly available machine, deGSM is able to handle very large genome sequence(s), e.g., the contigs (305 Gbp) and scaffolds (1.1 Tbp) recorded in GenBank database and Picea abies HTS dataset (9.7 Tbp). Moreover, deGSM also has faster or comparable construction speed compared with state-of-the-art approaches. With its high scalability and efficiency, deGSM has enormous potential in many large scale genomics studies. The deGSM is publicly available at: https://github.com/hitbc/deGSM.
随着高通量测序技术的快速发展和测序成本的逐渐降低,个体基因组测序已成为研究不同物种的基因型、变异情况和相关疾病的重要手段.然而,由于基因组上的大量重复序列和高变异区域,日益增大的测序数据量以及测序技术的局限等因素,如何准确且快速地将大量测序数据比对到参考基因组面临巨大挑战.阐述基于哈希思想的基因组数据的存储和索引方法.本文说明基于seed-and-extension思想的基本比对思路.本文提出一个基于de Bruijn图模型的索引结构DBG-index以及该索引的3层结构数据存储方式.分析该索引结构的特性并提出种子的基本操作方法.该索引结构利用图模型特性可以有效组织基因组上的重复序列,从而在整体上减少了候选种子数量并极大提高了比对速度.
The alignment of long-read RNA sequencing reads is non-trivial due to high sequencing errors and complicated gene structures. We propose deSALT, a tailored two-pass alignment approach, which constructs graph-based alignment skeletons to infer exons and uses them to generate spliced reference sequences to produce refined alignments. deSALT addresses several difficult technical issues, such as small exons and sequencing errors, which break through bottlenecks of long RNA-seq read alignment. Benchmarks demonstrate that deSALT has a greater ability to produce accurate and homogeneous full-length alignments. deSALT is available at: https://github.com/hitbc/deSALT .
Background Many genetic variants have been reported from sequencing projects due to decreasing experimental costs. Compared to the current typical paradigm, read mapping incorporating existing variants can improve the performance of subsequent analysis. This method is supposed to map sequencing reads efficiently to a graphical index with a reference genome and known variation to increase alignment quality and variant calling accuracy. However, storing and indexing various types of variation require costly RAM space. Methods Aligning reads to a graph model-based index including the whole set of variants is ultimately an NP-hard problem in theory. Here, we propose a variation-aware read alignment algorithm (VARA), which generates the alignment between read and multiple genomic sequences simultaneously utilizing the schema of the Landau-Vishkin algorithm. VARA dynamically extracts regional variants to construct a pseudo tree-based structure on-the-fly for seed extension without loading the whole genome variation into memory space. Results We developed the novel high-throughput sequencing read aligner deBGA-VARA by integrating VARA into deBGA. The deBGA-VARA is benchmarked both on simulated reads and the NA12878 sequencing dataset. The experimental results demonstrate that read alignment incorporating genetic variation knowledge can achieve high sensitivity and accuracy. Conclusions Due to its efficiency, VARA provides a promising solution for further improvement of variant calling while maintaining small memory footprints. The deBGA-VARA is available at: https://github.com/hitbc/deBGA-VARA .
Many genetic variants have been reported from sequencing projects due to decreasing experimental costs. Compared to the current typical paradigm, read mapping incorporating existing variants can improve the performance of subsequent analysis. However, storing and indexing various types of variation require costly RAM space. Aligning reads to a graph model-based index including the whole set of variants is ultimately an NP-hard problem in theory. This method is supposed to map sequencing reads efficiently to a graphical index with a reference genome and known variation to increase alignment quality and variant calling accuracy. Herein, we propose a variation-aware read alignment algorithm (VARA), which generates the alignment between read and multiple genomic sequences simultaneously utilizing the schema of the Landau-Vishkin algorithm. VARA dynamically extracts regional variants to construct a pseudo tree-based structure on-the-fly for seed extension without loading the whole genome variation into memory space. We developed the novel high-throughput sequencing read aligner deBGA-VARA by integrating VARA into deBGA. The deBGA-VARA is benchmarked both on simulated reads and the NA12878 sequencing dataset. The experimental results demonstrate that read alignment incorporating genetic variation knowledge can achieve high sensitivity and accuracy. Moreover, due to its efficiency, VARA provides a promising solution for further improvement of variant calling while maintaining small memory footprints. The deBGA-VARA is available at: https://github.com/hitbc/deBGA-VARA.
MOTIVATIONAs high-throughput sequencing (HTS) technology becomes ubiquitous and the volume of data continues to rise, HTS read alignment is becoming increasingly rate-limiting, which keeps pressing the development of novel read alignment approaches. Moreover, promising novel applications of HTS technology require aligning reads to multiple genomes instead of a single reference; however, it is still not viable for the state-of-the-art aligners to align large numbers of reads to multiple genomes.RESULTSWe propose de Bruijn Graph-based Aligner (deBGA), an innovative graph-based seed-and-extension algorithm to align HTS reads to a reference genome that is organized and indexed using a de Bruijn graph. With its well-handling of repeats, deBGA is substantially faster than state-of-the-art approaches while maintaining similar or higher sensitivity and accuracy. This makes it particularly well-suited to handle the rapidly growing volumes of sequencing data. Furthermore, it provides a promising solution for aligning reads to multiple genomes and graph-based references in HTS applications.AVAILABILITY AND IMPLEMENTATIONdeBGA is available at: https://github.com/hitbc/deBGA CONTACT: ydwang@hit.edu.cnSupplementary information: Supplementary data are available at Bioinformatics online.
Copy number variation (CNV) has been found to play an important role in human disease. Next-generation sequencing technology, including whole-genome sequencing (WGS) and whole-exome sequencing (WES), has become a primary strategy for studying the genetic basis of human disease. Several CNV calling tools have recently been developed on the basis of WES data. However, the comparative performance of these tools using real data remains unclear. An objective evaluation study of these tools in practical research situations would be beneficial. Here, we evaluated four well-known WES-based CNV detection tools (XHMM, CoNIFER, ExomeDepth, and CONTRA) using real data generated in house. After evaluation using six metrics, we found that the sensitive and accurate detection of CNVs in WES data remains challenging despite the many algorithms available. Each algorithm has its own strengths and weaknesses. None of the exome-based CNV calling methods performed well in all situations; in particular, compared with CNVs identified from high coverage WGS data from the same samples, all tools suffered from limited power. Our evaluation provides a comprehensive and objective comparison of several well-known detection tools designed for WES data, which will assist researchers in choosing the most suitable tools for their research needs.
Genome mapping is a procedure that could map a large number of short read data produced by High-throughput sequencing technologies to the human reference genome.And it has a great important meaning on analysis of expression of the genome,SNP site forecast and disease forecast.After a description of genome mapping background and the present situation of the research,this paper gives an introduction of basic structure of index concept and main method in mapping system.Then the paper emphatically puts forward the systematic analysis on index buiding principle based on BWT data structure,as well as presentation of main method of speeding up the sequence alignment.Finally,this paper gives main summaries of BWT technology,related method and principle involved in index constructing of mapping system,therefore reveals the essence of new index technology.