Extrachromosomal circular DNA (eccDNA) has emerged as a potential biomarker for disease due to its stable closed circular structure. However, the diagnostic utility of eccDNA remains underexplored. In this study, we demonstrate that the characteristics of eccDNA associated with genomic repetitive elements change in breast cancer patient tissues and plasma. These changes can serve as signatures for accurate cancer classification. We profiled eccDNA annotated to repeat elements across the genome in tissues and plasma, aggregating each repeat element to the superfamily and subfamily level. Our findings indicate that eccDNA associated with repetitive elements in cancer exhibits regular patterns of enrichment or depletion in specific elements, particularly at the family level. Additionally, these repeat element changes are present in different subtypes of breast cancer, correlated with varying hormone receptor expression. Although there are differences in the landscapes of eccDNA on repetitive elements between cancer tissues and paired plasma, the unique characteristics of eccDNA associated with repetitive sequences in the plasma of cancer patients facilitate better differentiation from normal individuals. These analyses reveal that changes in eccDNA associated with repeat sequences in human cancers can be used as diagnostic biomarkers for cancer patients.
Reconstructing the full-length sequence of extrachromosomal circular DNA (eccDNA) from short sequencing reads has proved challenging given the similarity of eccDNAs and their corresponding linear DNAs. Previous sequencing methods were unable to achieve high-throughput detection of full-length eccDNAs. Here we describe a new strategy that combined rolling circle amplification (RCA) and nanopore long-reads sequencing technology to generate full-length eccDNAs. We further developed a novel algorithm, called Full-Length eccDNA Detection (FLED), to reconstruct the sequence of eccDNAs. We used FLED to analyze seven human epithelial and cancer cell line samples and identified over 5,000 full-length eccDNAs per sample. The structures of identified eccDNAs were validated by both PCR and Sanger sequencing. Compared to other published nanopore-based eccDNA detectors, FLED exhibited higher sensitivity. In cancer cell lines, the genes overlapped with eccDNA regions were enriched in cancer-related pathways and cis-regulatory elements can be predicted in the up-stream or downstream of intact genes on eccDNA molecules, and the expressions of these cancer-related genes were dysregulated in tumor cell lines, indicating the regulatory potency of eccDNAs in biological processes. Our method takes advantage of nanopore long reads and enables unbiased reconstruction of full-length eccDNA sequences. FLED is imple-mented using Python3 which is freely available on GitHub (https://github.com/FuyuLi/FLED).
Background:There was a lacking of clinical diagnostic evidence in follow-up studies for reporting of secondary variants in 59 genes in American College of Medical Genetics and Genomics recommendations for reporting secondary findings and various strategies were applied to interpret the secondary variants. Results: Out of 1330 participants performed whole-exome sequencing, we identified 15 families with convincing clinical evidence. After Sanger validation and a comprehensive clinical follow-up, 10 families with both convincing clinical evidence and convincing genetic evidence of hereditary variants were found. Detailed clinical presentations and related clinical evidence were collected. Conclusions: Our research is a comprehensive follow-up study to identify secondary variants with convincing genetic and clinical evidence and it could help improve the strategy of screening actionable secondary variants and contribute to translation of genetic findings into medical practice.
Characterizing meiotic recombination rates across the genomes of nonhuman primates is important for understanding the genetics of primate populations, performing genetic analyses of phenotypic variation and reconstructing the evolution of human recombination. Rhesus macaques (Macaca mulatta) are the most widely used nonhuman primates in biomedical research. We constructed a high-resolution genetic map of the rhesus genome based on whole genome sequence data from Indian-origin rhesus macaques. The genetic markers used were approximately 18 million SNPs, with marker density 6.93 per kb across the autosomes. We report that the genome-wide recombination rate in rhesus macaques is significantly lower than rates observed in apes or humans, while the distribution of recombination across the macaque genome is more uniform. These observations provide new comparative information regarding the evolution of recombination in primates.
BACKGROUND:Current copy number variation (CNV) identification methods have rapidly become mature. However, the postdetection processes such as variant interpretation or reporting are inefficient. To overcome this situation, we developed REDBot as an automated software package for accurate and direct generation of clinical diagnostic reports for prenatal and products of conception (POC) samples. METHODS:We applied natural language process (NLP) methods for analyzing 30,235 in-house historical clinical reports through active learning, and then, developed clinical knowledge bases, evidence-based interpretation methods and reporting criteria to support the whole postdetection pipeline. RESULTS:Of the 30,235 reports, we obtained 37,175 CNV-paragraph pairs. For these pairs, the active learning approaches achieved a 0.9466 average F1-score in sentence classification. The overall accuracy for variant classification was 95.7%, 95.2%, and 100.0% in retrospective, prospective, and clinical utility experiments, respectively. CONCLUSION:By integrating NLP methods in CNVs postdetection pipeline, REDBot is a robust and rapid tool with clinical utility for prenatal and POC diagnosis.
Purpose To assess the clinical performance of an expanded noninvasive prenatal screening (NIPS) test (“NIPS-Plus”) for detection of both aneuploidy and genome-wide microdeletion/microduplication syndromes (MMS). Methods A total of 94,085 women with a singleton pregnancy were prospectively enrolled in the study. The cell-free plasma DNA was directly sequenced without intermediate amplification and fetal abnormalities identified using an improved copy-number variation (CNV) calling algorithm. Results A total of 1128 pregnancies (1.2%) were scored positive for clinically significant fetal chromosome abnormalities. This comprised 965 aneuploidies (1.026%) and 163 (0.174%) MMS. From follow-up tests, the positive predictive values (PPVs) for T21, T18, T13, rare trisomies, and sex chromosome aneuploidies were calculated as 95%, 82%, 46%, 29%, and 47%, respectively. For known MMS ( n = 32), PPVs were 93% (DiGeorge), 68% (22q11.22 microduplication), 75% (Prader–Willi/Angleman), and 50% (Cri du Chat). For the remaining genome-wide MMS ( n = 88), combined PPVs were 32% (CNVs ≥10 Mb) and 19% (CNVs <10 Mb). Conclusion NIPS-Plus yielded high PPVs for common aneuploidies and DiGeorge syndrome, and moderate PPVs for other MMS. Our results present compelling evidence that NIPS-Plus can be used as a first-tier pregnancy screening method to improve detection rates of clinically significant fetal chromosome abnormalities.
Tourette syndrome (TS) is a childhood-onset neuropsychiatric disorder characterized by repetitive motor movements and vocal tics. The clinical manifestations of TS are complex and often overlap with other neuropsychiatric disorders. TS is highly heritable; however, the underlying genetic basis and molecular and neuronal mechanisms of TS remain largely unknown. We performed whole-exome sequencing of a hundred trios (probands and their parents) with detailed records of their clinical presentations and identified a risk gene, ASH1L, that was both de novo mutated and associated with TS based on a transmission disequilibrium test. As a replication, we performed follow-up targeted sequencing of ASH1L in additional 524 unrelated TS samples and replicated the association (P value = 0.001). The point mutations in ASH1L cause defects in its enzymatic activity. Therefore, we established a transgenic mouse line and performed an array of anatomical, behavioral, and functional assays to investigate ASH1L function. The Ash1l+/− mice manifested tic-like behaviors and compulsive behaviors that could be rescued by the tic-relieving drug haloperidol. We also found that Ash1l disruption leads to hyper-activation and elevated dopamine-releasing events in the dorsal striatum, all of which could explain the neural mechanisms for the behavioral abnormalities in mice. Taken together, our results provide compelling evidence that ASH1L is a TS risk gene.
To assess the impact of genetic variation in regulatory loci on human health, we constructed a high-resolution map of allelic imbalances in DNA methylation, histone marks, and gene transcription in 71 epigenomes from 36 distinct cell and tissue types from 13 donors. Deep whole-genome bisulfite sequencing of 49 methylomes revealed sequence-dependent CpG methylation imbalances at thousands of heterozygous regulatory loci. Such loci are enriched for stochastic switching, which is defined as random transitions between fully methylated and unmethylated states of DNA. The methylation imbalances at thousands of loci are explainable by different relative frequencies of the methylated and unmethylated states for the two alleles. Further analyses provided a unifying model that links sequence-dependent allelic imbalances of the epigenome, stochastic switching at gene regulatory loci, and disease-associated genetic variation.
BACKGROUND: Next-generation sequencing is emerging as a viable alternative to chromosome microarray analysis for the diagnosis of chromosome disease syndromes. One next-generation sequencing methodology, copy number variation sequencing, has been shown to deliver high reliability, accuracy, and reproducibility for detection of fetal copy number variations in prenatal samples. However, its clinical utility as a first-tier diagnostic method has yet to be demonstrated in a large cohort of pregnant women referred for fetal chromosome testing. OBJECTIVE: We sought to evaluate copy number variation sequencing as a first-tier diagnostic method for detection of fetal chromosome anomalies in a general population of pregnant women with high-risk prenatal indications. STUDY DESIGN: This was a prospective analysis of 3429 pregnant women referred for amniocentesis and fetal chromosome testing for different risk indications, including advanced maternal age, high-risk maternal serum screening, and positivity for an ultrasound soft marker. Amniocentesis was performed by standard procedures. Amniocyte DNA was analyzed by copy number variation sequencing with a chromosome resolution of 0.1 Mb. Fetal chromosome anomalies including whole chromosome aneuploidy and segmental imbalances were independently confirmed by gold standard cytogenetic and molecular methods and their pathogenicity determined following guidelines of the American College of Medical Genetics for sequence variants. RESULTS: Clear interpretable copy number variation sequencing results were obtained for all 3429 amniocentesis samples. Copy number variation sequencing identified 3293 samples (96%) with a normal molecular karyotype and 136 samples (4%) with an altered molecular karyotype. A total of 146 fetal chromosome anomalies were detected, comprising 46 whole chromosome aneuploidies (pathogenic), 29 submicroscopic microdeletions/microduplications with known or suspected associations with chromosome disease syndromes (pathogenic), 22 other microdeletions/microduplications (likely pathogenic), and 49 variants of uncertain significance. Overall, the cumulative frequency of pathogenic/likely pathogenic and variants of uncertain significance chromosome anomalies in the patient cohort was 2.83% and 1.43%, respectively. In the 3 high-risk advanced maternal age, high-risk maternal serum screening, and ultrasound soft marker groups, the most common whole chromosome aneuploidy detected was trisomy 21, followed by sex chromosome aneuploidies, trisomy 18, and trisomy 13. Across all clinical indications, there was a similar incidence of submicroscopic copy number variations, with approximately equal proportions of pathogenic/likely pathogenic and variants of uncertain significance copy number variations. If karyotyping had been used as an alternate cytogenetics detection method, copy number variation sequencing would have returned a 1% higher yield of pathogenic or likely pathogenic copy number variations. CONCLUSION: In a large prospective clinical study, copy number variation sequencing delivered high reliability and accuracy for identifying clinically significant fetal anomalies in prenatal samples. Based on key performance criteria, copy number variation sequencing appears to be a well-suited methodology for first-tier diagnosis of pregnant women in the general population at risk of having a suspected fetal chromosome abnormality.
Large-scale, population-based genomic studies have provided a context for modern medical genetics. Among such studies, however, African populations have remained relatively underrepresented. The breadth of genetic diversity across the African continent argues for an exploration of local genomic context to facilitate burgeoning disease mapping studies in Africa. We sought to characterize genetic variation and to assess population substructure within a cohort of HIV-positive children from Botswana-a Southern African country that is regionally underrepresented in genomic databases. Using whole-exome sequencing data from 164 Batswana and comparisons with 150 similarly sequenced HIV-positive Ugandan children, we found that 13%-25% of variation observed among Batswana was not captured by public databases. Uncaptured variants were significantly enriched (p = 2.2 × 10-16) for coding variants with minor allele frequencies between 1% and 5% and included predicted-damaging non-synonymous variants. Among variants found in public databases, corresponding allele frequencies varied widely, with Botswana having significantly higher allele frequencies among rare (<1%) pathogenic and damaging variants. Batswana clustered with other Southern African populations, but distinctly from 1000 Genomes African populations, and had limited evidence for admixture with extra-continental ancestries. We also observed a surprising lack of genetic substructure in Botswana, despite multiple tribal ethnicities and language groups, alongside a higher degree of relatedness than purported founder populations from the 1000 Genomes project. Our observations reveal a complex, but distinct, ancestral history and genomic architecture among Batswana and suggest that disease mapping within similar Southern African populations will require a deeper repository of genetic variation and allelic dependencies than presently exists.
BACKGROUND: Intracranial aneurysm (IA) is usually a late-onset disease, affecting 1% to 3% of the general population and leading to lifethreatening subarachnoid hemorrhage. Genetic susceptibility has been implicated in IAs, but the causative genes remain elusive. METHODS: We performed next-generation sequencing in a discovery cohort of 20 Chinese IA patients. Bioinformatics filters were exploited to search for candidate deleterious variants with rare and low allele frequency. We further examined the candidate variants in a multiethnic sample collection of 86 whole exome sequenced unsolved familial IA cases from 3 previously published studies. RESULTS: We identified that the low-frequency variant c.4394C>A_p. Ala1465Asp (rs2298808) of ARHGEF17 was significantly associated with IA in our Chinese discovery cohort (P=7.3x10(-4); odds ratio=7.34). It was subsequently replicated in Japanese familial IA patients (P=0.039; odds ratio=4.00; 95% confidence interval=0.832-14.8) and was associated with IA in the large Chinese sample collection comprising 832 sporadic IA-affected and 599 control individuals (P=0.041; odds ratio=1.51; 95% confidence interval=1.02-Inf). When combining the sequencing data of all familial IA patients from 4 different ethnicities (ie, Chinese, Japanese, European American, and French-Canadian), we identified a significantly increased mutation burden for ARHGEF17 (21/106 versus 11/306; P=8.1x10(-7); odds ratio=6.6; 95% confidence interval=2.9-15.8) in cases as compared with controls. In zebrafish, arhgef17 was highly expressed in the brain blood vessel. arhgef17 knockdown caused blood extravasation in the brain region. Endothelial lesions were identified exclusively on cerebral blood vessels in the arhgef17-deficient zebrafish. CONCLUSIONS: Our results provide compelling evidence that ARHGEF17 is a risk gene for IA.
Whole-genome sequencing (WGS) allows for a comprehensive view of the sequence of the human genome. We present and apply integrated methodologic steps for interrogating WGS data to characterize the genetic architecture of 10 heart- and blood-related traits in a sample of 1,860 African Americans. In order to evaluate the contribution of regulatory and non-protein coding regions of the genome, we conducted aggregate tests of rare variation across the entire genomic landscape using a sliding window, complemented by an annotation-based assessment of the genome using predefined regulatory elements and within the first intron of all genes. These tests were performed treating all variants equally as well as with individual variants weighted by a measure of predicted functional consequence. Significant findings were assessed in 1,705 individuals of European ancestry. After these steps, we identified and replicated components of the genomic landscape significantly associated with heart- and blood-related traits. For two traits, lipoprotein(a) levels and neutrophil count, aggregate tests of low-frequency and rare variation were significantly associated across multiple motifs. For a third trait, cardiac troponin T, investigation of regulatory domains identified a locus on chromosome 9. These practical approaches for WGS analysis led to the identification of informative genomic regions and also showed that defined non-coding regions, such as first introns of genes and regulatory domains, are associated with important risk factor phenotypes. This study illustrates the tractable nature of WGS data and outlines an approach for characterizing the genetic architecture of complex traits.
Logistic regression and linear support vector machine(SVM)are the effective method to solve the problem of large-scale classification,but there has been no better research on their distributed implementation issues up to now.In recent years,the Spark platform based on the memory has been put forward,being the common framework in the mass data processing and analysis due to the inefficient willfulness of the distributed computing framework in the iterative algorithm.In this paper,the new quasi-Newton equation is used to solve the logistic regression and linear support vector machine and realized in the Spark framework.Experiments show that this method significantly improves the accuracy and efficiency of the large-scale classification problem.
The cost of Whole Genome Sequencing (WGS) has decreased tremendously in recent years due to advances in next-generation sequencing technologies. Nevertheless, the cost of carrying out large-scale cohort studies using WGS is still daunting. Past simulation studies with coverage at ~2x have shown promise for using low coverage WGS in studies focused on variant discovery, association study replications, and population genomics characterization. However, the performance of low coverage WGS in populations with a complex history and no reference panel remains to be determined.
Introduction: The metabolome is a collection of small molecules in a biologic sample, and may serve as biomarkers or predictors of heart disease. Whole genome sequence analysis offers the opportunity to investigate rare and low-frequency annotated variants across the human genome. We used whole genome sequence analysis to characterize the genetic architecture of the serum metabolome. Methods: Whole genome sequencing and measurement (chromotagraphy and mass spectroscopy) of 245 serum metabolites were done in 1,458 European Americans and 1,679 African Americans from the Atherosclerosis Risk in Communities (ARIC) study, and these data were used to perform a trans-ethnic meta-analysis. Common variants (MAF>5%) were analyzed individually using an additive genetic model. Rare and low-frequency protein-altering variants (MAF≤5%) were aggregated by genes. In order to determine the contribution of regulatory and non-protein coding regions of the genome, we conducted aggregate tests across the entire genome using a 4kb sliding window as well as in predefined regulatory elements, which includs enhancers, promoter, and 3’ and 5’ untranslated region of a gene. Results: We identified 119 significant associations between genetic variants and metabolite levels (significance threshold p<2.0*10 -10 for single variants, p<2.9х10 -10 for aggregate tests), of which 49 were novel, including genes involved in known Mendelian conditions, protein biological processes, and disease related pathways. Six genes ( DMGDH, AGA, ACY1, PRODH, DDC and CPS1 ) causing rare inborn errors of metabolism were associated with amino acid levels in the general population. A predicated regulatory variant in the AGA gene, encoding a protein involved in asparagine generation, was associated with serum asparagine levels independent of any coding variants in this gene. Seven genes ( ABCC2, PKD2L1, SLC10A1, FDX1, CYP3A43, UGT2B15 and SULT2A1 ) related to lipid-related metabolite levels were identified, whose gene products are involved in secretion, channeling and transportation. Analysis of regulatory regions unraveled associations between three steroid lipids and a member of the cytochrome P450 family, CYP3A43 . Five genes within the kinin-kallikrein pathway were identified to be related to small peptide levels, including KLKB1 , KNG1 , F12 , ACE and CPN1 . Variants in CPN1 , which is known to bind to fibrinogen, were associated with DSGEGDFXAEGGGVR, a peptide which is produced during fibrinogen to fibrin conversion. Conclusion: This study outlines an approach to characterize the genetic architecture of the human serum metabolome and shows that sequence variants affect multiple human metabolites. Using the principle of Mendelian randomization, the next step is to determine whether any of these metabolites are in causal pathways to disease.
Rhesus macaques ( Macaca mulatta ) are the most widely used nonhuman primate in biomedical research, have the largest natural geographic distribution of any nonhuman primate, and have been the focus of much evolutionary and behavioral investigation. Consequently, rhesus macaques are one of the most thoroughly studied nonhuman primate species. However, little is known about genome-wide genetic variation in this species. A detailed understanding of extant genomic variation among rhesus macaques has implications for the use of this species as a model for studies of human health and disease, as well as for evolutionary population genomics. Whole-genome sequencing analysis of 133 rhesus macaques revealed more than 43.7 million single-nucleotide variants, including thousands predicted to alter protein sequences, transcript splicing, and transcription factor binding sites. Rhesus macaques exhibit 2.5-fold higher overall nucleotide diversity and slightly elevated putative functional variation compared with humans. This functional variation in macaques provides opportunities for analyses of coding and noncoding variation, and its cellular consequences. Despite modestly higher levels of nonsynonymous variation in the macaques, the estimated distribution of fitness effects and the ratio of nonsynonymous to synonymous variants suggest that purifying selection has had stronger effects in rhesus macaques than in humans. Demographic reconstructions indicate this species has experienced a consistently large but fluctuating population size. Overall, the results presented here provide new insights into the population genomics of nonhuman primates and expand genomic information directly relevant to primate models of human disease.
Understanding the evolution of disease-associated mutations is fundamental to analyze pathogenetics of diseases. Mutation, recombination (by GC-biased gene conversion, gBGC), and selection have been known to shape the evolution of disease-associated mutations, but how these evolutionary forces work together is still an open question. In this study, we analyzed several human large-scale datasets (1000 Genomes, ESP6500, ExAC and ClinVar), and found that base-biased mutagenesis generates more GC→AT than AT→GC mutations, while gBGC promotes the fixation of AT→GC mutations to balance the impact of base-biased mutation on genome. Due to this effect of gBGC, purifying selection removes more deleterious AT→GC mutations than GC→AT from population, but many high-frequency (fixed and nearly fixed) deleterious AT→GC mutations are remained possibly due to high genetic load. As a special subset, disease-associated mutations follow this evolutionary rule, in which disease-associated GC→AT mutations are more enriched in rare mutations compared with AT→GC, while disease-associated AT→GC are more enriched in mutations with high frequency. Thus, we presented a base-biased evolutionary framework that explains the base-biased generation and accumulation of disease-associated mutations in human populations.
O1 The metabolomics approach to autism: identification of biomarkers for early detection of autism spectrum disorder