Alignment is the cornerstone of many long-read pipelines and plays an essential role in resolving structural variants (SVs). However, forced alignments of SVs embedded in long reads, inflexibility of integrating novel SVs models and computational inefficiency remain problems. Here, we investigate the feasibility of resolving long-read SVs with alignment-free algorithms. We ask: (1) Is it possible to resolve long-read SVs with alignment-free approaches? and (2) Does it provide an advantage over existing approaches? To this end, we implemented the framework named Linear, which can flexibly integrate alignment-free algorithms such as the generative model for long-read SV detection. Furthermore, Linear addresses the problem of compatibility of alignment-free approaches with existing software. It takes as input long reads and outputs standardized results existing software can directly process. We conducted large-scale assessments in this work and the results show that the sensitivity, and flexibility of Linear outperform alignment-based pipelines. Moreover, the computational efficiency is orders of magnitude faster.
Structural variants are a common cause of disease and contribute to a large extent to inter-individual variability, but their detection and interpretation remain a challenge. Here, we investigate 11 individuals with complex genomic rearrangements including germline chromothripsis by combining short- and long-read genome sequencing (GS) with Hi-C. Large-scale genomic rearrangements are identified in Hi-C interaction maps, allowing for an independent assessment of breakpoint calls derived from the GS methods, resulting in >300 genomic junctions. Based on a comprehensive breakpoint detection and Hi-C, we achieve a reconstruction of whole rearranged chromosomes. Integrating information on the three-dimensional organization of chromatin, we observe that breakpoints occur more frequently than expected in lamina-associated domains (LADs) and that a majority reshuffle topologically associating domains (TADs). By applying phased RNA-seq, we observe an enrichment of genes showing allelic imbalanced expression (AIG) within 100 kb around the breakpoints. Interestingly, the AIGs hit by a breakpoint (19/22) display both up- and downregulation, thereby suggesting different mechanisms at play, such as gene disruption and rearrangements of regulatory information. However, the majority of interpretable genes located 200 kb around a breakpoint do not show significant expression changes. Thus, there is an overall robustness in the genome towards large-scale chromosome rearrangements.
Objective Chronic pancreatitis (CP) is a potentially fatal disease of the exocrine pancreas, with no specific or effective approved therapies. Due to difficulty in accessing pancreas tissues, little is known about local immune responses or pathogenesis in human CP. We sought to characterise pancreatic immune responses using tissues derived from patients with different aetiologies of CP and non-CP organ donors in order to identify key signalling molecules associated with human CP. Design We performed single-cell level cellular indexing of transcriptomes and epitopes by sequencing and T-cell receptor (TCR) sequencing of pancreatic immune cells isolated from organ donors, hereditary and idiopathic patients with CP who underwent total pancreatectomy. We validated gene expression data by performing flow cytometry and functional assays in a second patient with CP cohort. Results Deep single-cell sequencing revealed distinct immune characteristics and significantly enriched CCR6+ CD4+ T cells in hereditary compared with idiopathic CP. In hereditary CP, a reduction in T-cell clonality was observed due to the increased CD4+ T (Th) cells that replaced tissue-resident CD8+ T cells. Shared TCR clonotype analysis among T-cell lineages also unveiled unique interactions between CCR6+ Th and Th1 subsets, and TCR clustering analysis showed unique common antigen binding motifs in hereditary CP. In addition, we observed a significant upregulation of the CCR6 ligand (CCL20) expression among monocytes in hereditary CP as compared with those in idiopathic CP. The functional significance of CCR6 expression in CD4+ T cells was confirmed by flow cytometry and chemotaxis assay. Conclusion Single-cell sequencing with pancreatic immune cells in human CP highlights pancreas-specific immune crosstalk through the CCR6-CCL20 axis, a signalling pathway that might be leveraged as a potential future target in human hereditary CP.
Segmental duplications are important for understanding human diseases and evolution. The challenge to distinguish allelic and duplication sequences has hindered their phased assembly as well as characterization of structural variant calls. Here we have developed a novel graph-based approach that leverages single nucleotide differences in overlapping reads to distinguish allelic and duplication sequences information from long read accurate PacBio HiFi sequencing. These differences enable to generate allelic and duplication-specific overlaps in the graph to spell out phased assembly used for structural variant calling. We have applied our method to three public genomes: CHM13, NA12878 and HG002. Our method resolved 86% of duplicated regions fully with contig N50 up to 79 kb and produced <800 structural variant phased calls, outperforming state-of-the-part SDA method in terms of all metrics. Furthermore, we demonstrate the importance of phased assemblies and variant calls to the biologically-relevant duplicated genes such as SMN1, SRGAP2C, NPY4R and FAM72A. Our phased assemblies and accurate variant calling specifically in duplicated regions will enable the study of the evolution and adaptation of various species.
Linking genomic variation to phenotypical traits remains a major challenge in evolutionary genetics. In this study, we use phylogenomic strategies to investigate a distinctive trait among mammals: the development of masculinizing ovotestes in female moles. By combining a chromosome-scale genome assembly of the Iberian mole, Talpa occidentalis, with transcriptomic, epigenetic, and chromatin interaction datasets, we identify rearrangements altering the regulatory landscape of genes with distinct gonadal expression patterns. These include a tandem triplication involving CYP17A1, a gene controlling androgen synthesis, and an intrachromosomal inversion involving the pro-testicular growth factor gene FGF9, which is heterochronically expressed in mole ovotestes. Transgenic mice with a knock-in mole CYP17A1 enhancer or overexpressing FGF9 showed phenotypes recapitulating mole sexual features. Our results highlight how integrative genomic approaches can reveal the phenotypic impact of noncoding sequence changes.
Islet yield is an important predictor of acceptable glucose control after total pancreatectomy with islet autotransplantation (TP-IAT). We assessed if pancreas volume calculated with preoperative MRI could assess islet yield and postoperative outcomes. We reviewed dynamic MRI studies from 154 adult TP-IAT patients (2009-2016), and associations between calculated volumes and digest islet equivalents (IEQs) were tested. In multivariate regression analysis, pancreas volume (P < .001) and preoperative HbA1c levels (P = .009) were independently associated with digest IEQs. The IEQ prediction formula was calculated according to each preoperative HbA1c level, (a) pancreas volume x 5800 for HbA1c >= 6.5, (b) pancreas volume x 10 000 for HbA1c >= 5.7/<6.5 and (iii) pancreas volume x 11 400 for HbA1c < 5.7. The formula was internally validated with 28 TP-IAT patients between 2017 and 2018 (r(2) = .657 andr(2) = .710 when restricted to 24 patients without prior pancreatectomy). An estimated IEQs/Body Weight (kg) >= 3700 predicted HbA1c <= 6.5 and insulin independence at 1 year after TP-IAT with 77% and 88% sensitivity and 55% and 43% specificity, respectively. The combination of pancreas volume and preoperative HbA1c levels may be useful to estimate islet yield. Estimated IEQs were reasonably sensitive to predict acceptable glucose control at 1 year.
Haplotype-resolved or phased genome assembly provides a complete picture of genomes and their complex genetic variations. However, current algorithms for phased assembly either do not generate chromosome-scale phasing or require pedigree information, which limits their application. We present a method named diploid assembly (DipAsm) that uses long, accurate reads and long-range conformation data for single individuals to generate a chromosome-scale phased assembly within 1 day. Applied to four public human genomes, PGP1, HG002, NA12878 and HG00733, DipAsm produced haplotype-resolved assemblies with minimum contig length needed to cover 50% of the known genome (NG50) up to 25 Mb and phased ~99.5% of heterozygous sites at 98–99% accuracy, outperforming other approaches in terms of both contiguity and phasing completeness. We demonstrate the importance of chromosome-scale phased assemblies for the discovery of structural variants (SVs), including thousands of new transposon insertions, and of highly polymorphic and medically important regions such as the human leukocyte antigen (HLA) and killer cell immunoglobulin-like receptor (KIR) regions. DipAsm will facilitate high-quality precision medicine and studies of individual haplotype variation and population diversity.
Structural variants (SVs) remain challenging to represent and study relative to point mutations despite their demonstrated importance. We show that variation graphs, as implemented in the vg toolkit, provide an effective means for leveraging SV catalogs for short-read SV genotyping experiments. We benchmark vg against state-of-the-art SV genotypers using three sequence-resolved SV catalogs generated by recent long-read sequencing studies. In addition, we use assemblies from 12 yeast strains to show that graphs constructed directly from aligned de novo assemblies improve genotyping compared to graphs built from intermediate SV catalogs in the VCF format.
Reconstructing haplotypes from sequencing data is one of the major challenges in genetics. Haplotypes play a crucial role in many analyses, including genome-wide association studies and population genetics. Haplotype reconstruction becomes more difficult for higher numbers of homologous chromosomes, as it is often the case for polyploid plants. This complexity is compounded further by higher heterozygosity, which denotes the frequent presence of variants between haplotypes. We have designed Ranbow, a new tool for haplotype reconstruction of polyploid genome from short read sequencing data. Ranbow integrates all types of small variants in bi- and multi-allelic sites to reconstruct haplotypes. To evaluate Ranbow and currently available competing methods on real data, we have created and released a real gold standard dataset from sweet potato sequencing data. Our evaluations on real and simulated data clearly show Ranbow's superior performance in terms of accuracy, haplotype length, memory usage, and running time. Specifically, Ranbow is one order of magnitude faster than the next best method. The efficiency and accuracy of Ranbow makes whole genome haplotype reconstruction of complex genome with higher ploidy feasible.
Motivation With the availability of new sequencing technologies, the generation of haplotype-resolved genome assemblies up to chromosome scale has become feasible. These assemblies capture the complete genetic information of both parental haplotypes, increase structural variant (SV) calling sensitivity and enable direct genotyping and phasing of SVs. Yet, existing SV callers are designed for haploid genome assemblies only, do not support genotyping or detect only a limited set of SV classes. Results We introduce our method SVIM-asm for the detection and genotyping of six common classes of SVs from haploid and diploid genome assemblies. Compared against the only other existing SV caller for diploid assemblies, DipCall, SVIM-asm detects more SV classes and reached higher F1 scores for the detection of insertions and deletions on two recently published assemblies of the HG002 individual. Availability and Implementation SVIM-asm has been implemented in Python and can be easily installed via bioconda. Its source code is available at github.com/eldariont/svim-asm. Contact vingron@molgen.mpg.de Supplementary information Supplementary data are available online.
Immune tolerance to allografts has been pursued for decades as an important goal in transplantation. Administration of apoptotic donor splenocytes effectively induces antigen-specific tolerance to allografts in murine studies. Here we show that two peritransplant infusions of apoptotic donor leukocytes under short-term immunotherapy with antagonistic anti-CD40 antibody 2C10R4, rapamycin, soluble tumor necrosis factor receptor and anti-interleukin 6 receptor antibody induce long-term (≥1 year) tolerance to islet allografts in 5 of 5 nonsensitized, MHC class I-disparate, and one MHC class II DRB allele-matched rhesus macaques. Tolerance in our preclinical model is associated with a regulatory network, involving antigen-specific Tr1 cells exhibiting a distinct transcriptome and indirect specificity for matched MHC class II and mismatched class I peptides. Apoptotic donor leukocyte infusions warrant continued investigation as a cellular, nonchimeric and translatable method for inducing antigen-specific tolerance in transplantation.
Motivation Structural variants are defined as genomic variants larger than 50bp. They have been shown to affect more bases in any given genome than SNPs or small indels. Additionally, they have great impact on human phenotype and diversity and have been linked to numerous diseases. Due to their size and association with repeats, they are difficult to detect by shotgun sequencing, especially when based on short reads. Long read, single molecule sequencing technologies like those offered by Pacific Biosciences or Oxford Nanopore Technologies produce reads with a length of several thousand base pairs. Despite the higher error rate and sequencing cost, long read sequencing offers many advantages for the detection of structural variants. Yet, available software tools still do not fully exploit the possibilities. Results We present SVIM, a tool for the sensitive detection and precise characterization of structural variants from long read data. SVIM consists of three components for the collection, clustering and combination of structural variant signatures from read alignments. It discriminates five different variant classes including similar types, such as tandem and interspersed duplications and novel element insertions. SVIM is unique in its capability of extracting both the genomic origin and destination of duplications. It compares favorably with existing tools in evaluations on simulated data and real datasets from PacBio and Nanopore sequencing machines. Availability and implementation The source code and executables of SVIM are available on Github: github.com/eldariont/svim. SVIM has been implemented in Python 3 and published on bioconda and the Python Package Index. Contact heller_d@molgen.mpg.de
BackgroundPancreatic fat may adversely affect -cell mass and function, possibly via local release of non-esterified fatty acids, and proinflammatory and vasoactive factors released by adipose tissue. However, the effects of intrapancreatic fat in patients with chronic pancreatitis undergoing total pancreatectomy with islet autotransplantation (TPIAT) have not been studied. This study investigated whether pancreatic fatty infiltration has a negative effect on metabolic outcomes following TPIAT. MethodsThe association between pancreatic fatty infiltration and diabetes outcomes was studied in 79 patients with low or high pancreatic fat content (LPF [n=53] and HPF [n=26], respectively) undergoing TPIAT. Pancreatic fatty infiltration was stratified using gross examinations during isolation and validated with histomorphometry of archived histology samples. ResultsFat area percentage in histology samples differed significantly between the LPF and HPF groups (2.1%4.3% vs 10.6%8.9%, respectively; P=0.0009). Insulin dependence was more common in the HPF group, whereas more patients in the LPF group were insulin independent or on partial insulin supplementation at 1year (P=0.022). Furthermore, 1- and 2-h glucose concentrations during mixed-meal tolerance tests were significantly higher in the HPF group (P=0.032 and 0.027, respectively) and -scores (a composite measure of islet function and metabolic control) were significantly greater in the LPF than HPF group (6.11.7 vs 4.62.0; P=0.034). ConclusionsPatients with HPF were more likely to be insulin dependent, with higher postprandial glucose excursion, suggesting that intrapancreatic fat may lead to -cell dysfunction with detrimental effects on diabetes outcomes after TPIAT.
Background: In TPIAT, current final islet product (IP) sterility testing identifies microbial contamination for rapid antibiotic prophylaxis. We sought to (1) quantify “bioburden” (BB) during manufacturing, (2) understand which processes may reduce BB, and (3) develop patient-specific strategies to lessen BB in IP. Methods: Pancreas preservation solution (PS) and IP sterility were assessed via BACTEC system, growth reported qualitatively (Y/N), and species identified by clinical laboratory. Quantitative data were collected by a Plate Count Method (PCM): PS and IP (n=142), and post-digestion Recombination Solution (RS) and pre-product/COBE (PC) supernatants (n=110) were incubated on TSA plates for 7 days. BB was defined as total CFU from PCM. Donor pancreas fibrosis was stratified as mild/moderate (1-8 out of 10) or severe (9-10).[SL1] To reduce BB, Gentamicin was added to RS (n=91). Results: BACTEC and PCM concordance was observed for PS and IP (p<.01), though Granulicatella species were not recoverable by PCM (n=7). In all cases (n=142), mean BB was higher in the severe fibrosis group at all sample time points [PS: 41,552 ± 12,574 vs 2,521 ± 1,364; RS: 42,792 ± 10,065 vs 3,188 ± 1,914; PC: 9,430 ± 2,630 vs 636 ± 361; and IP: 4,070 ± 1,067 vs 229 ± 91, (p<.05) (Figure 1)]. In the severe fibrosis group, mean BB was reduced from PS to PC (p<.05) and IP (p<.01) (Figure 1). Density gradient purification is significantly more likely in both mild and moderate cases vs severe (p<0.001) due to high tissue volume. BB is significantly lower in cases requiring purification for all timepoints (p<0.001). When excluding purified cases, fibrosis is still a predictive factor of increased BB in PS, RS, PC, and IP (p<0.01). In a pilot study, addition of gentamicin to the RS resulted in significant reduction IP sterility positive (n= 41 p<0.05). No significance was observed after 50 subsequent cases. Initial subset analysis (fibrosis, BB, positive PS, purification) were not significant. Conclusion: Although BB varies remarkably between patients, we characterized a subset with the highest IP BB (fibrosis ≥9, unpurified) where BB reduction may be inadequate. Cases requiring purification have lower BB, but not due to the purification process. BB reduction is inherit to islet manufacturing due to high media volumes and extensive washing steps. Current dilution during washing steps is ~1x106-fold from PS to IP; further dilution offers potential to reduce BB in high risk patients. A proposed novel method to reduce BB involves additional washes to dilute IP an added 1x104 and requires <20 minutes of processing. Further study is needed to identify high risk subsets and evaluate BB reduction strategies.References: 1. Berger MG, et al., Microbial contamination of transplant solutions during pancreatic islet autotransplants is not associated with clinical infection in a pediatric population, Pancreatology (2016), http://dx.doi.org/10.1016/j.pan.2016.03.019.
RNA-binding proteins (RBPs) play an important role in RNA post-transcriptional regulation and recognize target RNAs via sequence-structure motifs. The extent to which RNA structure influences protein binding in the presence or absence of a sequence motif is still poorly understood. Existing RNA motif finders either take the structure of the RNA only partially into account, or employ models which are not directly interpretable as sequence-structure motifs. We developed ssHMM, an RNA motif finder based on a hidden Markov model (HMM) and Gibbs sampling which fully captures the relationship between RNA sequence and secondary structure preference of a given RBP. Compared to previous methods which output separate logos for sequence and structure, it directly produces a combined sequence-structure motif when trained on a large set of sequences. ssHMM's model is visualized intuitively as a graph and facilitates biological interpretation. ssHMM can be used to find novel bona fide sequence-structure motifs of uncharacterized RBPs, such as the one presented here for the YY1 protein. ssHMM reaches a high motif recovery rate on synthetic data, it recovers known RBP motifs from CLIP-Seq data, and scales linearly on the input size, being considerably faster than MEMERIS and RNAcontext on large datasets while being on par with GraphProt. It is freely available on Github and as a Docker image.
Knut Reinert合作论文数Freie Universit?0?1t Berlin;Institut f??r Informatik2