We present a research profile of the Computational Biology Group at the Faculty of Mathematics, Physics and Informatics of Comenius University in Bratislava, Slovakia. Research in our group concentrates on the following areas: Computational prediction of plasmids in bacteria. Structural variant discovery from unassembled long reads using pangenome graphs. Statistical significance of genomic annotation co-localization. Viral variant quantification in waste water. Non-conventional yeast genome sequencing and analysis.
In most eukaryotes, chromosomal DNA terminates with tandem repeats of a short G-rich motif, such as the canonical TTAGGG sequence. The arrays of telomeric repeats are maintained by telomerase or by alternative lengthening of telomeres (ALT). Here, we report that nuclear chromosomes of several basidiomycetous yeasts classified into the order Microstromatales carry unusual telomeres. We demonstrate that instead of TTAGGG-like repeats these telomeres are composed of unique, tandemly repeated motifs, which are in most cases specific to a particular chromosomal end. In contrast to other basidiomycetes, the Microstromatales genomes lack orthologs coding for the telomerase catalytic subunit Est2 and a shelterin component Tpp1 indicating that noncanonical telomeric repeats are maintained by a telomerase-independent mechanism. We hypothesize that in a common ancestor of Microstromatales the loss of telomerase and Tpp1 was compensated by activation of an ALT-like mechanism, which promoted amplification of various motifs and formation of distinct telomeric repeat arrays at most chromosomal ends.
Pangenomes are becoming increasingly popular data structures for genomics analyses due to their ability to compactly represent the genetic diversity within populations. Constructing a pangenome graph, however, is still a time-consuming and expensive process. A promising approach for pangenome construction consists of progressively augmenting a pangenome graph with additional high-quality assemblies. Currently, there is no approach to augment a pangenome graph using unassembled reads from newly sequenced samples that does not require to align them and genotype the new individuals. In this work we present the first assembly-free and mapping-free approach for augmenting an existing pangenome graph using unassembled long reads from an individual not already present in the pangenome. Our approach consists of finding sample specific sequences in reads using efficient indexes, clustering reads corresponding to the same novel variant(s), and then building a consensus sequence to be added to the pangenome graph for each variant separately. Using simulated reads based on Human Pangenome Reference Consortium (HPRC) assemblies, we demonstrate the effectiveness of the proposed approach for progressively augmenting the pangenome with long reads, without the need for de novo assembly or predicting genetic variants of the new sample. The software is freely available at github.com/ldenti/palss. ### Competing Interest Statement The authors have declared no competing interest.
Structural variants (SVs) are medium and large-scale genomic alterations that shape phenotypic diversity and disease risk. Numerous methods have been proposed for discovering SVs, however their benchmarking has been inconsistent across studies, often resulting in contradictory findings. One of the main sources of conflicting evaluation re-sults is the lack of consistency in the SV callsets used as ground truth, ranging from curated callsets released by consortia to more recent approaches that construct callsets from high-quality telomere-to-telomere de novo haplotype assemblies. The discrepancies between benchmarks are further compounded by the choice of the reference genome (GRCh37, GRCh38, and T2T-CHM13), where using T2T-CHM13reveals a different deletion/insertion profile, indicating reduced reference bias. We evaluated the performance of several state-of-the-art SV discovery methods from long-read whole-genome sequencing data and observed substantial variation in their performance and rankings, depending on the choice of ground truth, reference genome, and genomic regions used for evaluation. Counter-intuitively, the more complete reference genome T2T-CHM13does not inherently solve the problem of SV benchmarking; instead it reveals the limitations of each detection method in complex genomic regions. The substantial variation in detection accuracy across different genomic regions calls for additional caution in downstream analyses and in drawing conclusions based on predicted SVs. These findings underscore the complexity of evaluating SV detection methods and highlight the need for careful consideration and, ideally, field-standard best practices when reporting performance metrics. ### Competing Interest Statement The authors have declared no competing interest. European Union, 101180581, 872539, 956229 National Science Foundation (NSF), DBI-2042518 Slovak Grant Agency, 1/0538/22, 1/0140/25 Ministero dell'Università e della Ricerca, 2022YRB97K
Comparing biological samples through sequencing is a core task in bioinformatics analyses such as variant detection, differential expression analysis, and epigenetic peak calling. Standard approaches typically rely on mapping newly sequenced reads to a reference genome. To avoid mapping and reference biases, k-mer-based approaches have been proposed as an alternative. Using this paradigm, our tool kdiff identifies genomic regions containing k-mers with differential abundances between samples. We demonstrate that our method effectively detects copy number variants in cancer genomes and remains robust against reference genome misassemblies. Additionally, we illustrate its utility in confirming telomere locations in noisy nanopore sequencing data. Our work demonstrates that alignment-free approaches can provide results comparable to standard alignment-based methods, while reducing the reference bias and significantly improving computational efficiency by leveraging fast k-mer counting tools.
Ribosomes are ribonucleoprotein complexes highly conserved across all domains of life. The size differences of ribosomal RNAs (rRNAs) can be mainly attributed to variable regions termed expansion segments (ESs) protruding out from the ribosomal surface. The ESs were found to be involved in a range of processes including ribosome biogenesis and maturation, translation, and co-translational protein modification. Here, we analyze the rRNAs of the yeasts from the Magnusiomyces/Saprochaete clade belonging to the basal lineages of the subphylum Saccharomycotina. We find that these yeasts are missing more than 400 nt from the 25S rRNA and 150 nt from the 18S rRNAs when compared to their canonical counterparts in Saccharomyces cerevisiae. The missing regions mostly map to ESs, thus representing a shift toward a minimal rRNA structure. Despite the structural changes in rRNAs, we did not identify dramatic alterations in the ribosomal protein inventories. We also show that the size-reduced rRNAs are not limited to the species of the Magnusiomyces/Saprochaete clade, indicating that the shortening of ESs happened independently in several other lineages of the subphylum Saccharomycotina.
An annotation is a set of genomic intervals sharing a particular function or property. Examples include genes or their exons, sequence repeats, regions with a particular epigenetic state, and copy number variants. A common task is to compare two annotations to determine if one is enriched or depleted in the regions covered by the other. We study the problem of assigning statistical significance to such a comparison based on a null model representing random unrelated annotations. To incorporate more background information into such analyses, we propose a new null model based on a Markov chain that differentiates among several genomic contexts. These contexts can capture various confounding factors, such as GC content or assembly gaps. We then develop a new algorithm for estimating p-values by computing the exact expectation and variance of the test statistic and then estimating the p-value using a normal approximation. Compared to the previous algorithm by Gafurov et al., the new algorithm provides three advances: (1) the running time is improved from quadratic to linear or quasi-linear, (2) the algorithm can handle two different test statistics, and (3) the algorithm can handle both simple and context-dependent Markov chain null models. We demonstrate the efficiency and accuracy of our algorithm on synthetic and real data sets, including the recent human telomere-to-telomere assembly. In particular, our algorithm computed p-values for 450 pairs of human genome annotations using 24 threads in under three hours. Moreover, the use of genomic contexts to correct for GC bias resulted in the reversal of some previously published findings.
We report the genome sequence of the pathogenic yeast Candida parapsilosis strain SR23 (CBS 7157) used in a number of experimental studies. The nuclear genome assembly consists of eight chromosome-sized contigs with a total size of 13.04 Mbp (N50 2.09 Mbp) and a G+C content of 38.7%.
We generalize a problem of finding maximum-scoring segment sets, previously studied by Csűrös (IEEE/ACM Transactions on Computational Biology and Bioinformatics, 2004, 1, 139–150), from sequences to graphs. Namely, given a vertex-weighted graph G and a non-negative startup penalty c, we can find a set of vertex-disjoint paths in G with maximum total score when each path’s score is its vertices’ total weight minus c. We call this new problem maximum-scoring path sets (MSPS). We present an algorithm that has a linear-time complexity for graphs with a constant treewidth. Generalization from sequences to graphs allows the algorithm to be used on pangenome graphs representing several related genomes and can be seen as a common abstraction for several biological problems on pangenomes, including searching for CpG islands, ChIP-seq data analysis, analysis of region enrichment for functional elements, or simple chaining problems.
An annotation is a set of genomic intervals sharing a particular function or property. Examples include genes, conserved elements, and epigenetic modifications. A common task is to compare two annotations to determine if one is enriched or depleted in the regions covered by the other. We study the problem of assigning statistical significance to such a comparison based on a null model representing two random unrelated annotations. To incorporate more background information into such analyses and avoid biased results, we propose a new null model based on a Markov chain which differentiates among several genomic contexts. These contexts can capture various confounding factors, such as GC content or sequencing gaps. We then develop a new algorithm for estimating p-values by computing the exact expectation and variance of the test statistic and then estimating the p-value using a normal approximation. Compared to the previous algorithm by Gafurov et al., the new algorithm provides three advances: (1) the running time is improved from quadratic to linear or quasi-linear, (2) the algorithm can handle two different test statistics, and (3) the algorithm can handle both simple and context-dependent Markov chain null models. We demonstrate the efficiency and accuracy of our algorithm on synthetic and real data sets, including the recent human telomere-to-telomere assembly. In particular, our algorithm computed p-values for 450 pairs of human genome annotations using 24 threads in under three hours. The use of genomic contexts to correct for GC-bias also resulted in the reversal of some previously published findings. Availability. The software is freely available at https://github.com/ fmfi-compbio/mcdp2 under the MIT licence. All data for reproducibility are available at https://github.com/fmfi-compbio/mcdp2 -reproducibility .
Lodderomyces beijingensis is an ascosporic ascomycetous yeast. In contrast to related species Lodderomyces elongisporus, which is a recently emerging human pathogen, L. beijingensis is associated with insects. To provide an insight into its genetic makeup, we investigated the genome of its type strain, CBS 14171. We demonstrate that this yeast is diploid and describe the high contiguity nuclear genome assembly consisting of eight chromosome-sized contigs with a total size of about 15.1 Mbp. We find that the genome sequence contains multiple copies of the mating type loci and codes for essential components of the mating pheromone response pathway, however, the missing orthologs of several genes involved in the meiotic program raise questions about the mode of sexual reproduction. We also show that L. beijingensis genome codes for the 3-oxoadipate pathway enzymes, which allow the assimilation of protocatechuate. In contrast, the GAL gene cluster underwent a decay resulting in an inability of L. beijingensis to utilize galactose. Moreover, we find that the 56.5 kbp long mitochondrial DNA is structurally similar to known linear mitochondrial genomes terminating on both sides with covalently closed single-stranded hairpins. Finally, we discovered a new double-stranded RNA mycovirus from the Totiviridae family and characterized its genome sequence.
Base calling in nanopore sequencing is a difficult and computationally intensive problem, typically resulting in high error rates. In many applications of nanopore sequencing, analysis of raw signal is a viable alternative. Dynamic time warping (DTW) is an important building block for raw signal analysis. In this paper, we propose several improvements to DTW class of algorithms to better account for specifics of nanopore signal modeling. We have implemented these improvements in a new signal-to-reference alignment tool Nadavca. We demonstrate that Nadavca alignments improve unsupervised methylation detection over Tombo. We also demonstrate that by providing additional information about the discriminative power of positions in the signal, an otherwise unsupervised method can approach the accuracy of supervised models. Availability and implementation Nadavca is available under MIT license at https://github.com/fmfi-compbio/nadavca . Nanopore sequencing data sets are available from ENA bioproject PRJEB64246. Jaminaea angkorensis reference genome assembly is available from Zenodo https://doi.org/10.5281/zenodo.8145315 .
In many bioinformatics applications the task is to identify biologically significant locations in an individual genome. In our work, we are interested in finding high-density clusters of such biologically meaningful locations in a graph representation of a pangenome, which is a collection of related genomes. Different formulations of finding such clusters were previously studied for sequences. In this work, we study an extension of this problem for graphs, which we formalize as finding a set of vertex-disjoint paths with a maximum score in a weighted directed graph. We provide a linear-time algorithm for a special class of graphs corresponding to elastic-degenerate strings, one of pangenome representations. We also provide a fixed-parameter tractable algorithm for directed acyclic graphs with a special path decomposition of a limited width.
Candida verbasci is an anamorphic ascomycetous yeast. We report the genome sequence of its type strain, 11-1055 (CBS 12699). The nuclear genome assembly consists of seven chromosome-sized contigs with a total size of 12.1 Mbp and has a relatively low G+C content (28.1%).
Identification of plasmids from sequencing data is an important and challenging problem related to antimicrobial resistance spread and other One-Health issues. We provide a new architecture for identifying plasmid contigs in fragmented genome assemblies built from short-read data. We employ graph neural networks (GNNs) and the assembly graph to propagate the information from nearby nodes, which leads to more accurate classification, especially for short contigs that are difficult to classify based on sequence features or database searches alone. We trained plASgraph2 on a data set of samples from the ESKAPEE group of pathogens. plASgraph2 either outperforms or performs on par with a wide range of state-of-the-art methods on testing sets of independent ESKAPEE samples and samples from related pathogens. On one hand, our study provides a new accurate and easy to use tool for contig classification in bacterial isolates; on the other hand, it serves as a proof-of-concept for the use of GNNs in genomics. Our software is available at https://github.com/cchauve/plasgraph2 and the training and testing data sets are available at https://github.com/fmfi-compbio/plasgraph2-datasets.
AbstractRibosomes are ribonucleoprotein complexes highly conserved across all domains of life. The size differences of ribosomal RNAs (rRNAs) can be mainly attributed to variable regions termed expansion segments (ESs) protruding out from the ribosomal surface. The ESs were found to be involved in a range of processes including ribosome biogenesis and maturation, translation, and co-translational protein modification. Here, we analyze the rRNAs of the yeasts from theMagnusiomyces/Saprochaeteclade belonging to the basal lineages of the subphylum Saccharomycotina. We find that these yeasts are missing more than 400 nt from the 25S rRNA and 150 nt from the 18S rRNAs when compared to their canonical counterparts inSaccharomyces cerevisiae. The missing regions mostly map to ESs, thus representing a shift toward a minimal rRNA structure. Despite the structural changes in rRNAs, we did not identify dramatic alterations of the ribosomal protein inventories. We also show that the size-reduced rRNAs are not limited to the species of theMagnusiomyces/Saprochaeteclade, indicating that the shortening of ESs happened independently in several other lineages of the subphylum Saccharomycotina.SignificanceExpansion segments are variable regions present in the ribosomal RNAs involved in the ribosome biogenesis and translation. Although some of them were shown to be essential, their functions and the evolutionary trajectories leading to their expansion and/or reduction are not fully understood. Here, we show that the yeasts from theMagnusiomyces/Saprochaeteclade have truncated expansion segments, yet the protein inventories of their ribosomes do not radically differ from the species possessing canonical ribosomal RNAs. We also show that the loss of expansion segments occurred independently in several phylogenetic lineages of yeasts pointing out their dispensable nature. The differences identified in yeast ribosomal RNAs open a venue for further studies of these enigmatic elements of the eukaryotic ribosome.
Motivation Short tandem repeats (STRs) are regions of a genome containing many consecutive copies of the same short motif, possibly with small variations. Analysis of STRs has many clinical uses, but is limited by technology mainly due to STRs surpassing the used read length. Nanopore sequencing, as one of long read sequencing technologies, produces very long reads, thus offering more possibilities to study and analyze STRs. Basecalling of nanopore reads is however particularly unreliable in repeating regions, and therefore direct analysis from raw nanopore data is required. Results Here we present WarpSTR, a novel method for characterizing both simple and complex tandem repeats directly from raw nanopore signals using a finite-state automaton and a search algorithm analogous to dynamic time warping. By applying this approach to determine the lengths of 241 STRs, we demonstrate that our approach decreases the mean absolute error of the STR length estimate compared to basecalling and STRique. Availability WarpSTR is freely available at https://github.com/fmfi-compbio/warpstr Contact jozef.sitarcik@uniba.sk
Abstract Motivation The analysis of bacterial isolates to detect plasmids is important due to their role in the propagation of antimicrobial resistance. In short-read sequence assemblies, both plasmids and bacterial chromosomes are typically split into several contigs of various lengths, making identification of plasmids a challenging problem. In plasmid contig binning, the goal is to distinguish short-read assembly contigs based on their origin into plasmid and chromosomal contigs and subsequently sort plasmid contigs into bins, each bin corresponding to a single plasmid. Previous works on this problem consist of de novo approaches and reference-based approaches. De novo methods rely on contig features such as length, circularity, read coverage, or GC content. Reference-based approaches compare contigs to databases of known plasmids or plasmid markers from finished bacterial genomes. Results Recent developments suggest that leveraging information contained in the assembly graph improves the accuracy of plasmid binning. We present PlasBin-flow, a hybrid method that defines contig bins as subgraphs of the assembly graph. PlasBin-flow identifies such plasmid subgraphs through a mixed integer linear programming model that relies on the concept of network flow to account for sequencing coverage, while also accounting for the presence of plasmid genes and the GC content that often distinguishes plasmids from chromosomes. We demonstrate the performance of PlasBin-flow on a real dataset of bacterial samples. Availability and implementation https://github.com/cchauve/PlasBin-flow.
Tomáš Vinař合作论文数Siepel Computational Genomics Lab,
Dept. of Biological Statistics and Computational Biology15