The ribosome, long considered an invariant actor of gene expression, recently emerged as a contributor to translational control through chemical modifications of ribosomal RNAs (rRNAs). These modifications are guided by small nucleolar RNAs (snoRNAs), which direct modifying enzymes to specific rRNA positions. Here, we report that lung adenocarcinoma (LUAD) cells that are resistant to tyrosine kinase inhibitors (TKIs) reshape both their translational program and rRNA 2’O-ribose methylation (2’Ome) profiles. EML4-ALK-positive LUAD cells that are resistant to crizotinib (ALK inhibitor) exhibit reduced global protein synthesis and a change in selective translation of mRNAs, encoding proteins previously linked to resistance. These cells show concomitant reduction in SNORD104 and 2’Ome at its associated 28S_Cm1327 position. Functional studies reveal that SNORD104 depletion abolishes 2’Ome at 28S_Cm1327 without affecting basal translation or cell viability, but enhances both under crizotinib exposure. Moreover, SNORD104 knockdown attenuates caspase activation and PARP cleavage during treatment, supporting reduced cell death. Altogether, our findings support a role for snoRNA-guided rRNA modification as a novel non-genomic mechanism in early adaptive responses to crizotinib therapy in LUAD.
Ancestral sequence reconstruction is an important task in bioinformatics, with applications ranging from protein engineering to the study of genome evolution. When sequences can only undergo substitutions, optimal reconstructions can be efficiently computed using well-known algorithms. However, accounting for indels in ancestral reconstructions is much harder. First, for biologically-relevant problem formulations, no polynomial-time exact algorithms are available. Second, multiple reconstructions are often equally parsimonious or likely, making it crucial to correctly display uncertainty in the results. Here, we consider a parsimony approach where only deletions are allowed, while addressing the aforementioned limitations. First, we describe an exact algorithm to obtain all the optimal solutions. The algorithm runs in polynomial time if only one solution is sought. Second, we show that all possible optimal reconstructions for a fixed node can be represented using a graph computable in polynomial time. While previous studies have proposed graph-based representations of ancestral reconstructions, this result is the first to offer a solid mathematical justification for this approach. Finally we provide arguments for the relevance of the deletion-only case for the general case.
A correlation is a binary vector that encodes all possible positions of overlaps of two words, where an overlap for an ordered pair of words (u, v) occurs if a suffix of u matches a prefix of v. As multiple pairs can have the same correlation, it is relevant to count how many pairs of words share the same correlation, depending on the alphabet size and word length n. We exhibit recurrences to compute the number of such pairs – which is termed population size – for any correlation; for this, we exploit a relationship between overlaps of two words and self-overlap of one word. This theorem allows us to compute the number of pairs with the longest overlap of a given length, solving two open questions Gabric raised in 2022. Finally, we also provide bounds for the asymptotic population ratio of any correlation. Given the importance of word overlaps in areas like combinatorics on words, bioinformatics, and digital communication, our results may ease analyses of algorithms for string processing, code design, or genome assembly.
Overlaps between words are crucial in many areas of computer science, such as code design, stringology, and bioinformatics. A self overlapping word is characterized by its periods and borders. A period of a word $u$ is the starting position of a suffix of $u$ that is also a prefix $u$, and such a suffix is called a border. Each word of length, say $n>0$, has a set of periods, but not all combinations of integers are sets of periods. Computing the period set of a word $u$ takes linear time in the length of $u$. We address the question of computing, the set, denoted $\Gamma_n$, of all period sets of words of length $n$. Although period sets have been characterized, there is no formula to compute the cardinality of $\Gamma_n$ (which is exponential in $n$), and the known dynamic programming algorithm to enumerate $\Gamma_n$ suffers from its space complexity. We present an incremental approach to compute $\Gamma_n$ from $\Gamma_{n-1}$, which reduces the space complexity, and then a constructive certification algorithm useful for verification purposes. The incremental approach defines a parental relation between sets in $\Gamma_{n-1}$ and $\Gamma_n$, enabling one to investigate the dynamics of period sets, and their intriguing statistical properties. Moreover, the period set of a word $u$ is the key for computing the absence probability of $u$ in random texts. Thus, knowing $\Gamma_n$ is useful to assess the significance of word statistics, such as the number of missing words in a random text.
Cancer onset and progression are known to be regulated by genetic and epigenetic events, including RNA modifications (a.k.a. epitranscriptomics). So far, more than 150 chemical modifications have been described in all RNA subtypes, including messenger, ribosomal, and transfer RNAs. RNA modifications and their regulators are known to be implicated in all steps of post-transcriptional regulation. The dysregulation of this complex yet delicate balance can contribute to disease evolution, particularly in the context of carcinogenesis, where cells are subjected to various stresses. We sought to discover RNA modifications involved in cancer cell adaptation to inhospitable environments, a peculiar feature of cancer stem cells (CSCs). We were particularly interested in the RNA marks that help the adaptation of cancer cells to suspension culture, which is often used as a surrogate to evaluate the tumorigenic potential. For this purpose, we designed an experimental pipeline consisting of four steps: (1) cell culture in different growth conditions to favor CSC survival; (2) simultaneous RNA subtype (mRNA, rRNA, tRNA) enrichment and RNA hydrolysis; (3) the multiplex analysis of nucleosides by LC-MS/MS followed by statistical/bioinformatic analysis; and (4) the functional validation of identified RNA marks. This study demonstrates that the RNA modification landscape evolves along with the cancer cell phenotype under growth constraints. Remarkably, we discovered a short epitranscriptomic signature, conserved across colorectal cancer cell lines and associated with enrichment in CSCs. Functional tests confirmed the importance of selected marks in the process of adaptation to suspension culture, confirming the validity of our approach and opening up interesting prospects in the field.
Motivation Seeking probabilistic motifs in a sequence is a common task to annotate putative transcription factor binding sites (TFBS). Useful motif representations include Position Weight Matrices (PWMs), dinucleotidic PWMs (di-PWMs), and Hidden Markov Models (HMMs). Dinucleotidic PWMs combine the simplicity of PWMs – a matrix form and a cumulative scoring function –, but also incoporate dependency between adjacent positions in the motif (unlike PWMs which disregard any dependency). For instance, to represent binding sites, the HOCOMOCO database provides di-PWM motifs derived from experimental data. Currently, two programs, SPRy-SARUS and MOODS, can search for di-PWMs in sequences. Results We propose a Python package, dipwmsearch , which provides an original and efficient algorithm for this task (it first enumerates matching words for the di-PWM, and then search them at once in the sequence even if it contains IUPAC codes). The user benefits from an easy installation via Pypi or conda , a documented Python interface, and reusable example scripts that smooth the use of di-PWMs. Availability and Implementation dipwmsearch is available at https://pypi.org/project/dipwmsearch/ and https://gite.lirmm.fr/rivals/dipwmsearch/ under Cecill license.
Gene expression is the synthesis of proteins from the information encoded on DNA. One of the two main steps of gene expression is the translation of messenger RNA (mRNA) into polypeptide sequences of amino acids. Here, by taking into account mRNA degradation, we model the motion of ribosomes along mRNA with a ballistic model where particles advance along a filament without excluded volume interactions. Unidirectional models of transport have previously been used to fit the average density of ribosomes obtained by the experimental ribo-sequencing (Ribo-seq) technique in order to obtain the kinetic rates. The degradation rate is not, however, accounted for and experimental data from different experiments are needed to have enough parameters for the fit. Here, we propose an entirely novel experimental setup and theoretical framework consisting in splitting the mRNAs into categories depending on the number of ribosomes from one to four. We solve analytically the ballistic model for a fixed number of ribosomes per mRNA, study the different regimes of degradation, and propose a criterion for the quality of the inverse fit. The proposed method provides a high sensitivity to the mRNA degradation rate. The additional equations coming from using the monosome (single ribosome) and polysome (arbitrary number) ribo-seq profiles enable us to determine all the kinetic rates in terms of the experimentally accessible mRNA degradation rate.
Background Protozoan parasites are known to attach specific and diverse group of proteins to their plasma membrane via a GPI anchor. In malaria parasites, GPI-anchored proteins (GPI-APs) have been shown to play an important role in host–pathogen interactions and a key function in host cell invasion and immune evasion. Because of their immunogenic properties, some of these proteins have been considered as malaria vaccine candidates. However, identification of all possible GPI-APs encoded by these parasites remains challenging due to their sequence diversity and limitations of the tools used for their characterization. Methods The FT-GPI software was developed to detect GPI-APs based on the presence of a hydrophobic helix at both ends of the premature peptide. FT-GPI was implemented in C ++and applied to study the GPI-proteome of 46 isolates of the order Haemosporida. Using the GPI proteome of Plasmodium falciparum strain 3D7 and Plasmodium vivax strain Sal-1, a heuristic method was defined to select the most sensitive and specific FT-GPI software parameters. Results FT-GPI enabled revision of the GPI-proteome of P. falciparum and P. vivax, including the identification of novel GPI-APs. Orthology- and synteny-based analyses showed that 19 of the 37 GPI-APs found in the order Haemosporida are conserved among Plasmodium species. Our analyses suggest that gene duplication and deletion events may have contributed significantly to the evolution of the GPI proteome, and its composition correlates with speciation. Conclusion FT-GPI-based prediction is a useful tool for mining GPI-APs and gaining further insights into their evolution and sequence diversity. This resource may also help identify new protein candidates for the development of vaccines for malaria and other parasitic diseases.
Finding the correct position of new sequences within an established phylogenetic tree is an increasingly relevant problem in evolutionary bioinformatics and metagenomics. Recently, alignment-free approaches for this task have been proposed. One such approach is based on the concept of phylogenetically-informative k-mers or phylo- k-mers for short. In practice, phylo- k-mers are inferred from a set of related reference sequences and are equipped with scores expressing the probability of their appearance in different locations within the input reference phylogeny. Computing phylo- k-mers, however, represents a computational bottleneck to their applicability in real-world problems such as the phylogenetic analysis of metabarcoding reads and the detection of novel recombinant viruses. Here we consider the problem of phylo- k-mer computation: how can we efficiently find all k-mers whose probability lies above a given threshold for a given tree node? We describe and analyze algorithms for this problem, relying on branch-and-bound and divide-and-conquer techniques. We exploit the redundancy of adjacent windows of the alignment to save on computation. Besides computational complexity analyses, we provide an empirical evaluation of the relative performance of their implementations on simulated and real-world data. The divide-and-conquer algorithms are found to surpass the branch-and-bound approach, especially when many phylo- k-mers are found.
MOTIVATION:Phylogenetic placement enables phylogenetic analysis of massive collections of newly sequenced DNA, when de novo tree inference is too unreliable or inefficient. Assuming that a high-quality reference tree is available, the idea is to seek the correct placement of the new sequences in that tree. Recently, alignment-free approaches to phylogenetic placement have emerged, both to circumvent the need to align the new sequences and to avoid the calculations that typically follow the alignment step. A promising approach is based on the inference of k-mers that can be potentially related to the reference sequences, also called phylo-k-mers. However, its usage is limited by the time and memory-consuming stage of reference data preprocessing and the large numbers of k-mers to consider.RESULTS:We suggest a filtering method for selecting informative phylo-k-mers based on mutual information, which can significantly improve the efficiency of placement, at the cost of a small loss in placement accuracy. This method is implemented in IPK, a new tool for computing phylo-k-mers that significantly outperforms the software previously available. We also present EPIK, a new software for phylogenetic placement, supporting filtered phylo-k-mer databases. Our experiments on real-world data show that EPIK is the fastest phylogenetic placement tool available, when placing hundreds of thousands and millions of queries while still providing accurate placements.AVAILABILITY AND IMPLEMENTATION:IPK and EPIK are freely available at https://github.com/phylo42/IPK and https://github.com/phylo42/EPIK. Both are implemented in C++ and Python and supported on Linux and MacOS.
Partial response to chemotherapy leads to disease resurgence. Upon treatment, a subpopulation of cancer cells, called drug-tolerant persistent cells, display a transitory drug tolerance that lead to treatment resistance 1,2 . Though drug-tolerance mechanisms remain poorly known, they have been linked to non-genomic processes, including epigenetics, stemness and dormancy 2–4 . 5-fluorouracil (5-FU), the most widely used chemotherapy in cancer treatment, is associated with resistance. While prescribed as an inhibitor of DNA replication, 5-FU alters all RNA pathways 5–9 . Here, we show that 5-FU treatment leads to the unexpected production of fluorinated ribosomes, exhibiting altered mRNA translation. 5-FU is incorporated into ribosomal RNAs of mature ribosomes in cancer cell lines, colorectal xenografts and human tumours. Fluorinated ribosomes appear to be functional, yet, they display a selective translational activity towards mRNAs according to the nature of their 5’-untranslated region. As a result, we found that sustained translation of IGF-1R mRNA, which codes for one of the most potent cell survival effectors, promoted the survival of 5-FU-treated colorectal cancer cells. Altogether, our results demonstrate that “man-made” fluorinated ribosomes favour the drug-tolerant cellular phenotype by promoting translation of survival genes. This could be exploited for developing novel combined therapies. By unraveling translation regulation as a novel gene expression mechanism helping cells to survive a drug-challenge, our study extends the spectrum of molecular mechanisms driving drug-tolerance.
The last decade has seen mRNA modification emerge as a new layer of gene expression regulation. The Fat mass and obesity-associated protein (FTO) was the first identified eraser of N6-methyladenosine (m6A) adducts, the most widespread modification in eukaryotic messenger RNA. This discovery, of a reversible and dynamic RNA modification, aided by recent technological advances in RNA mass spectrometry and sequencing has led to the birth of the field of epitranscriptomics. FTO crystallized much of the attention of epitranscriptomics researchers and resulted in the publication of numerous, yet contradictory, studies describing the regulatory role of FTO in gene expression and central biological processes. These incongruities may be explained by a wide spectrum of FTO substrates and RNA sequence preferences: FTO binds multiple RNA species (mRNA, snRNA and tRNA) and can demethylate internal m6A in mRNA and snRNA, N6,2 '-O-dimethyladenosine (m6Am) adjacent to the mRNA cap, and N1-methyladenosine (m1A) in tRNA. Here, we review current knowledge related to FTO function in healthy and cancer cells. In particular, we emphasize the divergent role(s) attributed to FTO in different tissues and subcellular and molecular contexts.
RNA translation has long been thought as a stable and uniform process by which a ribosome produces a protein encoded by the main Open Reading Frame (ORF) of an mRNA. Recently, growing evidence support incomplete correlation between RNA and protein abundance levels, the existence of alternative ORFs in numerous mammalian RNAs, and the involvement of ribosomes in gene expression regulation, thereby challenging previous views of translation. Ribosome profiling (aka Ribo-seq) has renewed the study of translation by enabling the mapping of translating ribosomes on the whole transcriptome using deep-sequencing. Despite increasing use of Ribo-seq, recent review articles conclude that flexible, interactive tools for mining such data are missing. As Ribo-seq protocols still evolve, flexibility is highly desirable for the end-user. Here we describe RNA-Ribo-Explorer (RRE) a stand-alone tool that fills this gap. With RRE, one can explore read-count profiles of RNAs obtained after mapping, compare them between conditions, and visualize the profiles of individual RNAs. Importantly, the user can mine the data by defining queries that combine several criteria to detect interesting subsets of RNAs. For instance, one can ask RRE to find all RNAs whose translation of UTR region compared to that of the main ORF has changed between two conditions. This feature seems useful for finding candidate RNAs whose translation status or processing has changed across conditions. RRE is a platform independent software and is freely available at https://gite.lirmm.fr/rivals/RRE/-/releases .
Cancer stem cells (CSCs) are a small but critical cell population for cancer biology since they display inherent resistance to standard therapies and give rise to metastases. Despite accruing evidence establishing a link between deregulation of epitranscriptome-related players and tumorigenic process, the role of messenger RNA (mRNA) modifications in the regulation of CSC properties remains poorly understood. Here, we show that the cytoplasmic pool of fat mass and obesity-associated protein (FTO) impedes CSC abilities in colorectal cancer through its N 6 ,2’-O-dimethyladenosine (m 6 A m ) demethylase activity. While m 6 A m is strategically located next to the m 7 G-mRNA cap, its biological function is not well understood and has not been addressed in cancer. Low FTO expression in patient-derived cell lines elevates m 6 A m level in mRNA which results in enhanced in vivo tumorigenicity and chemoresistance. Inhibition of the nuclear m 6 A m methyltransferase, PCIF1/CAPAM, fully reverses this phenotype, stressing the role of m 6 A m modification in stem-like properties acquisition. FTO-mediated regulation of m 6 A m marking constitutes a reversible pathway controlling CSC abilities. Altogether, our findings bring to light the first biological function of the m 6 A m modification and its potential adverse consequences for colorectal cancer management.
MOTIVATION:Novel recombinant viruses may have important medical and evolutionary significance, as they sometimes display new traits not present in the parental strains. This is particularly concerning when the new viruses combine fragments coming from phylogenetically distinct viral types. Here, we consider the task of screening large collections of sequences for such novel recombinants. A number of methods already exist for this task. However, these methods rely on complex models and heavy computations that are not always practical for a quick scan of a large number of sequences. RESULTS:We have developed SHERPAS, a new program to detect novel recombinants and provide a first estimate of their parental composition. Our approach is based on the precomputation of a large database of 'phylogenetically-informed k-mers', an idea recently introduced in the context of phylogenetic placement in metagenomics. Our experiments show that SHERPAS is hundreds to thousands of times faster than existing software, and enables the analysis of thousands of whole genomes, or long-sequencing reads, within minutes or seconds, and with limited loss of accuracy. AVAILABILITY AND IMPLEMENTATION:The source code is freely available for download at https://github.com/phylo42/sherpas. SUPPLEMENTARY INFORMATION:Supplementary data are available at Bioinformatics online.
The hierarchical overlap graph (HOG) is a graph that encodes overlaps from a given set P of n strings, as the overlap graph does. A best known algorithm constructs HOG in O(||P|| log n) time and O(||P||) space, where ||P|| is the sum of lengths of the strings in P. In this paper we present a new algorithm to construct HOG in O(||P||) time and space. Hence, the construction time and space of HOG are better than those of the overlap graph, which are O(||P|| + n^2).
A superstring of a set of strings correspond to a string which contains all the other strings as substrings. The problem of finding the Shortest Linear Superstring is a well-know and well-studied problem in stringology. We present here a variant of this problem, the Shortest Circular Superstring problem where the sought superstring is a circular string. We show a strong link between these two problems and prove that the Shortest Circular Superstring problem is NP-complete. Moreover, we propose a new conjecture on the approximation ratio of the Shortest Circular Superstring problem.
MOTIVATION:Phylogenetic placement (PP) is a process of taxonomic identification for which several tools are now available. However, it remains difficult to assess which tool is more adapted to particular genomic data or a particular reference taxonomy. We developed Placement Evaluation WOrkflows (PEWO), the first benchmarking tool dedicated to PP assessment. Its automated workflows can evaluate PP at many levels, from parameter optimization for a particular tool, to the selection of the most appropriate genetic marker when PP-based species identifications are targeted. Our goal is that PEWO will become a community effort and a standard support for future developments and applications of PP.AVAILABILITY AND IMPLEMENTATION:https://github.com/phylo42/PEWO.SUPPLEMENTARY INFORMATION:Supplementary data are available at Bioinformatics online.
Thierry Lecroq合作论文数LITIS EA 4108
Universite de Rouen5