Understanding how mutations affect protein function remains critical yet challenging, particularly for variants in clinical databases lacking experimental characterisation and for intrinsically disordered regions. Current computational approaches often operate as black boxes, providing predictions without sufficient transparency or quality assessment of the underlying data. Here we present ProteoCast, a user-friendly web server that predicts variant effects through evolutionary constraint analysis and structural context integration. ProteoCast provides a three-tier variant classification (impactful, mild, neutral) to help prioritise mutations for clinical interpretation and experimental validation. It incorporates multiple sequence alignment quality controls to ensure prediction reliability and flag positions with insufficient evolutionary information. Beyond single-variant classification, ProteoCast employs a novel segmentation approach based on mutational sensitivity to identify functional linear peptides in disordered regions. Interactive visualisations guide users through results interpretation, from variant-level predictions to protein-wide functional landscapes. Evaluation on 63,000 ClinVar variants demonstrates 77 % sensitivity and 87 % specificity for pathogenicity prediction, with performance maintained across species (85 % accuracy on Drosophila lethal mutations). ProteoCast successfully identifies twice as many functional motifs in intrinsically disordered regions compared to conservation-based phylogenetic methods. Predictions can be tuned to specific conformations, such as bound forms in protein complexes, for improved accuracy and interpretability. With its transparent, unsupervised methodology and computational efficiency (minutes per protein), ProteoCast democratises access to variant effect prediction and functional site discovery for the broader research community. The web server is freely available at: https://proteocast.ijm.fr/.
Dissecting the functional impact of genetic mutations is essential to advancing our understanding of genotype-phenotype relationships and identifying therapeutic targets. Despite progress in sequencing and genome editing technologies, proteome-wide mutation effect prediction remains challenging. Here we show that evolutionary information alone enables accurate prediction of mutation effects across entire proteomes. ProteoCast is a scalable and interpretable computational method that leverages protein sequence conservation to classify genetic variants and identify functionally important protein sites. We apply ProteoCast to the complete Drosophila melanogaster proteome (22,000 isoforms, 300 million mutations) and validate it against nearly 400,000 natural and experimental variants. It correctly classifies 85% of known lethal mutations as functionally impactful versus 13-18% of population variants. ProteoCast-guided genome editing experiments confirm these predictions. Moreover, ProteoCast successfully identifies functionally important protein modification sites and binding motifs. ProteoCast provides a publicly available resource and deployable pipeline for studying gene function and mutations in any organism.
RNase J (RNJ) is a ribonuclease found in bacteria, archaea, and plant chloroplasts, and plays diverse roles in RNA maturation and stability. Chloroplast RNJ is encoded by the nuclear RNJ locus and is essential for embryo maturation. Arabidopsis or tobacco plants depleted for RNJ accumulate massive amounts of double-stranded RNA (dsRNA), which interferes with translation and causes chlorosis. Land plant RNJ uniquely contains a C-terminal GT1 domain, a DNA-binding motif found in transcription factors. Here, we have used complementation of an Arabidopsis rnj mutant with versions of RNJ with a mutated or deleted GT1 domain to investigate its role in RNJ function. We show that in vitro, the recombinant GT1 domain binds both dsRNA and DNA, but not single-stranded nucleic acids, with no sequence specificity. Furthermore, while RNJ lacking GT1 binding complements the rnj mutant, these plants accumulate high levels of dsRNA as detected by immunolocalization and RNA-Seq. GT1 mutations also change RNJ solubility in vivo, suggesting that the GT1 domain is involved in localization within the plastid. Taken together, our results suggest that the GT1 domain plays a key role in dsRNA removal through localizing the enzyme and/or selectively binding the dsRNA substrate.
Nanopore sequencing of full-length cDNAs offers unprecedented details of the plastid RNA metabolism. After the generation of the nanopore reads, several bioinformatic steps are required to analyze the data. In this chapter, we describe in a few simple command lines the processing and mapping of the reads as well as the generation of virtual Northern blots as a simple and familiar way to visualize Nanopore sequencing data.
Global understanding of plastid gene expression has always been impaired by its complexity. RNA splicing, editing, and intercistronic processing create multiple transcripts isoforms that can hardly be resolved using traditional molecular biology techniques. During the last decade, the wide adoption of RNA-seq-based techniques has, however, allowed an unprecedented understanding of all the different steps of chloroplast gene expression, from transcription to translation. Current strategies are nonetheless unable to identify and quantify full length transcripts isoforms, a limitation that can now be overcome using Nanopore Sequencing. We here provide a complete protocol to produce, from total leaf RNA, cDNA libraries ready for Nanopore sequencing. While most Nanopore protocols take advantage of the mRNA polyA tail we here first ligate an RNA adapter to the 3' ends of the RNAs and use it to initiate the template switching reverse transcription. The cDNA is then prepared and indexed for use with the regular Oxford Nanopore v14 chemistry. This protocol is of particular interest to researchers willing to simultaneously study the multiple post-transcriptional processes prevalent in the chloroplast.
Given a time series in $R^n$ with a piecewise constant mean and independent noises, we propose an exact dynamic programming algorithm to minimize a least square criterion with a multiscale penalty promoting well-spread changepoints. Such a penalty has been proposed in Verzelen et al. (2020), and it achieves optimal rates for changepoint detection and changepoint localization. Our proposed algorithm, named Ms.FPOP, extends functional pruning ideas of Rigaill (2015) and Maidstone et al. (2017) to multiscale penalties. For large signals, $n \geq 10^5$, with relatively few real changepoints, Ms.FPOP is typically quasi-linear and an order of magnitude faster than PELT. We propose an efficient C++ implementation interfaced with R of Ms.FPOP allowing to segment a profile of up to $n = 10^6$ in a matter of seconds. Finally, we illustrate on simple simulations that for large enough profiles ($n \geq 10^4$) Ms.FPOP using the multiscale penalty of Verzelen et al. (2020) is typically more powerfull than FPOP using the classical BIC penalty of Yao (1989).
To fully understand gene regulation, it is necessary to have a thorough understanding of both the transcriptome and the enzymatic and RNA-binding activities that shape it. While many RNA-Seq-based tools have been developed to analyze the transcriptome, most only consider the abundance of sequencing reads along annotated patterns (such as genes). These annotations are typically incomplete, leading to errors in the differential expression analysis. To address this issue, we present DiffSegR - an R package that enables the discovery of transcriptome-wide expression differences between two biological conditions using RNA-Seq data. DiffSegR does not require prior annotation and uses a multiple changepoints detection algorithm to identify the boundaries of differentially expressed regions in the per-base log2 fold change. In a few minutes of computation, DiffSegR could rightfully predict the role of chloroplast ribonuclease Mini-III in rRNA maturation and chloroplast ribonuclease PNPase in (3 '/5 ')-degradation of rRNA, mRNA and tRNA precursors as well as intron accumulation. We believe DiffSegR will benefit biologists working on transcriptomics as it allows access to information from a layer of the transcriptome overlooked by the classical differential expression analysis pipelines widely used today. DiffSegR is available at https://aliehrmann.github.io/DiffSegR/index.html.
Plant mitochondria represent the largest group of respiring organelles on the planet. Plant mitochondrial messenger RNAs (mRNAs) lack Shine-Dalgarno-like ribosome-binding sites, so it is unknown how plant mitoribosomes recognize mRNA. We show that “mitochondrial translation factors” mTRAN1 and mTRAN2 are land plant–specific proteins, required for normal mitochondrial respiration chain biogenesis. Our studies suggest that mTRANs are noncanonical pentatricopeptide repeat (PPR)–like RNA binding proteins of the mitoribosomal “small” subunit. We identified conserved Adenosine (A)/Uridine (U)-rich motifs in the 5′ regions of plant mitochondrial mRNAs. mTRAN1 binds this motif, suggesting that it is a mitoribosome homing factor to identify mRNAs. We demonstrate that mTRANs are likely required for translation of all plant mitochondrial mRNAs. Plant mitochondrial translation initiation thus appears to use a protein-mRNA interaction that is divergent from bacteria or mammalian mitochondria.
In this work we derive new analytic expressions for fixation time in Wright-Fisher model with selection. The three standard cases for fixation are considered: fixation to zero, to one or both. Second order differential equations for fixation time are obtained by a simplified approach using only the law of total probability and Taylor expansions. The obtained solutions are given by a combination of exponential integral functions with elementary functions. We then state approximate formulas involving only elementary functions valid for small selection effects. The quality of our results are explored throughout an extensive simulation study. We show that our results approximate the discrete problem very accurately even for small population size (a few hundreds) and large selection coefficients.
Plastid gene expression involves many post-transcriptional maturation steps resulting in a complex transcriptome composed of multiple isoforms. Although short-read RNA-Seq has considerably improved our understanding of the molecular mechanisms controlling these processes, it is unable to sequence full-length transcripts. This information is crucial, however, when it comes to understanding the interplay between the various steps of plastid gene expression. Here, we describe a protocol to study the plastid transcriptome using nanopore sequencing. In the leaf of Arabidopsis thaliana, with about 1.5 million strand-specific reads mapped to the chloroplast genome, we could recapitulate most of the complexity of the plastid transcriptome (polygenic transcripts, multiple isoforms associated with post-transcriptional processing) using virtual Northern blots. Even if the transcripts longer than about 2500 nucleotides were missing, the study of the co-occurrence of editing and splicing events identified 42 pairs of events that were not occurring independently. This study also highlighted a preferential chronology of maturation events with splicing happening after most sites were edited.