Advances in sequencing technologies have revolutionised our ability to capture the complete RNA profile in tissue samples, allowing for comparative analyses of RNA levels between developmental processes, environmental responses, or treatments. However, quantifying changes in gene expression remains challenging, given inherent biological variability and technological limitations. To address this, we introduce a Bayesian framework for differential gene expression (DGE) analysis. Our framework unifies and streamlines a complex analysis, typically involving parameter estimations and multiple statistical tests, into a concise mathematical equation. This allows statistical evidence for differential expression to be computed rapidly and transparently. We show how this approach can be used to evaluate variabilty of individual genes between replicates. A comparison of our framework with existing tools revealed differences that can be explained by commonly employed thresholds in other packages. This motivated us to explore ranking genes based on their statistical evidence as opposed to a binary classification as DEGs. Our analysis leads us to advocate the use of Bayes factors within a rank-based approach. This framework offers enhanced computational efficiency and delivers a transparent way to analyse, interpret and communicate DGE results. ### Competing Interest Statement The authors have declared no competing interest.
The concept of 'Differentially Expressed Genes' (DEGs) is central to RNA-Seq studies, yet their identification suffers from reproducibility issues. This is largely a consequence of the inherent biological and technical variation that cannot be captured with small numbers of replicates. When thresholds for p-values and log fold changes are introduced, this variability can propagate an incomplete description of the data, leading to differing interpretations. Here, we compare traditional binary DEG classification with a rank-based method, grounded in Bayesian statistics, using a published yeast dataset comprising over 40 replicates. This analysis reveals how the choice of thresholds and number of replicates results in discrepancies between studies and potentially interesting genes being overlooked. Furthermore, by comparing wild-type with wild-type samples, we show how variability in gene expression can be mistaken for differential expression. Evaluating current practices for navigating the accuracy-error trade-off in the search for differentially expressed genes leads us to advocate rank-based methods and Bayesian statistics to mitigate the limitations of binary classifications and communicate uncertainty. ### Competing Interest Statement The authors have declared no competing interest.
Stomata regulate plant gas exchange via repeated turgor-driven changes of guard cell shape, thereby adjusting pore apertures. Grasses, which are among the most widespread plant families on the planet, are distinguished by their unique stomatal structure, which is proposed to have significantly contributed to their evolutionary and agricultural success. One component of their structure, which has received little attention, is the presence of a discontinuous adjoining cell wall of the guard cell pair. Here, we demonstrate the presence of these symplastic connections in a range of grasses and use finite element method simulations to assess hypotheses for their functional significance. Our results show that opening of the stomatal pore is maximal when the turgor pressure in dumbbell-shaped grass guard cells is equal, especially under the low pressure conditions that occur during the early phase of stomatal opening. By contrast, we demonstrate that turgor pressure differences have less effect on the opening of kidney-shaped guard cells, characteristic of the majority of land plants, where guard cell connections are rarely or not observed. Our data describe a functional mechanism based on cellular mechanics, which plausibly facilitated a major transition in plant evolution and crop development.
The mathematics of mass screening is crucial for understanding results from im-perfect detection methods in large datasets. Although this framework is well-established, it is not commonly applied in plant biology. Here, we view the iden-tification of messenger RNAs that travel over long distances in plants (mobile mRNAs) through the lens of mass screening statistics. RNA-Seq analyses have identified thousands of mobile mRNAs. Consideration of the detection accuracy and prevalence, however, cast doubt on these numbers. The presented method-ology is relevant to all areas of research where detection tests with less than 100% accuracy are applied to find rare events in large datasets.
Short-read RNA-seq studies of grafted plants have led to the proposal that thousands of messenger RNAs (mRNAs) move over long distances between plant tissues1–7, potentially acting as signals8–12. Transport of mRNAs between cells and tissues has been shown to play a role in several physiological and developmental processes in plants, such as tuberization13, leaf development14 and meristem maintenance15; yet for most mobile mRNAs, the biological relevance of transport remains to be determined16–19. Here we perform a meta-analysis of existing mobile mRNA datasets and examine the associated bioinformatic pipelines. Taking technological noise, biological variation, potential contamination and incomplete genome assemblies into account, we find that a high percentage of currently annotated graft-mobile transcripts are left without statistical support from available RNA-seq data. This meta-analysis challenges the findings of previous studies and current views on mRNA communication. This study reveals that a substantial number of transcripts that are currently annotated as graft-mobile lack statistical support from available RNA-seq data.
Sequencing technologies have revolutionised the field of molecular biology. We now have the ability to routinely capture the complete RNA profile in tissue samples. This wealth of data allows for comparative analyses of RNA levels at different times, shedding light on the dynamics of developmental processes, and under different environmental responses, providing insights into gene expression regulation and stress responses. However, given the inherent variability of the data stemming from biological and technological sources, quantifying changes in gene expression proves to be a statistical challenge. Here, we present a closed-form Bayesian solution to this problem. Our approach is tailored to the differential gene expression analysis of processed RNA-Seq data. The framework unifies and streamlines an otherwise complex analysis, typically involving parameter estimations and multiple statistical tests, into a concise mathematical equation for the calculation of Bayes factors. Using conjugate priors we can solve the equations analytically. For each gene, we calculate a Bayes factor, which can be used for ranking genes according to the statistical evidence for the gene's expression change given RNA-Seq data. The presented closed-form solution is derived under minimal assumptions and may be applied to a variety of other 2-sample problems.
Short-read RNA-Seq analyses of grafted plants have led to the proposal that large numbers of mRNAs move over long distances between plant tissues, acting as potential signals. The detection of transported transcripts by RNA-Seq is both experimentally and computationally challenging, requiring successful grafting, delicate harvesting, rigorous contamination controls and data processing approaches that can identify rare events in inherently noisy data. Here, we perform a meta-analysis of existing datasets and examine the associated bioinformatic pipelines. Our analysis reveals that technological noise, biological variation and incomplete genome assemblies give rise to features in the data that can distort the interpretation. Taking these considerations into account, we find that a substantial number of transcripts that are currently annotated as mobile are left without support from the available RNA-Seq data. Whilst several annotated mobile mRNAs have been validated, we cannot exclude that others may be false positives. The identified issues may impact also other RNA-Seq studies, in particular those using single nucleotide polymorphisms (SNPs) to detect variants. ### Competing Interest Statement The authors have declared no competing interest.
In Arabidopis a high number of distinct mRNAs move from shoot to root. We previously reported on the correlation of m5C-methylation and lack of mRNA transport in juvenile plants depending on the RNA methyltransferases DNMT2 NSUN2B . However, to our surprise we uncovered that lack of DNMT2 NSUN2B (writer) activity did not abolished transport of TCTP1 and HSC70.1 transcripts in flowering plants. We uncovered that transport of both transcripts is reinstated in dnmt2 nsun2b mutants after commitment to flowering. This finding suggests that additional factors are seemingly involved in regulating / mediating mRNA transport. In search of such candidates, we identified the two ALY2 and ALY4 nuclear mRNA export factors belonging to the ALYREF family as bona fide m5C readers mediating mRNA transport. We show that both proteins are allocated along the phloem and that they bind preferentially to mobile mRNAs. MST measurements indicate that ALY2 and ALY4 bind to mobile mRNAs with relative high affinity with ALY4 showing higher affinity towards m5C-methylated mobile mRNAs. An analysis of the graft-mobile transcriptome of juvenile heterografted-grafted wild type, dnmt2 nsun2b , aly2 and aly4 mutants revealed that the nuclear export factors are key regulators of mRNA transport. We suggest that depending on the developmental stage m5C methylation has a negative and positive regulatory function in mRNA transport and acts together with ALY2 and ALY4 to facilitate mRNA transport in both juvenile and flowering plants. ### Competing Interest Statement The authors have declared no competing interest.
Abstract Background: A popular method for the detection of graft-mobile mRNA on a genomic scale in plants is to perform RNA-Seq on heterografts comprising different ecotypes, species, or cultivars (types). Transcripts from one plant type that are detected in the other type can be assumed to have been transported across the graft junction. A necessary step is the ability to differentiate between transcripts from each plant type, which can be achieved based on known single nucleotide polymorphisms (SNPs) between types. One current challenge is to differentiate between RNA-Seq reads associated with specific SNPs and RNA-Seq errors. Here, we present baymobil, a Python package that implements a Bayesian framework for determining between graft-mobile transcripts and sequencing errors. Results: baymobil takes processed RNA-Seq data from homo- and heterografted plant samples as input and for each mRNA calculates a Bayes factor: the ratio of the evidence for transcripts having crossed the graft junction over the evidence for the data being a consequence of sequencing errors. These Bayes factors can be used to rank mRNAs based on the statistical support for their mobility. Furthermore, the package includes functions for the creation of simulated RNA-Seq datasets, which have been used to demonstrate that this approach outperforms existing mobility criteria and is a necessary step in the successful detection of graft-mobile mRNA. Conclusion: The baymobil Python package improves the accuracy of pipelines for the detection of graft-mobile mRNA based on SNP differences between heterografted types. It is openly and freely available via GitHub, and is easily installed with pip and pypi. A detailed tutorial with test data and results for the statistics and simulation package baymobil can be found on github.com/mtomtom/baymobil.
Summary The HSC70/HSP70 family of heat shock proteins are evolutionarily conserved chaperones involved in protein folding, protein transport, and RNA binding. Arabidopsis HSC70 chaperones are thought to act as housekeeping chaperones and as such are involved in many growth‐related pathways. Whether Arabidopsis HSC70 binds RNA and whether this interaction is functional has remained an open question. We provide evidence that the HSC70.1 chaperone binds its own mRNA via its C‐terminal short variable region (SVR) and inhibits its own translation. The SVR encoding mRNA region is necessary for HSC70.1 transcript mobility to distant tissues and that HSC70.1 transcript and not protein mobility is required to rescue root growth and flowering time of hsc70 mutants. We propose that this negative protein‐transcript feedback loop may establish an on‐demand chaperone pool that allows for a rapid response to stress. In summary, our data suggest that the Arabidopsis HSC70.1 chaperone can form a complex with its own transcript to regulate its translation and that both protein and transcript can act in a noncell‐autonomous manner, potentially maintaining chaperone homeostasis between tissues.
Plants are complex systems made up of many interacting components, ranging from architectural elements such as branches and roots, to entities comprising cellular processes such as metabolic pathways and gene regulatory networks. The collective behaviour of these components, along with the plant's response to the environment, give rise to the plant as a whole. Properties that result from these interactions and cannot be attributed to individual parts alone are called emergent properties, occurring at different time and spatial scales. Deepening our understanding of plant growth and development requires computational tools capable of handling a large number of interactions and a multiscale approach connecting properties across scales. There currently exist few methods able to integrate models across scales, or models capable of predicting new emergent plant properties. This perspective explores current approaches to modelling emergent behaviour in plants, with a focus on how current and future tools can handle multiscale plant systems.
Spikelets are the fundamental building blocks of Poaceae inflorescences and their development and branching patterns determine the various inflorescence architectures and grain yield of grasses. In wheat, the central spikelets produce the most and largest grains, while spikelet size gradually decreases acro- and basipetally, giving rise to the characteristic lanceolate shape of wheat spikes. The acropetal gradient correlates with the developmental age of spikelets, however the basal spikelets are developed first and the cause of their small size and rudimentary development is unclear. Here, we adapted G&T-seq, a low-input transcriptomics approach, to characterise gene expression profiles within spatial sections of individual spikes before and after the establishment of the lanceolate shape. We observed larger differences in gene expression profiles between the apical, central and basal sections of a single spike than between any section belonging to consecutive developmental timepoints. We found that SVP MADS-box transcription factors, including VRT-A2, are expressed highest in the basal section of the wheat spike and display the opposite expression gradient to flowering E-class SEP1 genes. Based on multi-year field trials and transgenic lines, we show that higher expression of VRT-A2 in the basal sections of the spike is associated with increased numbers of rudimentary basal spikelets. Our results, supported by computational modelling, suggest that the delayed transition of basal spikelets from vegetative to floral developmental programmes results in the lanceolate shape of wheat spikes. This study highlights the value of spatially resolved transcriptomics to gain new insights into developmental genetics pathways of grass inflorescences. One sentence summary Large transcriptional gradients exist within a wheat spike and are associated with rudimentary basal spikelet development, resulting in the characteristic lanceolate shape of wheat spikes.
Abstract Insect-vectored plant pathogens cause massive crop yield losses worldwide. Outbreaks often have patchy incidences on both spatial and temporal scales, which are thought to be driven by external factors such as the level of heterogeneity in plant hosts and vector dispersal. Mathematical modelling of Aster Yellows phytoplasma disease spread, a vector-borne plant pathogen, predicts under the assumption of well-mixed population dynamics only one possible stable outcome: all insect vectors end up being carriers of the pathogen. This outcome is in stark contrast to multi-year datasets revealing local co-existence of carrier and non-carrier vectors. We predict however dramatically different tripartite plant-vector-pathogen dynamics when explicitly taking space into account. Spatial simulations allow to explain both the stable coexistence of carrier and non-carrier insects and the large spatial and temporal variations in disease incidence, this co-existence being driven by emergent patterning. Moreover, simulations of the eco-evolutionary progression of the infection dynamics predict an evolutionary stable state that quantitatively matches field observations and prove that empirically observed complex patterning is the inevitable evolutionary outcome of this tripartite system. We infer that the known diversity in phytoplasma effector proteins could be an evolved mechanism for pathogen survival in fast-changing environments, such as agricultural areas.
Spikelets are the fundamental building blocks of Poaceae inflorescences, and their development and branching patterns determine the various inflorescence architectures and grain yield of grasses. In wheat (Triticum aestivum), the central spikelets produce the most and largest grains, while spikelet size gradually decreases acropetally and basipetally, giving rise to the characteristic lanceolate shape of wheat spikes. The acropetal gradient corresponds with the developmental age of spikelets; however, the basal spikelets are developed first, and the cause of their small size and rudimentary development is unclear. Here, we adapted G&T-seq, a low-input transcriptomics approach, to characterize gene expression profiles within spatial sections of individual spikes before and after the establishment of the lanceolate shape. We observed larger differences in gene expression profiles between the apical, central, and basal sections of a single spike than between any section belonging to consecutive developmental time points. We found that SHORT VEGETATIVE PHASE MADS-box transcription factors, including VEGETATIVE TO REPRODUCTIVE TRANSITION 2 (VRT-A2), are expressed highest in the basal section of the wheat spike and display the opposite expression gradient to flowering E-class SEPALLATA1 genes. Based on multi-year field trials and transgenic lines, we show that higher expression of VRT-A2 in the basal sections of the spike is associated with increased numbers of rudimentary basal spikelets. Our results, supported by computational modeling, suggest that the delayed transition of basal spikelets from vegetative to floral developmental programs results in the lanceolate shape of wheat spikes. This study highlights the value of spatially resolved transcriptomics to gain insights into developmental genetics pathways of grass inflorescences.
The long-distance transport of messenger RNAs (mRNAs) has been shown to be important for several developmental processes in plants. A popular method for identifying travelling mRNAs is to perform RNA-Seq on grafted plants. This approach depends on the ability to correctly assign sequenced mRNAs to the genetic background from which they originated. The assignment is often based on the identification of single-nucleotide polymorphisms (SNPs) between otherwise identical sequences. A major challenge is therefore to distinguish SNPs from sequencing errors. Here, we show how Bayes factors can be computed analytically using RNA-Seq data over all the SNPs in an mRNA. We used simulations to evaluate the performance of the proposed framework and demonstrate how Bayes factors accurately identify graft-mobile transcripts. The comparison with other detection methods using simulated data shows how not taking the variability in read depth, error rates and multiple SNPs per transcript into account can lead to incorrect classification. Our results suggest experimental design criteria for successful graft-mobile mRNA detection and show the pitfalls of filtering for sequencing errors or focusing on single SNPs within an mRNA.
Transport across membranes is critical for plant survival. Membranes are the interfaces at which plants interact with their environment. The transmission of energy and molecules into cells provides plants with the source material and power to grow, develop, defend, and move. An appreciation of the physical forces that drive transport processes is thus important for understanding the plant growth and development. We focus on the passive transport of molecules, describing the fundamental concepts and demonstrating how different levels of abstraction can lead to different interpretations of the driving forces. We summarize recent developments on quantitative frameworks for describing diffusive and bulk flow transport processes in and out of cells, with a more detailed focus on plasmodesmata, and outline open questions and challenges.
Molecular communication is key for multicellular organisms. In plants, the exchange of nutrients and signals between cells is facilitated by tunnels called plasmodesmata. Such transport processes in complex geometries can be simulated using particle-based approaches, these, however, are computationally expensive. Here, we evaluate the narrow escape problem as a framework for describing intercellular transport. We introduce a volumetric adjustment factor for estimating escape times from non-spherical geometries. We validate this approximation against full 3D stochastic simulations and provide results for a range of cell sizes and diffusivities. We discuss how this approach can be extended using recent results on multiple trap problems to account for different plasmodesmata distributions with varying apertures.
In biology, many research questions focus on uncovering the mechanisms that allow particles (molecules, proteins, etc.) to move from one location to another.Often these movements are from one domain to another and these domains are in some way contained.The "narrow escape problem" is a biophysics problem where the solution would provide the average time required for a Brownian particle to escape a bounded domain through a particular opening.Originally proposed by Holcman & Schuss (2004), solutions to this problem have been given and refined over the years (Schuss et al., 2007).Recently, solutions have been given (Kaye & Greengard, 2020) for more complex narrow escape problems, such as arbitrary escape pore patterning and size variation.However, these are provided without easily accessible implementations and confine the problem to a single container shape.