Predicting RNA splicing from genomic sequence is a crucial task for understanding gene regulation and interpreting genetic variation. Recent deep learning advancements have led to splicing prediction algorithms that achieve state-of-the-art performance compared to earlier models. However, owing to the limited interpretability of deep learning models, the predictive mechanisms of current splicing models remain poorly understood. Here we develop a framework to explain model prediction logic using interpretable distillation. Applying our framework, we find that RNA splicing prediction models suffer from pervasive confounders and blind spots, leading to poor performance on non-reference sequences. We find that splicing models recognize exons through surprisingly simple additive combinations of sequence motifs, including known splicing regulatory elements. Critically, our analysis also reveals that splicing models exploit genomic confounders unrelated to splicing and fail to adequately capture the effects of RNA structure, leading to systematic prediction errors. Our findings illuminate fundamental limitations of training models on genomic sequences and suggest ways to overcome them.
We prove that for any integral lattice $\mathcal{L} \subset \mathbb{R}^n$ (that is, a lattice $\mathcal{L}$ such that the inner product $\langle \mathbf{y}_1,\mathbf{y}_2 \rangle$ is an integer for all $\mathbf{y}_1, \mathbf{y}_2 \in \mathcal{L}$) and any positive integer $k$, \[ |\{ \mathbf{y} \in \mathcal{L} \ : \ \|\mathbf{y}\|^2 = k\}| \leq 2 \binom{n+2k-2}{2k-1} \; , \] giving a nearly tight reverse Minkowski theorem for integral lattices.
Nuclear speckles are membraneless organelles implicated in multiple RNA processing steps. In this work, we systematically characterize the sequence logic determining RNA localization to nuclear speckles. We find extensive similarities between the speckle localization code and the RNA splicing code, even for transcripts that do not undergo splicing. Specifically, speckle localization is enhanced by the presence of unspliced exon-like or intron-like sequence features. We demonstrate that interactions required for early spliceosomal complex assembly contribute to speckle localization. We also show that speckle localization of isolated endogenous exons is reduced by disease-associated single nucleotide variants. Finally, we find that speckle localization strongly correlates with splicing kinetics of splicing-competent constructs and is linked to the decision between exon inclusion and skipping. Together, these results suggest a model in which RNA speckle localization is associated with the formation of the early spliceosomal complex and enhances the efficiency of splicing reactions.
We show that $n$-bit integers can be factorized by independently running a quantum circuit with $\tilde{O}(n^{3/2})$ gates for $\sqrt{n}+4$ times, and then using polynomial-time classical post-processing. The correctness of the algorithm relies on a number-theoretic heuristic assumption reminiscent of those used in subexponential classical factorization algorithms. It is currently not clear if the algorithm can lead to improved physical implementations in practice.
Let $K$ be a convex body in $\mathbb{R}^n$, let $L$ be a lattice with covolume one, and let $\eta>0$. We say that $K$ and $L$ form an $\eta$-smooth cover if each point $x \in \mathbb{R}^n$ is covered by $(1 \pm \eta) vol(K)$ translates of $K$ by $L$. We prove that for any positive $\sigma, \eta$, asymptotically as $n \to \infty$, for any $K$ of volume $n^{3+\sigma}$, one can find a lattice $L$ for which $L, K$ form an $\eta$-smooth cover. Moreover, this property is satisfied with high probability for a lattice chosen randomly, according to the Haar-Siegel measure on the space of lattices. Similar results hold for random construction A lattices, albeit with a worse power law, provided the ratio between the covering and packing radii of $\mathbb{Z}^n$ with respect to $K$ is at most polynomial in $n$. Our proofs rely on a recent breakthrough by Dhar and Dvir on the discrete Kakeya problem.
Recent advances in ion mobility spectrometry have enabled the measurement of rotationally averaged collisional cross-sectional area (CCS) for millions of peptides as part of routine proteomic mass spectrometry workflows. One of the most striking findings in recent large ion mobility data sets is that CCS exhibits two distinct modes, most notably for charge 3+ peptides, with peptides predominantly exhibiting CCS in either the high or low mode. Here, using classical machine learning approaches, we identify that basic site positioning is a key sequence feature determining a peptide's CCS mode. Molecular dynamics simulations suggest that peptides in the high CCS mode tend to adopt more extended conformations and form charge-stabilized helical structures, whereas those in the low CCS mode adopt more compact, globular conformations. Further supporting this structural hypothesis, we provide evidence for preferential protonation near the C-terminus and uncover multiple position-dependent sequence determinants that all suggest the predominance of helix formation in the high CCS mode. Together, these findings will enable better integration of CCS measurements into protein identification and quantification pipelines, improving the performance of ion mobility-based proteomics.
We prove a conjecture due to Dadush, showing that if ℒ⊂ ℝ n is a lattice such that det(ℒ′) 1 for all sublattices ℒ′ ⊆ ℒ, then $$\sum_{ y ∈ℒ}^e -t 2||y||2 ≤3/2,$$ where t := 10(logn + 2). From this we also derive bounds on the number of short lattice vectors and on the covering radius.
RNA molecules often play critical roles in assisting the formation of membraneless organelles in eukaryotic cells. Yet, little is known about the organization of RNAs within membraneless organelles. Here, using super-resolution imaging and nuclear speckles as a model system, we demonstrate that different sequence domains of RNA transcripts exhibit differential spatial distributions within speckles. Specifically, we image transcripts containing a region enriched in SR protein binding motifs and another region enriched in hnRNP binding motifs. We show that these transcripts localize to the outer shell of speckles, with the SR motif-rich region localized closer to the speckle center relative to the hnRNP motif-rich region. Further, we identify that this intra-speckle RNA organization is driven by the strength of RNA-protein interactions inside and outside speckles. Our results hint at novel functional roles of nuclear speckles and likely other membraneless organelles in organizing RNA substrates for biochemical reactions.
Nuclear speckles, a type of membraneless nuclear organelle in higher eukaryotic cells, play a vital role in gene expression regulation. Using the reverse transcription-based RNA-binding protein binding sites sequencing (ARTR-seq) method, we study human transcripts associated with nuclear speckles. We identify three gene groups whose transcripts demonstrate different speckle localization properties and dynamics: stably enriched in nuclear speckles, transiently enriched in speckles at the pre-mRNA stage, and not enriched in speckles. Specifically, we find that stably-enriched transcripts contain inefficiently spliced introns. We show that nuclear speckles specifically facilitate splicing of speckle-enriched transcripts. We further reveal RNA sequence features contributing to transcript speckle localization, underscoring a tight interplay between genome organization, RNA cis-elements, and transcript speckle enrichment, and connecting transcript speckle localization with splicing efficiency. Finally, we show that speckles can act as hubs for the regulated retention of introns during cellular stress. Collectively, our data highlight a role of nuclear speckles in both co- and post-transcriptional splicing regulation.
In this review we give an overview of some recent "reverse Minkowski" results on the geometry of lattices. Such results provide upper bounds on the number of short vectors a lattice can have, assuming that it does not have any sublattice of low determinant. We also briefly describe the proof ideas, and mention some open questions.
Nuclear speckles represent a type of membraneless organelles in higher eukaryotic cells which are rich in snRNP components, certain splicing factors, including SR proteins (a family of RNA binding proteins (RBPs) named for containing regions with repetitive serine and arginine residues), polyA+ RNAs, and certain long noncoding RNAs (lncRNA). Multivalent interactions between RNAs and residing protein components drive the formation, as well as layered organization of components in nuclear speckles. Furthermore, RBPs have differential localization with respect to nuclear speckles; while SRSF1 proteins is enriched in nuclear speckles, heterogeneous nuclear ribonucleoprotein A1 (hnRNPA1) is either uniformly distributed in the nucleoplasm or slightly depleted from nuclear speckles. Using super-resolution imaging we show that the pre-mRNA transcripts containing SRSF1 motifs in the exon and hnRNPA1 motifs in the intron, are orientated in a way such that the regions containing the SRSF1 binding motifs are positioned relatively towards the interior of nuclear speckles and hnRNPA1 binding motifs are placed relatively towards the periphery of speckles. Knocking down of SRSF1 protein leads to the migration of RNA transcripts further towards the periphery of nuclear speckles, whereas knocking down of hnRNPA1 protein leads to the inward migration of RNA transcripts with respect to the speckle center. A similar effect of protein knockdown is seen on the endogenous nuclear speckle associated lncRNA, MALAT1. Our results collectively demonstrate that RNAs can adopt non-random orientation and positioning within the membraneless organelles. These observations might help explain the need for nuclear RNAs for structural integrity of nuclear speckles and the context dependent effect of SR and hnRNP proteins on splicing outcomes.
We prove that for any n∈N there is a convex body K⊆Rn whose surface area is at most n12+o(1), yet the translates of K by the integer lattice Zn tile Rn.
Electrospray ionization is a powerful and prevalent technique used to ionize analytes in mass spectrometry. The distribution of charges that an analyte receives (charge state distribution, CSD) is an important consideration for interpreting mass spectra. However, due to an incomplete understanding of the ionization mechanism, the analyte properties that influence CSDs are not fully understood. Here, we employ a machine learning-based high-throughput approach and analyze CSDs of hundreds of thousands of peptides. Interestingly, half of the peptides exhibit charges that differ from what one would naively expect (number of basic sites). We find that these peptides can be classified into two regimes-undercharging and overcharging-and that these two regimes display markedly different charging characteristics. Strikingly, peptides in the overcharging regime show minimal dependence on basic site count, and more generally, the two regimes exhibit distinct sequence determinants. These findings highlight the rich ionization behavior of peptides and the potential of CSDs for enhancing peptide identification.
Machine learning methods, particularly neural networks trained on large datasets, are transforming how scientists approach scientific discovery and experimental design. However, current state-of-the-art neural networks are limited by their uninterpretability: Despite their excellent accuracy, they cannot describe how they arrived at their predictions. Here, using an "interpretable-by-design" approach, we present a neural network model that provides insights into RNA splicing, a fundamental process in the transfer of genomic information into functional biochemical products. Although we designed our model to emphasize interpretability, its predictive accuracy is on par with state-of-the-art models. To demonstrate the model's interpretability, we introduce a visualization that, for any given exon, allows us to trace and quantify the entire decision process from input sequence to output splicing prediction. Importantly, the model revealed uncharacterized components of the splicing logic, which we experimentally validated. This study highlights how interpretable machine learning can advance scientific discovery.
Electrospray ionization is a powerful and prevalent technique used to ionize analytes in mass spectrometry. The distribution of charges that an analyte receives (charge state distribution, CSD) is an important consideration for interpreting mass spectra. However, due to an incomplete understanding of the ionization mechanism, the analyte properties that influence CSDs are not fully understood. Here, we employ a machine learning-based approach and analyze CSDs of hundreds of thousands of peptides. Interestingly, half of the peptides exhibit charges that differ from what one would naively expect (the number of basic sites). We find that these peptides can be classified into two regimes (undercharging and overcharging) and that these two regimes display markedly different charging characteristics. Notably, peptides in the overcharging regime show minimal dependence on basic site count, and more generally, the two regimes exhibit distinct sequence determinants. These findings highlight the rich ionization behavior of peptides and the potential of CSDs for enhancing peptide identification.
SummaryMachine learning methods, particularly neural networks trained on large datasets, are transforming how scientists approach scientific discovery and experimental design. However, current state-of-the-art neural networks are limited by their uninterpretability: despite their excellent accuracy, they cannot describe how they arrived at their predictions. Here, using an “interpretable-by-design” approach, we present a neural network model that provides insights into RNA splicing, a fundamental process in the transfer of genomic information into functional biochemical products. Although we designed our model to emphasize interpretability, its predictive accuracy is on par with state-of-the-art models. To demonstrate the model’s interpretability, we introduce a visualization that, for any given exon, allows us to trace and quantify the entire decision process from input sequence to output splicing prediction. Importantly, the model revealed novel components of the splicing logic, which we experimentally validated. This study highlights how interpretable machine learning can advance scientific discovery.
We obtain new upper bounds on the minimal density of lattice coverings of Euclidean space by dilates of a convex body K. We also obtain bounds on the probability (with respect to the natural Haar-Siegel measure on the space of lattices) that a randomly chosen lattice L satisfies that L+K is all of space. As a step in the proof, we utilize and strengthen results on the discrete Kakeya problem.
We prove that if ℒ⊂ℝ^n is a lattice such that (ℒ') ≥ 1 for all sublattices ℒ' ⊆ℒ, then ∑_𝐲∈ℒ 𝐲≠0 (𝐲^2+q)^-s≤∑_𝐳∈ℤ^n 𝐳≠0 (𝐳^2+q)^-s for all s > n/2 and all 0 ≤ q ≤ (2s-n)/(n+2), with equality if and only if ℒ is isomorphic to ℤ^n.
Yossi Azar合作论文数Blavatnik School of Computer Science, Tel-Aviv University9