All life depends on the reliable translation of RNA to protein according to complex interactions between translation machinery and RNA sequence features. While ribosomal occupancy and codon frequencies vary across coding regions, well-established metrics for computing coding potential of RNA do not capture such positional dependence. Here, we investigate positional bias in codon usage, which contextually accounts for the position of protein-coding signals embedded within coding regions. We demonstrate the existence of position-dependent patterns of codon frequency in the human transcriptome and describe these patterns using our POsition-Specific Codon Occurrence (POSCO) score that is more consistently associated with translation-initiating codons than other common sequence features. We further show that the patterns described by POSCO are not accounted for by other common scores, including position-dependent GC content, consensus sequences, and the presence of signal peptides in the translation product. More importantly, POSCO defines a spectrum of translational efficiency and local tRNA adaptation index (tAI). High POSCO scores correspond to the highest initial tAI values, which increase over the downstream length of the transcript, forming a translational highway. Meanwhile, low POSCO scores exhibit the lowest initial tAI values followed by a previously undescribed translational valley. An inverse correlation was found between POSCO score and ribosomal occupancy near the start codon. We also find that POSCO defines a spectrum of local folding energies, with high-POSCO transcripts showing the most stable folding immediately after the start codon. Finally, we examine the relationship between POSCO intensity and functional enrichment. We find that transcripts with start codons showing the highest POSCO are enriched for functions relating to development of musculoskeletal, cardiovascular, neurological, gastrointestinal, sensory, and other body systems. Furthermore, transcripts with high POSCO are depleted for functions related to immune response and detection of chemical stimulus. These findings lay important groundwork to improve our understanding of the regulation of translation, the calculation of coding potential, and the classification of RNA transcripts.
Single-cell proteomics enables direct measurement of cellular heterogeneity during dynamic biological processes, but its application to fragile and highly adherent neuronal models remains challenging. Here, we developed and applied an optimized single-cell proteomics workflow to characterize proteome remodeling during nerve growth factor (NGF)-induced differentiation of PC12 cells. To enable reliable single-cell analysis, we implemented gentle dissociation, antiaggregation strategies, and thermal inkjet-based cell dispensing, achieving high accuracy in single-cell isolation. Inclusion of n-dodecyl-β-d-maltoside (DDM) improved recovery of membrane-associated and low-solubility proteins. Coupled with LC-ion mobility-mass spectrometry, this workflow enabled quantification of 2,000-3,000 proteins per cell across the differentiation time course. Single-cell proteomic analysis revealed progressive and heterogeneous proteome remodeling during differentiation. While undifferentiated cells formed a relatively homogeneous population, later stages (Days 4-6) exhibited increased variability, including multimodal protein abundance distributions and separation into distinct subpopulations. Dimensionality reduction, clustering, and non-negative matrix factorization identified multiple coexisting proteomic states within the same time points, reflecting asynchronous differentiation trajectories. These subpopulations were characterized by coordinated differences in pathways related to intracellular trafficking, protein translation, cytoskeletal organization, and neuronal maturation. Comparison with bulk proteomics demonstrated that proteins associated with differentiated neuronal states, including those involved in neurite formation and structural remodeling, are underrepresented in population-averaged measurements but are enriched within specific single-cell subpopulations. Temporal and cluster-resolved analyses further revealed distinct protein expression trajectories, including early decreases in cell cycle and metabolic pathways and later increases in neuronal structural and regulatory proteins. Together, this study establishes an optimized workflow for single-cell proteomics of neuronal systems and demonstrates that NGF-induced PC12 differentiation proceeds through heterogeneous and divergent proteomic states that are not resolved by bulk analysis.
Triple-negative breast cancer (TNBC) is an aggressive breast cancer subtype associated with poor clinical outcomes and limited treatment options. The aryl hydrocarbon receptor (AHR) is a ligand-activated transcription factor that regulates xenobiotic metabolism through the induction of many cytochrome P450 enzymes such as CYP1A1. In this study, we investigated three 2-phenyl-imidazo[1,2-a]pyridine derivates (X19724, X19728, and X15695), previously developed as selective estrogen receptor degraders, for their activity in TNBC cell lines. All three compounds activated AHR-dependent transcriptional activity and exhibited selective cytotoxicity toward MDA-MB-468 cells among a panel of TNBC and nontumorigenic mammary epithelial cell lines. Nonresponsive TNBC cell lines lacked either detectable AHR expression or inducible CYP1A1 expression. In sensitive MDA-MB-468 cells, cytotoxicity was accompanied by membrane permeabilization. Genetic ablation or pharmacological antagonism of AHR, as well as inhibition of CYP1A1 activity, markedly attenuated compound-induced cytotoxicity. Proteomic analysis further revealed enrichment of pathways related to phospholipid synthesis and base excision repair, along with increased abundance of proteins associated with DNA damage responses. Finally, among the three compounds, X15695 exhibited the greatest metabolic stability following incubation with human liver microsomes. Collectively, these findings support a model in which activation of AHR-CYP1A1 axis contributes to selective cytotoxicity in TNBC cells, highlighting this pathway as a potential mechanistic vulnerability that can be exploited by imidazopyridine derivatives.
The aryl hydrocarbon receptor (AhR) is a ligand-activated transcription factor best known for mediating biological responses to a wide range of xenobiotics, such as dioxins and polycyclic aromatic hydrocarbons. Recently, AhR has emerged as an important player in cancer biology, with the potential for therapeutic applications through targeted modulation of its activity in specific cancer types. In this study, we report that 4,11-dichloro-BBQ (DiCl-BBQ), a benzimidazoisoquinoline, exhibits AhR-mediated antiproliferative activity in HepG2 hepatocellular carcinoma cells. DiCl-BBQ was found to decrease cell growth at nanomolar concentrations, and this antiproliferative effect persisted even after the compound's removal. Using inducible shRNA expression system, we demonstrated that the inhibitory effect of DiCl-BBQ was significantly reduced following AhR knockdown. Flow cytometric analysis revealed that DiCl-BBQ halted cell division and induced G1 cell cycle arrest in an AhR-dependent manner. Proteomic profiling identified the top four enriched pathways following DiCl-BBQ exposure: metabolism of RNA, translation, ribonucleoprotein complex biogenesis, and carboxylic acid metabolic processes. Notably, DiCl-BBQ caused a dramatic downregulation of translation-associated proteins, with this response diminished in AhR-depleted cells. Consistently, global protein synthesis was significantly repressed in DiCl-BBQ-treated cells. Together, these results indicate that DiCl-BBQ effectively inhibits HepG2 cells growth by inducing G1 cell cycle arrest and downregulating the protein translation machinery in an AhR-dependent manner.
The sequence of nucleotides that make up an RNA determines its structure, which determines its function. The RNA hairpin, also known as a stem-loop, is a ubiquitous and fundamental feature of RNA secondary structure. A common method of randomizing an RNA sequence is dinucleotide shuffling with the Altschul-Erickson algorithm, which preserves the dinucleotide content of the sequence. This algorithm generates randomized sequences by sampling Eulerian paths through the de Bruijn graph representation of the original sequence. We identified a subset of RNA hairpins in the bpRNA-1m meta-database that always form hairpins after repeated application of dinucleotide shuffling. We investigated these "unbreakable hairpins" and found several common properties. First, we found that unbreakable hairpins had on average similar folding energies compared to other hairpins of similar lengths, although they frequently contained ultra-stable hairpin loops. We found that they tend to be split by purines and pyrimidines on opposite sides of the stem. Furthermore, we found that this specific sequence feature restricts the number of distinct Eulerian paths through their de Bruijn graph representation, resulting in a small number of distinguishable dinucleotide-shuffled sequences. Beyond this algorithmic means of identification, these distinct sequences may have biological significance because we found that a significant percentage occur in a specific location of 16S ribosomal RNAs.
MOTIVATION:RNA secondary structure is often essential to function. Recent work has led to the development of high-throughput experimental probing methods for structure determination. Although structure is more conserved than primary sequence, much of the bioinformatics pipelines to connect RNA structure to function rely on nucleotide sequence alignments rather than structural similarity. There is a need to develop methods for secondary structure comparisons that are also fast and efficient to navigate the vast amounts of structural data. K-mer based similarity approaches are valued for their computational efficiency and have been applied for protein, DNA, and RNA primary sequences. However, these approaches have yet to be implemented for RNA secondary structure. RESULTS:Our method, bpRNA-CosMoS, fills this gap by using k-mers and length-weighted cosine similarity to compute similarity scores between RNA structures. bpRNA-CosMoS is built upon the bpRNA structure array, which represents the structural category of each nucleotide as a single-character structural code (e.g. hairpin=H, etc.). A structural comparison score is calculated through cosine similarity of the k-mer count vectors, generated from structure arrays. A major challenge with k-mer based methods is that they often ignore the length of the sequences being compared. We have overcome this with a length-weighted penalty that addresses cases of two RNAs of vastly different lengths. In addition, the use of "fuzzy counting" has added some optional flexibility to decrease the negative impact that small structural variations have on the similarity score. This results in a robust and efficient way to identify structural comparisons across large datasets. AVAILABILITY AND IMPLEMENTATION:The code and application guidelines of bpRNA-CosMoS are made available at github (https://github.com/BLasher113/bpRNA-CosMoS) and Zenodo (10.5281/zenodo.14715285).
RNA molecules adopt complex structures that perform essential biological functions across all forms of life, making them promising candidates for therapeutic applications. However, our ability to design new RNA structures remains limited by an incomplete understanding of their folding principles. While global metrics such as the minimum free energy are widely used, they are at odds with naturally occurring structures and incompatible with established design rules. Here, we introduce local stability compensation (LSC), a principle that RNA folding is governed by the local balance between destabilizing loops and their stabilizing adjacent stems, challenging the focus on global energetic optimization. Analysis of over 100,000 RNA structures revealed that LSC signatures are particularly pronounced in bulges and their adjacent stems, with distinct patterns across different RNA families that align with their biological functions. To validate LSC experimentally, we systematically analyzed thousands of RNA variants using DMS chemical mapping. Our results demonstrate that stem folding, as measured by reactivity, correlates with LSC (R2 = 0.458 for hairpin loops) and that instabilities show no significant effect on folding for distal stems. These findings demonstrate that LSC can be a guiding principle for understanding RNA function and for the rational design of custom RNAs.
Mounting an immune response requires energy, but how that energy is reallocated at the organismal level remains poorly understood. In Drosophila melanogaster, infection by a parasitoid wasp triggers a systemic metabolic switch known as immunometabolism which is characterized by a shift in metabolic activity and the redistribution of resources away from organismal development and toward the production of a cellular immune response. We identify the PDGF/VEGF (PVF) signaling pathway as a key initiator of this immunometabolic switch. Genetic manipulation of PVF signaling alters infection outcomes, modulates systemic metabolite profiles, and reveals a direct trade-off between immune function and development. Our findings establish the Drosophila-parasitoid wasp system as a genetically tractable model for understanding the molecular basis of immunometabolism.
In a modern society where artificial light sources rich in blue light (BL) are pervasive, concerns about the potential health impacts of BL on humans are growing. The damage BL causes to the human retina is well established, but its effects on non-retinal cells and the underlying mechanisms remain unclear. This study investigates the effects of phototoxic levels of BL on gene expression in adult Drosophila. We exposed Drosophila with genetically ablated eyes to continuous BL around the clock and assessed transcriptomic changes in non-retinal head tissues using RNAseq at three time points: 6, 10, and 14 days. Transcriptomic data revealed significant number of differentially expressed genes (DEGs) related to ribosomal biogenesis and energy metabolism pathways, including G6P metabolic process, fatty acyl-CoA biosynthesis, and glycolysis/ gluconeogenesis, due to BL exposure. To identify potential transcription factors (TFs) that may be involved in the BL-induced changes in gene expression, we utilized database of motifs to search the promoters and genomic sequences including introns of genes identified as DEGs. We determined that promoters for upregulated DEGs contain a DNA motif that is predicted to be bound by a TF with homology to a plant TF which is activated by light, encoded by CoRest. Further, a significant number of DEGs contained DNA motifs that could be targeted by another TF encoded by Xrp1, whose expression is significantly upregulated in BL. To test functional involvement of Xrp1 in BL response, we induced its overexpression in Drosophila neurons and showed that this significantly increased fly survival under BL, while knock out of Xrp1 accelerated mortality of BL-exposed flies. This study implicates potential positive feedback loop between CoRest and Xrp1, which is supported by previously published ChIPseq data, in response to blue light exposure and provides novel insights into the molecular mechanisms involved in mitigating its adverse effects.
The potential association of milk with childhood obesity has been widely debated and researched. Milk is known to contain many bioactive compounds as well as bovine exosomes rich in micro-RNA (miR) that can have effects on various cells, including stem cells. Among them, adipose stem cells (ASC) are particularly interesting due to their role in adipose tissue growth and, thus, obesity. The objective of this study was to evaluate the effect of milk consumption on miR present in circulating exosomes and the transcriptome of ASC in piglets. Piglets were supplemented for 11 weeks with 750 mL of whole milk (n = 6; M) or an isocaloric maltodextrin solution (n = 6; C). After euthanasia, ASC were isolated, quantified, and characterized. RNA was extracted from passage 1 ASC and sequenced. Exosomes were isolated and quantified from the milk and plasma of the pigs at 6-8 hours after milk consumption, and miRs were isolated from exosomes and sequenced. The transfer of exosomes from milk to porcine plasma was assessed by measuring bovine milk-specific miRs and mRNA in exosomes isolated from the plasma of 3 piglets during the first 6h after milk consumption. We observed a higher proportion of exosomes in the 80 nM diameter, enriched in milk, in M vs. C pigs. Over 500 genes were differentially expressed (DEG) in ASC isolated from M vs. C pigs. Bioinformatic analysis of DEG indicated an inhibition of the immune, neuronal, and endocrine systems and insulin-related pathways in ASC of milk-fed pigs compared with maltodextrin-fed pigs. Of the 900 identified miRs in porcine plasma exosomes, only 3 miRs were differentially abundant between the two groups and could target genes associated with neuronal functions. We could not detect exosomal miRs or mRNA transfer from milk to porcine-circulating plasma exosomes. Our data highlights the significant nutrigenomic role of milk consumption on ASC, a finding that does not appear to be attributed to miRs in bovine milk exosomes. The downregulation of insulin resistance and inflammatory-related pathways in the ASC of milk-fed pigs should be further explored in relation to milk and human health. In conclusion, the bioinformatic analyses and the absence of bovine exosomal miRs in porcine plasma suggest that miRs are not vertically transferred from milk exosomes.
Hop powdery mildew (PM) (Podosphaera macularis) causes substantial losses if left uncontrolled. Most resistant hop cultivars possess qualitative resistance based on R‐genes. One cultivar, Comet, has uncharacterized resistance that may be polygenic. This study focused on identifying genomic regions controlling PM resistance in Comet and ascertaining putative genetic mechanisms behind such resistance. A cross between Comet and susceptible male, USDA 64035M, was made. Offspring were screened for resistance under greenhouse conditions and genotyped using genotyping‐by‐sequencing. Genome‐wide analysis using mixed linear model analysis along with quantitative trait locus (QTL) analysis using either composite interval mapping or stepwise regression analyses was performed to identify QTLs. All analyses identified a region on chromosome 6 covering positions 308–314 Mb on the physical map. Analysis of the putative genes within this region identified 140 genes with 27 plant resistance‐like genes found in nine clusters. Six sulfur‐rich protein genes with homology to patatins, thionins, and agglutinins were identified in two clusters. Two glucan‐endo‐1,3‐beta‐glucosidase genes were identified bordering different R‐gene clusters. Finally, putative upregulators of transcription and stress‐response genes were identified. The 10 most highly associated single‐nucleotide polymorphisms for PM resistance were subsequently developed as KASP markers. The combination of R‐gene clusters, sulfur‐rich proteins, endo‐1,3‐beta‐glucosidase genes, and stress‐response genes may be responsible for resistance to PM in the cultivar Comet.
The prebiotic formation of RNA building blocks is well-supported experimentally, yet the emergence of sequence- and structure-specific RNA oligomers is generally attributed to biological selection via Darwinian evolution rather than prebiotic chemical selectivity. In this study, we used deep sequencing to investigate the partitioning of randomized RNA overhangs into ligated products by either splinted ligation or loop-closing ligation. Comprehensive sequence-reactivity profiles revealed that loop-closing ligation preferentially yields hairpin structures with loop sequences UNNG, CNNG, and GNNA (where N represents A, C, G, or U) under competing conditions. In contrast, splinted ligation products tended to be GC rich. Notably, the overhang sequences that preferentially partition to loop-closing ligation significantly overlap with the most common biological tetraloops, whereas the overhangs favoring splinted ligation exhibit an inverse correlation with biological tetraloops. Applying these sequence rules enables the high-efficiency assembly of functional ribozymes from short RNAs without template inhibition. Our findings suggest that the RNA tetraloop structures that are common in biology may have been predisposed and prevalent in the prebiotic pool of RNAs, prior to the advent of Darwinian evolution. We suggest that the one-step prebiotic chemical process of loop-closing ligation could have favored the emergence of the first RNA functions.
Ribosomes are information-processing macromolecular machines that integrate complex sequence patterns in messenger RNA (mRNA) transcripts to synthesize proteins. Studies of the sequence features that distinguish mRNAs from long noncoding RNAs (lncRNAs) may yield insight into the information that directs and regulates translation. Computational methods for calculating protein-coding potential are important for distinguishing mRNAs from lncRNAs during genome annotation, but most machine learning methods for this task rely on previously known rules to define features. Sequence-to-sequence (seq2seq) models, particularly ones using transformer networks, have proven capable of learning complex grammatical relationships between words to perform natural language translation. Seeking to leverage these advancements in the biological domain, we present a seq2seq formulation for predicting protein-coding potential with deep neural networks and demonstrate that simultaneously learning translation from RNA to protein improves classification performance relative to a classification-only training objective. Inspired by classical signal processing methods for gene discovery and Fourier-based image-processing neural networks, we introduce LocalFilterNet (LFNet). LFNet is a network architecture with an inductive bias for modeling the three-nucleotide periodicity apparent in coding sequences. We incorporate LFNet within an encoder-decoder framework to test whether the translation task improves the classification of transcripts and the interpretation of their sequence features. We use the resulting model to compute nucleotide-resolution importance scores, revealing sequence patterns that could assist the cellular machinery in distinguishing mRNAs and lncRNAs. Finally, we develop a novel approach for estimating mutation effects from Integrated Gradients, a backpropagation-based feature attribution, and characterize the difficulty of efficient approximations in this setting.
Predicting RNA degradation is a fundamental task in designing RNA-based therapeutic agents. Dual crowdsourcing efforts for dataset creation and machine learning were organized to learn biological rules and strategies for predicting RNA stability.
The nucleocapsid (N) protein of SARS-CoV-2 binds viral RNA, condensing it inside the virion, and phase separating with RNA to form liquid-liquid condensates. There is little consensus on what differentiates sequence-independent N-RNA interactions in the virion or in liquid droplets from those with specific genomic RNA (gRNA) motifs necessary for viral function inside infected cells. To identify the RNA structures and the N domains responsible for specific interactions and phase separation, we use the first 1,000 nt of viral RNA and short RNA segments designed as models for single-stranded and paired RNA. Binding affinities estimated from fluorescence anisotropy of these RNAs to the two-folded domains of N (the NTD and CTD) and comparison to full-length N demonstrate that the NTD binds preferentially to single-stranded RNA, and while it is the primary RNA-binding site, it is not essential to phase separation. Nuclear magnetic resonance spectroscopy identifies two RNA-binding sites on the NTD: a previously characterized site and an additional although weaker RNA-binding face that becomes prominent when binding to the primary site is weak, such as with dsRNA or a binding-impaired mutant. Phase separation assays of nucleocapsid domains with double-stranded and single-stranded RNA structures support a model where multiple weak interactions, such as with the CTD or the NTD's secondary face promote phase separation, while strong, specific interactions do not. These studies indicate that both strong and multivalent weak N-RNA interactions underlie the multifunctional abilities of N.
The structure of RNA can determine the function, and improvements in RNA secondary structure prediction can help in understanding the functions of RNA. Nucleotides in RNA sequences form base-pairing interactions in context-specific preferential behavior to help determine the secondary structure. Structure prediction algorithms have been developed to predict the secondary structure, including dynamic programming, and machine learning approaches. One of the central challenges in the prediction of secondary structure with deep learning is that these architectures are not good at bracketed structure prediction. To overcome this challenge, we present a deep learning approach for predicting secondary structure that uses an input predicted structure to provide a scaffolding for the structure prediction. We find that architectures using LSTM and self-attention-based transformer layers predict a strong baseline in the prediction of base pairs (F1 = 53.73), but significantly improves (F1 = 59.52) when predictions from dynamic programming methods are provided as input. Model interpretation shows that patterns of attention for different layers of the network are enriched for specific paired regions or regions that should be paired. Analysis of neural network models like this can shed light on possible missed interactions, and what other positions contribute most to output fixed positions.
Ribonucleic acid (RNA) is a polymeric molecule that is fundamental to biological processes, with structure being more highly conserved than primary sequence and often key to its function. Advances in RNA structure characterization have resulted in an increase in the number of accurate secondary structures. The task of uncovering common RNA structural motifs with a collective function through structural comparison, providing a level of similarity, remains challenging and could be used to improve RNA secondary structure databases and discover new RNA families. In this work, we present a novel secondary structure alignment method, bpRNA-align. bpRNA-align is a customized global structural alignment method, utilizing an inverted (gap extend costs more than gap open) and context-specific affine gap penalty along with a structural, feature-specific substitution matrix to provide similarity scores. We evaluate our similarity scores in comparison to other methods, using affinity propagation clustering, applied to a benchmarking data set of known structure types. bpRNA-align shows improvement in clustering performance over a broad range of structure types.
SummaryThe SARS-CoV-2 nucleocapsid phosphoprotein (CoV-N) contains two structured RNA binding domains, the NTD and the CTD, but the functional implications of two structurally independent binding domains remains unclear. The NTD/RNA interaction is well described but information about CTD/RNA binding is more limited. With a 1000 nucleotide fragment (g1-1000) of the SARS-CoV-2 genomic RNA, we map the interactions of both domains using NMR and show that CTD is sufficient for droplet formation. Further, we investigate CoV-N interactions with a primary RNA-binding site, g74-301, and a shuffled version with fewer paired nucleotides and higher entropy, sh-g74-301. Probed by microscopy and SEC-MALS, CoV-N forms droplets and large complexes faster with sh-g74-301. NMR titrations and mutagenesis studies confirm that NTD is responsible for the difference in binding patterns. These results suggest that CoV-N has preference for ssRNA, and that binding specificity is determined by NTD binding kinetics, while CTD is necessary for droplet formation.Funding: This work is funded by an NSF EAGER MCB 2034446. We also acknowledge the support of the Oregon State University NMR Facility funded by the National Institutes of Health, HEI Grant 1S10OD018518, and the M. J. Murdock Charitable Trust grant #2014162.Declaration of Interests: None to declare.
The transcriptome from lupulin glands and associated bracts from cone tissue of hop (Humulus lupulus) c.v. ‘Cascade’, during three stages of development: early, mid, and late or near-harvest, was sequenced. Significant increases were found in expression patterns of many genes involved in the biosynthesis of bitter acids, xanthohumol, and volatile secondary metabolites or “hop oils” during the middle stage of cone development. The biosynthesis of thiol precursors responsible for popular “tropical fruit” flavors in beer is not well known, but homologs of genes hypothesized to be involved in this process tend to be up-regulated during the late stage in hop cones. More research needs to be performed to describe the pathway of thiol precursor biosynthesis in hops. Hierarchical clustering revealed overlap of samples taken from each developmental stage, likely due to non-uniform ripening of cones on the plant. It is proposed that the mid-stage of cone development is critical for the development of important flavor-producing secondary metabolites in hops, and this is supported by previous research describing concentrations of secondary metabolites. It is hypothesized that abiotic stress during the mid-stage of cone development may be quite detrimental to the bitter acid concentrations ultimately found in hops.