Determining the diverse cellular states and their organization into cellular ecosystems that make up metastatic tumor is vital for elucidating the biological and prognostic diversity of cancer. However, large-scale studies profiling the clinical relevance of these cellular states and ecotypes are still lacking in metastatic cancers. In this study, we used EcoTyper, a machine learning framework, to comprehensively analyze transcriptomes from 2822 metastatic cancer patient samples covering 25 cancer types, enabling characterization of the fundamental cellular states and tumor ecosystems integral to metastatic cancer. We identified 45 distinct cellular states across 12 cell types and validated their robustness in validation cohorts. We observed that they differed in functional and prognostic associations. Survival analysis revealed that the clinically relevant cellular states, highlighting their promise as predictors of clinical outcomes. Functional enrichment analysis exhibited that the marker genes of cellular states were significantly enriched in cancer hallmark and immune-related pathways. In addition, our analysis identified five ecotypes associated with different clinical outcomes. Transcription factor enrichment analysis revealed key transcription factors (i.e. SPIB, SRF, and NR1D1) that were significantly associated with patient clinical outcomes. In conclusion, this study provided a high-resolution landscape of cellular states and ecosystems in metastatic tumors, offering new potential targets for the development of cancer treatment strategies and prognostic assessment.
BACKGROUND:RNA-RNA spatial interactions play crucial roles in various cellular processes, including gene transcriptional and post-transcriptional regulations. However, research on this area of RNA regulation remains limited in plants. RESULTS:Here, we adapt the global RNA-RNA interaction mapping method for plants and develop plant RNA in situ conformation sequencing, pRIC-seq, to generate comprehensive RNA-RNA spatial interaction maps for diploid and tetraploid cotton. We also perform global nuclear run-on followed by cap-selection assay, GRO-cap, and integrate multi-omics data to construct enhancer landscapes in these cotton species. Focusing on enhancer-promoter (E-P) RNA interactions, we find that tetraploid cotton, following polyploidy, innovates numerous novel E-P RNA interactions, thereby increasing its genomic regulatory complexity. Comparative analyses between wild-type and mutant fuzzless/lintless in tetraploid cotton reveal that RNA-RNA interactions, including E-P RNA interactions, play pivotal roles in fiber development. Our study also identifies short tandem repeats and transposable elements as potential mediators of E-P RNA interactions through base pairing within the cotton genome. Finally, integrating with genome-wide association studies (GWAS) and eQTLs from previous studies, we observe that our RNA-RNA interactions are significantly enriched near those functional mutation sites. Importantly, by using RAP-qPCR, we confirm that GWAS related enhancers interact with the promoters of protein-coding genes, explaining their regulatory mechanisms in fiber trait control. CONCLUSIONS:Our results provide the first genome-wide RNA-RNA interaction map in higher plants and offer valuable insights into the enhancer-regulated pathway and targets for future breeding studies.
Circular RNAs (circRNAs) serve as key post-transcriptional regulators in plant stress adaptation. Here, we comprehensively characterize the circRNA landscape of cotton (Gossypium arboreum) under multiple abiotic stresses condition using RNase R-enhanced sequencing. Through a stringent KNIFE-based algorithm pipeline, we identified 4,365 high-confidence circRNAs. Mechanistically, circRNA biogenesis was associated with long flanking introns and exhibited complex patterns of alternative splicing, revealing conserved production propensities but dynamic splice sites across species. Nuclear circRNAs frequently exhibited expression patterns decoupled from their host genes, typically in a stress-specific manner. By integrating miRNA-seq data, we constructed a circRNA-miRNA-mRNA regulatory network and found it centered mainly on cotton-specific miRNAs. Notably, we discovered an extraordinary dominance of chloroplast-derived circRNAs, accounting for over 80% of the total circRNAs repertoire. These chloroplast circRNAs clustered dominantly at clustering at the 3’ terminus of the photosynthetic gene psbA gene, suggesting a specialized post-transcriptional regulatory mechanism within the chloroplast. This study provides a high-resolution cotton circRNA atlas and highlights psbA-derived circRNAs as potential molecular targets for improving environmental resilience in crop.
Phosphite (P⁺ᴵᴵᴵ), an emerging reduced phosphorus (P) species of environmental significance, exerts complex but poorly characterized effects on algal communities. In this study, we systematically evaluated interspecific variability in P⁺ᴵᴵᴵ tolerance and recovery potential among four typical bloom-forming taxa by exposing each to a concentration gradient of P⁺ᴵᴵᴵ (0.07-14 mg L⁻¹), followed by phosphate (P⁺ⱽ) recovery. Results revealed species-specific responses to P⁺ᴵᴵᴵ were pronounced. During phosphite stress periods of 15-27 days (strain-dependent, see Methods), Chlamydomonas reinhardtii (C. reinhardtii) demonstrated exceptional P⁺ᴵᴵᴵ tolerance, with culture density gain inhibited by only 21.3% (final A₆₈₀ relative increase: 798% in phosphate control vs. 628% in 14 mg·L⁻¹ P⁺ᴵᴵᴵ) and intact cellular ultrastructure even at the highest concentration. In contrast, two strains of Chlorella pyrenoidosa (C. pyrenoidosa) exhibited severe growth inhibition-83.3% for C. pyrenoidosa-9 (final A₆₈₀ relative increase: 448% vs. 75%) and 84.6% for C. pyrenoidosa-10 (final A₆₈₀ relative increase: 344% vs. 53%)-along with extensive membrane rupture. Chlorella vulgaris (C. vulgaris) displayed transient adaptation at low P⁺ᴵᴵᴵ levels (0.07-0.7 mg·L⁻¹) but suffered serious oxidative damage at concentrations exceeding 0.7 mg·L⁻¹ , with density gain inhibited by 69.9% at 14 mg·L⁻¹ (final A₆₈₀ relative increase: 851% vs. 256%). Notably, P⁺ᴵᴵᴵ proved to be an ineffective as a bioavailable P source, as the intracellular P accumulation remained below 10% of that observed in P⁺ⱽ controls. During the subsequent P⁺ⱽ recovery phase, C. vulgaris achieved a rapid rebound, with cell densities increasing by 167%. Conversely, recovery in C. reinhardtii was delayed, attributable to its comparatively low P⁺ⱽ metabolic efficiency. Scanning electron microscopy (SEM) confirmed that P⁺ᴵᴵᴵ-induced cellular damage is partially reversible. However, exposure to 14 mg·L⁻¹ P⁺ᴵᴵᴵ caused reversible growth inhibition in sensitive species. Collectively, these results underscores the dual role of P+III as both a stressor and a potential modulator of algal blooms in eutrophic systems, capable of restructuring community structures. The findings enhance our understanding of reduced P bioavailability mechanisms and offer critical insights for ecological risk assessment and eutrophication management.
Understanding how cells commit to distinct fates over time is fundamental to elucidating the principles and mechanisms that govern organismal development, tissue regeneration and disease progression. Multimodal lineage tracing, which couples heritable lineage information with single-cell multi-omics, has revolutionized our ability to chart cellular dynamics and fate decisions at unprecedented resolution. However, the resulting datasets are inherently complex and heterogeneous, calling for sophisticated computational frameworks capable of transforming raw measurements into coherent biological insights. Here we comprehensively survey recent methodological advances that substantially expand the computational toolkit for analysing lineage-resolved, single-cell multi-omic data, enabling more accurate lineage reconstruction, trajectory inference, ancestral state estimation and identification of molecular programmes driving cell-state transitions. Emerging high-resolution lineage-tracing technologies and deep learning-based analytical models promise to further unlock the full potential of multimodal lineage tracing, offering an increasingly complete and quantitative view of cellular evolution in both health and disease.
Histone methylation is involved in a wide range of biological regulation in plants, and is conducted by three major components, including methyltransferases, demethylases, and histone readers. Compared with the other two components, research on histone readers is relatively limited. In this study, we demonstrate that OsSHH5 functions as an H3K9me1 reader to regulate rice disease resistance, tillering, and grain yield. Loss of OsSHH5 function significantly enhances both grain yield and disease resistance. Mechanistically, OsSHH5 recruits the H3K9 methyltransferase SGD733 and binds to H3K9me1, thereby maintaining H3K9me1 enrichment and facilitating gene silencing. In leaves, OsSHH5 interacts with the transcriptional factor HPY1 to target the resistance-related genes OsWAKg52 and OsWRKY81, maintaining their H3K9me1 levels and suppressing multiple PAMP-triggered immune responses, which ultimately reduces rice disease resistance. In tiller buds, OsSHH5 interacts with the transcriptional factor TCP19 to target the tillering-related gene OsNGR5, maintaining its H3K9me1 enrichment and inhibition of tillering, leading to reduced yield. Collectively, these findings reveal that OsSHH5 plays a vital role in integrating immune response, tillering, and grain yield in rice, providing new insights into the function of histone readers and offering a new strategy to improve rice yield and disease resistance.
Cotton is the world’s most important natural fiber crop and serves as an ideal model for studying plant genome evolution, cell differentiation, elongation, and cell wall biosynthesis. The first draft genome assembly for Gossypium raimondii, completed in 2012, marked the beginning of global efforts in studying cotton genomics. Over the past decade, the cotton research community has continued to assemble and refine the genomes for both wild and cultivated Gossypium species. With the accumulation of de novo genome assemblies and resequencing data across virous cotton populations, significant progress has been made in uncovering the genetic basis of key agronomic traits. Achieving the goal of cotton genomics-to-breeding (G2B) will require a deeper understanding of the spatiotemporal regulatory mechanisms involved in genome information storage and expression. We advocate for a cotton ENCODE project to systematically decode the functional elements and regulatory networks within the cotton genome. Technological advances, particularly on single-cell sequencing and high-resolution spatiotemporal omics, will be essential for elucidating these regulatory mechanisms. By integrating multi-omics data, genome editing tools, and artificial intelligence, these efforts will empower the genomics-driven strategies needed for future cotton G2B breeding.
Embryonic pattern formation and cell specification require precise cell division and cell cycle regulation. Splicing factors and the splicing of precursor mRNA (pre-mRNA) play significant roles in embryo development. However, how splicing factors control embryonic patterning via RNA splicing remains unclear. Here, we show that the mutation of SUPPRESSORS OF MEC-8 AND UNC-52 1 (SMU1), a conserved subunit of the spliceosomal B complex, causes compromised cell fate of the hypophysis and quiescent center (QC), failed embryonic root apical meristem (RAM) formation, as evidenced by altered WUSCHEL-RELATED HOMEOBOX 5 (WOX5) expression and perturbed auxin signaling. This results in smu1 embryo lethality. The splicing efficiency of three out of four CYCLIN-DEPENDENT KINASE ACTIVATOR (CAK) genes is decreased, leading to reduced protein levels in smu1 embryos. These CAK genes are required for hypophysis specification and embryonic RAM formation. SMU1 binds CAK transcripts in vitro and in vivo. Restoring the expression of either CAK gene partially rescues the defects in smu1 embryos, leading to the formation of QC-like cells, continued embryo development, and even the production of viable seeds. Our data suggest that SMU1 binds to CAK transcripts and promotes their splicing, enabling cell cycle progression to promote embryonic RAM formation.
E3 ligases are key enzymes required for protein degradation. Here, we identified a C3H2C3 RING domain-containing E3 ubiquitin ligase gene named GhATL68b. It is preferentially and highly expressed in developing cotton fiber cells and shows greater conservation in plants than in animals or archaea. The four orthologous copies of this gene in various diploid cottons and eight in the allotetraploid G. hirsutum were found to have originated from a single common ancestor that can be traced back to Chlamydomonas reinhardtii at about 992 million years ago. Structural variations in the GhATL68b promoter regions of G. hirsutum, G. herbaceum, G. arboreum, and G. raimondii are correlated with significantly different methylation patterns. Homozygous CRISPR-Cas9 knockout cotton lines exhibit significant reductions in fiber quality traits, including upper-half mean length, elongation at break, uniformity, and mature fiber weight. In vitro ubiquitination and cell-free protein degradation assays revealed that GhATL68b modulates the homeostasis of 2,4-dienoyl-CoA reductase, a rate-limiting enzyme for the β-oxidation of polyunsaturated fatty acids (PUFAs), via the ubiquitin proteasome pathway. Fiber cells harvested from these knockout mutants contain significantly lower levels of PUFAs important for production of glycerophospholipids and regulation of plasma membrane fluidity. The fiber growth defects of the mutant can be fully rescued by the addition of linolenic acid (C18:3), the most abundant type of PUFA, to the ovule culture medium. This experimentally characterized C3H2C3 type E3 ubiquitin ligase involved in regulating fiber cell elongation may provide us with a new genetic target for improved cotton lint production.
Brassica has long been of paramount importance to agriculture and human nutrition.Among the Brassica genus,Brassica rapa,the AA diploid,stands out as a prominent representative,valued for its versatility both as a vegetable and oil crop and its genetic proximity to other Brassica species.Economically,B.rapa plays an important role in the food industry,contributing to 12%of the global edible oil supply and 10%of vegetable yield[1].
Cellular responses to internal and external stimuli are orchestrated by intricate intracellular signaling pathways. To ensure an efficient and specific information flow, cells employ scaffold proteins as critical signaling organizers. With the ability to bind multiple signaling molecules, scaffold proteins can sequester signaling components within specific subcellular domains or modulate the efficiency of signal transduction. Scaffolds can also tune the output of signaling pathways by serving as regulatory targets. This review focuses on scaffold proteins associated with the plant GLYCOGEN SYNTHASE KINASE3-like kinase, BRASSINOSTEROID-INSENSITIVE2 (BIN2), that serves as a key negative regulator of brassinosteroid (BR) signaling. Here, we summarize current understanding of how scaffold proteins actively shape BR signaling outputs and cross-talk in plant cells via interactions with BIN2.
The abnormal development of urate granules in silkworm larvae leads to translucent mutants with a distinct transparent phenotype. Studies on such mutants are expected to enhance current understanding of uric acid metabolism. The hoarfrost translucent (oh) mutant exhibits a mottled, translucent larval integument due to the presence of smaller and irregularly shaped urate granules compared to wild-type individuals. Uric acid content in the silkworm larval integuments is significantly lower in the oh mutant. Using positional cloning, we successfully narrowed a ~ 180 kb region linked to the oh locus and identified the candidate gene, Bmpallidin, encoding the biosynthesis of lysosome-related organelles complex 1 subunit 6. Three alternative splicing isoforms were identified; Only isoform II was predicted to translate normally and was drastically reduced at both mRNA and protein levels in the oh mutant. Conversely, the non-functional isoform III showed slightly increased expression in the mutant. An 860 bp genomic sequence in the wild type was replaced by a 30 bp sequence in the mutant. Knockdown of Bmpallidin induced a translucent phenotype in first-instar larvae. These findings conclude that Bmpallidin is responsible for the oh mutant phenotype and plays a crucial role in urate granules formation in silkworms.
Cotton (Gossypium hirsutum) fibers are elongated single cells that rapidly accumulate cellulose during secondary cell wall (SCW) thickening, which requires cellulose synthase complex (CSC) activity. Here, we describe the CSC-interacting factor CASPARIAN STRIP MEMBRANE DOMAIN-LIKE1 (GhCASPL1), which contributes to SCW thickening by influencing CSC stability on the plasma membrane. GhCASPL1 is preferentially expressed in fiber cells during SCW biosynthesis and encodes a MARVEL domain protein. The ghcaspl1 ghcaspl2 mutant exhibited reduced plant height and produced mature fibers with fewer natural twists, lower tensile strength, and a thinner SCW compared to the wild type. Similarly, the Arabidopsis (Arabidopsis thaliana) caspl1 caspl2 double mutant showed a lower cellulose content and thinner cell walls in the stem vasculature than the wild type but normal plant morphology. Introducing the cotton gene GhCASPL1 successfully restored the reduced cellulose content of the Arabidopsis caspl1 caspl2 mutant. Detergent treatments, ultracentrifugation assays, and enzymatic assays showed that the CSC in the ghcaspl1 ghcaspl2 double mutant showed reduced membrane binding and decreased enzyme activity compared to the wild type. GhCASPL1 binds strongly to phosphatidic acid (PA), which is present in much higher amounts in thickening fiber cells compared to ovules and leaves. Mutating the PA-binding site in GhCASPL1 resulted in the loss of its colocalization with GhCesA8, and it failed to localize to the plasma membrane. PA may alter membrane structure to facilitate protein-protein interactions, suggesting that GhCASPL1 and PA collaboratively stabilize the CSC. Our findings shed light on CASPL functions and the molecular machinery behind SCW biosynthesis in cotton fibers.
Genomic analysis has revealed that the 1,637-Mb Gossypium arboreum genome contains approximately 81% transposable elements (TEs), while only 57% of the 735-Mb G. raimondii genome is occupied by TEs. In this study, we investigated whether there were unknown transcripts associated with TE or TE fragments and, if so, how these new transcripts were evolved and regulated. As sequence depths increased from 4 to 100 G, a total of 10,284 novel intergenic transcripts (intergenic genes) were discovered. On average, approximately 84% of these intergenic transcripts possibly overlapped with the long terminal repeat (LTR) insertions in the otherwise untranscribed intergenic regions and were expressed at relatively low levels. Most of these intergenic transcripts possessed no transcription activation markers, while the majority of the regular genic genes possessed at least one such marker. Genes without transcription activation markers formed their+1 and -1 nucleosomes more closely (only (117±1.4)bp apart), while twice as big spaces (approximately (403.5±46.0) bp apart) were detected for genes with the activation markers. The analysis of 183 previously assembled genomes across three different kingdoms demonstrated systematically that intergenic transcript numbers in a given genome correlated positively with its LTR content. Evolutionary analysis revealed that genic genes originated during one of the whole-genome duplication events around 137.7 million years ago (MYA) for all eudicot genomes or 13.7 MYA for the Gossypium family, respectively, while the intergenic transcripts evolved around 1.6 MYA, resultant of the last LTR insertion. The characterization of these low-transcribed intergenic transcripts can facilitate our understanding of the potential biological roles played by LTRs during speciation and diversifications.
Single-cell RNA sequencing (scRNA-seq) is a powerful approach for studying cellular differentiation, but accurately tracking cell fate transitions can be challenging, especially in disease conditions. Here we introduce PhyloVelo, a computational framework that estimates the velocity of transcriptomic dynamics by using monotonically expressed genes (MEGs) or genes with expression patterns that either increase or decrease, but do not cycle, through phylogenetic time. Through integration of scRNA-seq data with lineage information, PhyloVelo identifies MEGs and reconstructs a transcriptomic velocity field. We validate PhyloVelo using simulated data and Caenorhabditis elegans ground truth data, successfully recovering linear, bifurcated and convergent differentiations. Applying PhyloVelo to seven lineage-traced scRNA-seq datasets, generated using CRISPR–Cas9 editing, lentiviral barcoding or immune repertoire profiling, demonstrates its high accuracy and robustness in inferring complex lineage trajectories while outperforming RNA velocity. Additionally, we discovered that MEGs across tissues and organisms share similar functions in translation and ribosome biogenesis. A refined velocity model improves cell fate mapping with lineage-traced scRNA-seq data.
Inclusion bodies (IBs) of respiratory syncytial virus (RSV) are formed by liquid-liquid phase separation (LLPS) and contain internal structures termed “IB-associated granules” (IBAGs), where anti-termination factor M2-1 and viral mRNAs are concentrated. However, the mechanism of IBAG formation and the physiological function of IBAGs are unclear. Here, we found that the internal structures of RSV IBs are actual M2-1-free viral messenger ribonucleoprotein (mRNP) condensates formed by secondary LLPS. Mechanistically, the RSV nucleoprotein (N) and M2-1 interact with and recruit PABP to IBs, promoting PABP to bind viral mRNAs transcribed in IBs by RNA-recognition motif and drive secondary phase separation. Furthermore, PABP-eIF4G1 interaction regulates viral mRNP condensate composition, thereby recruiting specific translation initiation factors (eIF4G1, eIF4E, eIF4A, eIF4B and eIF4H) into the secondary condensed phase to activate viral mRNAs for ribosomal recruitment. Our study proposes a novel LLPS-regulated translation mechanism during viral infection and a novel antiviral strategy via targeting on secondary condensed phase.
BackgroundThe epidermis of cotton ovule produces fibers, the most important natural cellulose source for the global textile industry. However, the molecular mechanism of fiber cell growth is still poorly understood.ResultsHere, we develop an optimized protoplasting method, and integrate single-cell RNA sequencing (scRNA-seq) and single-cell ATAC sequencing (scATAC-seq) to systematically characterize the cells of the outer integument of ovules from wild type and fuzzless/lintless (fl) cotton (Gossypium hirsutum). By jointly analyzing the scRNA-seq data from wildtype and fl, we identify five cell populations including the fiber cell type and construct the development trajectory for fiber lineage cells. Interestingly, by time-course diurnal transcriptomic analysis, we demonstrate that the primary growth of fiber cells is a highly regulated circadian rhythmic process. Moreover, we identify a small peptide GhRALF1 that circadian rhythmically controls fiber growth possibly through oscillating auxin signaling and proton pump activity in the plasma membrane. Combining with scATAC-seq, we further identify two cardinal cis-regulatory elements (CREs, TCP motif, and TCP-like motif) which are bound by the trans factors GhTCP14s to modulate the circadian rhythmic metabolism of mitochondria and protein translation through regulating approximately one third of genes that are highly expressed in fiber cells.ConclusionsWe uncover a fiber-specific circadian clock-controlled gene expression program in regulating fiber growth. This study unprecedentedly reveals a new route to improve fiber traits by engineering the circadian clock of fiber cells.
Cotton is an irreplaceable economic crop currently domesticated in the human world for its extremely elongated fiber cells specialized in seed epidermis, which makes it of high research and application value. To date, numerous research on cotton has navigated various aspects, from multi-genome assembly, genome editing, mechanism of fiber development, metabolite biosynthesis, and analysis to genetic breeding. Genomic and 3D genomic studies reveal the origin of cotton species and the spatiotemporal asymmetric chromatin structure in fibers. Mature multiple genome editing systems, such as CRISPR/Cas9, Cas12 (Cpf1) and cytidine base editing (CBE), have been widely used in the study of candidate genes affecting fiber development. Based on this, the cotton fiber cell development network has been preliminarily drawn. Among them, the MYB-bHLH-WDR (MBW) transcription factor complex and IAA and BR signaling pathway regulate the initiation; various plant hormones, including ethylene, mediated regulatory network and membrane protein overlap fine-regulate elongation. Multistage transcription factors targeting CesA 4, 7, and 8 specifically dominate the whole process of secondary cell wall thickening. And fluorescently labeled cytoskeletal proteins can observe real-time dynamic changes in fiber development. Furthermore, research on the synthesis of cotton secondary metabolite gossypol, resistance to diseases and insect pests, plant architecture regulation, and seed oil utilization are all conducive to finding more high-quality breeding-related genes and subsequently facilitating the cultivation of better cotton varieties. This review summarizes the paramount research achievements in cotton molecular biology over the last few decades from the above aspects, thereby enabling us to conduct a status review on the current studies of cotton and provide strong theoretical support for the future direction.