Rationale: SPINOPHILIN (SPN, PPP1R9B) is an important tumor suppressor involved in the progression and malignancy of different tumors depending on its association with protein phosphatase 1 (PP1) and the ability of the PP1-SPN holoenzyme to dephosphorylate retinoblastoma (pRB). Methods: We performed a mutational analysis of SPN in human tumors, focusing on the region of interaction with PP1 and pRB. We explored the effect of the SPN-A566V mutation in an immortalized non-tumorigenic cell line of epithelial breast tissue, MCF10A, and in two different p53-mutated breast cancer cells lines, T47D and MDA-MB-468. Results: We characterized an oncogenic mutation of SPN found in human tumor samples, SPN-A566V, that affects both the SPN-PP1 interaction and its phosphatase activity. The SPN-A566V mutation does not affect the interaction of the PP1-SPN holoenzyme with pocket proteins pRB, p107 and p130, but it affects its ability to dephosphorylate them during G0/G1 and G1, indicating that the PP1-SPN holoenzyme regulates cell cycle progression. SPN-A566V also promoted stemness, establishing a connection between the cell cycle and stem cell biology via pocket proteins and PP1-SPN regulation. However, only cells with both SPN-A566V and mutant p53 have increased tumorigenic and stemness properties. Conclusions: SPN-A566V, or other equivalent mutations, could be late events that promote tumor progression by increasing the CSC pool and, eventually, the malignant behavior of the tumor.
The "one structure-one function" paradigm is central to most computational methods for predicting protein function. However, a subset of proteins known as shape-shifters can adopt multiple well-defined structures to perform distinct biological functions. These proteins, along with intrinsically disordered proteins (IDPs) that fold upon binding, challenge conventional views of protein functionality. Despite their biological significance, identifying fold-switching proteins remains difficult due to limited structural data and the limitations of current prediction methods.In this study, we present a computational protocol to identify potential shape-shifting proteins within structural databases. By integrating structural and sequence data, this approach enables the screening of the Protein Data Bank (PDB) for proteins exhibiting significant conformational variability.
The Ras protein superfamily comprises small GTPases that share a conserved G-domain but differ in flanking regions, regulation, and cellular roles. Because existing classifications rely mainly on G-domain phylogeny, this superfamily provides a useful test case for assessing whether protein language model embeddings recover biologically meaningful sequence organization consistent with established evolutionary classifications. Here, we analyzed a curated Ras superfamily dataset using three classification schemes: the classical five-family view, a G-domain phylogeny-based classification, and UniProtKB family annotation classification. We compared embeddings from multiple protein language models using supervised classification, unsupervised clustering, and residue-level ablation. Sequence-derived embeddings recovered known Ras superfamily organization across analyses. In supervised analyses, simple linear classifiers achieved high performance, indicating that Ras family and subfamily information is linearly accessible from sequence-derived embeddings. In unsupervised analyses, ESM-C layer 12 gave the strongest full-protein recovery of the a G-domain-based classification, whereas ProstT5 performed best for G-domain embeddings. Residue ablation identified recurrent candidate subfamily-informative regions both within and outside the G-domain, including signals mapping to structurally coherent regions associated with subfamily-specific regulatory or interaction-related functions. Together, these results indicate that protein language model embeddings provide an effective alignment-free representation of Ras family and subfamily organization, and can highlight candidate sequence regions associated with functional specialization from sequence alone.
The Gene Ontology is a central resource for representing biological knowledge, yet its internal structure is often treated as static-or as a black box-in computational analyses. Here, we examine 15 years of Gene Ontology evolution using network-based methods, revealing that Gene Ontology changes not only through incremental growth but also through punctuated, curator-driven restructuring. In particular, we document a major reorganization of the Cellular Component branch in 2019, where broad "part" terms were removed and the ontology was modularized into distinct domains for anatomical entities and protein-containing complexes. Semantic modularity aligns Gene Ontology with emerging frameworks such as the Common Anatomy Reference Ontology and Gene Ontology-Causal Activity Modeling, but also disrupts similarity metrics that rely solely on hierarchical proximity. More broadly, the restructuring of the cellular components branch consolidates a shift toward treating Gene Ontology as a multi-layer semantic network-a transformation rooted in a decade-long process of scientific and social consensus across institutions. These findings underscore the need for version-aware, multi-layer models to ensure reproducibility and interpretability-and to better represent biological function across compositional, spatial, and regulatory dimensions as ontologies continue to evolve.
Protein functional annotation is crucial in biology, but many protein-coding genes remain uncharacterized, especially in non-model organisms. FANTASIA (Functional ANnoTAtion based on embedding space SImilArity) integrates protein language models for large-scale functional annotation. Applied to ~1000 animal proteomes, FANTASIA predicts functions to virtually all proteins, including up to 50% that remained unannotated by traditional homology-based methods. This enables the discovery of novel gene functions, enhancing our understanding of molecular evolution and organismal biology. FANTASIA holds particular promise for functional discovery in non-model taxa, offering advantages over homology-based tools in sensitivity and generalizability. FANTASIA is available on GitHub at https://github.com/CBBIO/FANTASIA .
Protein function prediction is critical for a wide range of applications in biology, spanning from functional genomics to protein design and genome evolution, among others. However, accurately predicting protein function remains a longstanding challenge in computational biology, especially for non-model organisms. Traditional methods based on sequence similarity often fail to annotate a significant proportion of proteins. The emergence of protein language models has significantly improved this process, enabling more accurate and comprehensive functional annotation. In this work, we highlight how the ProtTrans language model outperforms other tools in per-protein annotation, offering a more precise approach to predicting protein function. We also introduce functional annotation based on embedding space similarity (FANTASIA; available at https://github.com/MetazoaPhylogenomicsLab/FANTASIA ), a tool developed to harness these advances for large-scale annotation of uncharacterized proteomes. We provide a detailed overview of how to use FANTASIA, interpret its outputs, and demonstrate its utility in three case studies: (a) enrichment analyses from transcriptomics data, (b) assigning novel functions to unannotated genes in model organisms, and (c) identifying genes involved in important functions in non-model organisms. These results demonstrate the potential of protein language models to advance functional annotation in diverse biological contexts.
ABSTRACTProtein language models have been tested and proved to be reliable when used on curated datasets but have not yet been applied to full proteomes. Accordingly, we tested how two different machine learning based methods performed when decoding functional information from the proteomes of selected model organisms. We found that protein Language Models are more precise and informative than Deep Learning methods for all the species tested and across the three gene ontologies studied, and that they better recover functional information from transcriptomics experiments. The results obtained indicate that these Language Models are likely to be suitable for large scale annotation and downstream analyses, and we recommend a guide for their use.
Cellular senescence connects aging and cancer. Cellular senescence is a common program activated by cells in response to various types of stress. During this process, cells lose their proliferative capacity and undergo distinct morphological and metabolic changes. Senescence itself constitutes a tumor suppression mechanism and plays a significant role in organismal aging by promoting chronic inflammation. Additionally, age is one of the major risk factors for developing breast cancer. Therefore, while senescence can suppress tumor development early in life, it can also lead to an aging process that drives the development of age-related pathologies, suggesting an antagonistic pleiotropic effect. In this work, we identified Rian/MEG8 as a potential biomarker connecting aging and breast cancer for the first time. We found that Rian/MEG8 expression decreases with age; however, it is high in mice that age prematurely. We also observed decreased MEG8 expression in breast tumors compared to normal tissue. Furthermore, MEG8 overexpression reduced the proliferative and stemness properties of breast cancer cells both in vitro and in vivo by activating apoptosis. MEG8 could exemplify the antagonistic pleiotropic theory, where senescence is beneficial early in life as a tumor suppression mechanism due to increased MEG8, resulting in fewer breast tumors at an early age. Conversely, this effect could be detrimental later in life due to aging and cancer, when MEG8 is reduced and loses its tumor-suppressive role.
Background Understanding how coding genes and their functions evolve over time is a key aspect of evolutionary biology. Protein coding genes poorly understood or characterized at the functional level may be related to important evolutionary innovations, potentially leading to incomplete or inaccurate models of evolutionary change, and limiting the ability to identify conserved or lineage-specific features. Homology-based methodologies often fail to transfer functional annotations in a large fraction of the coding gene repertoire in non-model organisms. This is particularly relevant in animals, where a large number of their coding genes yield no functional annotation.Results Here, we leverage homology, deep learning, and protein language models to investigate functional annotation in the ‘dark proteome’ (defined as the unknown functional landscape’) of ca. 1,000 gene repertoires of virtually all animal phyla, totaling ca. 23.2 million coding genes. We then explored the ‘dark proteome’ of all animal phyla revealing an enrichment in functions related to immune response, viral infection, response to stimuli, development, or signaling, among others. Furthermore, we provide an open-source pipeline - FANTASIA - to implement and benchmark these methodologies in any dataset.Conclusions Our results uncover the putative functions of poorly understood protein-coding genes across the Animal Tree of Life that were inaccessible before due to the limitations in homology inference, contributing to a more comprehensive understanding of the molecular basis of animal evolution, and providing a new tool for the functional annotation of protein-coding genes in newly generated genomes.### Competing Interest StatementThe authors have declared no competing interest.
Environmental testing of high-touch objects is a potential noninvasive approach for monitoring population-level trends of SARS-CoV-2 and other respiratory viruses within a defined setting. We aimed to determine the association between SARS-CoV-2 contamination on high-touch environmental surfaces, community level case incidence, and university student health data. Environmental swabs were collected from January 2022 to November 2022 from high-touch objects and surfaces from five locations on a large university campus in Florida, USA. RT-qPCR was used to detect and quantify viral RNA, and a subset of positive samples was analyzed by viral genome sequencing to identify circulating lineages. During the study period, we detected SARS-CoV-2 viral RNA on 90.7 % of 162 tested samples. Levels of environmental viral RNA correlated with trends in community-level activity and case reports from the student health center. A significant positive correlation was observed between the estimated viral gene copy number in environmental samples and the weekly confirmed cases at the university. Viral sequencing data from environmental samples identified lineages concurrently circulating in the local community and state based on genomic surveillance data. Further, we detected emerging variants in environmental samples prior to their identification by clinical genomic surveillance. Our results demonstrate the utility of viral monitoring on high-touch environmental surfaces for SARS-CoV-2 surveillance at a community level. In communities with delayed or limited testing facilities, immediate environmental surface testing may considerably inform epidemic dynamics.
Mutations of the androgen receptor (AR) associated with prostate cancer and androgen insensitivity syndrome may profoundly influence its structure, protein interaction network, and binding to chromatin, resulting in altered transcription signatures and drug responses. Current structural information fails to explain the effect of pathological mutations on AR structure-function relationship. Here, we have thoroughly studied the effects of selected mutations that span the complete dimer interface of AR ligand-binding domain (AR-LBD) using x-ray crystallography in combination with in vitro, in silico, and cell-based assays. We show that these variants alter AR-dependent transcription and responses to anti-androgens by inducing a previously undescribed allosteric switch in the AR-LBD that increases exposure of a major methylation target, Arg761. We also corroborate the relevance of residues Arg761 and Tyr764 for AR dimerization and function. Together, our results reveal allosteric coupling of AR dimerization and posttranslational modifications as a disease mechanism with implications for precision medicine.
Enveloped viruses depend on the host endoplasmic reticulum (ER) quality control (QC) machinery for proper glycoprotein folding. The endoplasmic reticulum quality control (ERQC) enzyme α-glucosidase I (α-GluI) is an attractive target for developing broad-spectrum antivirals. We synthesized 28 inhibitors designed to interact with all four subsites of the α-GluI active site. These inhibitors are derivatives of the iminosugars 1-deoxynojirimycin (1-DNJ) and valiolamine. Crystal structures of ER α-GluI bound to 25 1-DNJ and three valiolamine derivatives revealed the basis for inhibitory potency. We established the structure-activity relationship (SAR) and used the Site Identification by Ligand Competitive Saturation (SILCS) method to develop a model for predicting α-GluI inhibition. We screened the compounds against SARS-CoV-2 in vitro to identify those with greater antiviral activity than the benchmark α-glucosidase inhibitor UV-4. These host-targeting compounds are candidates for investigation in animal models of SARS-CoV-2 and for testing against other viruses that rely on ERQC for correct glycoprotein folding.
Introduction: Complex gastroschisis is one of the most frequent causes of short bowel syndrome (SBS) in pediatrics. Motility disorders, intestinal dilatation and hypoplasia are factors that difficult progression and intestinal rehabilitation. We present the case of an infant whom a venting jejunostomy aided in treatment. Methods: 5-month-old male infant, with antenatal diagnosis of gastroschisis, born at 34 weeks, with finding of associated intestinal atresia, ischemia and stenosis in the ring and extra abdominal intestine at birth, also absence of colon from cecum to descending portion, with remnant hypoplastic descending and sigmoid colon. Initial surgery was an intestinal resection leaving 20 cm of remaining dilated small intestine and a side-to-end jejuno-colonic anastomosis + jejunostomy. At 45 days of life, patient required remodeling of stoma due to jejunostomy prolapse. Prolapse recurred few days after the procedure. Since birth patient was managed by the pediatric intestinal failure team with parenteral nutrition and controlled and slow progress of enteral nutrition, due to high output and mucous prolapse of the stoma. At 80 days of life, the jejunostomy was surgical closed, leaving the intestine in continuity. Postoperative evolution was stationary, with poor oral tolerance, abdominal distension, biliary emesis and minimal progress despite having daily bowel movements. Contrast images showed dilation of the small intestine with passage of the contrast through the colon and patent anastomosis, without signs of mechanical obstruction. The group decided to place a percutaneous endoscopic jejunostomy tube with a 14 Fr endoscopic gastrostomy kit through an anterograde approach for ventilation, proximal to the jejuno-colonic anastomosis. Endoscopy showed the entire remnant intestine dilated with normal mucosa and a patent wide jejuno-colon anastomosis. There were no complications (Image 1). Results: After the procedure, the patient was managed with intermittent opening of the tube for decompression, with a good outcome, tolerating oral route with normal stools and without abdominal distension or vomiting (Image 2). Conclusions: Patients with intestinal failure and SBS due to gastroschisis born with intestinal dilatation and motility disorders represent a therapeutic challenge since lengthening/remodeling surgeries are questionable. The placement of a percutaneous venting jejunostomy is a less invasive procedure and may be helpful to promote enteral progression and rehabilitation of these patients, without completely excluding the distal intestine and without complications such as prolapse that surgical jejunostomies present.
α-Synuclein is a 140 amino-acid intrinsically disordered protein mainly found in the brain. Toxic α-synuclein aggregates are the molecular hallmarks of Parkinson’s disease. In vitro studies showed that α-synuclein aggregates in oligomeric structures of several 10th of monomers and into cylindrical structures (fibrils), comprising hundred to thousands of proteins, with polymorphic cross-β-sheet conformations. Oligomeric species, formed at the early stage of aggregation remain, however, poorly understood and are hypothezised to be the most toxic aggregates. Here, we studied the formation of wild-type (WT) and mutant (A30P, A53T, and E46K) dimers of α-synuclein using coarse-grained molecular dynamics. We identified two principal segments of the sequence with a higher propensity to aggregate in the early stage of dimerization: residues 36–55 and residues 66–95. The transient α-helices (residues 53–65 and 73–82) of α-synuclein monomers are destabilized by A53T and E46K mutations, which favors the formation of fibril native contacts in the N-terminal region, whereas the helix 53–65 prevents the propagation of fibril native contacts along the sequence for the WT in the early stages of dimerization. The present results indicate that dimers do not adopt the Greek key motif of the monomer fold in fibrils but form a majority of disordered aggregates and a minority (9–15%) of pre-fibrillar dimers both with intra-molecular and intermolecular β-sheets. The percentage of residues in parallel β-sheets is by increasing order monomer < disordered dimers < pre-fibrillar dimers. Native fibril contacts between the two monomers are present in the NAC domain for WT, A30P, and A53T and in the N-domain for A53T and E46K. Structural properties of pre-fibrillar dimers agree with rupture-force atomic force microscopy and single-molecule Förster resonance energy transfer available data. This suggests that the pre-fibrillar dimers might correspond to the smallest type B toxic oligomers. The probability density of the dimer gyration radius is multi-peaks with an average radius that is 10 Å larger than the one of the monomers for all proteins. The present results indicate that even the elementary α-synuclein aggregation step, the dimerization, is a complicated phenomenon that does not only involve the NAC region.
To identify potential miRNAs involved in the development/progression of type 2 diabetes (T2D) we performed data screening and integration of microRNA (miRNA) and messenger RNA (mRNA) expression profiles available in public data repositories. We retrieved two independent sequencing datasets: GSE52314 (small RNA-Seq) and GSE50244 (RNA-Seq). Then, we identified Differentially Expressed Genes (DEGs) and Differentially Expressed miRNAs (DEMs) from non-diabetic (ND) vs. people with T2D. The integrative analysis of identified DEGs and DEMs revealed 303 interactions involving 33 miRNAs and 256 genes. Functional enrichment analysis showed these interactions could play a functional role in carbohydrate, lipid, and nucleotide impaired metabolism. Some interactions we identified could also be relevant at an early stage of T2D. We also performed a similar integration analysis using a single study from a rat model, including both RNA-Seq and small RNA-Seq in the same pancreatic tissue: CRA000791. Interestingly, in this analysis one identified miRNA was also identified in the integrative analysis in human transcriptome: hsa-miR-27a-5p. Moreover, functional enrichment analysis showed that the target genes are mainly enriched in biological processes and pathways related to lipid and nucleotide metabolism.The miRNAs and interactions we identified could be relevant to gene expression regulation in T2D. Although further validation is necessary for extensive use in clinics, they provide new and useful knowledge to develop potential novel strategies for early diagnosis, prevention, and treatment of T2D.
Abnormal aggregation of amyloid β (Aβ) peptides into fibrils plays a critical role in the development of Alzheimer's disease. A two-stage "dock-lock" model has been proposed for the Aβ fibril elongation process. However, the mechanisms of the Aβ monomer-fibril binding process have not been elucidated with the necessary molecular-level precision, so it remains unclear how the lock phase dynamics leads to the overall in-register binding of the Aβ monomer onto the fibril. To gain mechanistic insights into this critical step during the fibril elongation process, we used molecular dynamics (MD) simulations with a physics-based coarse-grained UNited-RESidue (UNRES) force field and sampled extensively the dynamics of the lock phase process, in which a fibril-bound Aβ(9-40) peptide rearranged to establish the native docking conformation. Analysis of the MD trajectories with Markov state models was used to quantify the kinetics of the rearrangement process and the most probable pathways leading to the overall native docking conformation of the incoming peptide. These revealed a key intermediate state in which an intra-monomer hairpin is formed between the central core amyloidogenic patch 18VFFA21 and the C-terminal hydrophobic patch 34LMVG37. This hairpin structure is highly favored as a transition state during the lock phase of the fibril elongation. We propose a molecular mechanism for facilitation of the Aβ fibril elongation by amyloidogenic hydrophobic patches.
The glucocorticoid receptor (GR) is a ubiquitously expressed transcription factor that controls metabolic and homeostatic processes essential for life. Although numerous crystal structures of the GR ligand-binding domain (GR-LBD) have been reported, the functional oligomeric state of the full-length receptor, which is essential for its transcriptional activity, remains disputed. Here we present five new crystal structures of agonist-bound GR-LBD, along with a thorough analysis of previous structural work. We identify four distinct homodimerization interfaces on the GR-LBD surface, which can associate into 20 topologically different homodimers. Biologically relevant homodimers were identified by studying a battery of GR point mutants including crosslinking assays in solution, quantitative fluorescence microscopy in living cells, and transcriptomic analyses. Our results highlight the relevance of non-canonical dimerization modes for GR, especially of contacts made by loop L1-3 residues such as Tyr545. Our work illustrates the unique flexibility of GR's LBD and suggests different dimeric conformations within cells. In addition, we unveil pathophysiologically relevant quaternary assemblies of the receptor with important implications for glucocorticoid action and drug design.
α-Synuclein is an intrinsically disordered protein occurring in different conformations and prone to aggregate in β-sheet structures, which are the hallmark of the Parkinson disease. Missense mutations are associated with familial forms of this neuropathy. How these single amino-acid substitutions modify the conformations of wild-type α-synuclein is unclear. Here, using coarse-grained molecular dynamics simulations, we sampled the conformational space of the wild type and mutants (A30P, A53P, and E46K) of α-synuclein monomers for an effective time scale of 29.7 ms. To characterize the structures, we developed an algorithm, CUTABI (CUrvature and Torsion based of Alpha-helix and Beta-sheet Identification), to identify residues in the α-helix and β-sheet from Cα -coordinates. CUTABI was built from the results of the analysis of 14,652 selected protein structures using the Dictionary of Secondary Structure of Proteins (DSSP) algorithm. DSSP results are reproduced with 93% of success for 10 times lower computational cost. A two-dimensional probability density map of α-synuclein as a function of the number of residues in the α-helix and β-sheet is computed for wild-type and mutated proteins from molecular dynamics trajectories. The density of conformational states reveals a two-phase characteristic with a homogeneous phase (state B, β-sheets) and a heterogeneous phase (state HB, mixture of α-helices and β-sheets). The B state represents 40% of the conformations for the wild-type, A30P, and E46K and only 25% for A53T. The density of conformational states of the B state for A53T and A30P mutants differs from the wild-type one. In addition, the mutant A53T has a larger propensity to form helices than the others. These findings indicate that the equilibrium between the different conformations of the α-synuclein monomer is modified by the missense mutations in a subtle way. The α-helix and β-sheet contents are promising order parameters for intrinsically disordered proteins, whereas other structural properties such as average gyration radius, R g , or probability distribution of R g cannot discriminate significantly the conformational ensembles of the wild type and mutants. When separated in states B and HB, the distributions of R g are more significantly different, indicating that global structural parameters alone are insufficient to characterize the conformational ensembles of the α-synuclein monomer.
Polycyclic triterpenes are members of the terpene family produced by the cyclization of squalene. The most representative polycyclic triterpenes are hopanoids and sterols, the former are mostly found in bacteria, whereas the latter are largely limited to eukaryotes, albeit with a growing number of bacterial exceptions. Given their important role and omnipresence in most eukaryotes, contrasting with their scant representation in bacteria, sterol biosynthesis was long thought to be a eukaryotic innovation. Thus, their presence in some bacteria was deemed to be the result of lateral gene transfer from eukaryotes. Elucidating the origin and evolution of the polycyclic triterpene synthetic pathways is important to understand the role of these compounds in eukaryogenesis and their geobiological value as biomarkers in fossil records. Here, we have revisited the phylogenies of the main enzymes involved in triterpene synthesis, performing gene neighborhood analysis and phylogenetic profiling. Squalene can be biosynthesized by two different pathways containing the HpnCDE or Sqs proteins. Our results suggest that the HpnCDE enzymes are derived from carotenoid biosynthesis ones and that they assembled in an ancestral squalene pathway in bacteria, while remaining metabolically versatile. Conversely, the Sqs enzyme is prone to be involved in lateral gene transfer, and its emergence is possibly related to the specialization of squalene biosynthesis. The biosynthesis of hopanoids seems to be ancestral in the Bacteria domain. Moreover, no triterpene cyclases are found in Archaea, invoking a potential scenario in which eukaryotic genes for sterol biosynthesis assembled from ancestral bacterial contributions in early eukaryotic lineages.