DisProt (https://disprot.org/) is an open database integrating experimental evidence on intrinsically disordered proteins (IDPs), intrinsically disordered regions (IDRs), and their functions. Over the past two years, the database has grown over 20%, now comprising 3201 IDPs and 13 347 pieces of evidence, including over 1500 new structural state annotations and >1300 new function annotations. DisProt has systematically adopted the Minimum Information About Disorder Experiments (MIADE) guidelines, more than doubling annotations with experimental details and improving the interpretability of disorder-related experiments. The website has evolved into a hybrid knowledgebase and deposition system, introducing a Deposition Page that allows direct submissions by external users. Through BLAST-based homology propagation in MobiDB, DisProt disorder regions and linear interacting peptides have been extended from hundreds to hundreds of thousands of proteins across >11 000 organisms. This new release marks a paradigm shift by integrating computational predictions as valid evidence and introducing major updates and restructuring of the IDP Ontology, enhancing accuracy, interoperability, and semantic clarity. DisProt continues to support community engagement through training resources together with DisTriage, an AI-based literature triage tool, providing curators with regularly updated lists of prioritized publications.
Disordered proteins play essential roles in myriad cellular processes, yet their structural characterization remains a major challenge due to their dynamic and heterogeneous nature. We here present a community-driven initiative to address this problem by advocating a unified framework for determining conformational ensembles of disordered proteins. Our aim is to integrate state-of-the-art experimental techniques with advanced computational methods, including knowledge-based sampling, enhanced molecular dynamics, and machine learning models. The modular framework comprises three interconnected components: experimental data acquisition, computational ensemble generation, and validation. The systematic development of this framework will ensure the accurate and reproducible determination of conformational ensembles of disordered proteins. We highlight the open challenges necessary to achieve this goal, including force field accuracy, efficient sampling, and environmental dependency, advocating for collaborative benchmarking and standardized protocols.
Abstract The androgen receptor (AR) is a transcription factor whose overactivation is a primary driver of prostate cancer. Although AR interactions with several long non-coding RNAs (lncRNAs) have been implicated in castration-resistant prostate cancer, their underlying molecular mechanisms and functional consequences remain poorly understood. Here, we identify the N-terminal 37 residues of the intrinsically disordered AR N-terminal domain as the primary RNA-binding region that mediates selective interactions with the lncRNAs HOTAIR and SLNCR1. Residue-level mapping and mutational analysis define Y11, R13, and Q24 as key determinants of RNA recognition. This RNA-binding region partially overlaps with the F 23 QNLF 27 motif, previously shown to mediate N/C interdomain communication with the ligand-binding domain through folding upon binding. We observe a partner-dependent binding mode in which this motif remains dynamically disordered upon RNA binding. LncRNA binding promotes phase separation of the N-terminal domain, indicating that lncRNAs upregulated in late-stage prostate cancer may lower the threshold for AR condensate formation and contribute to ligand-independent AR signaling. LncRNAs modulate communication between the N-terminal and ligand-binding domains within condensates, suggesting that lncRNA binding may tune hormone-dependent full-length AR signaling. These findings provide a mechanistic framework for AR-lncRNA regulation, laying the groundwork for future therapeutic strategies against advanced prostate cancer. Graphical abstract
Abstract • Introduction. Prostate cancer “stemness” increases with tumor progression, predicts poor prognosis and contributes to therapy resistance. The cyclin D1 gene (CCND1), which encodes the regulatory subunit of a holoenzyme that phosphorylates RB and conveys transcriptional properties, is expressed in prostate cancer and augments growth of some castrate resistant prostate cancers. Clinical trials are targeting the kinase activity of cyclin D1 in prostate cancer. The functional significance of the cyclin D1 E domain remained to be further defined. We investigated a potential kinase-independent function of cyclin D1 in prostate cancer. • Methods. Analysis of patient gene expression, tumor histology, gene knockout transgenic mice, tissue culture stem cell assays, proteomic analysis. • Findings. We show cyclin D1 expression is increased in PCa tumor stroma compared with adjacent tissue. Using single cell sequencing cyclin D1 we identified cyclin D1 in cancer associated fibroblasts. In vivo, cell markers of PCa stemness including Trop2ICD activity. Deletion of both epithelial cell and stromal cyclin D1 further reduced markers of PCa stem cells suggesting a role for extra epithelial stromal cell cyclin D1 in the induction of PCa stemness. In vitro, cyclin D1 rescue of cyclin D1 deficient fibroblasts demonstrated cyclin D1 governed a secretome that augmented prostate cancer stemness, assessed by prostate cancer sphere size and number, and expression of pro-stemness chemokine receptors (CXCR2, CX3CR1). Mutational analysis showed the cyclin D1 heterotypic induction of stemness and chemokine receptor expression required an intrinsically disordered acidic rich “E domain” (AA 272-280) within the cyclin D1 carboxyl terminus. Integrated proteomics and ChIP Seq identified transcriptional targets governing cancer stemness and pro-tumorigenic inflammation. • Conclusions. Cyclin D1 augments prostate cancer stemness through heterotypic functions via an intrinsically disordered carboxyl terminal domain. Citation Format: Xuanmao Jiao, Danni Li, Rita Pancsa, Ritika Harish, Zhiping Li, Gideon Tolufashe, Hallgeir Rui, Beatrice S. Knudsen, Yanming Du, Hsin-Yao Tang, Peter Tompa, Richard G. Pestell. Stromal cyclin D1 promotes prostate cancer stemness via the transcriptional regulatory E domain [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2026; Part 1 (Regular Abstracts); 2026 Apr 17-22; San Diego, CA. Philadelphia (PA): AACR; Cancer Res 2026;86(7 Suppl):Abstract nr 2265.
AlphaFold2 changed structural biology by providing high-quality structure predictions for all possible proteins. Since its inception, a plethora of applications were built on AlphaFold2, expediting discoveries in virtually all areas related to protein science. In many cases, however, optimism seems to have made scientists forget about data leakage, a serious issue that needs to be addressed when evaluating machine learning methods. Here we provide a rigorous benchmark set that can be used in a broad range of applications built around AlphaFold2/3.
Intrinsically disordered proteins have key signalling and regulatory roles in cells and are frequently dysregulated in diseases such as cancer, neurodegeneration, inflammation and autoimmune disorders. Preventing the pathological functions mediated by structural disorder is crucial to successfully target proteins that drive transcription, biomolecular condensation and protein aggregation. However, owing to their heterogeneous, highly dynamic structural states, with ensembles of rapidly interconverting conformations, disordered proteins have been considered largely 'undruggable' by traditional approaches. Here, we review key developments of the field and suggest that the synergy of advanced experimental and computational approaches needs to be pursued to conquer this barrier in drug discovery.
BACKGROUND:Numerous cellular processes rely on biomolecular condensates formed through liquid-liquid phase separation (LLPS). Recently, it has become evident that somatic mutations can interfere with or over-activate the formation of phase-separated condensates. RESULTS:Here, we set out to systematically study the connection between cancer and biological condensation, specifically mapping the extent to which LLPS is affected in cancer and understanding the molecular pathomechanisms and therapeutic consequences of mutations affecting LLPS scaffolds. We identify both known and novel combinations of molecular functions that are specific to oncogenic fusion proteins and thus have a high potential for driving tumorigenesis. Protein regions driving condensate formation show an increased association with DNA- or chromatin-binding domains of transcription regulators within oncogenic fusion proteins, indicating a common molecular mechanism underlying several soft tissue sarcomas and hematologic malignancies where phase-separation-prone oncogenic fusion proteins form abnormal condensates along the DNA and thereby dysregulate gene expression programs. CONCLUSIONS:We find that proteins initiating LLPS are frequently implicated in somatic cancers, even surpassing their involvement in neurodegeneration. Our data shows that cancer-driving LLPS scaffolds tend to be potent oncogenes, giving rise to dominant phenotypes and lacking targeting options by current FDA-approved drugs. Finding the currently missing drugs to shut down oncogenic fusion proteins, to disrupt the condensation enabled by them, and to offset their downstream effects could provide cancer drugs widely applicable to diverse cancer incidences previously defying standard treatments.
In recent years, kinetics of aggregation and liquid-liquid phase separation (LLPS) of certain proteins have been extensively studied due to their link to neurodegenerative diseases such as amyotrophic lateral sclerosis and Parkinson's disease. Optical techniques including absorbance and fluorescence spectroscopy are widely used for these investigations. However, most commercial devices suffer from high sample volume (100 mu L/test) requirements. Optofluidic lab-on-chips (LOCs) provide accurate sampling and mixing in nanoliter scale sample volumes. State-of-the-art optofluidic chips have fabrication tolerances as low as the fiber core diameter, requiring monolithic fabrication and limiting the microfluidic design. We present the design and prototype fabrication of an optofluidic chip to monitor the absorbance and fluorescence of phase separating and aggregating protein samples. The design offers tolerances wide enough for assembly with a variety of microfluidic chips. The prototype LOC is compared with a commercial plate reader (BioTek Synergy MX) through internal validation of polystyrene beads (absorbance) and propidium iodide (fluorescence) dilution series. The LOC either matches or exceeds the plate reader in the limit of detection, linearity, and precision. Furthermore, LLPS and aggregation of G3BP1, a protein linked with neurodegenerative diseases, are demonstrated through absorbance and fluorescence (with thioflavin T dye), using the LOC and the commercial plate reader.
Peroxisome proliferator-activated receptor γ (PPARγ), which is expressed in a variety of malignancies, governs biological functions through transcriptional programs. Defining the molecular mechanisms governing the selection of canonical versus non-canonical PPARγ binding sequences may provide the opportunity to design regulators with distinct functions and side effects. Acetylation at K268/293 in mouse Pparγ2 participates in the regulation of adipose tissue differentiation, and the conserved lysine residues (K154/155) in mouse Pparγ1 governs lipogenesis in breast cancer cells. Herein, the PPARγ1 acetylated residues K154/155 were shown to be essential for oncogenic ErbB2 driven breast cancer growth and mammary tumor stem cell expansion in vivo. The induction of transcriptional modules governing growth factor signaling, lipogenesis, cellular apoptosis, and stem cell expansion were dependent upon K154/155. The acetylation status of the K154/155 residues determined the selection of genome-wide DNA binding sites, altering the selection from canonical to non-canonical (C/EBP) DNA sequence-specific binding. The gene signature reflecting the acetylation-dependent genomic occupancy in lipogenesis provided predictive value in survival outcomes of ErbB2+ breast cancer. The Pparγ1 acetylation site is critical for ErbB2-induced breast cancer tumor growth and may represent a relevant target for therapeutic coextinction.
Heterogeneous nuclear ribonucleoprotein A2/B1 (hnRNPA2B1) is a multifunctional RNA-binding protein involved in RNA maturation and mRNA transport. It has recently been shown to undergo liquid-liquid phase separation (LLPS), contributing to the assembly of membraneless organelles. Moreover, dysregulation of LLPS is associated with the formation of pathogenic protein aggregates, in which hnRNPA2B1 is frequently found. Despite its biological and pathological relevance, studies on the full-length protein remain limited due to its intrinsically disordered low-complexity domain, which renders hnRNPA2B1 highly aggregation-prone and difficult to purify. In this study, we report the successful expression and purification of full-length hnRNPA2B1 with high purity and minimal nucleic acid contamination. By optimizing buffer conditions, specifically ionic strength and pH, we maintained the protein in solution following cleavage of its solubility tag. Preliminary in vitro characterization under near-physiological conditions reveals that purified hnRNPA2B1 undergoes LLPS, forming dynamic liquid-like droplets that grow and mature into amorphous aggregates. Our approach provides a robust method for purifying hnRNPA2B1 suitable for LLPS and aggregation studies. This strategy may also be useful to purify other aggregation-prone, intrinsically disordered proteins.
The toxic effects of C9orf72- derived arginine-rich dipeptide repeats (R-DPRs) on cellular stress granules in amyotrophic lateral sclerosis (ALS) and frontotemporal dementia remain unclear at the molecular level. Stress granules are formed through the switch of Ras GTPase- activating protein- binding protein 1 (G3BP1) by RNA from a closed inactive state to an open activated state, driving the formation of the organelle by liquid-liquid phase separation (LLPS). We show that R-DPRs bind G3BP1 a thousand times stronger than RNA and initiate LLPS much more effectively. Their pathogenic effect is underscored by the slow transition of R- DPR-G3BP1 droplets to aggregated, ThS- positive states that can recruit ALS- linked proteins hnRNPA1, hnRNPA2, and TDP-43. Deletion constructs and molecular simulations show that R-DPR binding and LLPS are mediated via the negatively charged intrinsically disordered region 1 (IDR1) of the protein, allosterically regulated by its positively charged IDR3. Bioinformatic analyses point to the strong mechanistic parallels of these effects with the interaction of R-DPRs with nucleolar nucleophosmin 1 (NPM1) and underscore that R-DPRs interact with many other similar nucleolar and stress- granule proteins, extending the underlying mechanism of R-DPR toxicity in cells. Our results also highlight characteristic differences between the two R-DPRs, poly-GR and poly- PR, and suggest that the primary pathological target of poly-GR is not NPM1 in nucleoli, but G3BP1 in stress granules in affected cells.
Protein cis-regulatory elements (CREs) are regions that modulate the activity of a protein through intramolecular interactions. Kinases, pivotal enzymes in numerous biological processes, often undergo regulatory control via inhibitory interactions in cis. This study delves into the mechanisms of cis regulation in kinases mediated by CREs, employing a combined structural and sequence analysis. To accomplish this, we curated an extensive dataset of kinases featuring annotated CREs, organized into homolog families through multiple sequence alignments. Key molecular attributes, including disorder and secondary structure content, active and ATP-binding sites, post-translational modifications, and disease-associated mutations, were systematically mapped onto all sequences. Additionally, we explored the potential for conformational changes between active and inactive states. Finally, we explored the presence of these kinases within membraneless organelles and elucidated their functional roles therein. CREs display a continuum of structures, ranging from short disordered stretches to fully folded domains. The adaptability demonstrated by CREs in achieving the common goal of kinase inhibition spans from direct autoinhibitory interaction with the active site within the kinase domain, to CREs binding to an alternative site, inducing allosteric regulation revealing distinct types of inhibitory mechanisms, which we exemplify by archetypical representative systems. While this study provides a systematic approach to comprehend kinase CREs, further experimental investigations are imperative to unravel the complexity within distinct kinase families. The insights gleaned from this research lay the foundation for future studies aiming to decipher the molecular basis of kinase dysregulation, and explore potential therapeutic interventions.
The essential G 1 -cyclin, CCND1 , is frequently overexpressed in cancer, contributing to tumorigenesis by driving cell-cycle progression. D-type cyclins are rate-limiting regulators of G 1 -S progression in mammalian cells via their ability to bind and activate CDK4 and CDK6. In addition, cyclin D1 conveys kinase-independent transcriptional functions of cyclin D1. Here we report that cyclin D1 associates with H2B S14 via an intrinsically disordered domain (IDD). The same region of cyclin D1 was necessary for the induction of aneuploidy, induction of the DNA damage response, cyclin D1-mediated recruitment into chromatin, and CIN gene transcription. In response to DNA damage H2B S14 phosphorylation occurs, resulting in co-localization with γH2AX in DNA damage foci. Cyclin D1 ChIP seq and γH2AX ChIP seq revealed ~14% overlap. As the cyclin D1 IDD functioned independently of the CDK activity to drive CIN, the IDD domain may provide a rationale new target to complement CDK-extinction strategies.
Liquid-liquid phase separation (LLPS) is pivotal in forming biomolecular condensates, which are crucial in several biological processes. Intrinsically disordered regions (IDRs) are typically responsible for driving LLPS due to their multivalency and high content of charged residues that enable the establishment of electrostatic interactions. In our study, we examined the role of charge distribution in the condensation of the disordered N-terminal domain of human topoisomerase I (hNTD). hNTD is densely charged with oppositely charged residues evenly distributed along the sequence. Its LLPS behavior was compared with that of charge permutants exhibiting varying degrees of charge segregation. At low salt concentrations, hNTD undergoes LLPS. However, LLPS is inhibited by high concentrations of salt and RNA, disrupting electrostatic interactions. Our findings show that, in hNTD, moderate charge segregation promotes the formation of liquid condensates that are sensitive to salt and RNA, whereas marked charge segregation results in the formation of aberrant condensates. Although our study is based on a limited set of protein variants, it supports the applicability of the "stickers-and-spacers" model to biomolecular condensates involving highly charged IDRs. These results may help generate reliable models of the overall LLPS behavior of supercharged polypeptides.
Natural selection can drive organisms to strikingly similar adaptive solutions, but the underlying molecular mechanisms often remain unknown. Several amphibians have independently evolved highly adhesive skin secretions (glues) that support a highly effective antipredator defence mechanism. Here we demonstrate that the glue of the Madagascan tomato frog, Dyscophus guineti, relies on two interacting proteins: a highly derived member of a widespread glycoprotein family and a galectin. Identification of homologous proteins in other amphibians reveals that these proteins attained a function in skin long before glues evolved. Yet, major elevations in their expression, besides structural changes in the glycoprotein (increasing its structural disorder and glycosylation), caused the independent rise of glues in at least two frog lineages. Besides providing a model for the chemical functioning of animal adhesive secretions, our findings highlight how recruiting ancient molecular templates may facilitate the recurrent evolution of functional innovations.
The Protein Ensemble Database (PED) (URL: https://proteinensemble.org) is the primary resource for depositing structural ensembles of intrinsically disordered proteins. This updated version of PED reflects advancements in the field, denoting a continual expansion with a total of 461 entries and 538 ensembles, including those generated without explicit experimental data through novel machine learning (ML) techniques. With this significant increment in the number of ensembles, a few yet-unprecedented new entries entered the database, including those also determined or refined by electron paramagnetic resonance or circular dichroism data. In addition, PED was enriched with several new features, including a novel deposition service, improved user interface, new database cross-referencing options and integration with the 3D-Beacons network-all representing efforts to improve the FAIRness of the database. Foreseeably, PED will keep growing in size and expanding with new types of ensembles generated by accurate and fast ML-based generative models and coarse-grained simulations. Therefore, among future efforts, priority will be given to further develop the database to be compatible with ensembles modeled at a coarse-grained level.
The circumsporozoite protein (CSP) is the main surface antigen of the Plasmodium sporozoite (SPZ) and forms the basis of the currently only licensed anti-malarial vaccine (RTS,S/AS01). CSP uniformly coats the SPZ and plays a pivotal role in its immunobiology, in both the insect and the vertebrate hosts. Although CSP's N-terminal domain (CSPN) has been reported to play an important role in multiple CSP functions, a thorough biophysical and structural characterization of CSPN is currently lacking. Here, we present an alternative method for the recombinant production and purification of CSPN from Plasmodium falciparum (PfCSP(N)), which provides pure, high-quality protein preparations with high yields. Through an interdisciplinary approach combining in-solution experimental methods and in silico analyses, we provide strong evidence that PfCSP(N) is an intrinsically disordered region displaying some degree of compaction.
Tauopathies, a group of neurodegenerative disorders, are characterized by the abnormal aggregation of microtubule-associated Tau proteins in neurons and glial cells. The process of Tau proteins transitioning from soluble, intrinsically disordered monomers to disease-associated aggregates is still unclear. Investigating these molecular mechanisms requires the reconstitution of such processes in cellular and in vitro models using recombinant proteins at high purity and yield. However, the production of phase-separating or aggregation-prone recombinant proteins like Tau’s hydrophobic-rich domains or disease mutation-carrying variants on a large scale is highly challenging due to their limited solubility. To overcome this challenge, we have developed an improved strategy for expressing and purifying recombinant Tau proteins using the major ampullate spidroin-derived solubility tag (MaSp-NT*). This approach involves using NT* as a fusion tag to enhance the solubility and stability of expressed proteins by forming micelle-like particles within the cytosol of E. coli cells. We found that fusion with the NT* tag significantly increased the solubility and yield of highly hydrophobic and/or aggregation-prone Tau constructs. Our purification method for NT* fusion proteins yielded up to twenty-fold higher amounts than proteins purified using our novel tandem-tag (6xHis-SUMO-Tau-Heparin) purification system. This enhanced expression and yield were demonstrated with full-length Tau (hT40/Tau441), its particularly aggregation-prone repeat domain (Tau-MTBR), and Frontotemporal dementia (FTD)-associated mutant (Tau-P301L). These advancements offer promising avenues for the production of large quantities of Tau proteins suitable for in vitro experimental techniques such as nuclear magnetic resonance (NMR) spectroscopy without the need for a boiling step, bringing us closer to effective treatments for tauopathies.
An unambiguous description of an experiment, and the subsequent biological observation, is vital for accurate data interpretation. Minimum information guidelines define the fundamental complement of data that can support an unambiguous conclusion based on experimental observations. We present the Minimum Information About Disorder Experiments (MIADE) guidelines to define the parameters required for the wider scientific community to understand the findings of an experiment studying the structural properties of intrinsically disordered regions (IDRs). MIADE guidelines provide recommendations for data producers to describe the results of their experiments at source, for curators to annotate experimental data to community resources and for database developers maintaining community resources to disseminate the data. The MIADE guidelines will improve the interpretability of experimental results for data consumers, facilitate direct data submission, simplify data curation, improve data exchange among repositories and standardize the dissemination of the key metadata on an IDR experiment by IDR data sources.
Menin is a protein that is regulated via protein-protein interactions by different binding partners, such as mixed lineage leukemia protein (MLL) and androgen receptor (AR). We observed that menin-AR and menin-MLL interactions are regulated by concentration-dependent dimerization of menin, and its interaction with cancer-related AR. As a result of its oligomerization-dependent interaction with both AR and MLL, menin is recruited into AR-RNA and MLL-RNA condensates formed by liquid-liquid phase separation (LLPS), with different outcomes under AR-overexpression or MLL-overexpression conditions representing different cancer types. At high concentrations promoting menin dimerization, it inhibits MLL-RNA LLPS, while making AR-RNA condensates less dynamic, i.e., more gel-like. Regions of AR show both negative/positive cooperativity in menin binding. AR contains a specific menin-binding region (MBR) in its intrinsically disordered N-terminal domain (NTD), menin binding of which is inhibited by the adjacent DNA-binding domain (DBD), but facilitated by a hinge region located between its DBD and ligand-binding domain (LBD) as well as by N terminus of AR. Interestingly, the hinge region reduces the propensity of full-length AR to undergo LLPS in the presence of RNA, which is facilitated by an alternative hinge region present in the tumor-specific AR isoform, AR-v7. As both menin and MLL are recruited into AR-driven, functional cellular condensates aggravated in the case of AR-v7, we posit that the menin-AR-MLL system represents a fine-tuned condensate module of transcription regulation that is balanced toward the tumor-suppressor activity of menin. Our results suggest that this balance can be upset by prevalent oncogenic events, such as menin upregulation and/or AR-v7 overexpression, in cancer.