Compatibilized polymer blends are a complex yet versatile and widespread category of materials. When the components of a binary blend are immiscible, they are typically driven toward a macrophase-separated state, but with the introduction of electrostatic interactions, they can be either homogenized or shifted to microphase separation. However, both experimental and simulation approaches face significant challenges in efficiently exploring the vast design space of charge-compatibilized polymer blends, encompassing chemical interactions, architectural properties, and composition. In this work, we introduce a white-box machine learning approach integrated with polymer field theory to predict the phase behavior of these systems, which is significantly more accurate than conventional black-box machine learning approaches. The random phase approximation (RPA) calculation is used as a testbed to determine polymer phases. Instead of directly predicting the polymer phase output of RPA calculations from a large input space by a machine learning model, we build a parallel partial Gaussian process model to predict the most computationally intensive component of the RPA calculation that involves only polymer architecture parameters as inputs. This approach substantially reduces the computational cost of the RPA calculation across a vast input space with nearly 100% accuracy for the out-of-sample prediction, enabling rapid screening of polymer blend charge-compatibilization designs. More broadly, the white-box machine learning strategy offers a promising approach for the dramatic acceleration of polymer field-theoretic methods for mapping out polymer phase behavior.
Amyloid formation and liquid-liquid phase separation (LLPS) are two important phenomena in cellular biology, linked to both normal physiological functions and various pathologies. Here, we present a computational framework that scores amyloid propensities (amyloid-predict) or LLPS propensities (LLPS-predict) from protein language model embeddings, enabling rapid proteome-wide annotation of peptides and residues. amyloid-predict achieves classification performance that exceeds existing AI and physics-based tools on a hexapeptide benchmark while enabling substantially faster high-throughput screening; notably, amyloid-predict is sensitive to subtle mutational effects and is influenced by sequence patterning and context rather than amino acid composition alone. We apply these protein language model classifiers to all the IDRs in the human proteome and uncover several protein categories with significant enhancement in amyloid and/or LLPS propensity, suggesting insights into the biological roles of these protein categories. For example, signaling receptors, carbohydrate-binding proteins, and Ca2+ binding proteins are enriched in aggregation propensity, while mRNA-binding proteins, ribonucleoprotein complex, and nuclear matrix proteins are enriched in LLPS propensity. Interestingly, we observe patterns of both high amyloid and LLPS propensity in several amyloid-forming and prionic proteins. Together, these results provide side-by-side landscapes of LLPS and amyloid potential across the disordered human proteome while offering a rapid screening tool for basic biology, disease-mechanism studies, and rational design of peptide therapeutics.
Hydrophobicity governs a vast range of phenomena, from protein-protein interactions to nanomaterial assembly, and can be rigorously quantified by the dewetting free energy (Fdewet) of a molecule or surface. However, hydrophobicity remains widely treated as an additive property of amino acid identity, obscuring the fact that water's response is a collective property of the surface, shaped by curvature, chemical patterning, and neighboring residues. Direct calculation of Fdewet via specialized molecular simulations captures this collective behavior but is prohibitively slow, leaving in place broadly-used, decades-old sequence-based hydropathy scales that neglect the physics of solvation. Here we show that residue-level Fdewet can be predicted from a compact set of local water features (water structural signatures and residue-water potential energy) extracted from a brief and inexpensive all-atom simulation. We embed this insight in two models: HydroMap, which predicts Fdewet directly from water features, and FastHydroMap, a computationally inexpensive graph neural network surrogate trained on HydroMap that requires no solvent simulation. HydroMap and FastHydroMap capture context-dependent hydrophobicity that classical, sequence-only hydropathy scales miss. We demonstrate this across three protein systems: on an α-synuclein amyloid filament, strongly dewetting interfaces align with unassigned peptide densities, revealing hidden binding sites; in calmodulin, hydrophobicity redistributes upon Ca 2+ binding; and for Protein G, time-resolved hydrophobicity changes track the folding trajectory. Together, these models make Fdewet a computationally inexpensive descriptor for proteins, membranes, and other surfaces, enabling rapid scoring for materials design and a time-resolved view of dynamic hydrophobic-mediated processes such as protein folding. Significance Statement:Hydrophobicity, the tendency of surfaces to expel water, drives how proteins fold and how molecules recognize one another. For decades, it has widely been treated as a fixed property of an amino acid or chemical group, but water actually responds to the collective shape and chemistry of a surface, such as that presented by a protein, not to its components in isolation. Measuring this collective response from molecular simulation is rigorous but prohibitively slow. We show that it can instead be inferred from a compact set of features describing water structure and interactions near a surface, and we use this insight to build models that predict hydrophobicity rapidly and at residue resolution, enabling practical, physically grounded design of hydrophobic-mediated interactions.
Surfactant self-assembly in soft matter formulations spans a complex, multivariate design space, motivating the development of efficient and predictive computational modeling approaches to aid formulation design. However, conventional molecular simulation techniques, including all-atom molecular dynamics and coarse-grained methods, are limited by either accessible time and length scales or predictive accuracy in studying surfactant self-assembly. To address these challenges, we employ a multiscale methodology that uses small-scale all-atom simulations to parameterize statistical field-theoretic models via bottom-up coarse-graining, eliminating the need for experimental input. The resulting molecularly informed field theory is then sampled using self-consistent field theory calculations to efficiently predict self-assembly and phase behavior. We demonstrate this approach by constructing binary phase diagrams for cationic alkyl quaternary-ammonium surfactants (C16TAB, C16TAC, and C10TAB) in water across a range of temperatures, compositions, and salt concentrations. The model successfully captures all experimentally observed surfactant mesophases and reproduces the majority of phase transition orderings de novo. In addition, we show this approach provides a unified framework for predicting equilibrium properties such as mesostructure domain sizes, micelle aggregation numbers, and critical micelle concentrations with qualitative agreement to experimentally observed trends. This multiscale methodology has the potential to be integrated into high-throughput screening workflows for efficient prediction of phase diagrams in novel surfactant formulations.
Abstract Tau assembles into fibrillar aggregates that are pathological hallmarks of a group of neurodegenerative diseases collectively called tauopathies. Templated aggregation of naïve tau to seeding-competent fibrils that proceed from cell to cell is a key driver of prion-like progression of tauopathies. This study tests the hypothesis that tau, an intrinsically disordered protein (IDP), achieves in-register stacking to form seed-competent fibrils by a pinning action of tau to each other and/or the seed surface via a single dominant hotspot to avoid mismatch in tau stacking to fibrils. Structured solvation water has been proposed to be a signature of such hotspots at both the tau fibril-end surface and soluble tau monomers. Although jR2R3-P301L tau exhibits a heterogeneous hydration landscape in its intrinsically disordered state, with enhanced water structuring near the P301L mutation site, it is unclear whether a localized hotspot exists at the fibril end surface and surface water facilitates the initial contacts in templated aggregation. Using rapid 1 H- 15 N SOFAST-HMQC NMR to track seed-induced aggregation of jR2R3-P301L in real time, complemented by molecular dynamics simulation of fibril surface hydration, we identify a residue-specific pinning hotspot that is prone to dewetting followed by sequential folding and incorporation of the remaining segment in a two-step dock-and-lock process. Site-specific spin labeling further demonstrates that blocking this pinning hotspot disrupts templated aggregation, leading to shorter fibrils. The identification of a dominant pinning site will facilitate the rational design of binders to effectively disrupt fibril extension or serve as diagnostic or therapeutic strategies. Significance Statement Tau proteins must align and stack precisely with existing fibril ends to propagate pathology, yet the molecular signature that initiates and ensures this in-register alignment has been unclear. This study shows recruitment begins at a single, structurally well-defined contact site on the fibril surface that is enriched in release-prone hydration water. These findings reveal that water-release-prone hotspots, rather than the well-known amyloid-core regions alone, govern the initial steps in templating seeding and can hence be blocked, opening a new path for identifying therapeutic targets that could slow the progression of tau-related diseases.
Localized measurements of hydration dynamics surrounding poly(acrylic acid) (pAA) chains are executed using Overhauser dynamic nuclear polarization relaxometry to determine the role of polymer protonation state on the polymer solvation environment. Alterations in polymer concentration and solution pH uncovered three different regimes of local water dynamics ranging from bulk-water-like to subdued water diffusivity. These experiments, combined with molecular dynamics simulations, reveal that the presence of deprotonated carboxylic acid groups promotes a hydrophilic local environment and extended polymer chain configurations that together incentivize the incorporation of bulk-like water near the polymer. This results in a stretched polymer conformation that leads to faster observed hydration dynamics in the local hydration environment. Meanwhile, protonated carboxylic acid groups are expected to engage in intrapolymer hydrogen bonding between monomer groups, resulting in a more collapsed conformation and slower observed hydration dynamics. The importance of the carboxylic acid groups on mediating hydration behavior is further illustrated by their role in dictating the solvation environment around poly(acrylic acid-stat-(poly(ethylene glycol) methyl ether acrylate) (p(AA-stat-PEGMEA)) copolymers with varying monomer ratios. These local hydration measurements establish a connection between local water dynamics and polymer chain conformation and may provide additional molecular perspective into poly(acrylic acid)'s classification as a superabsorbent polymer.
Protein aggregation, impaired degradation, and immune activation are central hallmarks of neurodegenerative diseases, yet how these processes are coordinated remains unclear. Here, we identify Immune-Protein Degradation Bodies (I-PDBs), a previously unrecognized class of BAG2-driven, phase-separated organelles that integrate protein quality control with adaptive immunity. IFNγ induce I-PDB formation at the endoplasmic reticulum (ER), where they concentrate immunoproteasome components, MHC-I peptide-loading machinery, and ER-associated chaperones. I-PDBs redirect proteostatic cargo from centrosomal aggregation pathways to spatially restricted degradation sites optimized for antigenic peptide generation, coupling selective substrate clearance to CD8+ T cell engagement. Using a cellular model of aggregation-prone tau, we show that I-PDBs capture pathological tau fibrils at ER-microtubule interfaces and process them into potentially antigenic peptides, thus reducing the load of aggregation-prone tau peptides. We term this mechanism the Proteostasis-Associated Immune Relay (PAIR), establishing I-PDBs as critical hubs linking proteostasis to immune surveillance with broad implications for disease.
Polymer formulations are essential in diverse applications including personal care products, coatings, paints, adhesives, and plastic materials. Designing these formulations requires navigating large, complex design spaces, where phase and self-assembly behavior critically impact performance. The Flory-Huggins chi parameter, which quantifies segmental miscibility, is widely used to parametrize the excess free energy of mixing in formulation models. In this work, we introduce two data-efficient, top-down methods for estimating chi parameters using the Random Phase Approximation (RPA): (i) Boundary Nonlinear Regression (Boundary-NLR), which fits theoretical spinodal boundaries to experimental phase boundaries, and (ii) Surrogate Model Inverse Parameter Estimation (SMIPE), which uses a Gaussian Process Classifier to fit sparse phase maps via a surrogate model. Both methods allow rapid parametrization of polymer field-theoretic models without the need for additional experiments. We evaluate these approaches on data sets involving polymer-solvent-nonsolvent ternary mixtures and block copolymer-solvent systems, demonstrating their robustness to experimental noise and their relevance for real-world formulation design.
The interaction between amino acids (AAs) and hydration water is fundamental to protein folding and protein-protein interactions. Here, we proposed a hydrophobicity scale for AAs based on their computed free energetic cost of dewetting. This metric captures both entropic and enthalpic contributions of AA-water interactions and allows a systematic and intuitive classification of AAs. Using indirect umbrella sampling (INDUS), we rank individual AAs based on the relative magnitude of their dewetting free energies, from lowest (most hydrophobic) to highest (most hydrophilic). This new hydrophobicity scale is a starting point to evaluate different elements of water hydration behavior, and we focus here on the water structure and translational diffusivity of the hydration waters. While the latter is commonly used as a proxy for hydrophobicity, we show that its behavior is in fact nonmonotonic: hydrophobic residues show slow water diffusion due to highly structured hydration water networks, while highly hydrophilic residues have slow water diffusion due to strong hydrogen bonds with water despite less structured hydration networks. We extend our analysis of hydration properties to intrinsically disordered peptides with varied sequence patterning (sequences of proline/leucine and arginine/glutamic acid residues). We find that the hydration behavior of these peptides is highly context-dependent, with hydrophobic (hydrophilic) patches cooperatively enhancing hydrophobicity (hydrophilicity). These molecular insights of sequence-dependent hydration behaviors may be particularly impactful for the study of intrinsically disordered proteins implicated in liquid-liquid phase separation and aggregation, processes where AAs' hydration environments are complex and changing.
Intrinsically disordered proteins (IDPs) lack a stable 3D structure under physiological conditions, making them challenging to study and simulate. In this study, we compare the hydrophobicity and water-protein interactions of amino acids in three popular all-atom molecular dynamics (MD) force fields: amber03ws (a03ws), CHARMM36m (C36m), and a99SB-disp. Using the indirect umbrella sampling (INDUS) technique, we quantify the dewetting free energies of each amino acid in the force fields. Additionally, we analyze water structuring around the amino acids using the water triplet angle distribution and measure water diffusion in the hydration shells. Our results reveal that CHARMM36m has the lowest dewetting free energies, indicating higher amino acid hydrophobicity, while a99SB-disp exhibits the highest, suggesting lower hydrophobicity. Water diffusion is significantly slower in the hydration shells of a99SB-disp due to its unique water structuring (e.g., higher frequency of tetrahedral coordination), while there is much less of a water diffusion slowdown in a03ws and CHARMM36m. We show that these differences impact the behavior of an aggregation-prone tau fragment, jR2R3 P301L, in MD simulations. We find that CHARMM36m's propensity for dimer formation is attributed to its lower dewetting free energies, whereas a99SB-disp's higher-than-expected dimerization propensity is due to favorable, entropically driven changes in water structure upon peptide association. These findings underscore the importance of accurately modeling water-protein interactions for IDPs and protein-protein interactions as well as the sensitivity of these to the underlying force field. Our study suggests that dewetting free energies and water structuring metrics, such as the water triplet angle distribution, can be valuable for future force field development and for predicting phenomena related to water-protein interactions.
Complex fluids in confined geometries are found in numerous applications, including membranes, lubricants, and microelectronics. However, current computational approaches for studying these systems have a variety of shortcomings. Particle-based simulations are limited in accessible length and time scales, while the interaction parameters in field-theoretic approaches have no direct connections to specific chemistries. Here, we extend a multiscale framework that we earlier developed for bulk systems to address these challenges in confined polymer formulations. The methodology uses atomistic molecular dynamics simulations to parameterize coarse-grained field-theoretic models of confined fluids, which subsequently enable fast equilibration and the ability to surmount length scales inaccessible to particle-based simulation methods. We first use this workflow to study a model system consisting of a confined Gaussian fluid to validate and determine best practices for the coarse-graining methodology. Next, we demonstrate this methodology by applying it to an alkyl acrylic diblock copolymer and dodecane solution confined between α-iron oxide surfaces and examining the effect of diblock concentration and length on the structure of the adsorbed film. This approach has the potential to expedite the study of complex fluids in confined environments, bridging atomistic detail and mesoscale modeling with broad implications for materials design.
A general algorithm is introduced to compute single‐chain partition functions in field‐theoretic simulations of polymers with nested tree‐like topologies, including self‐consistent field theory simulations that invoke the mean‐field approximation. The algorithm is an extension of a method used in a number of recent studies on the phase behavior of bottlebrush block copolymers. In those studies, the computational cost of computing single‐chain partition functions is reduced by aggregating the statistical weight of degenerate side arms. By extending this method to chains with arbitrary degrees of branching, the computational cost is reduced to scale with the total length of unique segments in the chain instead of the total length/mass of the entire chain. The method is first validated on a model dendrimer system by comparing results to coarse‐grained molecular dynamics simulations and also demonstrate its advantage over more conventional approaches to compute single‐chain partition functions. The algorithm is subsequently used to analyze the phase behavior of a molecularly informed field‐theoretic model of poly(butyl acrylate)‐ graft ‐poly(dodecyl acrylate) (pBA‐ graft ‐pDDA) copolymers in a dodecane solvent. The methodology can help advance field‐theoretic investigations of branched polymers by leveraging degeneracy in the chain to reduce computational cost and avoid the need to develop architecture‐specific algorithms.
Tau forms fibrillar aggregates that are pathological hallmarks of a family of neurodegenerative diseases known as tauopathies. The synthetic replication of disease-specific fibril structures is a critical gap for developing diagnostic and therapeutic tools. This study debuts a strategy of identifying a critical and minimal folding motif in fibrils characteristic of tauopathies and generating seeding-competent fibrils from the isolated tau peptides. The 19-residue jR2R3 peptide (295 to 313) which spans the R2/R3 splice junction of tau, and includes the P301L mutation, is one such peptide that forms prion-competent fibrils. This tau fragment contains the hydrophobic VQIVYK hexapeptide that is part of the core of all known pathological tau fibril structures and an intramolecular counterstrand that stabilizes the strand–loop–strand (SLS) motif observed in 4R tauopathy fibrils. This study shows that P301L exhibits a duality of effects: it lowers the barrier for the peptide to adopt aggregation-prone conformations and enhances the local structuring of water around the mutation site to facilitate site-directed pinning and dewetting around sites 300-301 to achieve in-register stacking of tau to cross β-sheets. We solved a 3 Å cryo-EM structure of jR2R3-P301L fibrils in which each protofilament layer contains two jR2R3-P301L copies, of which one adopts a SLS fold found in 4R tauopathies and the other wraps around the SLS fold to stabilize it, reminiscent of the three- and fourfold structures observed in 4R tauopathies. These jR2R3-P301L fibrils are competent to template full-length 4R tau in a prion-like manner.
Molecular insight into amyloid aggregation is crucial for understanding the details of protein fibril nucleation and growth, which play a significant role in a wide range of proteinopathies. The length and time scales for fibrillization make its computational study an intrinsically multiscale problem, necessitating the use of coarse-grained modeling. A wide variety of coarse-grained models for peptides have been proposed, often parametrized with a combination of top-down and bottom-up approaches. Here, we present a predictive, sequence-transferable bottom-up coarse-grained model, systematically developed using only information from atomistic simulations by applying an extended-ensemble relative entropy minimization technique. The resulting model is capable of accurately recovering conformational properties of peptides constructed from a reduced alphabet of amino acids, of predicting secondary structures of isolated and interacting peptides from their sequences alone, and of simulating aggregation of peptides that have been experimentally characterized as amyloidogenic. Finally, we couple such coarse-grained simulations with a genetic algorithm to characterize the sequence space of the reduced alphabet and identify features of sequences for which ordered fibrillar states are both thermodynamically favorable and kinetically accessible.
Current developments in the precise synthesis of sequence-controlled polymers allow for new opportunities in designing materials with finely tunable properties. In particular, polypeptoids offer a robust platform for sequence-specific polymers that can be produced at gram scale and offer a range of sidechain chemistries that far exceed those of polypeptides and natural protein-based biopolymers. However, the vast chemical design space of polypeptoids demands high-throughput screening, which is not yet synthetically feasible. Moreover, the lack of large structural and property databases limits the development of AI-based predictive models. These challenges highlight the need for systematic, physics-based computational methods to understand and predict how sequence impacts the polypeptoid structure and material properties. Here, we create a multiscale simulation workflow to develop bottom-up coarse-grained (CG) peptoid models using the relative entropy approach, to create a library of peptoid monomers suitable for studying the CG models of a wide range of sequences in both long-chain and multi-chain simulations. Using a representative subset of peptoid chemistries, we validate the resulting CG models by comparison with all-atom simulations and experimental end-to-end distance measurements measured through double electron-electron resonance spectroscopy. This approach is encouraging for polymer platforms that lack large databases as it offers a bottom-up framework to navigate the vast sequence and chemistry space of sequence-defined polymers, enabling molecular-level insight and in silico screening of peptoid-based materials.
A critical discovery of the past decade is that tau protein fibrils adopt disease-specific hallmark structures in each tauopathy. The faithful generation of synthetic fibrils adopting hallmark structures that can serve as targets for developing diagnostic and/or therapeutic strategies remains a grand challenge. We report on a rational design of synthetic fibrils built of a short peptide that adopts a critical structural motif in tauopathy fibrils found in Alzheimer's Disease (AD) and Chronic Traumatic Encephalopathy (CTE). They serve as minimal prions with exquisite seeding competency, in vitro and in tau biosensor cells, for recruiting tau constructs ten times larger its size en route to AD or CTE fibril structures. We demonstrate that the generation of AD and CTE-like fibril structures is dramatically catalyzed in the presence of mini-AD prions and further influenced by salt composition in solution. Double Electron-Electron Resonance studies confirmed the preservation of AD-like folds across multi-generational seeding. Fibrils formed with the full AD/CTE-like core show strong seeding competency, with their templating effect dominating over the choice of salt composition that tunes the initial selection of AD- and CTE-like fibril populations. The mini-AD prions serve as a potent catalyst with templating capabilities that offer a novel strategy to design pathological tau fibril models.
Cellulose acetate (CA), a prominent water-soluble derivative of cellulose, is a promising biodegradable ingredient that has applications in films, membranes, fibers, drug delivery, and more. In this work, we present a molecularly informed field-theoretic model for CA to explore its phase behavior in aqueous solutions. By integrating atomistic details into large-scale field-theoretic simulations via the relative entropy coarse-graining framework, our approach enables efficient calculations of CA’s miscibility window as a function of the degree of substitution (DS) of cellulose hydroxyl groups with acetate side chains. This allows us to capture the intricate phase behavior of CA, particularly its unique miscibility at intermediate substitution, without relying on experimental input. Additionally, the model directly probes CA solution behavior specific to the relative DS at C2, C3, and C6 alcohol sites, providing insights for the rational design of water-soluble CA for diverse applications. This work demonstrates a promising integration of molecularly informed field theories, complementing wet-lab experimentation, for engineering the next-generation polymeric materials with precisely tailored properties.
Solution formulations involving polymers are the basis for a wide range of products spanning consumer care, therapeutics, lubricants, adhesives, and coatings. These multicomponent systems typically show rich self-assembly and phase behavior that are sensitive to even small changes in chemistry and composition. Longstanding computational efforts have sought techniques for predictive modeling of formulation structure and thermodynamics without experimental guidance, but the challenges of addressing the long time scales and large length scales of self-assembly while maintaining chemical specificity have thwarted the emergence of general approaches. As a consequence, current formulation design remains largely Edisonian. Here, we present a multiscale modeling approach that accurately predicts, without any experimental input, the complete temperature-concentration phase diagram of model diblock polymers in solution, as established postprediction through small-angle X-ray scattering. The methodology employs a strategy whereby atomistic molecular dynamics simulations is used to parametrize coarse-grained field-theoretic models; simulations of the latter then easily surmount long equilibration time scales and enable rigorous determination of solution structures and phase behavior. This systematic and predictive approach, accelerated by access to well-defined block copolymers, has the potential to expedite in silico screening of novel formulations to significantly reduce trial-and-error experimental design and to guide selection of components and compositions across a vast range of applications.
PEO restructures water near the polymer, reducing free volume and slowing local water.
Design of next-generation membranes requires a nanoscopic understanding of the effect of biologically inspired heterogeneous surface chemistries and topologies (roughness) on local water and solute behavior. In particular, the rejection of small, neutral solutes, such as boric acid, poses a heretofore unsolved challenge. In prior work, a computational inverse design technique using an evolutionary optimization successfully uncovered new surface design strategies for optimized transport of water over solutes in smooth, model pores consisting of two surface chemistries. However, extending such an approach to more complex (and realistic) scenarios involving many surface chemistries as well as surface roughness is challenging due to the expanded design space. In this work, we develop a new approach that uses active learning to optimize in a reduced feature space of surface group interactions, finding parameters that lead to their assembly into ordered, optimal patterns. This approach rapidly identifies novel surface functionalizations that maximize the difference in water and boric acid transport through the nanopore. Moreover, we find that the roughness of the nanopore wall, independent of its chemistry, can be leveraged to enhance transport selectivity: oscillations in the pore wall diameter optimally inhibit boric acid transport by creating energetic wells from which the solute must escape to transport down the pore. This proof-of-concept demonstrates the potential for active learning strategies, in concert with molecular simulations, to rapidly navigate complex design spaces of aqueous interfaces and is promising as a tool for engineering water-mediated surface interactions for a broad range of applications.