Condensates formed by oppositely charged intrinsically disordered proteins provide model systems for understanding how transient electrostatic interactions govern structure and dynamics in biomolecular assemblies. Here we investigate a nearly charge-neutral condensate composed of 50 Prothymosin alpha (ProTalpha) and 40 Histone H1 molecules using a single-bead-per-residue coarse-grained model combining the HPS hydropathy model for disordered regions with a Go model for the globular domain of Histone H1 under NPT conditions at pressures from 2 to 12 bar. We find that chain dimensions, including the radius of gyration (Rg), end-to-end distance (Ree), and their ratio R, are insensitive to pressure, indicating that chain conformations remain largely unchanged over the pressure range studied. Histone H1 exhibits systematically larger values of R than ProTalpha because of its globular-core plus disordered-tail architecture. Translational diffusion coefficients decrease monotonically with pressure, from approximately 0.22 to 0.06 nm^2/ns, with substantial chain-to-chain heterogeneity comparable to the mean diffusion coefficient. Chain relaxation follows a stretched exponential with beta less than 1 that decreases with pressure. ProTalpha relaxation times of approximately 12 to 40 ns obey Rouse scaling, whereas Histone H1 deviates because of the internal constraint imposed by its globular domain. ProTalpha-Histone H1 contact lifetimes of approximately 0.43 to 0.56 ns are much shorter than the Rouse relaxation time, placing the system firmly in the fast-exchange regime where transient electrostatic contacts renormalize chain friction rather than acting as permanent cross-links, consistent with the moderate stretching exponent beta of approximately 0.55 to 0.70 observed across all pressures.
Recent advances in machine learning (ML) are accelerating scientific discovery across various disciplines of life sciences. In proteomics, ML models show great promise for studying intrinsically disordered proteins (IDPs)–highly dynamic proteins that lack stable three-dimensional structures, yet play key roles in many biological functions. These models are useful for fast predictions of IDP properties and substantially reducing simulation costs. Nonetheless, ML models often generalize poorly when the training and test distributions differ. In this work, we address this by designing domain-adaptation (DA)–based ML models via adversarial learning for robust IDP conformational property prediction under realistic distribution shifts, e.g., differences in sequence length and net charge. Our approach enables training on IDP sequences that are inexpensive to simulate while transferring reliably to harder regimes, thereby cutting simulation costs. For example, models trained on short-sequence simulations accurately predict properties of long sequences. Experiment results across multiple IDP datasets show substantial gains in generalization to unseen domains, reducing prediction errors by 40–80% in several scenarios.
We use a combination of Brownian dynamics (BD) simulation results and deep learning (DL) strategies for the rapid identification of large structural changes caused by missense mutations in intrinsically disordered proteins (IDPs). We used ∼6500 IDP sequences from MobiDB database of length 20-300 to obtain gyration radii from BD simulation on a coarse-grained single-bead amino acid model (HPS2 model) used by us and others [Dignon, G. L. PLoS Comput. Biol. 2018, 14, e1005941,Tesei, G. Proc. Natl. Acad. Sci. U.S.A. 2021, 118, e2111696118,Seth, S. J. Chem. Phys. 2024, 160, 014902] to generate the training sets for the DL algorithm. Using the gyration radii ⟨Rg⟩ of the simulated IDPs as the training set, we develop a multilayer perceptron neural net (NN) architecture that predicts the gyration radii of 33 IDPs previously studied by using BD simulation with 97% accuracy from the sequence and the corresponding parameters from the HPS model. We now utilize this NN to predict gyration radii of every permutation of missense mutations in IDPs. Our approach successfully identifies mutation-prone regions that induce significant alterations in the radius of gyration when compared to the wild-type IDP sequence. We further validate the prediction by running BD simulations on the subset of identified mutants. The neural network yields a (104-106)-fold faster computation in the search space for potentially harmful mutations. Our findings have substantial implications for rapid identification and understanding of diseases related to missense mutations in IDPs and for the development of potential therapeutic interventions. The method can be extended to accurate predictions of other mutation effects in disordered proteins.
We introduce a novel hybrid machine learning (ML) framework to predict the radius of gyration and other conformational properties of intrinsically disordered proteins (IDPs). Our model integrates sequence information with physical features derived from a coarse-grained model validated by experimental data. Specifically, we combine hidden states from sequence-based models with 23 physical features projected into a shared latent space, and apply an attention mechanism that assigns weights to each residue to highlight the most informative regions of the sequence. This attention-guided fusion significantly improves predictive accuracy across multiple metrics, including mean absolute percentage error and mean squared error, while also enhancing confidence in the predictions. We trained and evaluated our models on Brownian dynamics (BD) simulation results for approximately 7000 IDPs from the MobiDB database (each with >99% disorder score). We find that sequence-based models consistently outperform feature-only models, with the GRU achieving the best performance among sequence-only approaches. Moreover, combining sequence and feature information further improves accuracy across all architectures, with the hybrid biGRU model delivering the best overall predictive performance. SHAP analysis reveals the relative importance of physical features, offering model explainability, and guiding feature selection. Notably, using a small number of top features often reduces model complexity and improves generalization. Furthermore, an integrated gradient analysis reveals that in addition to the length of the IDPs, the three parameters (sequence charge and hydropathy decoration parameters (SCD and SHD), and charge asymmetry parameter f*) play a key role in the predictions of ML. Our framework provides a fast, interpretable, and scalable tool for predicting IDP behavior, enabling efficient initial screening prior to costly molecular simulations.
We report simulation studies of 33 single intrinsically disordered proteins (IDPs) using coarse-grained bead-spring models where interactions among different amino acids are introduced through a hydropathy matrix and additional screened Coulomb interaction for the charged amino acid beads. Our simulation studies of two different hydropathy scales (HPS1, HPS2) [Dignon et al., PLoS Comput. Biol. 14, e1005941 (2018); Tesei et al. Proc. Natl. Acad. Sci. U. S. A. 118, e2111696118 (2021)] and the comparison with the existing experimental data indicate an optimal interaction parameter ϵ = 0.1 and 0.2 kcal/mol for the HPS1 and HPS2 hydropathy scales. We use these best-fit parameters to investigate both the universal aspects as well as the fine structures of the individual IDPs by introducing additional characteristics. (i) First, we investigate the polymer-specific scaling relations of the IDPs in comparison to the universal scaling relations [Bair et al., J. Chem. Phys. 158, 204902 (2023)] for the homopolymers. By studying the scaled end-to-end distances ⟨RN2⟩/(2Lℓp) and the scaled transverse fluctuations l̃⊥2=⟨l⊥2⟩/L, we demonstrate that IDPs are broadly characterized with a Flory exponent of ν ≃ 0.56 with the conclusion that conformations of the IDPs interpolate between Gaussian and self-avoiding random walk chains. Then, we introduce (ii) Wilson charge index (W) that captures the essential features of charge interactions and distribution in the sequence space and (iii) a skewness index (S) that captures the finer shape variation of the gyration radii distributions as a function of the net charge per residue and charge asymmetry parameter. Finally, our study of the (iv) variation of ⟨Rg⟩ as a function of salt concentration provides another important metric to bring out finer characteristics of the IDPs, which may carry relevant information for the origin of life.
We report a novel method based on the current blockade (CB) characteristics obtained from a dual nanopore device that can determine DNA barcodes with near-perfect accuracy using a Brownian dynamics simulation strategy. The method supersedes our previously reported velocity correction algorithm (S. Seth and A. Bhattacharya, RSC Advances, 11:20781-20787, 2021), taking advantage of the better measurement of the time-of-flight (TOF) protocol offered by the dual nanopore setup. We demonstrate the efficacy of the method by comparing our simulation data from a coarse-grained model of a polymer chain consisting of 2048 excluded volume beads of diameter 𝜎 = 24 bp using with those obtained from experimental CB data from a 48,500 bp λ-phage DNA, providing a 48500 2400 ≅ 24 base pair resolution in simulation. The simulation time scale is compared to the experimental time scale by matching the simulated time-of-flight (TOF) velocity distributions with those obtained experimentally (Rand et al., ACS Nano, 16:5258-5273, 2022). We then use the evolving coordinates of the dsDNA and the molecular features to reconstruct the current blockade characteristics on the fly using a volumetric model based on the effective van der Waal radii of the species inside and in the immediate vicinity of the pore. Our BD simulation mimics the control-zoom-in-logic to understand the origin of the TOF distributions due to the relaxation of the out-of-equilibrium conformations followed by a reversal of the electric fields. The simulation algorithm is quite general and can be applied to differentiate DNA barcodes from different species.
We study the universal aspects of polymer conformations and transverse fluctuations for a single swollen chain characterized by a contour length L and a persistence length ℓp in two dimensions (2D) and three dimensions (3D) in the bulk, as well as in the presence of excluded volume (EV) particles of different sizes occupying different area/volume fractions. In the absence of the EV particles, we extend the previously established universal scaling relations in 2D [Huang et al., J. Chem. 140, 214902 (2014)] to include 3D and demonstrate that the scaled end-to-end distance ⟨RN2⟩/(2Lℓp) and the scaled transverse fluctuation ⟨l⊥2⟩/L as a function of L/ℓp collapse onto the same master curve, where ⟨RN2⟩ and ⟨l⊥2⟩ are the mean-square end-to-end distance and transverse fluctuations. However, unlike in 2D, where the Gaussian regime is absent due to the extreme dominance of the EV interaction, we find that the Gaussian regime is present, albeit very narrow in 3D. The scaled transverse fluctuation in the limit L/ℓp ≪ 1 is independent of the physical dimension and scales as ⟨l⊥2⟩/L∼(L/ℓp)ζ-1, where ζ = 1.5 is the roughening exponent. For L/ℓp ≫ 1, the scaled fluctuation scales as ⟨l⊥2⟩/L∼(L/ℓp)ν-1, where ν is the Flory exponent for the corresponding spatial dimension (ν2D = 0.75 and ν3D = 0.58). When EV particles of different sizes for different area or volume fractions are added into 2D and 3D systems, our results indicate that the crowding density either does not or does only weakly affect the universal scaling relations. We discuss the implications of these results in living matter by showing the experimental result for a dsDNA on the master plot.
We report current blockade (CB) characteristics of molecular motifs residing on a model dsDNA using electrokinetic Brownian dynamics (EKBD) and study the role of the valence of the counterions as the dsDNA translocates through a solitary nanopore (NP) driven by an electric field. We explicitly incorporate all the charges on the DNA backbone, co- and counter-ions, and investigate CB characteristics of two charged sidechain motifs exactly. Our simulation brings out the details of binding and unbinding of the counter-ions and the time dependent counter-ion condensation on the translocating DNA for mono- and di-valent salt conditions. An important and less intuitive finding is that the drop in the conventional (positive) current through the pore is due to the condensation of the counter-ions on the translocating DNA and not so much due to drop in the co-ions passing through the pore. This finding aligns with previous studies conducted by Tanaka et al. [Phys. Rev. Lett. 94, 148103 (2005)], Cui [J. Phys. Chem B 114, 2015 (2010)], and Holm et al. [Phys. Rev. Lett. 112, 018101 (2014)]. We further find that this condensation is larger for the divalent ions leading to a slowing down of the translocation speed and yielding a longer dwell time for the motifs. Finally, we use the exact CB characteristics from the EKBD simulation to reconstruct the same CB characteristics using a volumetric ansatz on the segment inside and in the vicinity of the pore using on the ordinary BD model without the explicit presence of co- and counter-ions. Refinement of this ansatz will allow us to obtain the CB characteristics for longer genome fragments using low-cost ordinary BD simulation.
We study DNA translocation through a double nanopore system subject to a net bias using Brownian dynamics simulation on a model system. We consider the limit d LR < < L, where d LR is the distance between the pores and L = Nσ is the contour length of the chain consisting of N monomers of diameter σ . In this limit, we generalize a scaling ansatz for the mean first passage time, originally proposed for the driven translocation through a single nanopore, for the double nanopore system and demonstrate its validity using simulation data. The simulation data enables us to extract the pore friction as a function of the chain stiffness. The method can be used to determine the mean first passage time 〈 τ 〉 for longer chains difficult to extract from BD simulation.
Recent experiments demonstrated that knots in single molecule dsDNA can be formed by compression in a nanochannel. In this manuscript, we further elucidate the underlying molecular mechanisms by carrying out a compression experiment in silico, where an equilibrated coarse-grained double-stranded DNA confined in a square channel is pushed by a piston. The probability of forming knots is a non-monotonic function of the persistence length and can be enhanced significantly by increasing the piston speed. Under compression knots are abundant and delocalized due to a backfolding mechanism from which chain-spanning loops emerge, while knots are less frequent and only weakly localized in equilibrium. Our in silico study thus provides insights into the formation, origin and control of DNA knots in nanopores.
We report Brownian dynamics simulation results with the specific goal to identify key parameters controlling the experimentally measurable characteristics of protein tags on a dsDNA construct translocating through a double nanopore setup. First, we validate the simulation scheme in silico by reproducing and explaining the physical origin of the asymmetric experimental dwell time distributions of the oligonucleotide flap markers on a 48 kbp long dsDNA at the left and the right pore. We study the effect of the electric field inside and beyond the pores, critical to discriminate the protein tags based on their effective charges and masses revealed through a generic power-law dependence of the average dwell time at each pore. The simulation protocols monitor piecewise dynamics at a sub-nanometer length scale and explain the disparate velocity using the concepts of nonequilibrium tension propagation theory. We further justify the model and the chosen simulation parameters by calculating the Péclet number which is in close agreement with the experiment. We demonstrate that our carefully chosen simulation strategies can serve as a powerful tool to discriminate different types of neutral and charged tags of different origins on a dsDNA construct in terms of their physical characteristics and can provide insights to increase both the efficiency and accuracy of an experimental dual-nanopore setup.
DNA capture with high fidelity is an essential part of nanopore translocation. We report several important aspects of the capture process and subsequent translocation of a model DNA polymer through a solid-state nanopore in the presence of an extended electric field using the Brownian dynamics simulation that enables us to record statistics of the conformations at every stage of the translocation process. By releasing the equilibrated DNAs from different equipotentials, we observe that the capture time distribution depends on the initial starting point and follows a Poisson process. The field gradient elongates the DNA on its way toward the nanopore and favors a successful translocation even after multiple failed threading attempts. Even in the limit of an extremely narrow pore, a fully flexible chain has a finite probability of hairpin-loop capture, while this probability decreases for a stiffer chain and promotes single file translocation. Our in silico studies identify and differentiate characteristic distributions of the mean first passage time due to single file translocation from those due to translocation of different types of folds and provide direct evidence of the interpretation of the experimentally observed folds [M. Gershow and J. A. Golovchenko, Nat. Nanotechnol. 2, 775 (2007) and Mihovilovic et al., Phys. Rev. Lett. 110, 028102 (2013)] in a solitary nanopore. Finally, we show a new finding-that a charged tag attached at the 5' end of the DNA enhances both the multi-scan rate and the uni-directional translocation (5' → 3') probability that would benefit the genomic barcoding and sequencing experiments.
We report an accurate method to determine DNA barcodes from the dwell time measurement of protein tags (barcodes) along the DNA backbone using Brownian dynamics simulation of a model DNA and use a recursive theoretical scheme which improves the measurements to almost 100% accuracy. The heavier protein tags along the DNA backbone introduce a large speed variation in the chain that can be understood using the idea of non-equilibrium tension propagation theory. However, from an initial rough characterization of velocities into “fast” (nucleotides) and “slow” (protein tags) domains, we introduce a physically motivated interpolation scheme that enables us to determine the barcode velocities rather accurately. Our theoretical analysis of the motion of the DNA through a cylindrical nanopore opens up the possibility of its experimental realization and carries over to multi-nanopore devices used for barcoding.
We report an accurate method to determine DNA barcodes from the dwell time measurement of protein tags (barcodes) along the DNA backbone using Brownian dynamics simulation of a model DNA and use a recursive theoretical scheme which improves the measurements to almost 100% accuracy. The heavier protein tags along the DNA backbone introduce a large speed variation in the chain that can be understood using the idea of non-equilibrium tension propagation theory. However, from an initial rough characterization of velocities into "fast" (nucleotides) and "slow" (protein tags) domains, we introduce a physically motivated interpolation scheme that enables us to determine the barcode velocities rather accurately. Our theoretical analysis of the motion of the DNA through a cylindrical nanopore opens up the possibility of its experimental realization and carries over to multi-nanopore devices used for barcoding.
DNA capture with high fidelity is an essential part of nanopore translocation. We report several important aspects of the capture process and subsequent translocation of a model DNA polymer through a solid-state nanopore in presence of an extended electric field using the Brownian dynamics simulation that enables us to record statistics of the conformations at every stage of the translocation process. By releasing the equilibrated DNAs from different equipotentials, we observe that the capture time distribution depends on the initial starting point and follows a Poisson process. The field gradient elongates the DNA on its way towards the nanopore and favors a successful translocation even after multiple failed threading attempts. Even in the limit of an extremely narrow pore, a fully flexible chain has a finite probability of hairpin-loop capture while this probability decreases for a stiffer chain and promotes single file translocation. Our studies for the first time identify and differentiate characteristic distributions of the mean first passage time due to single file translocation from those due to translocation of different types of folds and provide direct evidences of the interpretation of the experimentally observed folds [M. Gershow et al., Nat. Nanotech. 2, 775 (2007) & M. Mihovilovic et al. Phys. Rev. Letts. 110, 028102 (2013)] in solitary nanopores.
We report Brownian dynamics (BD) simulation results for a coarse-grained (CG) model semi-flexible polymer threading through two nanopores. Particularly we study a “tug-of-war” situation where equal and opposite forces are applied on each pore to avoid folds for the polymer segment in between the pores. We calculate mean first passage times (MFPT) through the left and the right pores and show how the MFPT decays as a function of the off-set voltage between the pores. We present results for several bias voltages and chain stiffness. Our BD simulation results validate recent experimental results and offer avenues to further explore various aspects of multi-pore translocation problem using BD simulation strategies which we believe will provide insights to design new experiments.
The potential of a double nanopore system to determine DNA barcodes has been demonstrated experimentally. By carrying out Brownian dynamics simulation on a coarse-grained model DNA with protein tag (barcodes) at known locations along the chain backbone, we demonstrate that due to large variation of velocities of the chain segments between the tags, it is inevitable to under/overestimate the genetic lengths from the experimental current blockade and time of flight data. We demonstrate that it is the tension propagation along the chain's backbone that governs the motion of the entire chain and is the key element to explain the non uniformity and disparate velocities of the tags and DNA monomers under translocation that introduce errors in measurement of the length segments between protein tags. Using simulation data we further demonstrate that it is important to consider the dynamics of the entire chain and suggest methods to accurately decipher barcodes. We introduce and validate an interpolation scheme using simulation data for a broad distribution of tag separations and suggest how to implement the scheme experimentally.