We present a multi-objective binder design paradigm based on instruction fine-tuning and direct preference optimization (DPO) of autoregressive protein language models (pLMs). Multiple design objectives are encoded in the language model through direct optimization on expert curated preference sequence datasets comprising preferred and dispreferred distributions. We show the proposed alignment strategy enables ProtGPT2 to effectively design binders conditioned on specified receptors and a drug developability criterion. Generated binder samples demonstrate median isoelectric point (pI) improvements by 17%-60%.
We present a scalable strategy for development of mesh-free hybrid neuro-symbolic partial differential equation solvers based on existing mesh-based numerical discretization methods. Particularly, this strategy can be used to efficiently train neural network surrogate models of partial differential equations by (i) leveraging the accuracy and convergence properties of advanced numerical methods, solvers, and preconditioners, as well as (ii) better scalability to higher order PDEs by strictly limiting optimization to first order automatic differentiation. The presented neural bootstrapping method (hereby dubbed NBM) is based on evaluation of the finite discretization residuals of the PDE system obtained on implicit Cartesian cells centered on a set of random collocation points with respect to trainable parameters of the neural network. Importantly, the conservation laws and symmetries present in the bootstrapped finite discretization equations inform the neural network about solution regularities within local neighborhoods of training points. We apply NBM to the important class of elliptic problems with jump conditions across irregular interfaces in three spatial dimensions. We show the method is convergent such that model accuracy improves by increasing number of collocation points in the domain and predonditioning the residuals. We show NBM is competitive in terms of memory and training speed with other PINN-type frameworks. The algorithms presented here are implemented using \texttt{JAX} in a software package named \texttt{JAX-DIPS} (https://github.com/JAX-DIPS/JAX-DIPS), standing for differentiable interfacial PDE solver. We open sourced \texttt{JAX-DIPS} to facilitate research into use of differentiable algorithms for developing hybrid PDE solvers.
We present a highly scalable strategy for developing mesh-free neuro-symbolic partial differential equation solvers from existing numerical discretizations found in scientific computing. This strategy is unique in that it can be used to efficiently train neural network surrogate models for the solution functions and the differential operators, while retaining the accuracy and convergence properties of state-of-the-art numerical solvers. This neural bootstrapping method is based on minimizing residuals of discretized differential systems on a set of random collocation points with respect to the trainable parameters of the neural network, achieving unprecedented resolution and optimal scaling for solving physical and biological systems.
We propose a novel composite framework to find unknown fields in the context of inverse problems for partial differential equations (PDEs). We blend the high expressibility of deep neural networks as universal function estimators with the accuracy and reliability of existing numerical algorithms for partial differential equations as custom layers in semantic autoencoders. Our design brings together techniques of computational mathematics, machine learning and pattern recognition under one umbrella to incorporate domain-specific knowledge and physical constraints to discover the underlying hidden fields. The network is explicitly aware of the governing physics through a hard-coded PDE solver layer in contrast to most existing methods that incorporate the governing equations in the loss function or rely on trainable convolutional layers to discover proper discretizations from data. This subsequently focuses the computational load to only the discovery of the hidden fields and therefore is more data efficient. We call this architecture Blended inverse-PDE networks (hereby dubbed BiPDE networks) and demonstrate its applicability for recovering the variable diffusion coefficient in Poisson problems in one and two spatial dimensions, as well as the diffusion coefficient in the time-dependent and nonlinear Burgers' equation in one dimension. We also show that the learned hidden parameters are robust to added noise on input data. (C) 2021 Elsevier Inc. All rights reserved.
We present a theoretical framework to model the electric response of cell aggregates. We establish a coarse representation for each cell as a combination of membrane and cytoplasm dipole moments. Then we compute the effective conductivity of the resulting system, and thereafter derive a Fokker-Planck partial differential equation that captures the time-dependent evolution of the distribution of induced cellular polarizations in an ensemble of cells. Our model predicts that the polarization density parallel to an applied pulse follows a skewed t-distribution, while the transverse polarization density follows a symmetric t-distribution, which are in accordance with our direct numerical simulations. Furthermore, we report a reduced order model described by a coupled pair of ordinary differential equations that reproduces the average and the variance of induced dipole moments in the aggregate. We extend our proposed formulation by considering fractional order time derivatives that we find necessary to explain anomalous relaxation phenomena observed in experiments as well as our direct numerical simulations. Owing to its time-domain formulation, our framework can be easily used to consider nonlinear membrane effects or intercellular couplings that arise in several scientific, medical and technological applications.
The influence of galaxy cluster environment on the kinematics of the stripped globular clusters Globular clusters (GCs) contain relic information about the star formation history of their host galaxies. However, the development of a theoretical model for formation, evolution and disruption of GCs in a cosmological context is still at its infancy. Lack of a theoretical model to relate kinematics and distribution of the intracluster GCs to the underlying environmental processes in the history of the galaxy cluster has become a pressing issue in observational astrophysics in recent years. This is particularly due to the observation of tens of thousands of GCs in the intracluster environment of the Virgo as well as Coma galaxy clusters. In the postprocessing of IllustrisTNG, the most recent cosmological and magneto-hydrodynamical simulation, we implemented theoretical models of formation and disruption of the intracluster GCs generated in dwarf elliptical galaxies (dEs), which are the most abundant galaxies in cluster environments. By implementing dark matter particle-tagging scheme, we obtained the kinematics and distribution of stripped GCs around dEs at z = 0. Then, we retained the origin imprinted in these GCs’ phase-space properties (positions and velocities) which could be used to design observational strategies to identify the full population of native GCs born in dE galaxies.
In this chapter, following the previous one, we briefly present the modern approach to real-space renormalization group (RG) theory based on tensor network formulations which was developed during the last two decades. The aim of this sequel is to suggest a novel framework based on tensor networks in order to find the fixed points of complex systems via coarse-graining. The main result of RG is that it provides a systematic way to study the collective dynamics of a large ensemble of elements that interact according to a complex underlying network topology. RG explicitly seeks the fixed points of the complex system in the space of interactions and unravels the universality class of the complex system as well as calculates a plethora of important observables. We hope that tensor networks can particularly pave the way for better understanding of the sustainable interdependent networks (Amini et al., Sustainable interdependent networks: from theory to application, 2018) through proposing efficient computational strategies and discovering insightful features of the network behaviors.
We introduce a numerical framework that enables unprecedented direct numerical studies of the electropermeabilization effects of a cell aggregate at the meso-scale. Our simulations qualitatively replicate the shadowing effect observed in experiments and reproduce the time evolution of the impedance of the cell sample in agreement with the trends observed in experiments. This approach sets the scene for performing homogenization studies for understanding the effect of tissue environment on the efficiency of electropermeabilization. We employ a forest of Octree grids along with a Voronoi mesh in a parallel environment that exhibits excellent scalability. We exploit the electric interactions between the cells through a nonlinear phenomenological model that is generalized to account for the permeability of the cell membranes. We use the Voronoi Interface Method (VIM) to accurately capture the sharp jump in the electric potential on the cell boundaries. The case study simulation covers a volume of (1mm)3 with more than 27,000 well-resolved cells with a heterogeneous mix of morphologies that are randomly distributed throughout a spheroid region.
We introduce an approach for simulating epitaxial growth by use of an island dynamics model on a forest of quadtree grids, and in a parallel environment. To this end, we use a parallel framework introduced in the context of the level-set method. This framework utilizes: discretizations that achieve a second-order accurate level-set method on non-graded adaptive Cartesian grids for solving the associated free boundary value problem for surface diffusion; and an established library for the partitioning of the grid. We consider the cases with: irreversible aggregation, which amounts to applying Dirichlet boundary conditions at the island boundary; and an asymmetric (Ehrlich–Schwoebel) energy barrier for attachment/detachment of atoms at the island boundary, which entails the use of a Robin boundary condition. We provide the scaling analyses performed on the Stampede supercomputer and numerical examples that illustrate the capability of our methodology to efficiently simulate different aspects of epitaxial growth. The combination of adaptivity and parallelism in our approach enables simulations that are several orders of magnitude faster than those reported in the recent literature and, thus, provides a viable framework for the systematic study of mound formation on crystal surfaces.
Complex networks are composed of nodes (entities) and edges (connections) with any arbitrary topology. There may also exist multiple types of interactions among these nodes and each node may admit different states in each of its interactions with its neighbors. Understanding complex networks dwells on understanding their structure and function. However, current representations model nodes as single-state entities that are connected to each other differently and treat their dynamics separately with some differential equations. Alternatively, a unified framework might be accessible using the tensor network representation that is already utilized in physics communities. In a sequel of chapters we introduce tensor network representation and renormalization as an alternative framework to explore the universal behaviors of complex systems. We hope that tensor networks can particularly pave the way for better understanding of the sustainable interdependent networks (Amini et al., Sustainable interdependent networks: from theory to application, 2018) through proposing efficient computational strategies and discovering insightful features of the network behaviors.
Galaxy clusters contain a large population of low mass dwarf elliptical galaxies whose exact origin is unclear: their colors, structural properties and kinematics differ substantially from those of dwarf irregulars in the field. We use the Illustris cosmological simulation to study differences in the assembly paths of dwarf galaxies (3e8 < M_*/M_sun < 1e10) according to their environment. We find that cluster dwarfs achieve their maximum total and stellar mass on average ~ 8 and ~ 4.5 Gyr ago (or redshifts z = 1.0 and z = 0.4, respectively), around the time of infall into the clusters. In contrast, field dwarfs not subjected to environmental stripping, reach their maximum mass at redshift z = 0. This different assembly history naturally produces a color bimodality, with blue isolated dwarfs and redder cluster dwarfs exhibiting negligible star-formation today. The cessation of star formation happens over median times 3.5-5 Gyr depending on stellar mass, and shows a large scatter (~ 1-8 Gyr), with the lower values associated with starburst events that occur at infall through the virial radius or pericentric passages. We argue that such starbursts together with the early assembly of cluster dwarfs can provide a natural explanation for the higher specific frequency of globular clusters (GCs) in cluster dwarfs, as found observationally. We present a simple model for the formation and stripping of GCs that supports this interpretation. The origin of dwarf ellipticals in clusters is, therefore, consistent with an environmentally-driven evolution of field dwarf irregulars. However the z = 0 field analogs of cluster dwarf progenitors have today stellar masses a factor ~ 3 larger --a difference arising from the early truncation of star formation in cluster dwarfs.