
This study introduces a novel approach for the unrestricted de novo design of transition metal catalysts, leveraging the power of genetic algorithms (GAs) and density functional theory (DFT) calculations. By focusing on the Suzuki reaction, known for its significance in forming carbon-carbon bonds, we demonstrate the effectiveness of fragment-based and graph-based genetic algorithms in identifying novel ligands for palladium-based catalytic systems. Our research highlights the capability of these algorithms to generate ligands with desired thermodynamic properties, moving beyond the restriction of enumerated chemical libraries. Limitations in the applicability of machine learning models are overcome by calculating thermodynamic properties from first principle. The inclusion of synthetic accessibility scores further refines the search, steering it towards more practically feasible ligands. Through the examination of both palladium and alternative transition metal catalysts like copper and silver, our findings reveal the algorithms' ability to uncover unique catalyst structures within the target energy range, offering insights into the electronic and steric effects necessary for effective catalysis. This work not only proves the potential of genetic algorithms in the cost-effective and scalable discovery of new catalysts but also sets the stage for future exploration beyond predefined chemical spaces, enhancing the toolkit available for catalyst design.
Linear correlation coefficients were calculated between the reported Young’s modulus (YM) values and non-covalent interactions within cellulose-oligolignol complexes, considering the composition of an efficient adhesive formulation previously reported. A paradigmatic relationship was observed. Molecular complexes of oligolignols with cellulose Iβ were modeled using hybrid quantum mechanics/molecular mechanics (QM/MM) computations to obtain wavefunctions at the interaction region. Subsequently, a study of non-covalent interactions (NCI) based on the atoms in molecules (AIM) theory was implemented, utilizing graphics processing units (GPUs) for calculations. Our findings indicate that non-covalent interactions control the forces associated to adhesive-cellulose contacts, primarily through X-H···O hydrogen bonds, which promote the adhesion of oligolignols on cellulose Iβ. Results indicate that the adhesion strength projected from larger YM values cannot be described solely by the number of stronger hydrogen bonds nor by the number of the weak interactions but by the entire contributions of specific interactions. Thus, significant linear correlations were observed between reported values of Young’s modulus and the molecular interactions observed, rendering the influence of oligolignol structure on the adhesion phenomenon in our cellulose Iβ crystallite model. These observations promote the NCI and AIM analysis in a new framework to design adhesive formulations.
The equilibrium stability of a protein is determined by its amino acid sequence and the solution conditions, such as temperature, pH and presence of chemical denaturant. The stability of a single protein in two identical solutions can nonetheless differ if other macromolecules, termed cosolutes or crowders, are present in one of the solutions at concentrations high enough to occupy a substantial fraction of the solution volume. This effect, due to the presence of the crowders, decreases or increases the stability depending on the interactions between the protein and crowders. Hard-core steric repulsions, which are responsible for the reduction in free volume, are expected to entropically stabilize the protein while attractive interactions can be destabilizing. Here we use a coarse-grained protein model to assess the impact of different types of crowder-protein interactions on the stability of a 35-amino acid model sequence folding into a helical bundle. We find that, for the same interaction strength and concentration, spherical crowders with a hydrophobic character are more destabilizing than crowders interacting nonspecifically with the protein. However, the two types of interactions differ in the degree of association between crowders and protein. At an interaction strength for which the attractive interactions roughly counteracts the stabilizing hard-core repulsions, the nonspecific interactions lead to much stronger crowder-protein association than the hydrophobic interactions. Additionally, we study crowders in the form of polypeptide chains, which are capable of hydrogen bonding with the protein. These peptide crowders have a destabilizing effect even at relatively low crowder concentrations, especially if the sequence of the peptide crowders includes hydrophobic amino acids. Our findings emphasize the importance of the interplay between different types of attractive crowder-protein interactions and entropic effects in determining the net effect on protein stability.
Synthetic biology aims to engineer biological circuits, which often involve gene expression. A particularly promising group of regulatory elements are riboswitches because of their versatility with respect to their targets, but early synthetic designs were not as attractive because of a reduced dynamic range with respect to protein regulators. Only recently, the creation of toehold switches helped overcome this obstacle by also providing an unprecedented degree of orthogonality. However, a lack of automated design and optimization tools prevents the widespread and effective use of toehold switches in high throughput experiments. To address this, we developed Toeholder, a comprehensive open-source software for toehold design and in silico comparison. Toeholder takes into consideration sequence constraints from experimentally tested switches, as well as data derived from molecular dynamics simulations of a toehold switch. We describe the software and its in silico validation results, as well as its potential applications and impacts on the management and design of toehold switches.
Ignition delay times (IDT) for stoichiometric propane (C 3 H 8 ) diluted with nitrogen were measured in a shock tube facility under reflected shock wave conditions at pressures ranging from 1 to 10 atm and temperatures between 850 and 1500 K. The experiments were limited to a maximum pressure of 10 atm due to the facility’s constraints. In addition, numerical simulations were conducted using several detailed kinetic mechanisms at pressures from 1 to 30 atm and three equivalence ratios ( φ = 0.5, 1, and 2) to provide comparative insights. The results indicated that IDT decreases as pressure increases, with a more significant reduction observed between 1 and 10 atm compared to 10 to 30 atm. While most models exhibited similar trends and minimal discrepancies, the GRI Mech 3.0 mechanism demonstrated a slower prediction of ignition delay times at temperatures below 1250 K. In contrast, the POLIMI model exhibited a relatively faster prediction at temperatures above 1250 K, with the deviation between the two models becoming more pronounced as pressure increased. A comparative analysis revealed that the experimental predictions of propane autoignition behavior were in good agreement with the results obtained using the ARAMCO 3.0 mechanism. To further understand the chemistry governing the autoignition process of C 3 H 8 , a sensitivity analysis was performed for a stoichiometric mixture at three distinct temperatures (850 K, 1200 K, and 1550 K).
A high priority of the World Health Organization (WHO) is the study of drugs against Pseudomonas aeruginosa , which has developed antibiotic resistance. In this order, recent research is analyzing biomaterials and metal oxide nanoparticles, such as chitosan (QT) and TiO 2 (NT), which can transport molecules with biological activity against bacteria, to propose them as drug carrier candidates. In the present work, 10 modified benzofuran-isatin molecules were studied through computational simulation using density functional theory (DFT) and molecular docking assays against Hfq and LpxC (proteins of P. aeruginosa ). The results show that the ligand efficiency of commercial drugs C-CP and C-AZI against Hfq is low compared with the best-designed molecule MOL-A. However, we highlight that the influence of NT promotes a better interaction of some molecules, where MOL-E generates a better interaction by 0.219 kcal/mol when NT is introduced in Hfq, forming the system Hfq-NT (Target-NT). Similar behavior is observed in the LpxC target, in which MOL-J is better at 0.072 kcal/mol. Finally, two pharmacophoric models for Hfq and LpxC implicate hydrophobic and aromatic-hydrophobic fragments.
This study leverages a graph-based genetic algorithm (GB-GA) for the design of efficient nitrogen-fixing catalysts as alternatives to the Schrock catalyst, with the aim to improve the energetics of key reaction steps. Despite the abundance of nitrogen in the atmosphere, it remains largely inaccessible due to its inert nature. The Schrock catalyst, a molybdenum-based complex, offered a breakthrough but its practical application is limited due to low turnover numbers and energetic bottlenecks. The genetic algorithm in our study explores the chemical space for viable modifications of the Schrock catalyst, evaluating each modified catalyst’s fitness based on reaction energies of key catalytic steps and synthetic accessibility. Through a series of selection and optimization processes, we obtained fully converged catalytic cycles for 20 molecules at the B3LYP level of theory. From these results, we identified three promising molecules, each demonstrating unique advantages in different aspects of the catalytic cycle. This study offers valuable insights into the potential of generative models for catalyst design. Our results can help guide future work on catalyst discovery for the challenging nitrogen fixation process.
Lead (Pb) is a pervasive contaminant and poses a serious threat to living beings. The present study aims at batch and fixed bed column scale potential of commercial compost (CCB) and peanut shells biosorbents (PSB) for the sequestration of Pb from contaminated aqueous systems. The PSB and CCB were characterized with FTIR, SEM and Brunauer Emmett-Teller (BET) to get insight of the adsorption behavior of both materials. Fixed bed column scale experiments were performed at steady state flow (2.5 and 5.0 mL/min), initial Pb concentrations (25 and 50 mg/L) and dosage of each adsorbent (3.0 and 6.0 g/column). Columns packed (15.9 cm2) with PSB and CCB have revealed excellent adsorption of Pb with PSB as compared with CCB. The total volume of injected contaminated water was 1,500 mL and 3,000 mL at 2.5 and 5.0 mL/min, respectively while total bed volume number was 157. A series of batch experiments with CCB and PSB was conducted at adsorbent dosage (1.25–5.0 g/L), initial Pb level (25–100 mg/L), interaction time (0–180 min) and solution pH (4–10) at room temperature. Batch scale results revealed that PSB removed 92% Pb from water at 25 mg Pb/L concentration as compared with CCB (79%). The presence of competing ions in groundwater showed less Pb removal as compared with synthetic water. The experimental data were simulated with equilibrium isothermal models: Langmuir, Freundlich, and kinetic models: pseudo first order, pseudo second order and intra-particle diffusion. The Freundlich and pseudo second order models better described the equilibrium and kinetic experimental data, respectively with maximum sorption of 42.5 mg/g by PSB which is also evident from FTIR functional groups and SEM results. While equilibrium sorption of Pb onto CCB was equally explained by Freundlich and Langmuir models. These findings indicate that PSB could be an active and ecofriendly biosorbent for the sequestration of metals from contaminated aqueous systems.
Carnosine (CAR) and anserine (ANS) are histidine-containing dipeptides that show the therapeutic properties and protective abilities against diabetes and cognitive deficit. Both dipeptides are rich in meat products and have been used as a supplement. However, in humans, both compounds have a short half-life due to the rapid degradation by dizinc carnosinase 1 (CN1) which is a hurdle for its therapeutic application. To date, a comparative study of carnosine- and anserine-CN1 complexes is limited. Thus, in this work, molecular dynamics (MD) simulations were performed to explore the binding of carnosine and anserine to CN1. CN1 comprises 2 chains (Chains A and B). Both monomers are found to work independently and alternatingly. The displacement of Zn2+ pair is found to disrupt the substrate binding. CN1 employs residues from the neighbour chain (H235, T335, and T337) to form the active site. This highlights the importance of a dimer for enzymatic activity. Anserine is more resistant to CN 1 than carnosine because of its bulky and dehydrated imidazole moiety. Although both dipeptides can direct the peptide oxygen to the active Zn2+ which can facilitate the catalytic reaction, the bulky methylated imidazole on anserine promotes various poses that can retard the hydrolytic activity in contrast to carnosine. Anserine is likely to be the temporary competitive inhibitor by retarding the carnosine catabolism.
Photon capture by chlorophylls and other chromophores in light-harvesting complexes and photosystems is the driving force behind the light reactions of photosynthesis. Excitation of photosystem II allows it to receive electrons from the water-oxidizing oxygen-evolution complex and to transfer them to an electron-transport chain that generates a transmembrane electrochemical gradient and ultimately reduces plastocyanin, which donates its electron to photosystem I. Subsequently, excitation of photosystem I leads to electron transfer to a ferredoxin which can either reduce plastocyanin again (in so-called “cyclical electron-flow”) and release energy for the maintenance of the electrochemical gradient, or reduce NADP+ to NADPH. Although photons in the far-red (700–750 nm) portion of the solar spectrum carry enough energy to enable the functioning of the photosynthetic electron-transfer chain, most extant photosystems cannot usually take advantage of them due to only absorbing light with shorter wavelengths. In this work, we used computational methods to characterize the spectral and redox properties of 49 chlorophyll derivatives, with the aim of finding suitable candidates for incorporation into synthetic organisms with increased ability to use far-red photons. The data offer a simple and elegant explanation for the evolutionary selection of chlorophylls a, b, c, and d among all easily-synthesized singly-substituted chlorophylls, and identified one novel candidate (2,12-diformyl chlorophyll a) with an absorption peak shifted 79 nm into the far-red (relative to chlorophyll a) with redox characteristics fully suitable to its possible incorporation into photosystem I (though not photosystem II). chlorophyll d is shown by our data to be the most suitable candidate for incorporation into far-red utilizing photosystem II, and several candidates were found with red-shifted Soret bands that allow the capture of larger amounts of blue and green light by light harvesting complexes.
Measurements of nonlinear optical (NLO) properties of different binary mixtures having carbon disulfide (CS2) as the common component, namely CS2-acetone, CS2-cyclopentanone, CS2-toluene, and CS2-carbon tetrachloride (CCl4), are carried out by using the z-scan technique. Open-aperture z-scan (OAZS) and close-aperture z-scan (CAZS) experiments are performed to determine the nonlinear absorption coefficient (β) and nonlinear refractive index (n2) of all binary liquid mixtures at various compositions of the components by employing a pulsed, high repetition rate (HRR) femtosecond laser. Also, we were able to use the flowing liquid to measure NLO properties in the CS2-acetone binary mixture to remove the cumulative thermal effects produced due to the pulsed HRR laser light. Nonlinear refractive index (n2) values are found to be influenced by the weak dipole-induced dipole intermolecular interactions between the nonpolar CS2 and polar acetone as well as cyclopentanone of the respective binary mixtures. On the contrary n2 values are not found to be affected by the intermolecular interactions in CS2-toluene and CS2-CCl4 binary mixtures. In comparison, the nonlinear absorption coefficient (β) values are not found to be affected by the same in all different sets of binary mixtures.
We test our meta-molecular dynamics (MD) based approach for finding low-barrier (<30 kcal/mol) reactions (SciPost Chem. 2021, 1, 003) on uni- and bimolecular reactions extracted from the barrier dataset developed by Grambow et al. (Scientific Data 2020, 7, 137). For unimolecular reactions the meta-MD simulations identify 25 of the 26 products found by Grambow et al., while the subsequent semiempirical screening eliminates an additional four reactions due to at an overestimation of the reaction energies or estimated barrier heights relative to DFT. In addition, our approach identifies an additional 36 reactions not found by Grambow et al., 10 of which are <30 kcal/mol. For bimolecular reactions the meta-MD simulations identify 19 of the 20 reactions found by Grambow et al., while the subsequent semiempirical screening eliminates an additional reaction. In addition, we find 34 new low-barrier reactions. For bimolecular reactions we found that it is necessary to ”encourage” the reactants to go to previously undiscovered products, by including products found by other MD simulations when computing the biasing potential as well as decreasing the size of the molecular cavity in which the MD occurs, until a reaction is observed. We also show that our methodology can find the correct products for two reactions that are more representative of those encountered in synthetic organic chemistry. The meta-MD hyperparameters used in this study thus appears to be generally applicable to finding low-barrier reactions.
Protein engineers conventionally use tools such as Directed Evolution to find new proteins with better functionalities and traits. More recently, computational techniques and especially machine learning approaches have been recruited to assist Directed Evolution, showing promising results. In this article, we propose POET, a computational Genetic Programming tool based on evolutionary computation methods to enhance screening and mutagenesis in Directed Evolution and help protein engineers to find proteins that have better functionality. As a proof-of-concept, we use peptides that generate MRI contrast detected by the Chemical Exchange Saturation Transfer contrast mechanism. The evolutionary methods used in POET are described, and the performance of POET in different epochs of our experiments with Chemical Exchange Saturation Transfer contrast are studied. Our results indicate that a computational modeling tool like POET can help to find peptides with 400% better functionality than used before.
The substitution of Ile to Val at residue 117 (I117V) of neuraminidase (NA) reduces the susceptibility of the A/H5N1 influenza virus to oseltamivir (OTV). However, the molecular mechanism by which the I117V mutation affects the intermolecular interactions between NA and OTV has not been fully elucidated. In this study, we performed molecular dynamics (MD) simulations to analyze the characteristic conformational changes that contribute to the reduced binding affinity of NA to OTV after the I117V mutation. The results of MD simulations revealed that after the I117V mutation in NA, the changes in the secondary structure around the mutation site had a noticeable effect on the residue interactions in the OTV-binding site. In the case of the WT NA-OTV complex, the positively charged side chain of R118, located in the β-sheet region, frequently interacted with the negatively charged side chain of E119, which is an amino acid residue in the OTV-binding site. This can reduce the electrostatic repulsion of E119 toward D151, which is also a negatively charged residue in the OTV-binding site, so that both E119 and D151 simultaneously form hydrogen bonds with OTV more frequently, which greatly contributes to the binding affinity of NA to OTV. After the I117V mutation in NA, the side chain of R118 interacted with the side chain of E119 less frequently, likely because of the decreased tendency of R118 to form a β-sheet structure. As a result, the electrostatic repulsion of E119 toward D151 is greater than that of the WT case, making it difficult for both E119 and D151 to simultaneously form hydrogen bonds with OTV, which in turn reduces the binding affinity of NA to OTV. Hence, after the I117V mutation in NA, influenza viruses are less susceptible to OTV because of conformational changes in residues of R118, E119, and D151 around the mutation site and in the binding site.
Background Intrinsically disordered proteins (IDPs) have been shown to exhibit cryoprotective activity toward other cellular enzymes without any obvious conserved sequence motifs. This study investigated relationships between the physical properties of several human genome-derived IDPs and their cryoprotective activities. Methods Cryoprotective activity of three human-genome derived IDPs and their truncated peptides toward lactate dehydrogenase (LDH) and glutathione S-transferase (GST) was examined. After the shortest cryoprotective peptide was defined (named FK20), cryoprotective activity of all-D-enantiomeric isoform of FK20 (FK20-D) as well as a racemic mixture of FK20 and FK20-D was examined. In order to examine the lack of increase of thermal stability of the target enzyme, the CD spectra of GST and LDH in the presence of a racemic mixture of FK20 and FK20-D at varying temperatures were measured and used to estimate Tm. Results Cryoprotective activity of IDPs longer than 20 amino acids was nearly independent of the amino acid length. The shortest IDP-derived 20 amino acid length peptide with sufficient cryoprotective activity was developed from a series of TNFRSF11B fragments (named FK20). FK20, FK20-D, and an equimolar mixture of FK20 and FK20-D also showed similar cryoprotective activity toward LDH and GST. Tm of GST in the presence and absence of an equimolar mixture of FK20 and FK20-D are similar, suggesting that IDPs’ cryoprotection mechanism seems partly from a molecular shielding effect rather than a direct interaction with the target enzymes.
A bio-based Silica/Calcium Carbonate (CS–SiO 2 /CaCO 3 ) nanocomposite was synthesized in this study using waste eggshells (ES) and rice husks (RH). The adsorbents (ESCaCO 3 , RHSiO 2 and, CS-SiO 2 /CaCO 3 ) characterized using XRD show crystallinity associated with the calcite and quartz phase. The FTIR of ESCaCO 3 shows the CO −2 3 group of CaCO 3, while the spectra of RHSiO2 majorly show the siloxane bonds (Si–O–Si) in addition to the asymmetric and symmetric bending mode of SiO 2 . The spectra for Chitosan (CS) show peaks corresponding to the C=O vibration mode of amides, C–N stretching, and C–O stretching. The CS–SiO 2 /CaCO 3 nanocomposite shows the spectra pattern associated with ESCaCO 3 and RHSiO 2. The FESEM micrograph shows a near monodispersed and spherical CS–SiO 2 /CaCO 3 nanocomposite morphology, with an average size distribution of 32.15 ± 6.20 nm. The corresponding EDX showed the representative peaks for Ca, C, Si, and O. The highest removal efficiency of phenol over the adsorbents was observed over CS–SiO 2 /CaCO 3 nanocomposite compared to other adsorbents. Adsorbing 84–89% of phenol in 60–90 min at a pH of 5.4, and a dose of 0.15 g in 20 ml of 25 mg/L phenol concentration. The result of the kinetic model shows the adsorption processes to be best described by pseudo-second-order. The highest correlation coefficient ( R 2 ) of 0.99 was observed in CS-SiO 2 /CaCO 3 nanocomposite, followed by RHSiO 2 and ESCaCO 3 . The result shows the equilibrium data for all the adsorbents fitting well to the Langmuir isotherm model, and follow the trend CS-SiO 2 /CaCO 3 > ESCaCO 3 > RHSiO 2 . The Langmuir equation and Freundlich model in this study show a higher correlation coefficient ( R 2 = 0.9912 and 0.9905) for phenol adsorption onto the CS–SiO 2 /CaCO 3 nanocomposite with a maximum adsorption capacity ( q m ) of 14.06 mg/g compared to RHSiO 2 (10.64 mg/g) and ESCaCO 3 (10.33 mg/g). The results suggest good monolayer coverage on the adsorbent’s surface (Langmuir) and heterogeneous surfaces with available binding sites (Freundlich).
The dihydroazulene/vinylheptafulvene (DHA/VHF) thermocouple is a promising candidate for thermal heat batteries that absorb and store solar energy as chemical energy without the need for insulation. However, in order to be viable the energy storage capacity and lifetime of the high energy form (i.e. the free energy barrier to the back reaction) of the canonical parent compound must be increased significantly to be of practical use. We use semiempirical quantum chemical methods, machine learning, and density func- tional theory to virtually screen over 230 billion substituted DHA molecules to identify promising candidates. We identify a molecule with a predicted energy density of 0.38 kJ/g, which is significantly larger than the 0.14 kJ/g computed for the parent compound. The free energy barrier to the back reaction is 11 kJ/mol higher than the parent com- pound, which should correspond to a half-life of about 10 days - 4 months. This is considerably longer than the 3-39 hours (depending on solvent) observed for the parent compound and sufficiently long for many practical applications. Our paper makes two main important contributions: 1) a novel and generally applicable methodological approach that makes screening of huge libraries for properties involving chemical reactivity with modest computational resources, and 2) a clear demonstration that the storage capacity of the DHA/VHF thermocouple cannot be increased to >0.5 kJ/g by combining simple substituents.
A graph-based genetic algorithm (GA) is used to identify molecules (ligands) with high absolute docking scores as estimated by the Glide software, starting from randomly chosen molecules from the ZINC database, for four different targets: Bacillus subtilis chorismate mutase (CM), human β2-adrenergic G protein-coupled receptor (β2AR), the DDR1 kinase domain (DDR1), and β-cyclodextrin (BCD). By the combined use of functional group filters and a score modifier based on a heuristic synthetic accessibility (SA) score our approach identifies between ca 500 and 6000 structurally diverse molecules with scores better than known binders by screening a total of 400,000 molecules starting from 8000 randomly selected molecules from the ZINC database. Screening 250,000 molecules from the ZINC database identifies significantly more molecules with better docking scores than known binders, with the exception of CM, where the conventional screening approach only identifies 60 compounds compared to 511 with GA+Filter+SA. In the case of β2AR and DDR1 the GA+Filter+SA approach finds significantly more molecules with docking scores lower than -9.0 and -10.0. The GA+Filters+SA docking methodology is thus effective in generating a large and diverse set of synthetically accessible molecules with very good docking scores for a particular target. An early incarnation of the GA+Filter+SA approach was used to identify potential binders to the COVID-19 main protease and submitted to the early stages of the COVID Moonshot project, a crowd-sourced initiative to accelerate the development of a COVID antiviral.
We attempt to explain why search algorithms can find molecules with particular properties in an enormous chemical space (ca 1060 molecules) by considering only a tiny subset (typically 103−6 molecules). Using a very simple example, we show that the number of potential paths that the search algorithms can follow to the target is equally vast. Thus, the probability of randomly finding a molecule that is on one of these paths is quite high and from here a search algorithm can follow the path to the target molecule. A path is defined as a series of molecules that have some non-zero quantifiable similarity (score) with the target molecule and that are increasingly similar to the target molecule. The minimum path length from any point in chemical space to the target corresponds is on the order of 100 steps, where a step is the change of and atom- or bond-type. Thus, a perfect search algorithm should be able to locate a particular molecule in chemical space by screening on the order of 100s of molecules, provided the score changes incrementally. We show that the actual number for a genetic search algorithm is between 100 and several millions, and depending on the target property and its dependence on molecular changes, the molecular representation, and the number of solutions to the search problem.
Mathematical models of the dynamics of infectious disease transmission are used to forecast epidemics and assess mitigation strategies. We reveal that the classic Susceptible-Infectious-Recovered (SIR) epidemic model resembles a dynamic model of a batch reactor carrying out an autocatalytic reaction with catalyst deactivation. This analogy between disease transmission and chemical reactions enables the cross-pollination of ideas between epidemic and chemical kinetic modeling.