This paper presents an effort for development of predictive quantitative structure-activity relationship (QSAR) models to assess mutagenicity of a set of 95 aromatic and heteroaromatic amines using newly developed set of Laplacian matrix-based molecular descriptors applying graph convolution. In order to calculate these descriptors, we used (1) properties of isolated atoms, (2) properties of the atoms, which arise from their inclusion in molecular structures, and (3) properties derived from the graphs that represent the molecules. The QSAR models developed are based on counter-propagation artificial neural network algorithm. The optimization was performed in an automated manner using genetic algorithm for (1) selection of the most suitable descriptors as well as for (2) selection of the most suitable parameters of the neural networks. Results presented in this study show that this new type of descriptors, reported in this paper, encode structural information capable of predicting TA98 Ames’ mutagenicity of the amines.
Cell membranes exhibit specific structural and chirality properties influencing their biological behavior and functionality. Artemia salina endothelial-like cell membranes, structurally simpler, provide insights into fundamental cellular structures, whereas human endothelial cell membranes represent complex, specialized tissues essential for understanding advanced vascular functions. This study aims to compare the structural and chiral properties of Artemia salina endothelial-like cell membranes and human endothelial cell membranes through computational molecular-level modeling, evaluating potential histological and biological implications. Membrane models for Artemia salina and human endothelial cells were developed using Protein Data Bank (PDB) structures. Computational descriptors, including radius of gyration (Rg), solvent-accessible surface area (SASA), geometric asymmetry index (GAI), chiral moment (CM), fractal dimension (FD), and additional chirality indices (SOC, HCI, ACI, CAI, ME, RDF) were calculated to assess membrane complexity, structural asymmetry, and chirality. Significant structural divergences between Artemia salina and human endothelial membranes were identified. Artemia membranes exhibited lower values of Rg, SASA, and chirality metrics, indicating simpler, more symmetrical structures. In contrast, human endothelial membranes displayed elevated structural complexity, pronounced asymmetry, higher chirality indices, and more significant structural heterogeneity, consistent with their specialized physiological functions. Principal Component Analysis (PCA) further highlighted clear structural clustering distinctions between the two models. The comparative analysis underscores fundamental structural and functional divergences between Artemia salina and human endothelial cell membranes. Artemia membranes represent simplified, uniform cellular arrangements optimized for fundamental physiological roles, while human endothelial membranes exhibit complex architectures, structural specialization, and significant chirality essential for dynamic vascular functionalities. These computational descriptors offer potential diagnostic biomarkers for evaluating endothelial functionality and pathological states.
Chirality is a pervasive and functionally critical feature of biological macromolecules, yet its distributed and emergent forms remain poorly quantified in complex systems such as membrane proteins. We present Chirobiophore, a novel paradigm for capturing biochirality across scales—from atomic geometries to global structural asymmetries. Unlike traditional stereochemical metrics, Chirobiophore employs a multidimensional model-independent vector comprising Local Tetrahedral Asymmetry (LTA), Helical Path Curvature (HPC), Asymmetric Environment Score (AES), Directional Density Profile (DDP), Leaflet Asymmetry Index (LAI), and Orientation Twist Score (OTS). This framework enables coordinate-invariant comparisons of structurally diverse proteins in a continuous chirality space. We demonstrate its application to canonical, GPCR, and topologically complex membrane proteins, revealing distinct chirality signatures and functional clustering. Furthermore, we map Chirobiophore descriptors to tissue-level asymmetry indices, providing a bridge between molecular structure and morphogenetic patterning. Chirobiophore offers a unified, extensible platform for structural biology, synthetic design, and developmental modeling of chirality.
The development of chirality descriptors for quantitative chirality structure–activity relationship (QCSAR) modeling has always attracted attention, owing to the importance of chiral molecules in pharmaceutical, agriculture, food, and fragrance industries, and environmental toxicology. The utility of a multidimensional space of novel relative chirality indices (RCIs) in the QCSAR modeling of twenty CCR2 antagonists is reported upon in this paper. The numerical characterization of chirality by the RCI approach gives a large pool of chirality descriptors with different degrees of mutual correlation (the correlation coefficient among the computed descriptors varied from 0.02 to 0.99). In the present study, the final data set contains 198 chirality descriptors for each of the twenty CCR2 antagonist molecules, providing a multidimensional space for modeling. The data reduction using principal component analysis resulted in the extraction of eight principal components (PCs). The linear regression using the principal component scores (PCSs) resulted in a three-predictor prediction model with good statistics: R2 = 0.823; Adj R2 = 0.790. The regression models were rebuilt using the chirality descriptors (RCIs) that are most correlated with each of the scores (PCSs) of the three principal components. The R2 value for the regression models with three RCIs as the predictors is 0.742 and the five-fold cross validation, Rcv2, is 0.839. The new chirality descriptors, namely, the RCIs calculated using a different weighting scheme, provide a multidimensional space of chirality descriptors for a set of chiral molecules, and such a multidimensional chirality space is a powerful tool to build quantitative chiral structure–activity relationship (QCSAR) models.
This article reports the development of a set of new molecular descriptors derived from convolution using the Laplacians of molecular graphs and their line graphs. These descriptors have been applied in quantitative structure-activity relationship (QSAR) studies to predict the toxicity of 69 benzene derivatives and the aqueous solubility of a diverse dataset of 375 drug-like structures, using multivariate linear regression and a nonlinear machine learning algorithm known as counter-propagation artificial neural networks. The descriptors are developed using atomic properties of only the non-hydrogen atoms in the molecule. Using this approach, we developed a total of 54 new graph convolution-based descriptors. Results indicate that the newly definedinvariants provide a new set of molecular descriptors for the characterization of molecular structures and QSAR studies
This chapter discusses the authors' mathematical proteomics methods for characterization of high-dimensional proteomics data. Three discrete mathematics-based methods for two-dimensional gel electrophoresis (2-DE) data are discussed: graph invariant approach, information-theoretic biodescriptors technique, and spectrum-like representation of projection of three-dimensional 2-DE data. A similarity-based approach is used to reduce the number of relevant descriptors for the nanotoxicoproteomics section, and then important selected descriptors are used to compare the proteomics patterns using principal component analysis and self-organizing maps.
This chapter gives a detailed presentation of the theoretical background and computational approaches to the utility of alignment-free sequence descriptors and multidimensional variable reduction methods in the characterization and visualization of biological sequence data. The utility of such novel methods developed by the authors of this chapter is shown using data on case studies of severe acute respiratory syndrome, Middle East respiratory syndrome, Coronavirus disease-2019, and Zika viruses.
Currently, we are witnessing the emergence of big data in various fields including the biomedical and natural sciences. The size of chemoinformatics and bioinformatics databases is increasing every day. This gives us both challenges and opportunities. This chapter discusses the mathematical methods used in these fields both for the generation and analysis of such data. It is emphasized that proper use of robust statistical and machine learning methods in the analysis of the available big data may facilitate both hypothesis-driven and discovery-oriented research.
Owing to the homochirality among the α-amino acids, the building blocks, chiral environmental prevails within the structure of biological macromolecules namely, the proteins, receptors, and enzymes. This results in chiral distinction of the ligands such as drug molecules and toxicants by the biological targets. Chiral distinction of enantiomers is not only important in the biological activity of enantiomers but also in pharmacokinetics and metabolism of chemicals. The molecular descriptors that are based only on the connectivity of atoms cannot differentiate the enantiomers and diastereomers. In order to model the differential activity of enantiomers and diastereomers, molecular descriptors capable of encoding the difference in spatial arrangements of atoms and groups around a chiral center are needed. In this paper we report a modified approach that enables to compute a large pool of chirality descriptors for a given set of molecules. The new chirality descriptors can differentiate enantiomers and diastereomers. Application of the new chiral descriptors in structure–activity modelling of bioactivity of chiral molecules is illustrated for the dopamine (DA) D2 and opiate σ receptor affinities of seven pairs of enantiomers of 3-(3-hydroxyphenyl)piperidines.
During an emergency, such as a pandemic in which time and resources are extremely scarce, it is important to find effective and rapid solutions when searching for possible treatments. One possibility in this regard is the repurposing of available “on the market” drugs. This is a proof of the concept study showing the potential of a collaboration between two research groups, engaged in computer-aided drug design and control of viral infections, for the development of early strategies to combat future pandemics. We describe a QSAR (quantitative structure activity relationship) based repurposing study on molecular topology and molecular docking for identifying inhibitors of the main protease (Mpro) of SARS-CoV-2, the causative agent of COVID-19. The aim of this computational strategy was to create an agile, rapid, and efficient way to enable the selection of molecules capable of inhibiting SARS-CoV-2 protease. Molecules selected through in silico method were tested in vitro using human coronavirus 229E as a surrogate for SARS-CoV-2. Three strategies were used to screen the antiviral activity of these molecules against human coronavirus 229E in cell cultures, e.g., pre-treatment, co-treatment, and post-treatment. We found >99% of virus inhibition during pre-treatment and co-treatment and 90–99% inhibition when the molecules were applied post-treatment (after infection with the virus). From all tested compounds, Molport-046-067-769 and Molport-046-568-802 are here reported for the first time as potential anti-SARS-CoV-2 compounds.
INTRODUCTION:Coronaviruses comprise a group of enveloped, positive-sense single-stranded RNA viruses that infect humans as well as a wide range of animals. The study was performed on a set of 573 sequences belonging to SARS, MERS and SARS-CoV-2 (CoVID-19) viruses. The sequences were represented with alignment-free sequence descriptors and analyzed with different chemometric methods: Euclidean/Mahalanobis distances, principal component analysis and self-organizing maps (Kohonen networks). We report the cluster structures of the data. The sequences are well-clustered regarding the type of virus; however, some of them show the tendency to belong to more than one virus type.BACKGROUND:This is a study of 573 genome sequences belonging to SARS, MERS and SARS-- CoV-2 (CoVID-19) coronaviruses.OBJECTIVES:The aim was to compare the virus sequences, which originate from different places around the world.METHODS:The study used alignment free sequence descriptors for the representation of sequences and chemometric methods for analyzing clusters.RESULTS:Majority of genome sequences are clustered with respect to the virus type, but some of them are outliers.CONCLUSION:We indicate 71 sequences, which tend to belong to more than one cluster.
The design for vaccines using in silico analysis of genomic data of different viruses has taken many different paths, but lack of any precise computational approach has constrained them to alignment methods and some alignment-free techniques. In this work, a precise computational approach has been established wherein two new mathematical parameters have been suggested to identify the highly conserved and surface-exposed regions which are spread over a large region of the surface protein of the virus so that one can determine possible peptide vaccine candidates from those regions. The first parameter, w, is the sum of the normalized values of the measure of surface accessibility and the normalized measure of conservativeness, and the second parameter is the area of a triangle formed by a mathematical model named 2D Polygon Representation. This method has been, therefore, used to determine possible vaccine targets against SARS-CoV-2 by considering its surface-situated spike glycoprotein. The results of this model have been verified by a parallel analysis using the older approach of manually estimating the graphs describing the variation of conservativeness and surface-exposure across the protein sequence. Furthermore, the working of the method has been tested by applying it to find out peptide vaccine candidates for Zika and Hendra viruses respectively. A satisfactory consistency of the model results with pre-established results for both the test cases shows that this in silico alignment-free analysis proposed by the model is suitable not only to determine vaccine targets against SARS-CoV-2 but also ready to extend against other viruses.
BACKGROUND:In this report, we consider a data set, which consists of 310 Zika virus genome sequences taken from different continents, Africa, Asia and South America. The sequences, which were compiled from GenBank, were derived from the host cells of different mammalian species (Simiiformes, Aedes opok, Aedes africanus, Aedes luteocephalus, Aedes dalzieli, Aedes aegypti, and Homo sapiens). METHODS:For chemometrical treatment, the sequences have been represented by sequence descriptors derived from their graphs or neighborhood matrices. The set was analyzed with three chemometrical methods: Mahalanobis distances, principal component analysis (PCA) and self organizing maps (SOM). A good separation of samples with respect to the region of origin was observed using these three methods. RESULTS:Study of 310 Zika virus genome sequences from different continents. To characterize and compare Zika virus sequences from around the world using alignment-free sequence comparison and chemometrical methods. CONCLUSION:Mahalanobis distance analysis, self organizing maps, principal components were used to carry out the chemometrical analyses of the Zika sequence data. Genome sequences are clustered with respect to the region of origin (continent, country). Africa samples are well separated from Asian and South American ones.
We consider a novel approach to mathematically define a graphing method to represent amino acid sequences of proteins in two-dimensional plane and characterize them numerically. The amino acids are represented by their relative magnitude of their hydrophobicity. Each amino acid is compared with a vector and moves in relative direction which generates a graph. Applications are shown in Zaire Ebola Virus to conclude how this plotting can be more useful than base sequence plotting. Also, superimposition graph of SARS and SARS-CoV-2 shows that these sequences are strongly related and Various other applications are shown too to explain it's fruitfulness.
With the increasing frequency of viral epidemics, vaccines to augment the human immune response system have been the medium of choice to combat viral infections. The tragic consequences of the Zika virus pandemic in South and Central America a few years ago brought the issues into sharper focus. While traditional vaccine development is time-consuming and expensive, recent advances in information technology, immunoinformatics, genetics, bioinformatics, and related sciences have opened the doors to new paradigms in vaccine design and applications. Peptide vaccines are one group of the new approaches to vaccine formulation. In this chapter, we discuss the various issues involved in the design of peptide vaccines and their advantages and shortcomings, with special reference to the Zika virus for which no drugs or vaccines are as yet available. In the process, we outline our work in this field giving a detailed step-by-step description of the protocol we follow for such vaccine design so that interested researchers can easily follow them and do their own designing. Several flowcharts and figures are included to provide a background of the software to be used and results to be anticipated.
Nenad Trinajstic合作论文数University of Zagreb5