Elucidating the role of the hydrophobic groove in mineralocorticoid receptor (MR)–spironolactone interactions is important for structure-based drug design, receptor modulation, and the development of more selective MR antagonists. Despite the clinical importance of spironolactone, the contribution of the hydrophobic groove, particularly residues M807, F829, M845, C849, and M852, remains underexplored. Here, we demonstrate through molecular dynamic simulations that these hydrophobic residues, together with polar residue N770, stabilize the thioacetyl moiety of spironolactone. Binding free energy calculations of the hydrophobic groove, both with the complete binding site and with the groove alone, demonstrate the impact of the groove’s hydrophobicity along with the polar residues N770, Q776, and R817. Simulation results, supported by statistical analysis, highlight the groove’s structural and energetic significance. Site-directed mutagenesis targeting residues F829, M845, and C849 further clarifies their role in the binding mechanism, offering insights for rational drug design and biomarker development. The crystal structure of the MR–spironolactone complex (PDB ID: 3VHU) was retrieved and mutated using COOT. Mutant complexes were constructed and subjected to 1 μs molecular dynamics simulations using GROMACS. Binding free energies were calculated via MM/PBSA. Residue–ligand interactions were analyzed from MD trajectories using LigPlot + and GROMACS tools. Statistical significance of residue contributions was assessed using ANOVA, comparing polar and hydrophobic residue mutations across simulated complexes.
Helicobacter pylori infections represent a significant global health challenge, particularly due to the increasing prevalence of antibiotic resistance. This study aims to identify novel therapeutic targets by characterising H. pylori Histidyl tRNA synthetase (HpHRS), an enzyme essential for protein biosynthesis. We successfully cloned, expressed, and purified HpHRS using affinity chromatography and size exclusion chromatography techniques. Circular dichroism spectroscopy revealed conformational changes in HpHRS upon substrate binding, with a notable decrease in α-helical content and an increase in β-sheet structures. Fluorescence quenching studies confirmed the binding of L-histidine and ATP to the enzyme's active site. Virtual screening of the ZINC database identified two potential inhibitors, ZINC39960778 and ZINC30878996, which exhibited higher affinities than the natural substrate. Molecular dynamics simulations conducted over 100 ns demonstrated stable protein-ligand interactions, with the HRS-Histidyl Adenylate Monophosphate (HAM) complex showing the highest stability. Free energy landscape analysis indicated greater conformational freedom in ligand-bound complexes than apoproteins. These findings provide valuable insights into the structure and function of HpHRS, contributing to the development of new antibacterial agents targeting this crucial enzyme in H. pylori. Given the growing concern regarding antibiotic resistance, this study presents a promising avenue for designing new therapeutic strategies against H. pylori infections. Further research should focus on the experimental validation of the identified inhibitors and optimisation to develop potent and selective HpHRS inhibitors as potential antibacterial candidates.
Autoimmunity has been explored in various viral infections, and its relevance to respiratory viruses deserves more attention, especially its immune derangement during these infections, which could potentially trigger relapse and induction of many new cases. Our study aimed to utilize publicly available transcriptomic respiratory viral datasets of rhinovirus, influenza virus, respiratory syncytial virus, and COVID-19 to understand their autoimmune activation. Antibodies produced against the autoantigens associated with respiratory viruses resulted in the identification of three biomarker genes: TRIM21, ELANE, and CTSG. These genes are reported to be involved in the pathways of neuroactive ligand-receptor interaction, neutrophil extracellular trap formation, apoptosis, amebiasis, renin-angiotensin system, and lysosome, commonly triggering the systemic lupus erythematosus (SLE) pathway in genetically susceptible SLE patients. These results emphasize that the key genes are enriched mainly in the immune system process linking SLE pathogenesis. Literature sources suggest that the biomarkers induce autoreactivity through bystander activation and molecular mimicry which results in aberrant B-cell activation and the formation of neutrophil extracellular traps leading to autoimmunity. Thus, these key biomarkers indicate a new direction for early diagnosis, risk assessment, and treatment of respiratory virus infections and SLE pathogenesis.
ABSTRACTBovine serum albumin (BSA) plays a crucial role as a carrier protein in plasma, binding various ligands, including drugs. Understanding the interaction between BSA and saquinavir, an antiretroviral drug, is essential for predicting its pharmacokinetics and pharmacodynamics. We employed spectroscopic approaches, including circular dichroism spectrometry and fluorescence spectroscopy, to investigate the binding of saquinavir to BSA. CD studies revealed conformational changes upon saquinavir mesylate binding, and the complex was stable up to 45°C during thermal denaturation. Saquinavir quenched the intrinsic fluorescence of BSA, indicating static quenching due to complex formation. Additionally, molecular docking simulations were performed to elucidate the favored binding site and interactions. The molecular docking results revealed that Subdomains IIA and IIB, which are proximal to Sudlow Site I, are the principal binding sites for the antiviral drug saquinavir. The ligand‐bound pose of BSA also revealed that residue Trp213, which is adjacent to saquinavir, further validated the results of the fluorescence quenching assay, suggesting that residue Trp213 is quenched upon binding with saquinavir. MD simulations allowed us to explore the dynamic behavior of the BSA–saquinavir complex over time. We observed conformational fluctuations, solvent exposure, flexibility of binding pockets, free energy landscape, and binding energy. This study enhances our understanding of drug–protein interactions and contributes to drug development and optimization.
This study aims to investigate the comparative binding pattern of TTHA1873 and its mutants (R55A and R138A) with DNA through molecular docking and molecular dynamics (MD) simulations. The docking results suggests that the Wild type (WT-TTHA1873), R55A, R138A and double mutant R55A/R138A having docking scores of -225.80 kcal/mol, -209.81 kcal/mol, -197.53 kcal/mol, -195.55 kcal/mol respectively and WT-TTHA1873 has more significant binding capability with DNA in comparison to mutants. The MD analysis revealed that the WT-TTHA1873 demonstrated stable interactions with DNA and exhibited a reduced conformational space compared to the mutants. By examining the atomic interactions, it was observed that significant variations in the hydrogen bonding pattern between WT-TTHA1873 and its mutants while interacting with DNA resulted in structural anomalies in the mutants and differences in DNA-binding specificity. The calculated binding free energies imply more stability of the WT-TTHA1873-DNA complex, while the mutants showed lesser binding affinity toward its interacting partner, double-stranded DNA. It is apparent that substituting single mutation R55A and R138A on TTHA1873 abolishes their DNA-binding ability. The present study portrays the critical role of R55 and R138 from TTHA1873 as likely involved in DNA binding.
In recent years, several experimental evidences suggest that amino acid repeats are closely linked to many disease conditions, as they have a significant role in evolution of disordered regions of the polypeptide segments. Even though many algorithms and databases were developed for such analysis, each algorithm has some caveats, like limitation on the number of amino acids within the repeat patterns and number of query protein sequences. To this end, in the present work, a new method called the internal sequence repeats across multiple protein sequences (ISRMPS) is proposed for the first time to identify identical repeats across multiple protein sequences. It also identifies distantly located repeat patterns in various protein sequences. Our method can be applied to study evolutionary relationships, epitope mapping, CRISPR-Cas sequencing methods, and other comparative analytical assessments of protein sequences.
Myxobacteria are predatory bacteria with antimicrobial activity, utilizing complex mechanisms to kill their prey and assimilate their macromolecules. Having large genomes encoding hundreds of secondary metabolites, hydrolytic enzymes and antimicrobial peptides, these organisms are widely studied for their antibiotic potential. MyxoPortal is a comprehensive genomic database hosting 262 genomes of myxobacterial strains. Datasets included provide genome annotations with gene locations, functions, amino acids and nucleotide sequences, allowing analysis of evolutionary and taxonomical relationships between strains and genes. Biosynthetic gene clusters are identified by AntiSMASH, and dbAMP-generated antimicrobial peptide sequences are included as a resource for novel antimicrobial discoveries, while curated datasets of CRISPR/Cas genes, regulatory protein sequences, and phage associated genes give useful insights into each strain's biological properties. MyxoPortal is an intuitive open-source database that brings together application-oriented genomic features that can be used in taxonomy, evolution, predation and antimicrobial research. MyxoPortal can be accessed at http://dicsoft1.physics.iisc.ac.in/MyxoPortal/.Database URL: http://dicsoft1.physics.iisc.ac.in/MyxoPortal/.Graphical Abstract
Although many studies have addressed the significance of interactions in the dimeric structures of proteins, there is no dedicated database of these interactions. To this end, the Molecular Interactions in Protein Dimer Structure (MIPDS) database has been developed; it is an open-access repository containing 60 298 3D structures of dimeric proteins sourced from the Protein Data Bank. This helps researchers comprehend the types of interaction, which include those mediated by water, small molecules or ligands and direct interactions, in 3D structures at the molecular level. The database is accessible through a user-friendly interface, where users can conduct searches based on PDB accession number, interaction type and geometric parameters. It can be viewed in textual and graphical formats using the plug-in JSmol. MIPDS is updated weekly using programmed scripts to incorporate newly released dimeric structures and analyses of their interaction types. The database is intended for the scientific community working in structural biology, structural bioinformatics, drug discovery and development. MIPDS is freely accessible to users worldwide at http://dicsoft1.physics.iisc.ac.in/mipds.
Zika virus (ZIKV) and Dengue virus (DENV) infections cause severe disease in humans and are significant socio-economic burden worldwide. These flavivirus infections are difficult to diagnose serologically due to antigenic overlap. The phylogenetic analysis shows that ZIKV clusters with DENVs at a higher node of the phylogenetic tree with significant genomic and structural similarity. Our study aims to identify gene biomarkers for the classification of Dengue and Zika viral infections using machine learning algorithms and bioinformatics analysis. The gene expression count matrix for single-cell RNA sequencing dataset GSE110496 was analyzed using binary classifiers, namely Logistic regression, Support Vector Machines, Random Forest, and Decision trees. The GSE110496 dataset represents a unique study of the transcriptional and translational dynamics of DENV and ZIKV infections at 4-, 12-, 24-, and 48-h time points for human hepatoma (Huh7) cells. Out of which 24-h time point has been analyzed in this study, at the optimal threshold of viral molecules. Feature selection was performed using two different approaches Random Forest Classifier (RFC) for gene ranking and Recursive Feature Elimination (RFE). Out of which RFE, showed more accuracy and precision. The classification accuracy of 89.4
The present investigation illustrates the conceptualization, synthesis, crystallographic analysis, and computational assessment of a new racetam derivative with a pyrrolidone ring as the pharmacophore. The compound demonstrates drug-like characteristics similar to LEV, an approved anti-epileptic drug, indicating its potential for developing novel epilepsy drugs. Molecular docking and molecular dynamic simulations were used to evaluate the binding affinity between the compound and SV2A, a protein-ligand complex. The protein-ligand complex attained structural equilibrium in the final 50 nanoseconds of the 200 nanosecond molecular dynamics simulation. The analysis reveals that regions with higher flexibility are primarily located in the extramembrane regions of the protein. Intermolecular contact analysis reveals hydrogen bonding and hydrophobic interactions as the primary types of interactions. The Molecular Mechanics Poisson-Boltzmann Surface Area (MMPBSA) calculation highlights the energetic aspects of ligand binding and the participation of important residues in the binding pocket. Unique interactions like those involving bifurcated hydrogen bonding, and novel pi-anion interaction makes the study significant. Quantum chemical calculations for the compound (LIG) done using DFT corroborates the protein-ligand interactions on the basis of Molecular Electrostatic Potential (MEP) maps. The study establishes the structure-activity relationship (SAR) of the newly developed pyrrolidone-based compound, identifying it as a promising lead molecule for epilepsy treatment.
The insulin superfamily proteins (ISPs), in particular, insulin, IGFs and relaxin proteins are key modulators of animal physiology. They are known to have evolved from the same ancestral gene and have diverged into proteins with varied sequences and distinct functions, but maintain a similar structural architecture stabilized by highly conserved disulphide bridges. The recent surge of sequence data and the structures of these proteins prompted a need for a comprehensive analysis, which connects the evolution of these sequences (427 sequences) in the light of available functional and structural information including representative complex structures of ISPs with their cognate receptors. This study reveals (a) unusually high sequence conservation of IGFs (>90 % conservation in 184 sequences) and provides a possible structure-based rationale for such high sequence conservation; (b) provides an updated definition of the receptor-binding signature motif of the functionally diverse relaxin family members (c) provides a probable non-canonical C-peptide cleavage site in a few insulin sequences. The high conservation of IGFs appears to represent a classic case of resistance to sequence diversity exerted by physiologically important interactions with multiple partners. We also propose a probable mechanism for C-peptide cleavage in a few distinct insulin sequences and redefine the receptor-binding signature motif of the relaxin family. Lastly, we provide a basis for minimally modified insulin mutants with potential therapeutic application, inspired by concomitant changes observed in other insulin superfamily protein members supported by molecular dynamics simulation.
Protein dynamics linked to numerous biomolecular functions, such as ligand binding, allosteric regulation, and catalysis, must be better understood at the atomic level. Reactive atoms of key residues drive a repertoire of biomolecular functions by flipping between alternate conformations or conformational substates, seldom found in protein structures. Probing such sparsely sampled alternate conformations would provide mechanistic insight into many biological functions. We are therefore interested in evaluating the instance of amino acids adopted alternate conformations, either in backbone or side-chain atoms or in both. Accordingly, over 70000 protein structures appear to contain alternate conformations only 'A' and 'B' for any atom, particularly the instance of amino acids that adopted alternate conformations are more for Arg, Cys, Met, and Ser than others. The resulting protein structure analysis depicts that amino acids with alternate conformations are mainly found in the helical and β-regions and are often seen in high-resolution X-ray crystal structures. Furthermore, a case study on human cyclophilin A (CypA) was performed to explain the pre-existing intrinsic dynamics of catalytically critical residues from the CypA and how such intrinsic dynamics perturbed upon Ser99Thr mutation using molecular dynamics simulations on the ns-μs timescale. Simulation results demonstrated that the Ser99Thr mutation had impaired the alternate conformations or the catalytically productive micro-environment of Phe113, mimicking the experimentally observed perturbation captured by X-ray crystallography. In brief, a deeper comprehension of alternate conformations adopted by the amino acids may shed light on the interplay between protein structure, dynamics, and function.
Studying the relationship between sequences and their corresponding three-dimensional structure assists structural biologists in solving the protein-folding problem. Despite several experimental and in-silico approaches, still understanding or decoding the three-dimensional structures from the sequence remains a mystery. In such cases, the accuracy of the structure prediction plays an indispensable role. To address this issue, an updated web server (CSSP-2.0) has been created to improve the accuracy of our previous version of CSSP by deploying the existing algorithms. It uses input as probabilities and predicts the consensus for the secondary structure as a highly accurate three-state Q3 (helix, strand, and coil). This prediction is achieved using six recent top-performing methods: MUFOLD-SS, RaptorX, PSSpred v4, PSIPRED, JPred v4, and Porter 5.0. CSSP-2.0 validation includes datasets involving various protein classes from the PDB, CullPDB, and AlphaFold databases. Our results indicate a significant improvement in the accuracy of the consensus Q3 prediction. Using CSSP-2.0, crystallographers can sort out the stable regular secondary structures from the entire complex structure, which would aid in inferring the functional annotation of hypothetical proteins. The web server is freely available at https://bioserver3.physics.iisc.ac.in/cgi-bin/cssp-2/
Cancer is a multigene and widespread disease. Increasing drug resistance leads to the development of new therapeutic targets. Recent research indicates that various cellular components called Stress Granules (SGs) are engaged in the cancer-related signaling pathway. The phosphoinositol-3-kinase (PI3K)/AKT/mammalian target of rapamycin (mTOR) signaling, considered a master regulator in cancer, has been shown through genomic profiling studies to play a key role in Esophageal Cancer (EC). In this study, we performed the in-silico analysis of an RNA sequencing dataset to investigate the effects of omipalisib, a PI3K/mTOR inhibitor, on EC cell lines. Our objective was to identify novel molecular targets, particularly stress granule-related proteins, that contribute to drug resistance in EC. Using computational approaches including differential gene expression analysis, pathway analysis, and functional enrichment, we examined the transcriptomic changes in response to omipalisib treatment. Our analysis revealed downregulation of the PI3K/mTOR signaling pathway and upregulation of compensatory pathways such as FOXO and JAK-STAT signaling in response to omipalisib. Notably, we identified 16 stress granule-related proteins that were significantly upregulated, suggesting their potential role in drug resistance mechanisms. These findings provide new insights into the molecular mechanisms underlying drug resistance in EC and highlight potential novel targets for therapeutic intervention. Currently, EC is limited by the number of potential drugs for treatment, poor prognosis, and is prone to chemotherapeutic resistance to existing clinically proven drugs. Our computational analysis offers valuable insights into targeting stress granules for cancer drug discovery, potentially enhancing the development of new therapeutic strategies for EC. These results provide a strong foundation for future experimental validation and drug development efforts aimed at overcoming resistance to EC treatment. Received: 1 July 2024 | Revised: 20 September 2024 | Accepted: 20 November 2024 Conflicts of Interest The authors declare that they have no conflicts of interest to this work. Data Availability Statement The data that support the findings of this study are openly available in the NCBI Gene Expression Omnibus (GEO) database under accession number GSE143462 at https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE143462. Author Contribution Statement Vinod Jangid: Conceptualization, Methodology, Software, Validation, Formal analysis, Data curation, Writing - original draft, Visualization. Chandrasekar Narayanan Rahul: Conceptualization, Methodology, Validation, Investigation,Writing - original draft. Aarthi Rashmi B: Validation, Writing - review & editing. Kanagaraj Sekar:Validation, Writing - review & editing, Supervision, Project administration.
Syringic acid (SA) is an active carcinogenesis inhibitor; however, the low bioavailability and unstable functional groups hinder its activity. Here, a chemically synthesized novel SA analog (SA10) is evaluated for its anticancer activity using in-vitro and in-silico studies. K562 cell line study revealed that SA10 had shown a higher rate of inhibition (IC50 = 50.40 μg/mL) than its parental compound, SA (IC50 = 96.92 μg/mL), at 50 μM concentration. The inhibition ratio was also been evaluated by checking the expression level of NFkB and Bcl-2 and showing that SA10 has two-fold increase in the inhibitory mechanism than SA. This result demonstrates that SA10 acts as an NFkB inhibitor and an apoptosis inducer. Further, molecular docking and simulation have been performed to get insights into the possible inhibitory mechanism of SA and SA10 on NFkB at the atomistic level. The molecular docking results exemplify that both SA and SA10 bind to the active site of NFkB, thereby interfering with the association between DNA and NFkB. SA10 exhibits a more robust binding affinity than SA and is firmly docked well into the interior of the NFkB, as confirmed by MM-PBSA calculations. In a nutshell, the Benzimidazole scaffold containing SA10 has shown more NFkB inhibitory activity in K562 cells than SA, which could be helpful as an ideal therapeutic NFkB inhibitor for treating cancers.
Background:For years now, cancer treatments have entailed tried-and-true methods. Yet, oncologists and clinicians recommend a series of surgeries, chemotherapy, and radiation therapy. Yet, even amidst these treatments, the number of deaths due to cancer increases at an alarming rate. The prognosis of cancer patients is influenced by mutations, age, and various cancer stages. However, the association between these variables is unclear. Methods: The present work adopts a machine learning technique—k-nearest neighbor; for both regression and classification tasks, regression for predicting the survival time of oral cancer patients, and classification for classifying the patients into one of the predefined oral cancer stages. Two cross-validation approaches—hold-out and k-fold methods—have been used to examine the prediction results. Results: The experimental results show that the k-fold method performs better than the hold-out method, providing the least mean absolute error score of 0.015. Additionally, the model classifies patients into a valid group. Of the 429 records, 97 (out of 106), 99 (out of 119), 95 (out of 113), and 77 (out of 91) were classified to its correct label as stages – 1, 2, 3, and 4. The accuracy, recall, precision, and F-measure for each classification group obtained are 0.84, 0.85, 0.85, and 0.84. Conclusions: The study showed that aged patients with a higher number of mutations than young patients have a higher risk of short survival. Senior patients with a more significant number of mutations have an increased risk of getting into the last cancer stage
To regulate biological activity in humans, the Notch signaling pathway (NSP) plays an essential role in a wide array of cellular development and differentiation process. In recent years, many studies have reported that aberrant activation of Notch is associated with the tumor process; but no appropriate database exists to fill this significant gap. To address this, we created a pioneering database NCSp, which is open access and comprises intercommunicating pathways and related protein mutations. This allows scientists to understand better the cause of single amino acid mutations in proteins. Therefore, NCSp provides information on the predicted functional effect of human protein mutations, which aids in understanding the importance of mutations linked to the Notch crosstalk signaling pathways in cancerous and non-cancerous systems. This database might be helpful for therapeutic mutation analysis, molecular biology, and structural biology researchers. The NCSp database can be accessed through https://bioserver3.physics.iisc.ac.in/cgi-bin/nccspd/ .
The Dengue virus M protein is a 75 amino acid polypeptide with two helical transmembranes (TM). The TM domain oligomerizes to form an ion channel, facilitating viral release from the host cells. The M protein has a critical role in the virus entry and life cycle, making it a potent drug target. The oligomerization of the monomeric protein was studied using ab initio modeling and molecular dynamics (MD) simulation in an implicit membrane environment. The representative structures obtained showed pentamer as the most stable oligomeric state, resembling an ion channel. Glutamic acid, threonine, serine, tryptophan, alanine, isoleucine form the pore-lining residues of the pentameric channel, conferring an overall negative charge to the channel with approximate length of 51.9 Å. Residue interaction analysis (RIN) for M protein shows that Ala94, Leu95, Ser112, Glu124, and Phe155 are the central hub residues representing the physicochemical interactions between domains. The virtual screening with 165 different ion channel inhibitors from the ion channel library shows monovalent ion channel blockers, namely lumacaftor, glipizide, gliquidone, glisoxepide, and azelnidipine to be the inhibitors with high docking scores. Understanding the three-dimensional structure of M protein will help design therapeutics and vaccines for Dengue infection.
Background: Crohn's disease (CD) is a chronic idiopathic inflammatory bowel disease affecting the entire gastrointestinal tract from the mouth to the anus. These patients often experience a period of symptomatic relapse and remission. A 20 - 30% symptomatic recurrence rate is reported in the first year after surgery, with a 10% increase each subsequent year. Thus, surgery is done only to relieve symptoms and not for the complete cure of the disease. The determinants and the genetic factors of this disease recurrence are also not well-defined. Therefore, enhanced diagnostic efficiency and prognostic outcome are critical for confronting CD recurrence.Methods: We analysed ileal mucosa samples collected from neo-terminal ileum six months after surgery (M6=121 samples) from Crohn's disease dataset (GSE186582). The primary aim of this study is to identify the potential genes and critical pathways in post-operative recurrence of Crohn's disease. We combined the differential gene expression analysis with Recursive feature elimination (RFE), a machine learning approach to get five critical genes for the postoperative recurrence of Crohn's disease. The features (genes) selected by different methods were validated using five binary classifiers for recurrence and remission samples: Logistic Regression (LR), Decision tree classifier (DT), Support Vector Machine (SVM), Random Forest classifier (RF), and K-nearest neighbor (KNN) with 10-fold cross-validation. We also performed weighted gene co-expression network analysis (WGCNA) to select specific modules and feature genes associated with Crohn's disease postoperative recurrence, smoking, and biological sex. Combined with other biological interpretations, including Gene Ontology (GO) analysis, pathway enrichment, and protein-protein interaction (PPI) network analysis, our current study sheds light on the in-depth research of CD diagnosis and prognosis in postoperative recurrence.Results: PLOD2, ZNF165, BOK, CX3CR1, and ARMCX4, are the important genes identified from the machine learning approach. These genes are reported to be involved in the viral protein interaction with cytokine and cytokine receptors, lysine degradation, and apoptosis. They are also linked with various cellular and molecular functions such as Peptidyl-lysine hydroxylation, Central nervous system maturation, G protein-coupled chemoattractant receptor activity, BCL-2 homology (BH) domain binding, Gliogenesis and negative regulation of mitochondrial depolarization. WGCNA identified a gene co-expression module that was primarily involved in mitochondrial translational elongation, mitochondrial translational termination, mitochondrial translation, mitochondrial respiratory chain complex, mRNA splicing via spliceosome pathways, etc.; Both the analysis result emphasizes that the mitochondrial depolarization pathway is linked with CD recurrence leading to oxidative stress in promoting inflammation in CD patients.Conclusion: These key genes serve as the novel diagnostic biomarker for the postoperative recurrence of Crohn's disease. Thus, among other treatment options present until now, these biomarkers would provide success in both diagnosis and prognosis, aiming for a long-lasting remission to prevent further complications in CD.
—Cancer is often caused by missense mutations, where a single nucleotide substitution leads to an amino acid change and affects protein function. This study proposes a novel machine learning (ML) approach to calculate missing values in the tp53 database for three computational methods: SIFT, Provean, and Mutassessor scores. The computed values are compared with those obtained from the imputation method. Using these values, an ML classification model trained on 80,406 samples achieves an accuracy of 85%, while the impute method achieves 75%. The scores and statistics are used to classify samples into five classes: Benign, likely pathogenic, possibly pathogenic, pathogenic, and a variant of uncertain significance. Additionally, a comparative analysis is conducted on 58,444 samples, evaluating six ML techniques. The accuracy obtained by each of these is mentioned alongside the algorithm: logistic regression (89%), k-nearest neighbor (99%), decision tree (95%), random forest (99.8%), support vector machine with the polynomial kernel (91%), support vector machine with RBF kernel (84%), and deep neural networks (98.2%). These results demonstrate the effectiveness of the proposed ML approach for pathogenicity prediction.