
Abstract Bayesian optimization (BO) is widely used for reaction optimization, but under small-data conditions, regression of absolute objective values may not align with the practical goal of prioritizing promising experiments. Here, we compared ranking-based BO (RankBO) combined with Thompson sampling (Rank-TS) with regression-based BO using two reaction benchmarks. Rank-TS showed a clear advantage on Direct Pd-catalyzed arylation in optimization performance, global ranking quality, and recovery of high-yielding conditions. For Suzuki–Miyaura coupling, which comprised 12 substrate combinations, its improvement was small in the pooled analysis. However, in the equal-weight macro analysis across the 12 substrate-defined search spaces, Rank-TS showed higher ranking quality and faster high-yield-condition recovery. These results indicate that Rank-TS is particularly useful for prioritizing reaction conditions within a fixed-substrate combination under small-data conditions, and that its relative performance depends on how the chemical search space and ranking task are defined.
Abstract Accurate prediction of equilibrium partition coefficients of organic compounds in biological and environmental media is critical for evaluating their environmental fate and bioaccumulation potential, as well as for guiding drug design and toxicology studies. Linear solvation energy relationships using Abraham solute descriptors (ASD-LSERs) have achieved remarkable success in characterizing such complex biological and environmental partitioning systems. However, the limited availability of high-quality experimental descriptors, particularly for structurally complex compounds, hinders their practical application and further constrains the accuracy of descriptor predictive models. Here, we propose a linear solvation energy relationship (4SD-LSER) using descriptors derived from logarithmic n-hexadecane–air, n-octanol–air, and water–air partition coefficients, along with the topological McGowan molar volume. To evaluate its performance, 1,849 experimental partition coefficients for 779 neutral compounds across 12 biologically and environmentally relevant systems were compiled. The 4SD-LSER was calibrated and exhibited good descriptive power for these systems. Remarkably, when combined with appropriate fragment-based or machine learning-based descriptor prediction approaches, the 4SD-LSER achieved prediction errors largely within ±0.5 log units for structurally simple compounds and within ±1.0 log unit for more complex compounds (e.g., pesticides, pharmaceuticals, and flame retardants), exhibiting state-of-the-art accuracy, especially for complex compounds. This study demonstrates that models originally developed for well-characterized solvent systems to predict partition coefficients or solvation free energies can be readily extended to biological and environmental systems via LSER. Its performance is poised to improve further with advances in theoretical and machine-learning approaches.
Molecular representations that capture both chemical information and property-relevant differences are essential for reliable molecular property prediction in drug discovery and toxicity assessments. Fragment-level SMILES representation learning is effective at modeling substructure semantics, but structural reconstruction objectives alone do not explicitly encode the descriptor-derived physicochemical information needed for downstream prediction, especially in small data and scaffold-split settings. To address this limitation, we propose PKP-DiffMol, a physicochemical knowledge-prompt encoding and latent diffusion framework built upon the pretrained SMI-EDITOR encoder. PKP-DiffMol improves fragment-level molecular representations by using numerical and semantic physicochemical knowledge. The numerical physicochemical prior branch encodes RDKit descriptors as continuous quantitative priors, while the physicochemical knowledge prompt encoder transforms the same descriptors into semantic physicochemical representations by using a Transformer-based text encoder. The two complementary views are then integrated through hierarchical physicochemical knowledge fusion to construct a chemically enriched, fused latent space. In this space, a quality-controlled conditional latent diffusion module learns label-conditioned distributions, generates synthetic molecular latent representations, and retains reliable samples through a Mahalanobis distance quality gate. Experiments on seven MoleculeNet classification datasets under scaffold splitting show that PKP-DiffMol increases the mean ROC-AUC from 77.80% for the SMI-EDITOR backbone to 80.53% and obtains the best result on five of the seven datasets. The largest gains are observed on BACE (+6.91%), MUV (+4.29%), and SIDER (+3.92%), showing strong predictive performance across bioactivity prediction, virtual screening, and adverse drug reaction prediction tasks. Further analyses indicate that numerical and semantic physicochemical knowledge improves class separability and the correspondence between representation distances and RDKit descriptor distances, while Mahalanobis-filtered label-conditioned latent augmentation supports prediction-head refinement.
Abstract The fast multipole method (FMM) for computing Coulombic forces in molecular dynamics (MD) simulations offers benefits for a variety of systems, including nonperiodic configurations such as molecules in vacuum and large-scale systems where many other approaches encounter limitations. Because FMM parameters govern both simulation accuracy and computational performance, and because the optimal solution is highly system-dependent, we have developed FMM-AutoOpt for the automated benchmarking of FMM parameters with GROMACS. Relying solely on standard Python libraries and GROMACS FMM (Kohnke, B.; et al.J. Chem. Theory Comput.2020, 16, 6938–6949.), the tool is versatile and flexible, facilitating rapid parameter screening to ensure optimal performance and high simulation accuracy.
Abstract Recently, a new category of machine learning approaches for tabular data has emerged: tabular foundation models (TFM), based on in-context learning. A TFM is a neural network (usually a transformer) pretrained primarily on synthetic data. Its input is an entire data set: features and labels for training records, along with features for the target records. This study comprehensively benchmarked these methods for predicting the properties of small organic molecules. Several TFM were compared with multiple machine learning methods across 11 data sets (regression, random and structure-aware splits, up to 10,000 molecules each). The evaluation showed that TFM consistently outperform XGBoost, CatBoost, multilayer perceptrons, and other descriptor-based methods, even with careful hyperparameter selection for the baselines. TFM also demonstrate accuracy on par with or better than graph-based methods, including those pretrained on chemical data. Uni-Mol2, a pretrained deep neural network operating on 3D atomic coordinates, slightly outperforms TFM in some experiments. However, this comparison deliberately disfavors TFM, as they do not use any chemical pretraining and rely on a minimalistic set of 2D molecular descriptors without feature engineering. Some further improvement in results for relatively large data sets and random splits is achieved using retrieval: for each test molecule, the 500 closest neighbors (by Tanimoto similarity) are selected from the training set, and TFM inference is performed on this local subset. Overall, TFM (particularly the TabPFN family) are highly promising for predicting the properties of small molecules.
Abstract G-quadruplexes (G4) are noncanonical nucleic acid secondary structures formed by guanine-rich sequences that can form higher-order multimers under appropriate conditions and represent attractive targets for multivalent molecular recognition. Here, we propose a rational ligand design strategy based on multivalency, in which ligand molecules are covalently linked into polymer-like macromolecules. Using coarse-grained molecular dynamics, we show that these “polyligands” intercalate G4 multimers more efficiently than their monomeric counterparts, achieving strong stabilization through cooperative effects, even when individual ligands exhibit low intrinsic binding affinity. Polyligand intercalation induces structural changes in G4 multimers, modifying their size and stiffness. Using a quantitative polymer theory framework, we connect these structural changes with the thermal stabilization of G4 multimers. Polyligands represent a general design principle for G4-targeting molecular stabilizers, offering enhanced G4 stabilization via programmable linker length and backbone rigidity.
Abstract Predicting the impact of single-point mutations on protein thermodynamic stability is crucial for protein engineering of therapeutic and industrial applications. By effectively capturing the three-dimensional structural information of proteins and the spatial physical environment of each residue, the inverse folding models (IFMs) upon fine-tuning, such as ThermoMPNN, achieved state-of-the-art performance in predicting thermostability changes in proteins caused by mutations. However, IFMs are limited in their capacity to capture protein deep evolutionary information, whereas protein language models (pLMs) excel. Here, we present SA-MPNN, a lightweight, end-to-end hybrid framework that dynamically integrates the protein sequence representations from a protein language model (ESM2) into the ThermoMPNN architecture to improve protein stability prediction by combining evolutionary representations with geometric structural embeddings. By evaluating various feature fusion strategies, we selected a self-attention-based integration mechanism to effectively combine the two modalities. Trained on the large-scale Megascale data set, SA-MPNN achieved modest but consistent gains over ThermoMPNN on various benchmark data sets, with particularly noticeable improvements in several correlation analysis and screening-oriented evaluations. Finally, wet-lab validation was performed on the top-ranking variants of Acetivibrio thermocellusβ-glucosidase (AtBgl1A) as a case study. The experimental results demonstrated that multiple designed mutants exhibited improved thermostability, and the optimal variant, GC20, achieved a melting temperature (Tm) of 76.98 °C, representing a 5.97 °C increase over the wild-type, thereby supporting the practical applicability of SA-MPNN in protein engineering.
Abstract Ephrin receptors (Ephs) are receptor tyrosine kinases that regulate cellular growth, differentiation, and motility. EphA2, often overexpressed in cancer, is notable for its ligand-independent activation, which drives pro-oncogenic signaling distinct from the canonical, ligand-dependent pathway that restricts cell movement. While ligand binding induces extracellular clustering, kinase activation depends on dimerization within the transmembrane (TM) region. EphA1 and EphA2 differ substantially in their function, likely due to differences in their TM, but more so in their juxtamembrane (JM), and membrane-proximal fibronectin type III (FN1/FN2) domains. How these latter two regions modulate dimerization has not yet been investigated. To address this, we performed extensive coarse-grained (CG) simulations using Martini3 in an anionic POPC/PS/PIP2 model of the plasma membrane. Both receptors formed stable TM dimers, though EphA1 favored a symmetric AXXXGXXXG-centered interface, whereas EphA2 favored an alternate leucine zipper interface. All protein constructs sampled multiple configurations, reflecting substantial intrinsic variability. AlphaFold3 does not yield reliable predictions and cannot account for the effects of the lipid bilayer. In the CG simulations, basic residues in the JM region remained membrane-bound, and the EphA2 FN domain displayed sustained PIP2 interactions, consistent with previous observations. Notably, constructs with the FN2 domain alone (a fragment relevant to Alzheimer's Disease) restricted TM association in both receptors, whereas inclusion of the second FN1 domain restored dimerization but produced receptor-specific extracellular interfaces. These differences arise from distinct FN1-FN2 linker flexibilities, a different level of sequence conservation of EphA1 compared to several other Ephs and FN-domain membrane contacts, which together shape TM geometry and lipid engagement. Our results predict how TM and TM-proximal elements cooperatively tune dimerization in Eph receptors. This work offers a molecular rationale for the divergent activation behaviors of EphA1 and EphA2 and provides testable models relevant to cancer and neurodegeneration biology as well as Eph-driven signaling.
Abstract Protein–protein interactions (PPIs) regulate essential cellular processes and represent an important class of therapeutic targets; however, discovering effective modulators of PPIs remains a formidable challenge. Although deep learning approaches have been widely explored for PPI modulator discovery, many rely on simplified representations that obscure interchain boundary information and fine-grained PPI–modulator interaction (PPIMI) patterns, limiting their robustness under distribution shifts. To address this challenge, we introduce TvTPPIMI, a framework that leverages learnable boundary tokens to encode partner-aware boundary information and models PPIMI at atom–residue resolution. In a case study targeting the AURKA–TPX2 interaction, TvTPPIMI prioritized putative modulatory candidates from a small-molecule screening library. Structure-based docking, attention analysis, multireplica molecular dynamics simulations, MM/GBSA binding free-energy estimation, and noncovalent interaction analyses provided post hoc physical support for the stable AURKA binding of selected candidates, highlighting CE02–6266 as the most favorable compound among the tested hits. Together, these results suggest that TvTPPIMI provides a generalizable computational framework with coarse-grained, attention-based interpretive cues for PPIMI prediction and can be integrated with structure- and dynamics-based analyses to support PPI modulator discovery.
Abstract Specific recognition between T-cell receptors (TCRs) and peptide–major histocompatibility complexes (pMHCs) is central to adaptive immunity, yet accurate prediction of TCR–pMHC specificity remains challenging. Existing models mainly rely on sequence features or isolated molecular structures, limiting their ability to capture interface-level determinants within the ternary recognition complex. Here, we constructed the multimodal TCR–pMHC ternary complex (MM-TCR) data set, integrating paired TCR–pMHC sequences, V/J gene annotations, and modeled TCR–pMHC complex structures refined by short molecular dynamics-based relaxation. Based on MM-TCR, we developed TCRspec, an interpretable multimodal framework combining sequence embeddings, gene-usage features, and complex-level structural representations. Under a stringent CD-HIT TCR-cluster-disjoint split, TCRspec achieved an average AUROC of 0.896 and AUPRC of 0.882 across seven antigen-specific test data sets, outperforming representative baseline models. Cross-validation and ablation analyses confirmed the contribution of ternary complex structural information and MD-refined structures. In independent OOD peptide–TCR systems, TCRspec retained discriminative performance and identified model-inferred peptide positions associated with TCR recognition, providing a structure-informed framework for TCR specificity prediction.
Abstract This review addresses a critical need in the emerging intersection of digital chemistry and material properties by critically evaluating the current capabilities and future prospects of applying machine learning (ML) to ionic liquids (ILs) for thermal energy storage (TES) applications. We examine available data sets for key thermophysical properties relevant to TES, including thermal conductivity, heat capacity, density, viscosity, surface tension, melting point, enthalpy, CO2 solubility, and toxicity. Particular emphasis is placed on chemical space coverage, biases, and practical limitations. We compare different modeling approaches, spanning from conventional QSPR/QSAR and group-contribution methods to contemporary ML approaches, including kernel models, ensemble methods, graph neural networks (GNNs), and deep learning architectures. The impact of molecular representations and descriptor choices and validation strategies in determining model robustness and transferability is examined. To structure this assessment, we introduce a three-tier property-selection framework for IL-based TES and apply a six-criterion evaluation checklist consistently across all nine reviewed properties. The analysis is supported by literature-summary tables that consolidate data sets, chemical space coverage, modeling approaches, and validation practices across different property domains. Finally, we outline recommendations for data set standardization, open sharing of data and code, and the establishment of benchmark validation protocols to improve reproducibility and accelerate the discovery of ILs for TES applications.
Abstract Molecular property prediction is a key task in AI-driven drug discovery, yet the prevalence and impact of label imbalances in molecular property regression remain poorly understood. Through a systematic benchmark of widely used molecular property data sets, we show that target values are often highly imbalanced and that prediction errors are consistently concentrated in sparsely represented regions of the label space. We further evaluate representative imbalance-learning approaches developed for general regression tasks and find that their effectiveness on molecular data sets is limited. To address this challenge, we propose interval-aware mixture-of-experts (IA-MoE), a plug-in framework that partitions the continuous target space into intervals and promotes expert specialization across different regions of the label distribution. IA-MoE can be seamlessly integrated with diverse molecular encoders without modifying their underlying architectures. Experiments on five molecular property benchmarks and four graph neural network backbones demonstrate that IA-MoE consistently improves the predictive performance in underrepresented regions while maintaining or improving overall accuracy. Our findings establish label imbalance as an important challenge in molecular property regression and highlight interval-aware expert learning as an effective strategy for addressing this challenge.
Abstract Accurate prediction of protein–ligand binding poses is essential for structure-based drug discovery. In recent years, computational approaches, particularly molecular docking, have generated large numbers of candidate ligand poses. However, existing scoring functions remain limited in their ability to distinguish near-native poses from incorrect docking decoys. In this work, we propose UniPoseScore, a Graphormer3D-based framework for ligand pose scoring and refinement. The model simultaneously predicts pose RMSD and atomic displacement vectors to guide ligand pose refinement. Across four benchmark data sets, UniPoseScore showed consistently strong performance and better generalization, with a particularly clear advantage in RMSD correlation. Visualization of the learned embeddings further suggests that the model captures structural features correlated with pose accuracy. On the CASF-2016 benchmark, UniPoseScore achieves a refinement success rate of 74.3%, highlighting the effectiveness of UniPoseScore in improving ligand binding poses and its potential applications in drug discovery. The UniPoseScore is available at https://github.com/April-Zhangjq/UniPoseScore.
Abstract Fub7 is a rare pyridoxal 5′-phosphate (PLP)-dependent enzyme that can subsequently catalyze the γ-elimination of O-acetyl-l-homoserine (OAH) and Michael addition with n-valeraldehyde (NVA) to synthesize 5-alkyl-pipecolic acid, which shows significant potential in synthetic chemistry. In this study, we constructed the computational models and performed MD and QM/MM calculations to illuminate the complete catalytic cycle of Fub7, which contains five reaction stages, including external aldimine formation, γ-elimination, Michael addition, reformation of the internal aldimine, and final cyclization and dehydration, of which γ-elimination and Michael addition correspond to relatively high energy barriers. In addition, according to our test calculations, Fub7 can also catalyze the β-elimination. During the catalysis, the terminal amino of Lys211 exhibits significant conformational changes and was found to play multiple roles. On one hand, Lys211 forms internal aldimine with the cofactor PLP. On the other hand, Lys211 functions as a catalytic base/acid or as a bridge to mediate a series of proton transfer in the γ-elimination and Michael addition. Tyr109 only participates in the Michael addition. The deprotonated Tyr109 acts as a base to abstract the α-proton of NVA to generate the carbanion nucleophile, which is the key for Michael addition. These findings may provide useful information for understanding the catalysis of PLP-dependent enzymes and designing biocatalysts for γ-substitution.
Abstract Mimotopes selected by phage display have been used for antibody epitope mapping for nearly four decades, but residue-level recovery of functional epitopes has remained unreliable. This is not a problem of algorithms but of assumptions: mimotopes recapitulate native epitope recognition through compensatory interaction networks that preserve binding energetics, without conserving sequence or surface geometry. Predictors trained on sequence similarity or static interface geometry therefore miss the residues that drive recognition. We present a physics-based workflow that refines antibody-bound mimotope conformations by microsecond-scale molecular dynamics and maps the resulting structures onto the cognate antigen using Folddisco. Across four antibody–antigen systems with crystallographically defined epitopes, the workflow reaches 60 to 100% precision in epitope localization and recovers 100% of mutagenesis- or structurally validated functional hotspots, compared with at most 25% for three widely used benchmarks (EpiSearch, ClusPro, SEPPA-mAb). For the clinically relevant ChiLob 7/4 system, it further identifies a minimal tetrapeptide (PWVP) sufficient for antibody recognition, validated by Western blot and immunoprecipitation. By grounding epitope prediction in experimentally selected binding information and energy-resolved conformational refinement, the workflow connects phage display to mechanistic structural insight.
Abstract The detailed mechanism of halohydrin dehalogenase (HHDH)-catalyzed ring-expansion reaction between spiro-epoxides and the nucleophile OCN– to form spiro-oxazolidinones was investigated using molecular docking, molecular dynamics (MD) simulations, and quantum mechanics/molecular mechanics (QM/MM) calculations. Molecular docking and molecular mechanics/Poisson–Boltzmann surface area (MM-PBSA) results revealed that the R-configurational substrate exhibited significantly superior binding affinity compared to the S-configurational one, validating the stable binding mode within the enzyme′s active pocket. The fundamental pathway initiates with the nucleophilic attack of OCN– accompanied by proton transfer, which is followed by ring closure coupled with a proton transfer to form a five-membered ring. A total of four possible selective nucleophilic attack pathways were evaluated, and the most energetically favorable pathway associated with the R-configurational product has an energy barrier of 17.1 kcal/mol, which is close to the experimental value of 18.1 kcal/mol derived from kinetic parameters. This study provides valuable theoretical insights for the future engineering of HHDHs for chiral spiro-oxazolidinone synthesis.
Abstract Domain-specific language models are attractive because they can be tailored to suit their users. However, the high-end computational resources needed to pretrain large language models (LLMs) by conventional methods tend to prohibit their development. Furthermore, there is a dearth of large, domain-specific, question-answering (QA) data sets to finetune LLMs for prompt engineering, even when computational resources can be found to pretrain LLMs. Our paper presents a method for autogenerating a data set of 168,080 domain-specific QA pairs relating to magnetic materials. We show how this can be employed to finetune LLMs for the downstream task of Question-Answering (QA) for the magnetic materials domain. We use the bidirectional encoder representations from transformers (BERT) architecture as our benchmark LLM to compare the relative performance of 6 LLMs. These include 3 vanilla (BERT-base-cased) LLMs that have been finetuned on various QA data sets (i) our domain-specific QA data set (MagQA); (ii) the Stanford Question-Answering Data set (SQuAD v2); (iii) our QA data set mixed with SQuAD v2. The other 3 (BERT-base-cased) LLMs have been subjected to domain-adaptive pretraining on a corpus of 97,308 magnetic papers prior to their finetuning on the same combination of QA data sets (i)-(iii). As well as our QA data set autogeneration pipeline, we show how our work can be extrapolated to the larger architectures of RoBERTa-base and DeBERTa-base. We also undertake an ablation study of the finetuning process using four sizes of QA data sets to gain insights into the quantity of data that are needed for sufficient domain-specification. The highest performing LLM was MagBERT_MagQA_Mixed, a vanilla BERT-base-cased LLM that had been finetuned on the combination of our autogenerated MagQA data set and SQuAD v2, achieving an F1 score of 78.43% and an exact-match score of 72.84% when tested on a manually annotated data set on magnetic materials. These results demonstrate that domain-specific BERT models only need to be finetuned from vanilla BERT models (i.e. no domain-adaptive pretraining is needed), pending the availability of sufficiently large, high-quality, domain-specific QA data sets that this work shows how to autogenerate.
Abstract An accurate description of halogen bonding in biomolecules remains a challenge in modern biochemistry, as exemplified by the flavin-dependent human dehalogenase, iodotyrosine deiodinase. Analyses of the static crystal structure and kinetics data suggested that halogen bonding is absent in the enzyme’s active site, which does not preclude its transient formation along the binding or unbinding pathway. No studies have yet explored iodotyrosine deiodinase’s dynamic behavior from the perspective of halogen bonding. Here, extensive equilibrium and nonequilibrium molecular dynamics (MD) simulations together with quantum-mechanical calculations are performed to elucidate how the cooperative interplay of transient halogen bonds and stable hydrogen bonds governs substrate recognition and retention. An extra-point charge model is employed to describe σ-hole of the halogen to capture its charge anisotropy in classical simulation. The isotropic steered molecular dynamics simulations and dynamic time warping algorithm identify dominant unbinding pathways for I-Tyr, Cl-Tyr, and I-Phenol. Substrate release kinetics is controlled by a stepwise disruption of key hydrogen bonds and salt bridges with K182, Y161, E157, and the flavin cofactor. Molecular dynamics simulation-derived geometries were screened, representative poses were chosen based on structural criteria, and halogen-bond-like interactions were validated through natural bond orbital and energy-decomposition analyses. Geometries consistent with halogen-bond-like interactions (0.5–2.0 kcal/mol) are observed for I-Tyr and I-Phenol near the protein surface. In contrast, Cl-Tyr interactions identified solely by structural criteria largely correspond to false positives, underscoring the necessity of quantum-mechanical validation for reliably characterizing halogen-bond-like interactions. These early, short-lived halogen bonds help orient the native substrate toward a possible alternate site, triggering a communication network that grows with strong hydrogen bonds during substrate recognition. By integrating dynamic simulations with quantum-mechanical validation, this work reconciles conflicting views on halogen bonding in human iodotyrosine deiodinase, revealing transient geometries consistent with halogen-bond-like interactions alongside persistent hydrogen bonds.
Abstract Atropisomerism, a form of chirality arising from restricted rotation around single bonds, holds significant pharmacological importance. Yet, the automated identification and representation of atropisomerism in molecular databases and cheminformatics tools remain limited. Identifying bonds that lead to stable atropisomers typically involves labor-intensive experiments or costly quantum chemical scans. In this work, we present an extension of the molecule identifier MolBar that automatically classifies each sp2–sp2 bond in a three-dimensional molecular structure as either rapidly convertible or configurationally stable (rotational barrier ≥ 20 kcal/mol), specifying the absolute configuration for the latter. Our methodology uses two distinct classifiers: a physically motivated classifier for steric hindrance and a message-passing graph neural network for electronic effects. Integrated into MolBar, this module enables unsupervised, large-scale annotation of atropisomeric bonds along with their absolute configurations. This advancement facilitates the representation of atropisomerism in databases and supports atropisomerism-aware machine learning in drug discovery. A new version of MolBar (1.2.0) is available on the Python Package Index (PyPI) and can be installed using the pip command: pip install molbar.
Traditional carcinogenicity assessment relies on animal experiments, which are costly, time-consuming, and difficult to extrapolate across species. Consequently, the carcinogenic potential of most chemical pollutants remains uncharacterized, posing significant risks to human health. Here, we combined 2356 compounds from public databases with six molecular fingerprints, five base machine learning models, and ensemble strategies to develop PolluCarc-MFSE (a multifingerprint stacking ensemble learning model using Klekota-Roth, extended-connectivity fingerprint (ECFP), and MACCS (Molecular ACCess System) fingerprints) as the optimal predictive model. This model was then applied to predict the carcinogenicity of 179 petrochemical pollutants with unknown carcinogenic potential, as listed in the emission standard of pollutants for the petroleum chemistry industry (GB 31571-2015, 2024 amendment). Among 213 carcinogenic compounds (130 predicted by the model and 83 identified via dataset intersection), the major types were high-ring-number polycyclic aromatic hydrocarbons, halogenated unsaturated hydrocarbons, and multisubstituted aromatics. To assess biological relevance, disease-associated targets for ten pollutant-linked cancers were retrieved from GeneCards. Using a novel multimetric scoring approach in network analysis, we identified 29 hub targets (e.g., JAK2, EGFR, TP53, MYC) that exhibit stable binding to the predicted carcinogens. Pathway enrichment confirmed involvement in cancer-related, PI3K-Akt, and JAK-STAT signaling pathways, while Disease Ontology linked these targets to liver, breast, and lung malignancies. This integrated "data-model-mechanism" framework provides a practical tool for prioritizing high-risk chemical pollutants and demonstrates how computational toxicology can support regulatory decision-making and the adoption of New Approach Methodologies in chemical risk assessment.