
Traditionally recognized for guiding rRNA modifications, small nucleolar RNAs (snoRNAs) are increasingly appreciated as key regulators of drug response. However, snoRNA and drug association data remain limited, and computational approaches capable of capturing their complex, high-order relationships are scarce. Here, we present HMHLVI, a hybrid multi-view hypergraph learning with variational inference framework that integrates sequence, structural, and association information to predict snoRNA-drug response associations. HMHLVI constructs three higher-order networks called attribute, topology, and association to model intrinsic features, structural dependencies, and known interactions, respectively. By employing shared and modality-specific hypergraph convolutional encoders together with an adaptive temperature-regulated multi-view attention mechanism, the framework effectively learns both common and view-specific biomolecular representations. In addition, a variational autoencoder is introduced to model the higher-order association network and to capture latent interaction patterns in a low-dimensional probabilistic space, enhancing robustness and denoising capability. Across three cross-validation settings, including random zero, multi-column zero, and multi-row zero, HMHLVI consistently outperformed other state-of-the-art models. System-level validations including functional enrichment, molecular docking, thermodynamic analysis, and clinical survival assessment confirmed its biological relevance. Key regulatory snoRNAs such as SNORD43 and SNORD116 were identified, and an integrated network linking drugs, diseases, snoRNAs, and target genes was constructed. Notably, SCARNA6 was predicted to modulate docetaxel response in esophageal squamous cell carcinoma, and higher SCARNA6 expression was associated with favorable overall survival in an exploratory Kaplan-Meier analysis (log-rank p = 0.0029).
Early detection of lung cancer using low-dose computed tomography (LDCT) is critical for improving patient outcomes. In this study, we propose a RepViT-based lightweight multimodal fusion network (Rep-LMFNet) framework, which integrates LDCT imaging and structured clinical information for risk prediction. The framework incorporates adaptive multi-scale module (ASMM), dynamic channel recalibration module (DCRM), adaptive feature enhancement module (AFEM), and hierarchical multimodal fusion module (HMFM) to improve feature representation and multimodal interaction. Experimental results under stratified five-fold cross-validation demonstrate that the proposed framework achieves an area under the receiver operating characteristic curve (AUC) of 0.877 (95
Recent advances in single-cell immune profiling enable simultaneous measurement of T cell receptors (TCRs) and transcriptomes, offering unprecedented opportunities to characterize adaptive immunity at single-cell resolution. However, the substantial heterogeneity between these modalities poses challenges for effective integration, and existing approaches often rely on simplistic alignment assumptions that may overlook modality-specific biological signals. To address this challenge, we introduce TransTCR, a multimodal learning framework that integrates optimal transport (OT) with contrastive learning to align TCR and transcriptomic representations. Specifically, TransTCR first extracts informative features from each modality using pretrained foundation models, then performs OT-based projection into a shared latent space to achieve distribution-aware alignment. A bidirectional contrastive objective further refines instance-level correspondence by maximizing agreement between the paired TCR–RNA profiles. Extensive experiments demonstrate that TransTCR substantially outperforms existing single-modality and multimodal baselines across antigen specificity recognition tasks, achieving state-of-the-art performance in intra-donor antigen specificity prediction, inter-donor antigen specificity prediction, and clustering evaluations. Overall, TransTCR provides a powerful computational tool for integrating and analyzing multimodal T cell data, facilitating deeper insight into adaptive immune responses.
Uncovering genes that drive cancer is fundamental to elucidating the mechanisms underlying cancer development and to advancing cancer research. Recent years have witnessed the emergence of cancer driver gene identification from multi-omics data as a key research area, facilitated by the rapid progress of high-throughput molecular technologies. Although numerous algorithms have been proposed for cancer driver gene discovery, the precise identification of these genes continues to pose a challenge owing to the lack of labeled data. This study presents SDMGAE, a Self-supervised Dual Masked Graph AutoEncoder-based method for cancer driver gene identification. This framework integrates two components: a self-supervised graph learning module and a driver gene prediction module. During the self-supervised graph learning phase, nodes and edges of protein–protein interaction (PPI) networks are masked separately to consider both node and structural information. Subsequently, the graph autoencoder is employed to reconstruct the PPI network without using labelled data. In the driver gene prediction stage, we employ the pre-trained graph neural network encoder to obtain the embeddings, which are then processed through the logistic regression to generate prediction outcomes. To evaluate the effectiveness of SDMGAE, we performed benchmarking experiments across 10 distinct types of cancer data. Experimental outcomes reveal that SDMGAE exhibits improved performance in cancer driver gene detection compared with state-of-the-art methods.
Staphylococcus aureus (S. aureus) is a well-recognized pathogen known for its multi-drug resistance and diverse virulence mechanisms. Its ability to grow biofilms on implanted medical devices enhances its antimicrobial resistance (AMR) and virulence. Despite its clinical relevance, the underlying genetic basis of S. aureus biofilm formation remains insufficiently characterized, particularly regarding key biofilm-associated genes (BAGs) and their regulatory contributions. This study presents a two-part integrative approach to identify genetic determinants of biofilm formation in 178 Egyptian, clinical S. aureus isolates. The framework integrates a genome-wide association study (GWAS) module with a learning-based classification module. GWAS was conducted using a linear mixed model, while logistic regression was the best-performing model in binary and multiclass classification. Integrating both modules, we identified 20 BAGs as promising determinants of biofilm formation. Protein–protein interaction network and pathway enrichment analyses revealed their involvement in biofilm-related pathways. Of the identified BAGs, nine genes have direct links to biofilm formation in S. aureus or other bacteria, while the rest are linked to AMR, nutrient acquisition, and cell division. This study presents a robust framework for biofilm genomics research, uncovering 20 candidate BAGs that span diverse biological functions and capture the multi-faceted nature of biofilm formation in S. aureus.
Accurate segmentation of medical volumes in magnetic resonance imaging is essential for the exploration of organ structures. Despite the impressive performance of supervised learning in medical image segmentation, its reliance on the large availability of labeled datasets poses a challenge due to the expertise and effort required for data acquisition. Semi-supervised learning (SSL) methods extract additional information from unlabeled datasets without involvement of expert annotation, and many approaches have been proposed to assist with this task. However, they often yield satisfactory results in the central regions but do not perform well in the edge regions due to low contrast edges and ambiguous tissue intensities at the boundary. Meanwhile, standard loss functions are biased towards large regions and cause under-penalizing of boundary errors. In this paper, we propose a novel framework that effectively leverages unlabeled data to improve segmentation performance in cardiac structures. Firstly, we design a mutual learning module with multiple decoders to obtain different predicted probabilities through feature perturbations. Secondly, we apply a novel consistency constraint between labeled and unlabeled data by a dual fine-grained boundary loss that provide global characteristics-based guidance from the transition of the boundary region and an edge-aware uncertainty loss. We evaluate the proposed framework on two publicly available cardiac datasets, including the automated cardiac diagnosis challenge (ACDC) dataset under a 2D setting and the left atrium (LA) dataset under a 3D setting, and compare it with seven recent state-of-the-art semi-supervised methods. Through extensive qualitative visualization and quantitative experiments under standard semi-supervised settings, we demonstrate the effectiveness of the proposed approach and its superior segmentation performance compared with existing methods. Furthermore, clinically relevant cardiac measurements derived from the segmentation results are evaluated to analyze the clinical relevance of the proposed framework. To support reproducibility and future research, our code is publicly available at: https://github.com/waqasanwaar/SSL4MIS_PMEUL .
Drug–drug interactions (DDIs) can compromise therapeutic efficacy and patient safety, making accurate computational prediction highly important in drug discovery and clinical decision support. We propose Multi-GraphDDI, a structure-only framework that predicts DDIs without relying on external biological networks. In this model, three complementary molecular fingerprints, namely extended-connectivity fingerprints (ECFP4), PubChem fingerprints, and pharmacophore fingerprints, are encoded as three grayscale channels and fused into a single image representation, while a parallel branch transforms the two-dimensional molecular graph into a topology-aware embedding through a five-layer residual graph isomorphism network (GIN). A bidirectional feature-interaction module together with four-head cross-attention is then used to align the image-based and graph-based representations, and the fused features are further used to estimate interaction scores. On ChCh-Miner, ZhangDDI, and DeepDDI, Multi-GraphDDI achieved AUC/AUPR/F1 scores of 0.9986/0.9998/0.9730, 0.9858/0.9633/0.8855, and 0.9922/0.9920/0.9598, respectively, outperforming competing methods. These results indicate that integrating heterogeneous structural cues through coarse- and fine-grained feature interaction provides an effective and scalable solution for DDI prediction.
Combination drug therapy is an effective approach to combating drug resistance and enhancing therapeutic efficacy in complex diseases such as cancer. Nevertheless, discovering synergistic drug pairs remains difficult because of the enormous combinatorial possibilities and the context-dependent, nonlinear interactions among drugs and cellular systems. Although recent computational methods have advanced drug synergy prediction, they often fail to preserve pharmacologically meaningful molecular substructures, capture global properties of complex molecular graphs, or model the varying importance of genes across cellular contexts. In this study, we propose GATESynergy, a novel deep learning framework for predicting drug synergy that combines a hierarchical gated gene-aware encoder (HiGate) with a molecular global–local aggregator (MoGLA). Specifically, MoGLA integrates local graph message-passing with global self-attention to jointly model functional motifs and long-range structural dependencies, overcoming the inability of conventional GNNs to simultaneously preserve pharmacologically meaningful substructures and capture global molecular topology. HiGate employs a residual gene-wise gating mechanism that adaptively weights genes according to their contextual relevance, yielding context-specific cell-line representations that address the neglect of gene-level importance in existing encoders. Furthermore, we propose a multi-head interactive additive attention module that uses global query summarization to efficiently fuse drug-drug-cell line representations and capture diverse synergistic interaction patterns. Extensive benchmarking results demonstrate that GATESynergy consistently surpasses current state-of-the-art methods. Moreover, a case study on 42 FDA-approved drugs further validates the effectiveness of GATESynergy in discovering novel synergistic drug combinations. The source data and code are available at https://github.com/coding-in-github/GATESynergy .
Breast cancer remains a major global health burden, underscoring the urgent need for reliable early detection strategies. Exosomes, as mediators of intercellular communication, have shown promise in early tumor screening through Raman spectroscopy and gene expression profiling in pancreatic and colorectal cancers. However, the application of exosomal gene expression profiles for breast cancer prediction remains largely unexplored. Exosomal mRNA profiles were obtained from exoRBase 3.0 (242 breast cancer, 244 healthy controls). Sample sex was inferred using XIST and UTY expression, yielding 337 female samples for analysis. A nested cross-validation framework (20 repetitions, fivefold) was implemented, with differential expression analysis and feature selection performed exclusively within each training fold to prevent information leakage. Ten machine learning classifiers were evaluated on an independent held-out test set. Model performance was assessed using accuracy, precision, recall, and F1-score. Feature selection demonstrated high stability (average Jaccard score 0.5912), with 9 genes consistently selected across all 100 iterations and a set of robust feature genes was identified. Among classifiers, xgbTree achieved the best performance (AUC 0.992, accuracy 0.970, F1 0.979) on the independent test set, supporting exosomal mRNA profiles as a promising non-invasive approach for the early breast cancer detection.
PIWI-interacting RNAs (piRNAs) are an important class of non-coding RNA molecules in epigenetic regulation. It plays a crucial role in maintaining genomic stability and inhibiting transposable elements, and have been proven to participate in various diseases by regulating gene expression and influencing signaling pathways. Traditional biological experimental methods have limitations such as low throughput, long cycles, and high costs, making them difficult to meet the requirements of large-scale systematic screening. In this study, we develop a predictive framework named PiDA-DVLSA. We integrate autoencoder, dual graph transformer, and multi-head self-attention mechanisms, and construct an end-to-end multimodal deep learning system. We use autoencoder to perform nonlinear dimensionality reduction and denoising on piRNA sequence features and disease phenotype semantic features, and extract potential representations with strong discriminative ability. Then, we use graph transformers to model the high-order topological relationships between nodes in isomorphic similar graphs, and input heterogeneous graph transformers to learn complex cross-entity interaction patterns in heterogeneous networks. Finally, we achieve adaptive fusion of multi-source information through multi-head self-attention mechanisms. PiDA-DVLSA performs excellently on the benchmark dataset, with AUC and AUPR reach 0.9437 and 0.9195, respectively, significantly outperform eight mainstream algorithms. In independent case validations for breast cancer, clioblastoma, and Alzheimer disease, our model successfully predicts multiple biologically significant potential associations, further confirming its practicality and effectiveness in real scientific research scenarios and providing a solid computational basis for future precision diagnostic and therapeutic applications. PiDA-DVLSA is freely available at https://github.com/zhaoqi106/PiDA-DVLSA .
Drug-target binding affinity (DTA) is central to computer-aided drug design. Although biochemical assays yield accurate measurements, their high cost and inefficiency limit scalability. Computational approaches, particularly deep learning, have recently shown promising progress in accelerating DTA prediction and reducing experimental burden. However, existing methods rarely incorporate biochemical knowledge, which limits their performance. To address this issue, we propose a deep learning framework termed DTANet+ for drug-target affinity prediction, which is inspired by the biochemical properties of drugs and targets. Specifically, given that functional groups and peptide chains, which vary significantly in length, are crucial in determining drug and target properties, we propose a Kernel-diverse Feature Extraction Block (KFEB) and Cross-Scale Interaction Module (CSIM) to extract drug and target features. Inspired by the crucial role of drug-target binding sites in precise affinity prediction, we introduce the Drug-Target Interaction Module (DTIM) to explore drug-target relationships and focus on binding sites. Besides, we design a Multi-modal Fusion Module (MFM) to effectively integrate the complementary information across different modalities and conduct prediction. Extensive experiments on the KIBA and Davis datasets show that DTANet+ consistently outperforms existing methods, achieving higher Concordance Index (CI) and lower Mean Squared Error (MSE). These results confirm the effectiveness of DTANet+ and highlight its potential for drug repositioning and personalized medicine. The source code is publicly available at: https://github.com/202324131016T/DTANet_plus .
With the rapid development of artificial intelligence and medical image analysis, MRI-based automated diagnosis has provided an effective approach for Alzheimer's disease (AD) assessment. To improve the performance of MRI-based AD classification, this study proposes an AD diagnosis model termed High-level CBAM-ResNet34. First, T1-weighted structural MRI data are preprocessed using a unified pipeline and further converted into two-dimensional slices for model training. Then, a ResNet34-based classification framework is constructed, in which the convolutional block attention module (CBAM) is introduced into the high-level feature stage to enhance discriminative feature representation. In addition, to better adapt to inter-dataset differences, the negative-class weight parameter is selected according to the empirical performance on each dataset during training. Experimental results on the ADNI dataset show that the proposed model achieved an AUC of 0.8757 and an accuracy of 0.8160, outperforming several representative comparison models in overall performance. Ablation and comparison experiments further verified the effectiveness of the proposed design, and external validation on the open access series of imaging studies 1 (OASIS 1) dataset demonstrated its generalization ability. These results indicate that the proposed model is effective for MRI-based AD diagnosis and provides a useful reference for computer-aided neuroimaging analysis.
Predicting drug-target affinity (DTA) plays a pivotal role in drug discovery and repurposing. While existing computational approaches predominantly rely on 1D sequences or 2D structural data, they often fail to fully capture the intricate nature of molecular interactions. To address this limitation, we propose Deep3D-DTA, a novel tri-modal deep learning framework that integrates 1D sequence semantics, 2D graph topology, and 3D spatial geometry complementary representations for both drugs and target proteins. The proposed architecture offers three key advancements: First, it employs a pre-trained protein language model to encode amino acid sequences, effectively capturing long-range sequential dependencies. Second, it constructs precise 3D structural representations by computing interatomic distances and bond angles, enabling accurate modeling of the spatial conformations of both proteins and compounds. Third, it leverages a hybrid feature extraction module that combines graph neural networks with multi-head attention mechanisms to learn hierarchical structural patterns. Extensive experiments on three widely used benchmark datasets (Davis, KIBA, and Metz) demonstrate that Deep3D-DTA significantly outperforms state-of-the-art methods in DTA prediction. These results highlight its potential as a robust and reliable computational tool for accelerating drug discovery and reducing development costs through more accurate affinity prediction.
Recent advances in spatial transcriptomics (ST) have enabled the extraction of gene expression patterns while retaining spatial context. Identifying spatial domains is crucial for ST research. However, most existing spatial domain recognition methods cannot capture more complex relationships between gene expression profiles and spatial information. To bridge this gap, we propose a novel self-supervised learning framework named STNMAE for identifying spatial domains using a neighbor-aware multi-view masked graph autoencoder. Specifically, to fully exploit the dependencies between local neighbor information and globally similar expressions, we first construct multiple neighbor views with distinct similarity measures based on the gene expression profiles and spatial information. Additionally, we utilize a feature-masked encoder to extract more expressive embeddings. Then, STNMAE learns multiple view-unique embeddings through a multi-view autoencoder. Furthermore, the framework also uses regularization techniques through the latent representation prediction module to avoid overfitting and reduce the direct effect of input features. We apply STNMAE on seven ST datasets with different resolutions across distinct platforms. Finally, extensive evaluation confirms STNMAE's superiority over current state-of-the-art methods, indicating a substantial improvement in ST data analysis.
Integrating single-cell RNA sequencing (scRNA-seq) with spatial transcriptomics (ST) enables the projection of cell-type-resolved transcriptional programs onto tissue architecture. However, existing integration methods are often unstable because spot-level inference is performed directly in high-dimensional gene space, where extreme sparsity, measurement noise, and strong multicollinearity among marker genes amplify the estimation variance. As a result, inferred cell type proportions may be dominated by a small subset of genes, making them highly sensitive to noise and systematically distorting rare or low-abundance cell types. Here, we present ST-LDAW, which is a computational framework explicitly designed to address these challenges. ST-LDAW combines probabilistic topic modeling with damped weighted least squares optimization to enhance robustness at both the representation and inference levels. Topic-based modeling reduces dimensionality and mitigates gene-level noise by capturing coherent transcriptional programs, whereas damped weighting constrains the influence of unstable or low-confidence features, preventing variance inflation and overfitting during deconvolution. Benchmarking of simulated spatial mixtures demonstrated that ST-LDAW achieved a recall rate of 94% and an accuracy of 80%, surpassing existing regression-based and mapping-based methods in terms of sensitivity and precision. These results highlight ST-LDAW's ability to reliably identify cell types in complex, sparse datasets and its robust performance in handling rare or low-abundance cell types. Application to breast cancer ST data further reveals the subtype-specific cellular composition, functional heterogeneity, intercellular communication patterns, and key epithelial hub genes.
Lysine crotonylation (Kcr), as an emerging post-translational modification, plays a crucial role in core life activities such as chromatin dynamics and gene expression. To address the current limitations of Kcr site detection techniques, including high experimental costs, complex procedures, and high false-positive rates, as well as the poor generalization performance of existing computational models caused by limited training data and class imbalance, this study proposes an innovative intelligent recognition framework named MVFAN-Kcr. The system integrates multi-view feature fusion and attention mechanisms to synergistically enhance the prediction accuracy and robustness of Kcr site identification. Physicochemical property features are combined with global sequence semantic information derived from the ESM-2 protein language model to construct a fused feature representation that captures both local physicochemical information and global contextual information. To optimize computational efficiency, feature selection is performed using analysis of variance. To effectively address the problem of imbalanced data distribution, a stratified undersampling strategy based on the chi-square test is developed. In addition, a convolutional neural network combined with attention is designed to efficiently extract local sequence patterns and enhance the representation of key features. Under a rigorous evaluation framework based on five-fold cross-validation and an independent test set, MVFAN-Kcr demonstrates excellent predictive performance. The method significantly outperforms baseline approaches in terms of accuracy, achieving an ACC value of 79.18%, while the area under the ROC curve reaches 0.8618. SHAP analysis and gradient-based analysis reveal a strong dependence on pKa-related properties, charge characteristics, and locally enriched small side-chain residues, and indicate that the model can nonlinearly integrate high-order sequence semantic information for effective identification of Kcr sites. Overall, MVFAN-Kcr combines data balancing strategies, multi-view feature fusion, and attention mechanisms to achieve high accuracy, robustness, and interpretability, providing an effective tool for protein Kcr site prediction. The data and code are available on https://github.com/Lilyjoys/MVFAN-Kcr , and a free web platform is provided at http://www.mvfan-kcr.com .
DNA methylation is a covalent modification of cytosine and adenine bases that regulates gene expression and underlies diverse biological processes and diseases. Existing computational methods often rely on fixed-scale feature extraction or static positional encodings, limiting their ability to model both fine-grained sequence motifs and long-range dependencies across multiple methylation chemistries. We present MeDiCNet, a unified deep-learning framework that combines multi-scale dynamic convolution with enhanced positional attention to predict N6-methyladenine, 5-hydroxymethylcytosine and N4-methylcytosine sites. MeDiCNet encodes nucleotide identity, extracts local patterns via a dynamic convolution module, captures global context through a Transformer encoder into which positional information is injected via rotary and clipped relative position embeddings, and adaptively fuses these feature streams through a gated fusion module for final classification. We evaluated MeDiCNet on seventeen benchmark datasets spanning bacteria, fungi, plants and mammals. Compared with other methods, MeDiCNet improved overall accuracy (ACC) by up to 8.1% and Matthews correlation coefficient (MCC) by up to 0.10. For example, it achieved 94.82% accuracy on the F. vesca 6mA dataset and the area under the ROC (AUC) curve above 0.98 on the Mus musculus 5hmC dataset. Crucially, rigorous analysis confirms that MeDiCNet recovers biologically authentic motifs with high fidelity in an unsupervised manner, while requiring only about 16% of the parameters of comparable large language models. These results demonstrate MeDiCNet's ability to capture complex local and global sequence features, providing a robust, efficient, and interpretable tool for large-scale, cross-type epigenomic analysis.
Protein phosphorylation, a pivotal post-translational modification mechanism, plays essential roles in cellular signaling and disease regulation. While O-phosphorylation has been extensively investigated, the biological significance of N-phosphorylation has only recently gained attention, with its study hindered by inherent challenges including site instability and detection limitations. To address these challenges, we present NphosNet, a deep learning framework with four technical innovations. First, we constructed a novel, class-imbalanced N-phosphorylation dataset comprising pH-913, pK-2060, and pR-1700 subsets. For comprehensive feature representation, we developed a hybrid embedding strategy combining amino acid tokenization with positional encoding, enhanced by ProtT5 and EMBER2 pre-trained model embeddings to capture deep semantic information from protein sequences. Architecturally, we introduced a three-branch framework integrating Transformer modules, optimized xLSTM blocks, CNN components, and spatial attention-enhanced ResNet units for multi-dimensional feature extraction. A novel weighted three-channel cross-attention mechanism was specifically designed for effective feature fusion across branches. Comparative evaluations demonstrate NphosNet's superior performance in N-phosphorylation site prediction (pH/pK/pR), achieving AUC values of 0.9227, 0.9099, and 0.9377 respectively, significantly outperforming existing methods. This advancement provides a robust computational tool for elucidating N-phosphorylation mechanisms in cellular processes and disease pathogenesis.
Identifying protein-protein interaction (PPI) sites is crucial for predicting protein function, uncovering disease mechanisms, and designing drugs. Experimental methods for PPI site identification are often costly and time-consuming, necessitating the development of efficient computational approaches. However, existing methods still face significant challenges in balancing high accuracy with computational efficiency. To address these limitations, we propose ProtFormer-Site, a novel PPI site prediction framework that integrates large protein language models (ESM2 and SaProt) with a parameter-efficient fine-tuning strategy (LoRA). We introduce a specialized ProtFormer backbone featuring a recycling mechanism to iteratively refine residue-level interaction features. The framework includes two variants: a sequence-only model and a structure-enhanced model, catering to different data availabilities. ProtFormer-Site demonstrated outstanding performance on three benchmark datasets, achieving Matthews correlation coefficient (MCC) improvements ranging from 22.4% to 61.5% compared to state-of-the-art methods. Furthermore, ProtFormer-Site demonstrates exceptional scalability, maintaining significantly lower log-transformed inference times across varying sequence lengths compared to state-of-the-art methods. Its computational efficiency makes it uniquely suited for large-scale, high-throughput prediction tasks. These results indicate that ProtFormer-Site offers a robust, accurate, and computationally efficient solution for PPI site prediction.
Triple-negative breast cancer (TNBC) is a biologically aggressive subtype of breast cancer marked by high heterogeneity and poor prognosis. Copper metabolism has been implicated in TNBC progression, but its functional contributions remain insufficiently defined. In this study, we analyzed transcriptomic data from 229 TNBC and adjacent normal samples from The Cancer Genome Atlas (TCGA) to identify 26 differentially expressed copper metabolism-related genes (DEGs-CM). A nine-gene Cox model (HEPHL1, COX7A1, COX4I2, JUN, MAPT, MT1A, AOC3, DCT, AOC2) demonstrated robust prognostic value, with time-dependent AUCs of 0.88, 0.84, and 0.80 at 1, 3, and 5 years. Functional enrichment analyses revealed epithelial-mesenchymal transition (EMT) and angiogenesis pathways enriched in high-risk groups. Four genes (AOC3, COX4I2, COX7A1, JUN) were further identified as copper metabolism-related metastasis genes (CMMRGs) through correlation with metastasis-associated programs. Based on these genes, machine learning classifiers were developed to predict TNBC presence and lymph node metastasis. Classifiers trained on the full dataset achieved consistently high performance, with most models showing AUCs greater than 0.97 (random forest, XGBoost, and AdaBoost classifier) even when using reduced gene panels (26-, 9-, and 4-gene sets), demonstrating stable classification across gene panels. In contrast, performance declined notably in metastasis-specific classification, largely due to the limited number of labeled metastatic samples. Among these, the 9-gene panel yielded the highest test AUCs across most models (gradient boosting machine AUC 0.64), suggesting that it may provide an optimal balance between model complexity and discriminative power, while also highlighting a key limitation related to the restricted sample size and incomplete clinical annotations in the metastasis-specific dataset. Single-cell RNA sequencing confirmed fibroblast-specific enrichment of CMMRGs and associated EMT signatures, suggesting a mechanistic link between copper metabolism, stromal remodeling, and metastasis. These results establish a copper-centered molecular framework for TNBC diagnosis and metastasis prediction, supporting the translational potential of copper metabolism-related genes in clinical applications.