Nucleosome positioning plays a central role in chromatin organization and gene regulation, yet its accurate computational prediction remains challenging. This study introduces a Hybrid Dilated Gated Separable Convolutional Neural Network (HDGS-Net), which integrates dilated convolution, gated convolution, and depthwise separable convolution to achieve continuous prediction of in vitro nucleosome occupancy at single-base resolution across the entire Saccharomyces cerevisiae genome. On benchmark datasets, HDGS-Net attained an average Pearson correlation coefficient of 0.87, outperforming conventional methods and demonstrating excellent cross-chromosome generalization capability. Sequence analysis confirms that DNA dinucleotide physical properties dominate nucleosome positioning, with AT-rich sequences inhibiting binding and GC-rich sequences promoting binding. Analysis of transcription start regions verifies that flanking nucleosome sequence features are highly conserved across different chromatin environments, supporting the universal regulatory role of sequence preference. Cross-species analysis demonstrates that the guiding efficacy of DNA sequence on nucleosome positioning varies among species, showing quantitatively decreasing contributions in Caenorhabditis elegans, Saccharomyces cerevisiae, and Schizosaccharomyces pombe. This study provides a high-accuracy predictive tool for investigating dynamic nucleosome positioning.
MOTIVATION:Recent advances in single-cell sequencing have transformed precise measurement of gene expression at cellular resolution, enabling unprecedented dissection of cellular heterogeneity and intricate biological processes. The accumulation of multi-omics data offers new avenues for cell clustering-a critical foundation for cell-type identification and downstream analyses. However, substantial challenges persist in simultaneously achieving effective integration of complementary information in multi-omics data and their appropriate weight allocation. RESULTS:Here, we propose an Adaptive Multi-View clustering framework with the Information Bottleneck principle to solve the multi-omics data clustering task (named scAMVIB). The proposed model could learn multi-view omics representations that capture both inter-omics associations and omics-specific patterns, with the adaptive weight allocation. Specifically, multi-view data comprise two components: (i) the integrated omics feature matrix derived from the similarity network fusion strategy and (ii) omics-specific representations from distinct platforms. These inputs are processed through a multi-view information bottleneck clustering framework that leverages cross-view complementarity to enhance representations. View weights are adaptively assigned via maximum entropy regularization, proportional to their information content. The final cell partitions are obtained through sequential iterative optimization. Comprehensive experiments across multiple datasets demonstrate that scAMVIB has strong competitiveness in clustering while maintaining biological interpretability.
Accurate prediction of drug-target binding affinity (DTA) is crucial for accelerating drug discovery and development. Although deep learning-based approaches have demonstrated remarkable success in DTA prediction, current methods usually suffer from two key limitations: (1) inadequate utilization of multi-scale feature information, leading to suboptimal representations of drugs and targets; and (2) challenges in effectively balancing features across scales for robust feature fusion. In this study, we propose MFCLDTA, a novel deep learning-based framework for DTA prediction through the multi-scale feature contrastive learning mechanism. The proposed approach captures drug and target features across three scales: molecular sequence, molecular structure, and affinity graph. Through the multi-scale contrastive learning mechanism, MFCLDTA maximizes mutual information between different scales while ensuring robust feature alignment. Integration of these multi-scale representations achieves superior DTA prediction performance. Comprehensive benchmarks on Davis and KIBA datasets demonstrate MFCLDTA consistently outperforms state-of-the-art baseline approaches. Ablation studies confirm the critical importance of both multi-scale feature integration and contrastive learning in enhancing prediction accuracy. The source code of MFCLDTA is available at https://github.com/ssaixiansheng/MFCLDTA.
Cancer drug response (CDR) prediction is crucial for advancing precision medicine. Advanced computational predictors extracted CDR patterns from multimodal features of cell lines and drugs under the guidance of known CDRs. However, most existing methods struggle to extract biologically meaningful and generalizable CDR representations due to insufficient semantic alignment among the multimodal features. In addition, semantic inconsistencies between multimodal and topological features further hinder the predictive accuracy of CDRs. To address the challenges, a novel Reinforcement Learning-driven Adaptive Synergistic Optimization Network-CDR (RLASON-CDR) is put forward for CDR prediction. RLASON-CDR first constructs multimodal CDR representations aligned within and across drugs and cell lines. Next, it captures high-order topological representations of CDRs from the cell line-drug response network. Finally, a reinforcement learning-based network is proposed to adaptively explore potential CDR patterns by synergistically optimizing these representations instead of semantic fusion. Extensive evaluations demonstrate that RLASON-CDR consistently outperforms existing methods and is robust for predicting unknown CDRs. Gradient attribution analysis further reveals that RLASON-CDR identifies key modality contributions and uncovers biologically response patterns. Furthermore, biological significance analysis indicates that RLASON-CDR can effectively reveal mechanisms of drug response and provide valuable guidance for precision therapy in clinical applications.
Understanding long noncoding RNA (lncRNA) function is essential for revealing molecular mechanisms and developing effective therapies for complex diseases, as lncRNAs play important regulatory roles in many disease-related biological processes. However, existing lncRNA function predictors struggle to extract discriminative features from multimodal omics data and to model the semantic and topological structure of the gene ontology (GO), which severely limits their ability to achieve biologically meaningful and functionally informative predictions. To address these challenges, we propose a novel framework for lncRNA function prediction, namely MiCLSAO. Firstly, MiCLSAO utilizes multiview cross-contrastive learning with attention mechanisms to extract highly discriminative lncRNA features from diverse omics similarity networks. Secondly, graph convolutional networks are applied to learn initial features of GO terms, while multiscale topological and semantic relationships are incorporated to adaptively refine term representations. Finally, an lncRNA function predictor is developed by dynamically integrating the representations of lncRNAs and GO terms using a Kolmogorov-Arnold network. Extensive experiments demonstrate that MiCLSAO consistently outperforms state-of-the-art methods across multiple metrics, with significant capability to recover known functions and uncover novel ones. Moreover, MiCLSAO demonstrates remarkable practical utility and potential value by providing more functionally informative annotations for lncRNAs.
Accurately predicting cancer drug response (CDR) is crucial for personalized cancer therapies and drug repositioning. Efficient CDR prediction requires to integrate multimodal data including sequences, structures, multilevel omics, and diverse biological networks of drugs and cell lines, to capture intricate underlying patterns. So we proposed TransGCDR, a novel method that integrates cross-modal multilevel homogeneous and heterogeneous features to derive highly discriminative embeddings, thereby enhancing CDR prediction. TransGCDR first learns structural representations of drugs using a Transformer from fingerprint substructures and Graph Convolutional Networks for molecular graphs. Meanwhile, it utilizes a Graph Attention Network to extract cell line representations from graphs integrating multilevel omics data, including gene expression, copy number variation, and somatic mutations. Next, it generates homogeneous embeddings for drugs and cell lines from these cross-modal representations while extracting drug- and cell linecentric heterogeneous contextual embeddings from prior CDRs. These embeddings are then integrated using a Residual Attention Graph Convolutional Network to obtain comprehensive representations. Finally, an MLP-based model is trained on the learned embeddings to predict CDRs. Extensive evaluations demonstrate that TransGCDR outperforms state-of-the-art methods in CDR prediction across various metrics. Ablation studies confirm that every module enhances the precise CDR characterization from the multimodal and multilevel data, which is key to the success of TransGCDR. Moreover, case studies highlight the significant clinical potential of TransGCDR to predict CDRs.
Motivation: Computational drug repositioning is a vital path to improve efficiency of drug discovery, which aims to find potential Drug-Disease Associations (DDAs) to develop new effects of the existing drugs. Many approaches detected novel DDAs from heterogenous network which integrates similar drugs, similar diseases and the known DDAs. However, sparsity of the known DDAs and intrinsic synergic relations on representations of drugs and diseases in the heterogenous network are still the main challenges for DDAs prediction. Results: To address the problems, a novel drug repositioning approach is proposed here. Firstly, a drug similar network is constructed by Gaussian similarity kernel fusion of multisource drug similarities. Likewise, a disease similar network is generated by the same strategy. Secondly, the known DDAs network is extended by a bi-random walk algorithm from the above-mentioned similar networks. Meanwhile, representations of drugs and diseases are learned from their similar networks through graph convolutions and then a DDAs network is induced from the representations. Finally, to discover latent DDAs, the inductive DDAs network is refined iteratively by collaborating with the extended known DDAs network. Comprehensive experimental results show that our method outperforms several state-of-the-art methods for predicting DDAs on indicators including precision, recall, F1-score, MCC, ROC and AUPR. Moreover, case studies suggest that our method is highly effective in practices. The success of our method may be attributed to three aspects: (1) reliable similar relationships of drugs and diseases; (2) enhanced connectivity of the heterogenous network; (3) reasonable collaborative induction on DDAs network. Our method is freely available at https://github.com/BioMLab/DRCLN.
This study achieved cancer type and survival time prediction by transforming transcriptomic features into feature maps and employing deep learning models. Using transcriptomic data from 27 cancer types and survival data from 10 types in the TCGA database, a pan-cancer transcriptomic feature map was constructed through data cleaning, feature extraction, and visualization. Using Inception network and gated convolutional modules yielded a pan-cancer classification accuracy of 91.8 %. Additionally, by extracting 31 differential genes from different cancer feature maps, an interaction network diagram was drawn, identifying two key genes, ANXA5 and ACTB. These genes are potential biomarkers related to cancer progression, angiogenesis, metastasis, and treatment resistance. Survival prediction analysis on 10 cancer types, combined with feature maps and data amplification, cancer survival prediction accuracy reached from 0.75 to 0.91. This transcriptomic feature map provides a novel approach for cancer omics analysis, to facilitate personalized treatments and reflecting individual differences.
Single-cell RNA sequencing (scRNA-seq) analysis is capable of elucidating cell heterogeneity and diversity, making cell-level biological research possible. Cell type clusterization is one of the main goals of scRNA-seq analysis. However, existing methods lack a flexible aggregation mechanism to fuse node attribute features and graph structural features adaptively. They also ignore multi-scale information embedded in different layers of networks. Here, we propose AGAC (Attention-driven Graph Attentional Clustering), a unified computational framework that applies hierarchical feature aggregation with mixed attentional mechanisms to scRNA-seq data clustering. Firstly, AGAC learns the attribute features of cells through graph attention autoencoder and the graph structural features among cells through graph neural network (GNN) respectively. Secondly, AGAC utilizes the heterogeneous intelligent fusion module to fuse these two kinds of features dynamically, as well as the multi-scale intelligent fusion module to aggregate the multi-scale information embedded in different layers of the GNN adaptively. In addition, AGAC adopts a dual self-supervision module for end-to-end training and synchronous optimization. Experimental results on 13 real scRNA-seq datasets demonstrate that AGAC is more effective than existing baseline methods because it takes into account the discriminative information in the network comprehensively and generates clustering results directly. The source code of AGAC can be downloaded at https://github.com/zzyqh/AGAC.
N6-methyladenosine (m6A) is a key epitranscriptomic marker enriched in long noncoding RNAs (lncRNAs) that is closely involved in complex disease mechanisms. Although accurate detection of m6A sites in lncRNAs is essential for understanding disease mechanisms, the development of effective computational predictors remains challenging due to the limited number of annotated sites. Moreover, most existing predictors are specifically designed for messenger RNAs (mRNAs) based on abundant mRNA-specific knowledge, yet they exhibit limited generalizability to lncRNAs. Given the similarities between mRNAs and lncRNAs, a transferable framework capable of leveraging their shared features is critical for advancing m6A site prediction in lncRNAs. To address this challenge, we propose DSNm6A, a deep learning framework that learns cross-RNA transferable sequence representations for effective lncRNA m6A site detection. To comprehensively capture patterns and signals of m6A sites, lncRNA and mRNA sequences are first encoded from complementary multiple facets, including One-Hot encoding, nucleotide physicochemical properties and cumulative frequency, and position-specific propensity. Based on these sequence encodings, a domain separation network integrating CNN, Bi-LSTM, and BERT modules is then employed to explicitly disentangle domain-invariant features shared between mRNAs and lncRNAs from their domain-specific counterparts. The shared features are finally fed into a fully connected layer for accurate lncRNA m6A sites prediction. Cross-validation and independent test results demonstrate that DSNm6A consistently outperforms existing methods across nearly all performance metrics, attributed to its superior capacity to learn transferable m6A-related features across RNA types. In addition, DSNm6A exhibits strong robustness and generalization across species.
Prediction of drug-target interactions (DTIs) is one of the crucial steps for drug repositioning. Identifying DTIs through bio-experimental manners is always expensive and time-consuming. Recently, deep learning-based approaches have shown promising advancements in DTI prediction, but they face two notable challenges: (i) how to explicitly capture local interactions between drug-target pairs and learn their higher-order substructure embeddings; (ii) How to filter out redundant information to obtain effective embeddings for drugs and targets. Results: In this study, we propose a novel approach, termed DSANIB, to infer potential interactions between drugs and targets. DSANIB comprises two primary components: (1) DSAN component: The Inter-view Attention Network Module explicitly learns the local interactions between drugs and targets, while the Intra-view Attention Network Module aggregates information from local interaction features to obtain their higher-order substructure embeddings. (2) Information Bottleneck (IB) component: DSANIB adopts the IB strategy, which could retain relevant information while minimizing the redundant features to obtain their discriminative representations. Extensive experimental results demonstrate that DSANIB outperforms other SOTA prediction models. In addition, visualization of drug and target embeddings learned through DSANIB could provide interpretable insights for the prediction results.
Interactions between long non-coding RNAs (lncRNAs) and microRNAs (miRNAs) play an important role in the development of complex human diseases by collaboratively regulating gene transcription and expression. Therefore, identifying lncRNA-miRNA interactions (LMIs) is essential for diagnosing and treating complex human diseases. Because identifying LMIs with wet experiments is time-consuming and labor-intensive, some computational methods have been developed to infer LMIs. However, these approaches excel at utilizing single-modal information but struggle to integrate multimodal data from lncRNAs and miRNAs, which is essential for uncovering complex patterns in LMIs, ultimately limiting their performance. Therefore, this article proposes a novel multimodal contrastive representation learning model (MCRLMI) for LMI predictions. The model fully integrates multi-source similarity information and sequence encodings of lncRNAs and miRNAs. It leverages a graph convolutional network (GCN) and a Transformer to capture local neighborhood structural features and long-distance dependencies, respectively, enabling the collaborative modeling of structural and semantic information. Subsequently, to effectively integrate multimodal characteristics with encoded information, a multichannel attention mechanism and contrastive learning are introduced to fuse the extracted features. Finally, a Kolmogorov-Arnold Network (KAN) is trained with the optimized embeddings to predict LMIs. Extensive experiments show that the proposed MCRLMI consistently outperforms existing methods. Moreover, case studies further validate the potential of MCRLMI to identify novel LMIs in practical applications.
Identifying disease-associated microRNAs (miRNAs) could help understand the deep mechanism of diseases, which promotes the development of new medicine. Recently, network-based approaches have been widely proposed for inferring the potential associations between miRNAs and diseases. However, these approaches ignore the importance of different relations in meta-paths when learning the embeddings of miRNAs and diseases. Besides, they pay little attention to screening out reliable negative samples which is crucial for improving the prediction accuracy. In this study, we propose a novel approach named MGCNSS with the multi-layer graph convolution and high-quality negative sample selection strategy. Specifically, MGCNSS first constructs a comprehensive heterogeneous network by integrating miRNA and disease similarity networks coupled with their known association relationships. Then, we employ the multi-layer graph convolution to automatically capture the meta-path relations with different lengths in the heterogeneous network and learn the discriminative representations of miRNAs and diseases. After that, MGCNSS establishes a highly reliable negative sample set from the unlabeled sample set with the negative distance-based sample selection strategy. Finally, we train MGCNSS under an unsupervised learning manner and predict the potential associations between miRNAs and diseases. The experimental results fully demonstrate that MGCNSS outperforms all baseline methods on both balanced and imbalanced datasets. More importantly, we conduct case studies on colon neoplasms and esophageal neoplasms, further confirming the ability of MGCNSS to detect potential candidate miRNAs. The source code is publicly available on GitHub https://github.com/15136943622/MGCNSS/tree/master
Long non-coding RNAs are a class of essential non-coding RNAs with a length of more than 200 nts. Recent studies have indicated that lncRNAs have various complex regulatory functions, which play great impacts on many fundamental biological processes. However, measuring the functional similarity between lncRNAs by traditional wet-experiments is time-consuming and labor intensive, computational-based approaches have been an effective choice to tackle this problem. Meanwhile, most sequences-based computation methods measure the functional similarity of lncRNAs with their fixed length vector representations, which could not capture the features on larger k-mers. Therefore, it is urgent to improve the predict performance of the potential regulatory functions of lncRNAs. In this study, we propose a novel approach called MFSLNC to comprehensively measure functional similarity of lncRNAs based on variable k-mer profiles of nucleotide sequences. MFSLNC employs the dictionary tree storage, which could comprehensively represent lncRNAs with long k-mers. The functional similarity between lncRNAs is evaluated by the Jaccard similarity. MFSLNC verified the similarity between two lncRNAs with the same mechanism, detecting homologous sequence pairs between human and mouse. Besides, MFSLNC is also applied to lncRNA-disease associations, combined with the association prediction model WKNKN. Moreover, we also proved that our method can more effectively calculate the similarity of lncRNAs by comparing with the classical methods based on the lncRNA-mRNA association data. The detected AUC value of prediction is 0.867, which achieves good performance in the comparison of similar models.
Long non-coding RNAs (lncRNAs) play important roles by regulating proteins in many biological processes and life activities. To uncover molecular mechanisms of lncRNA, it is very necessary to identify interactions of lncRNA with proteins. Recently, some machine learning methods were proposed to detect lncRNA-protein interactions according to the distribution of known interactions. The performances of these methods were largely dependent upon: (1) how exactly the distribution of known interactions was characterized by feature space; (2) how discriminative the feature space was for distinguishing lncRNA-protein interactions. Because the known interactions may be multiple and complex model, it remains a challenge to construct discriminative feature space for lncRNA-protein interactions. To resolve this problem, a novel method named DFRPI was developed based on deep autoencoder and marginal fisher analysis in this paper. Firstly, some initial features of lncRNA-protein interactions were extracted from the primary sequences and secondary structures of lncRNA and protein. Secondly, a deep autoencoder was exploited to learn encode parameters of the initial features to describe the known interactions precisely. Next, the marginal fisher analysis was employed to optimize the encode parameters of features to characterize a discriminative feature space of the lncRNA-protein interactions. Finally, a random forest-based predictor was trained on the discriminative feature space to detect lncRNA-protein interactions. Verified by a series of experiments, the results showed that our predictor achieved the precision of 0.920, recall of 0.916, accuracy of 0.918, MCC of 0.836, specificity of 0.920, sensitivity of 0.916 and AUC of 0.906 respectively, which outperforms the concerned methods for predicting lncRNA-protein interaction. It may be suggested that the proposed method can generate a reasonable and effective feature space for distinguishing lncRNA-protein interactions accurately. The code and data are available on https://github.com/D0ub1e-D/DFRPI.
The measurement of gene functional similarity plays a critical role in numerous biological applications, such as gene clustering, the construction of gene similarity networks. However, most existing approaches still rely heavily on traditional computational strategies, which are not guaranteed to achieve satisfactory performance. In this study, we propose a novel computational approach called GOGCN to measure gene functional similarity by modeling the Gene Ontology (GO) through Graph Convolutional Network (GCN). GOGCN is a graph-based approach that performs sufficient representation learning for terms and relations in the GO graph. First, GOGCN employs the GCN-based knowledge graph embedding (KGE) model to learn vector representations (i.e., embeddings) for all entities (i.e., terms). Second, GOGCN calculates the semantic similarity between two terms based on their corresponding vector representations. Finally, GOGCN estimates gene functional similarity by making use of the pair-wise strategy. During the representation learning period, GOGCN promotes semantic interaction between terms through GCN, thereby capturing the rich structural information of the GO graph. Further experimental results on various datasets suggest that GOGCN is superior to the other state-of-the-art approaches, which shows its reliability and effectiveness.
MOTIVATION:In recent years, a large number of biological experiments have strongly shown that miRNAs play an important role in understanding disease pathogenesis. The discovery of miRNA-disease associations is beneficial for disease diagnosis and treatment. Since inferring these associations through biological experiments is time-consuming and expensive, researchers have sought to identify the associations utilizing computational approaches. Graph Convolutional Networks (GCNs), which exhibit excellent performance in link prediction problems, have been successfully used in miRNA-disease association prediction. However, GCNs only consider 1st-order neighborhood information at one layer but fail to capture information from high-order neighbors to learn miRNA and disease representations through information propagation. Therefore, how to aggregate information from high-order neighborhood effectively in an explicit way is still challenging.RESULTS:To address such a challenge, we propose a novel method called mixed neighborhood information for miRNA-disease association (MINIMDA), which could fuse mixed high-order neighborhood information of miRNAs and diseases in multimodal networks. First, MINIMDA constructs the integrated miRNA similarity network and integrated disease similarity network respectively with their multisource information. Then, the embedding representations of miRNAs and diseases are obtained by fusing mixed high-order neighborhood information from multimodal network which are the integrated miRNA similarity network, integrated disease similarity network and the miRNA-disease association networks. Finally, we concentrate the multimodal embedding representations of miRNAs and diseases and feed them into the multilayer perceptron (MLP) to predict their underlying associations. Extensive experimental results show that MINIMDA is superior to other state-of-the-art methods overall. Moreover, the outstanding performance on case studies for esophageal cancer, colon tumor and lung cancer further demonstrates the effectiveness of MINIMDA.AVAILABILITY AND IMPLEMENTATION:https://github.com/chengxu123/MINIMDA and http://120.79.173.96/.
为确保冷链物流在缩减成本的同时满足农产品品质需求,以异构数据为基础,提出一种农产品冷链物流节点部署方法.基于可扩展标记语言文档,利用数据源模块、转变模块、集成模块以及应用模块,构建异构数据集成模型;根据集成的物流异构数据,架构冷链物流节点部署模型,采用粒子群优化算法进行求解,获取节点方位,通过分析影响寻优性能的极大速度与加速常数等指标参数取值范围,完善农产品冷链物流节点部署结果.仿真结果表明,所提方法的设计方案的配送成本与需求量满足度最优,具有较好的有效性.
DNA N6-Methyladenine (6mA) is a common epigenetic modification, which plays some significant roles in the growth and development of plants. It is crucial to identify 6mA sites for elucidating the functions of 6mA. In this article, a novel model named i6mA-vote is developed to predict 6mA sites of plants. Firstly, DNA sequences were coded into six feature vectors with diverse strategies based on density, physicochemical properties, and position of nucleotides, respectively. To find the best coding strategy, the feature vectors were compared on several machine learning classifiers. The results suggested that the position of nucleotides has a significant positive effect on 6mA sites identification. Thus, the dinucleotide one-hot strategy which can describe position characteristics of nucleotides well was employed to extract DNA features in our method. Secondly, DNA sequences of Rosaceae were divided into a training dataset and a test dataset randomly. Finally, i6mA-vote was constructed by combining five different base-classifiers under a majority voting strategy and trained on the Rosaceae training dataset. The i6mA-vote was evaluated on the task of predicting 6mA sites from the genome of the Rosaceae, Rice, and Arabidopsis separately. In Rosaceae, the performances of i6mA-vote were 0.955 on accuracy (ACC), 0.909 on Matthew correlation coefficients (MCC), 0.955 on sensitivity (SN), and 0.954 on specificity (SP). Those indicators, in the order of ACC, MCC, SN, SP, were 0.882, 0.774, 0.961, and 0.803 on Rice while they were 0.798, 0.617, 0.666, and 0.929 on Arabidopsis. According to the indicators, our method was effectiveness and better than other concerned methods. The results also illustrated that i6mA-vote does not only well in 6mA sites prediction of intraspecies but also interspecies plants. Moreover, it can be seen that the specificity is distinctly lower than the sensitivity in Rice while it is just the opposite in Arabidopsis. It may be resulted from sequence similarity among Rosaceae, Rice and Arabidopsis.
Preeclampsia (PE) is a maternal disease that causes maternal and child death. Treatment and preventive measures are not sound enough. The problem of PE screening has attracted much attention. The purpose of this study is to screen placental mRNA to obtain the best PE biomarkers for identifying patients with PE. We use Limma in the R language to screen out the 48 differentially expressed genes with the largest differences and used correlation-based feature selection algorithms to reduce the dimensionality and avoid attribute redundancy arising from too many mRNA samples participating in the classification. After reducing the mRNA attributes, the mRNA samples are sorted from large to small according to information gain. In this study, a classifier model is designed to identify whether samples had PE through mRNA in the placenta. To improve the accuracy of classification and avoid overfitting, three classifiers, including C4.5, AdaBoost, and multilayer perceptron, are used. We use the majority voting strategy integrated with the differentially expressed genes and the genes filtered by the best subset method as comparison methods to train the classifier. The results show that the classification accuracy rate has increased from 79% to 82.2%, and the number of mRNA features has decreased from 48 to 13. This study provides clues for the main PE biomarkers of mRNA in the placenta and provides ideas for the treatment and screening of PE.