Molecular Docking is a critical task in structure-based virtual screening. Recent advancements have showcased the efficacy of diffusion-based generative models for blind docking tasks. However, these models do not inherently estimate protein-ligand binding strength thus cannot be directly applied to virtual screening tasks. Protein-ligand scoring functions serve as fast and approximate computational methods to evaluate the binding strength between the protein and ligand. In this work, we introduce normalized mixture density network (NMDN) score, a deep learning (DL)-based scoring function learning the probability density distribution of distances between protein residues and ligand atoms. The NMDN score addresses limitations observed in existing DL scoring functions and performs robustly in both pose selection and virtual screening tasks. Additionally, we incorporate an interaction module to predict the experimental binding affinity score to fully utilize the learned protein and ligand representations. Finally, we present an end-to-end blind docking and virtual screening protocol named DiffDock-NMDN. For each protein-ligand pair, we employ DiffDock to sample multiple poses, followed by utilizing the NMDN score to select the optimal binding pose, and estimating the binding affinity using scoring functions. Our protocol achieves an average enrichment factor of 4.96 on the LIT-PCBA data set, proving effective in real-world drug discovery scenarios where binder information is limited. This work not only presents a robust DL-based scoring function with superior pose selection and virtual screening capabilities but also offers a blind docking protocol and benchmarks to guide future scoring function development.
With the potential application of biomass-derived feedstock upgradation to sustainable aviation fuels, it is essential to enhance the hydroisomerization performance of ZSM-22 zeolite while improving its resistance to residual oxygen-containing compounds. As the defect sites in the ZSM-22 zeolite, the abundant Si-OH groups are closely related to the catalytic performance and stability, serving as the main attack sites for the generated water. In this work, liquid-mediated defect-healing treatment is performed to heal Si-OH to Si-O-Si, leading to the enhancement of the crystallinity and pore connectivity without affecting the Si/Al, micropore volume, and morphology and preventing the micropores blockage and dealumination caused by conventional silylation and silication procedures. The outcome of the declined Si-OH groups is the reduction of the Lewis acid site without altering the Br & Oslash;nsted acidity. In the n-dodecane hydroisomerization, the catalyst obtained by the liquid-mediated defect-healing treatment route shows an increased conversion and isomer yield compared to the parent and healed catalysts prepared by other healing methods. This is mainly due to enhanced confinement of the micropore void, resulting in decreased apparent activation energy and reduced yield to multi-branched isomers prone to cracking. Furthermore, the healed catalyst exhibits improved resistance and structure stability in the hydroisomerization of feedstocks containing butanol. The work provides a prospective application of ZSM-22 zeolite in the hydroisomerization for complex and severe reactants through the essential Si-OH healing method.
Advancing the conventional hydrothermal and ionic liquid-assisted synthesis methodologies to produce ZSM-22 zeolites with similar characteristics yet notably distinct defect sites holds substantial significance in elucidating their influence on shape selectivity and hydroisomerization performance. This study investigated the impact of the initial gel ratio and synthesis parameters on the ionic liquid-assisted crystallization process and the crystal size of the ZSM-22 zeolite. The NH4F feed exhibits a minimal effect on the crystallization process, in contrast to the beneficial effect of improved ionic liquid and decreased water feed on substantially diminishing the crystal size of the resulting zeolites, which exhibit a morphology similar to that of the ZSM-22 zeolite obtained through the conventional synthesis route. Furthermore, the investigation into the crystallization process revealed that both ionic liquids and high temperatures play indispensable roles in accelerating the nucleation and crystal growth of ZSM-22 zeolite. When applied in the n-dodecane hydroisomerization, the nanosized ZSM-22 zeolites prepared by ionic liquid-assisted route exhibit remarkable isomer yield (above 72%), surpassing the conventional and large-size ZSM-22 zeolites. After the product distribution analysis, it is evident that the fewer Si-OH defect sites and the correspondingly strict confinement environment effectively limit the formation of the multibranched intermediates prone to crack, which ultimately improves the hydroisomerization performance, intrinsically reflecting the shape selectivity of the 10-membered ring zeolite by eliminating the interference of the framework defect sites. This work emphasizes the exceptional hydroisomerization performance of the ZSM-22 zeolite, highlighting the significance of its efficient synthesis and wide application.
Drug repositioning offers promising prospects for accelerating drug discovery by identifying potential drug-disease associations (DDAs) for existing drugs and diseases. Previous methods have generated meta-path-augmented node or graph embeddings for DDA prediction in drug-disease heterogeneous networks. However, these approaches rarely develop end-to-end frameworks for path instance-level representation learning as well as the further feature selection and aggregation. By leveraging the abundant topological information in path instances, more fine-grained and interpretable predictions can be achieved. To this end, we introduce deep multiple instance learning into drug repositioning by proposing a novel method called MilGNet. MilGNet employs a heterogeneous graph neural network (HGNN)-based encoder to learn drug and disease node embeddings. Treating each drug-disease pair as a bag, we designed a special quadruplet meta-path form and implemented a pseudo meta-path generator in MilGNet to obtain multiple meta-path instances based on network topology. Additionally, a bidirectional instance encoder enhances the representation of meta-path instances. Finally, MilGNet utilizes a multi-scale interpretable predictor to aggregate bag embeddings with an attention mechanism, providing predictions at both the bag and instance levels for accurate and explainable predictions. Comprehensive experiments on five benchmarks demonstrate that MilGNet significantly outperforms ten advanced methods. Notably, three case studies on one drug (Methotrexate) and two diseases (Renal Failure and Mismatch Repair Cancer Syndrome) highlight MilGNet's potential for discovering new indications, therapies, and generating rational meta-path instances to investigate possible treatment mechanisms. The source code is available at https://github.com/gu-yaowen/MilGNet.
Drug resistance in Mycobacterium tuberculosis (Mtb) is a significant challenge in the control and treatment of tuberculosis, making efforts to combat the spread of this global health burden more difficult. To accelerate anti-tuberculosis drug discovery, repurposing clinically approved or investigational drugs for the treatment of tuberculosis by computational methods has become an attractive strategy. In this study, we developed a virtual screening workflow that combines multiple machine learning and deep learning models, and 11 576 compounds extracted from the DrugBank database were screened against Mtb. Our screening method produced satisfactory predictions on three data-splitting settings, with the top predicted bioactive compounds all known antibacterial or anti-TB drugs. To further identify and evaluate drugs with repurposing potential in TB therapy, 15 screened potential compounds were selected for subsequent computational and experimental evaluations, out of which aldoxorubicin and quarfloxin showed potent inhibition of Mtb strain H37Rv, with minimal inhibitory concentrations of 4.16 and 20.67 μM/mL, respectively. More inspiringly, these two compounds also showed antibacterial activity against multidrug-resistant TB isolates and exhibited strong antimicrobial activity against Mtb. Furthermore, molecular docking, molecular dynamics simulation, and the surface plasmon resonance experiments validated the direct binding of the two compounds to Mtb DNA gyrase. In summary, our effective comprehensive virtual screening workflow successfully repurposed two novel drugs (aldoxorubicin and quarfloxin) as promising anti-Mtb candidates. The verification results provide useful information for the further development and clinical verification of anti-TB drugs.
Computational drug repositioning, through predicting drug-disease associations (DDA), offers significant potential for discovering new drug indications. Current methods incorporate graph neural networks (GNN) on drug-disease heterogeneous networks to predict DDAs, achieving notable performances compared to traditional machine learning and matrix factorization approaches. However, these methods depend heavily on network topology, hampered by incomplete and noisy network data, and overlook the wealth of biomedical knowledge available. Correspondingly, large language models (LLMs) excel in graph search and relational reasoning, which can possibly enhance the integration of comprehensive biomedical knowledge into drug and disease profiles. In this study, we first investigate the contribution of LLM-inferred knowledge representation in drug repositioning and DDA prediction. A zero-shot prompting template was designed for LLM to extract high-quality knowledge descriptions for drug and disease entities, followed by embedding generation from language models to transform the discrete text to continual numerical representation. Then, we proposed LLM-DDA with three different model architectures (LLM-DDANode Feat, LLM-DDADual GNN, LLM-DDAGNN-AE) to investigate the best fusion mode for LLM-based embeddings. Extensive experiments on four DDA benchmarks show that, LLM-DDAGNN-AE achieved the optimal performance compared to 11 baselines with the overall relative improvement in AUPR of 23.22%, F1-Score of 17.20%, and precision of 25.35%. Meanwhile, selected case studies of involving Prednisone and Allergic Rhinitis highlighted the model's capability to identify reliable DDAs and knowledge descriptions, supported by existing literature. This study showcases the utility of LLMs in drug repositioning with its generality and applicability in other biomedical relation prediction tasks.
Background China’s older population is facing serious health challenges, including malnutrition and multiple chronic conditions. There is a critical need for tailored food recommendation systems. Knowledge graph–based food recommendations offer considerable promise in delivering personalized nutritional support. However, the integration of disease-based nutritional principles and preference-related requirements needs to be optimized in current recommendation processes. Objective This study aims to develop a knowledge graph–based personalized meal recommendation system for community-dwelling older adults and to conduct preliminary effectiveness testing. Methods We developed ElCombo, a personalized meal recommendation system driven by user profiles and food knowledge graphs. User profiles were established from a survey of 96 community-dwelling older adults. Food knowledge graphs were supported by data from websites of Chinese cuisine recipes and eating history, consisting of 5 entity classes: dishes, ingredients, category of ingredients, nutrients, and diseases, along with their attributes and interrelations. A personalized meal recommendation algorithm was then developed to synthesize this information to generate packaged meals as outputs, considering disease-related nutritional constraints and personal dietary preferences. Furthermore, a validation study using a real-world data set collected from 96 community-dwelling older adults was conducted to assess ElCombo’s effectiveness in modifying their dietary habits over a 1-month intervention, using simulated data for impact analysis. Results Our recommendation system, ElCombo, was evaluated by comparing the dietary diversity and diet quality of its recommended meals with those of the autonomous choices of 96 eligible community-dwelling older adults. Participants were grouped based on whether they had a recorded eating history, with 34 (35%) having and 62 (65%) lacking such data. Simulation experiments based on retrospective data over a 30-day evaluation revealed that ElCombo’s meal recommendations consistently had significantly higher diet quality and dietary diversity compared to the older adults’ own selections (P<.001). In addition, case studies of 2 older adults, 1 with and 1 without prior eating records, showcased ElCombo’s ability to fulfill complex nutritional requirements associated with multiple morbidities, personalized to each individual’s health profile and dietary requirements. Conclusions ElCombo has shown enhanced potential for improving dietary quality and diversity among community-dwelling older adults in simulation tests. The evaluation metrics suggest that the food choices supported by the personalized meal recommendation system surpass autonomous selections. Future research will focus on validating and refining ElCombo’s performance in real-world settings, emphasizing the robust management of complex health data. The system’s scalability and adaptability pinpoint its potential for making a meaningful impact on the nutritional health of older adults.
Pancreatic ductal adenocarcinoma (PDAC) is one of the leading causes of cancer-related death. Therefore, we intend to explore novel strategies against PDAC. The exosomes-based biomimetic nanoparticle is an appealing candidate served as a drug carrier in cancer treatment, due to its inherit abilities. In the present study, we designed dasatinib-loaded hybrid exosomes by fusing human pancreatic cancer cells derived exosomes with dasatinib-loaded liposomes, followed by characterization for particle size (119.9 ± 6.10 nm) and zeta potential (-11.45 ± 2.24 mV). Major protein analysis from western blot techniques reveal the presence of exosome marker proteins CD9 and CD81. PEGylated hybrid exosomes showed pH-sensitive drug release in acidic condition, benefiting drug delivery to acidic cancer environment. Dasatinib-loaded hybrid exosomes exhibited significantly higher uptake rates and cytotoxicity to parent PDAC cells by two-sample t-test or by one-way ANOVA analysis of variance, as compared to free drug or liposomal formulations. The results from our computational analysis demonstrated that the drug-likeness, ADMET, and protein-ligand binding affinity of dasatinib are verified successfully. Cancer derived hybrid exosomes may serve as a potential therapeutic candidate for pancreatic cancer treatment.
构建基于药物多特征融合的药物疾病关联预测模型,为药物知识发现提供新思路.借助药物的化学结构、药物-副作用关联、药物-靶标关联的3个特征,构建融合的药物综合相似度及基于MeSH的疾病语义相似度特征表示方法.利用图卷积神经网络模型抽取药物-疾病图数据特征信息,构建基于多特征融合的药物疾病关联预测模型(MFFGCN),进而实现未知的药物疾病关联发现.利用269种药物、598种疾病及其之间的18 416种关联关系,对药物疾病存在的未知关联进行预测,借助AUC、AUPR、准确率、灵敏度、召回率、F1等多个评价指标进行评价.结果表明,多特征融合的药物疾病关联预测方法的AUC指标为0.866 2,较单一特征的平均预测指标最大相对提升为2.48%,较4种代表性基线方法的指标最大相对提升为1.67%;AUPR指标为0.341 2,较单一特征预测结果最大相对提升为1.67%,较4种代表性基线方法提升27.49%.对预测结果中药物-疾病预测关联得分中排名前10的组合及阿霉素为例的单一药物预测组合进行文献研究验证、临床治疗验证,同样证明MFFGCN在未知的药物疾病关联预测上表现良好,能有效地发现药物的新适应症,为药物重定位提供方法借鉴和理论依据.
BACKGROUND:Medical disputes are a global public health issue that is receiving increasing attention. However, studies investigating the relationship between hospital legal construction and medical disputes are scarce. The development of a multicenter model incorporating machine learning (ML) techniques for the individualized prediction of medical disputes would be beneficial for medical workers. OBJECTIVE:This study aimed to identify predictors related to medical disputes from the perspective of hospital legal construction and the use of ML techniques to build models for predicting the risk of medical disputes. METHODS:This study enrolled 38,053 medical workers from 130 tertiary hospitals in Hunan province, China. The participants were randomly divided into a training cohort (34,286/38,053, 90.1%) and an internal validation cohort (3767/38,053, 9.9%). Medical workers from 87 tertiary hospitals in Beijing were included in an external validation cohort (26,285/26,285, 100%). This study used logistic regression and 5 ML techniques: decision tree, random forest, support vector machine, gradient boosting decision tree (GBDT), and deep neural network. In total, 12 metrics, including discrimination and calibration, were used for performance evaluation. A scoring system was developed to select the optimal model. Shapley additive explanations was used to generate the importance coefficients for characteristics. To promote the clinical practice of our proposed optimal model, reclassification of patients was performed, and a web-based app for medical dispute prediction was created, which can be easily accessed by the public. RESULTS:Medical disputes occurred among 46.06% (17,527/38,053) of the medical workers in Hunan province, China. Among the 26 clinical characteristics, multivariate analysis demonstrated that 18 characteristics were significantly associated with medical disputes, and these characteristics were used for ML model development. Among the ML techniques, GBDT was identified as the optimal model, demonstrating the lowest Brier score (0.205), highest area under the receiver operating characteristic curve (0.738, 95% CI 0.722-0.754), and the largest discrimination slope (0.172) and Youden index (1.355). In addition, it achieved the highest metrics score (63 points), followed by deep neural network (46 points) and random forest (45 points), in the internal validation set. In the external validation set, GBDT still performed comparably, achieving the second highest metrics score (52 points). The high-risk group had more than twice the odds of experiencing medical disputes compared with the low-risk group. CONCLUSIONS:We established a prediction model to stratify medical workers into different risk groups for encountering medical disputes. Among the 5 ML models, GBDT demonstrated the optimal comprehensive performance and was used to construct the web-based app. Our proposed model can serve as a useful tool for identifying medical workers at high risk of medical disputes. We believe that preventive strategies should be implemented for the high-risk group.
Ligand-based virtual screening (LBVS) is a promising approach for rapid and low-cost screening of potentially bioactive molecules in the early stage of drug discovery. Compared with traditional similarity-based machine learning methods, deep learning frameworks for LBVS can more effectively extract high-order molecule structure representations from molecular fingerprints or structures. However, the 3D conformation of a molecule largely influences its bioactivity and physical properties, and has rarely been considered in previous deep learning-based LBVS methods. Moreover, the relative bioactivity benchmark dataset is still lacking. To address these issues, we introduce a novel end-to-end deep learning architecture trained from molecular conformers for LBVS. We first extracted molecule conformers from multiple public molecular bioactivity data and consolidated them into a large-scale bioactivity benchmark dataset, which totally includes millions of endpoints and molecules corresponding to 954 targets. Then, we devised a deep learning-based LBVS called EquiVS to learn molecule representations from conformers for bioactivity prediction. Specifically, graph convolutional network (GCN) and equivariant graph neural network (EGNN) are sequentially stacked to learn high-order molecule-level and conformer-level representations, followed with attention-based deep multiple-instance learning (MIL) to aggregate these representations and then predict the potential bioactivity for the query molecule on a given target. We conducted various experiments to validate the data quality of our benchmark dataset, and confirmed EquiVS achieved better performance compared with 10 traditional machine learning or deep learning-based LBVS methods. Further ablation studies demonstrate the significant contribution of molecular conformation for bioactivity prediction, as well as the reasonability and non-redundancy of deep learning architecture in EquiVS. Finally, a model interpretation case study on CDK2 shows the potential of EquiVS in optimal conformer discovery. The overall study shows that our proposed benchmark dataset and EquiVS method have promising prospects in virtual screening applications.
Introduction: Exploring the potential efficacy of a drug is a valid approach for drug development with shorter development times and lower costs. Recently, several computational drug repositioning methods have been introduced to learn multi-features for potential association prediction. However, fully leveraging the vast amount of information in the scientific literature to enhance drug-disease association prediction is a great challenge.Methods: We constructed a drug-disease association prediction method called Literature Based Multi-Feature Fusion (LBMFF), which effectively integrated known drugs, diseases, side effects and target associations from public databases as well as literature semantic features. Specifically, a pre-training and fine-tuning BERT model was introduced to extract literature semantic information for similarity assessment. Then, we revealed drug and disease embeddings from the constructed fusion similarity matrix by a graph convolutional network with an attention mechanism.Results: LBMFF achieved superior performance in drug-disease association prediction with an AUC value of 0.8818 and an AUPR value of 0.5916.Discussion: LBMFF achieved relative improvements of 31.67% and 16.09%, respectively, over the second-best results, compared to single feature methods and seven existing state-of-the-art prediction methods on the same test datasets. Meanwhile, case studies have verified that LBMFF can discover new associations to accelerate drug development. The proposed benchmark dataset and source code are available at: https://github.com/kang-hongyu/LBMFF.
BACKGROUND:The growing number of patients visiting pediatric emergency departments could have a detrimental impact on the care provided to children who are triaged as needing urgent attention. Therefore, it has become essential to continuously monitor and analyze the admissions and waiting times of pediatric emergency patients. Despite the significant challenge posed by the shortage of pediatric medical resources in China's health care system, there have been few large-scale studies conducted to analyze visits to the pediatric emergency room. OBJECTIVE:This study seeks to examine the characteristics and admission patterns of patients in the pediatric emergency department using electronic medical record (EMR) data. Additionally, it aims to develop and assess machine learning models for predicting waiting times for pediatric emergency department visits. METHODS:This retrospective analysis involved patients who were admitted to the emergency department of Children's Hospital Capital Institute of Pediatrics from January 1, 2021, to December 31, 2021. Clinical data from these admissions were extracted from the electronic medical records, encompassing various variables of interest such as patient demographics, clinical diagnoses, and time stamps of clinical visits. These indicators were collected and compared. Furthermore, we developed and evaluated several computational models for predicting waiting times. RESULTS:In total, 183,024 eligible admissions from 127,368 pediatric patients were included. During the 12-month study period, pediatric emergency department visits were most frequent among children aged less than 5 years, accounting for 71.26% (130,423/183,024) of the total visits. Additionally, there was a higher proportion of male patients (104,147/183,024, 56.90%) compared with female patients (78,877/183,024, 43.10%). Fever (50,715/183,024, 27.71%), respiratory infection (43,269/183,024, 23.64%), celialgia (9560/183,024, 5.22%), and emesis (6898/183,024, 3.77%) were the leading causes of pediatric emergency room visits. The average daily number of admissions was 501.44, and 18.76% (34,339/183,204) of pediatric emergency department visits resulted in discharge without a prescription or further tests. The median waiting time from registration to seeing a doctor was 27.53 minutes. Prolonged waiting times were observed from April to July, coinciding with an increased number of arrivals, primarily for respiratory diseases. In terms of waiting time prediction, machine learning models, specifically random forest, LightGBM, and XGBoost, outperformed regression methods. On average, these models reduced the root-mean-square error by approximately 17.73% (8.951/50.481) and increased the R2 by approximately 29.33% (0.154/0.525). The SHAP method analysis highlighted that the features "wait.green" and "department" had the most significant influence on waiting times. CONCLUSIONS:This study offers a contemporary exploration of pediatric emergency room visits, revealing significant variations in admission rates across different periods and uncovering certain admission patterns. The machine learning models, particularly ensemble methods, delivered more dependable waiting time predictions. Patient volume awaiting consultation or treatment and the triage status emerged as crucial factors contributing to prolonged waiting times. Therefore, strategies such as patient diversion to alleviate congestion in emergency departments and optimizing triage systems to reduce average waiting times remain effective approaches to enhance the quality of pediatric health care services in China.
The identification of gene-disease associations plays an important role in the exploration of pathogenic mech-anisms and therapeutic targets. Computational methods have been regarded as an effective way to discover the potential gene-disease associations in recent years. However, most of them ignored the combination of abundant genetic, therapeutic information, and gene-disease network topology. To this end, we re-organized the current gene-disease association benchmark dataset by extracting the newest gene-disease associations from the OMIM database. Then, we developed a multi-graph representation learning-based ensemble model, named MGREL to predict gene-disease associations. MGREL integrated two feature generation channels to extract gene and disease features, including a knowledge extraction channel which learned high-order representations from genetic and therapeutic information, and a graph learning channel which acquired network topological representations through multiple advanced graph representation learning methods. Then, an ensemble learning method with 5 machine learning models was used as the classifier to predict the gene-disease association. Comprehensive ex-periments have demonstrated the significant performance achieved by MGREL compared to 5 state-of-the-art methods. For the major measurements (AUC = 0.925, AUPR = 0.935), the relative improvements of MGREL compared to the suboptimal methods are 3.24%, and 2.75%, respectively. MGREL also achieved impressive improvements in the challenging tasks of predicting potential associations for unknown genes/diseases. In addition, case studies implied potential applications for MGREL in the discovery of potential therapeutic targets.
BACKGROUND:Urinary tract infection (UTI) is one of major nosocomial infections significantly affecting the outcomes of immobile stroke patients. Previous studies have identified several risk factors, but it is still challenging to accurately estimate personal UTI risk.AIM:To develop predictive models for UTI risk identification for immobile stroke patients.METHODS:Research data were collected from our previous multicentre study. Derivation cohort included 3982 immobile stroke patients collected from November 1st, 2015 to June 30th, 2016; external validation cohort included 3837 patients collected from November 1st, 2016 to July 30th, 2017. Six machine learning models and an ensemble learning model were derived, based on 80% of derivation cohort, and effectiveness was evaluated with the remaining 20%. Shapley additive explanation values were used to determine feature importance and examine the clinical significance of prediction models.FINDINGS:In all, 2.59% (103/3982) patients were diagnosed with UTI in derivation cohort, 1.38% (53/3837) in external cohort. The ensemble learning model performed the best in area under the receiver operating characteristic (ROC) curve in internal validation (82.2%); second best in external validation (80.8%). In addition, the ensemble learning model performed the best sensitivity in both internal and external validation sets (80.9% and 81.1%, respectively). Seven UTI risk factors (pneumonia, glucocorticoid use, female sex, mixed cerebrovascular disease, increased age, prolonged length of stay, and duration of catheterization) were also identified.CONCLUSION:This ensemble learning model demonstrated promising performance. Future work should continue to develop a more concise scoring tool based on machine learning models and prospectively examining the model in practical use, thus improving clinical outcomes.
Computational drug repositioning is an effective way to find new indications for existing drugs, thus can accelerate drug development and reduce experimental costs. Recently, various deep learning-based repurposing methods have been established to identify the potential drug-disease associations (DDA). However, effective utilization of the relations of biological entities to capture the biological interactions to enhance the drug-disease association prediction is still challenging. To resolve the above problem, we proposed a heterogeneous graph neural network called REDDA (Relations-Enhanced Drug-Disease Association prediction). Assembled with three attention mechanisms, REDDA can sequentially learn drug/disease representations by a general heterogeneous graph convolutional network-based node embedding block, a topological subnet embedding block, a graph attention block, and a layer attention block. Performance comparisons on our proposed benchmark dataset show that REDDA outperforms 8 advanced drug-disease association prediction methods, achieving relative improvements of 0.76% on the area under the receiver operating characteristic curve (AUC) score and 13.92% on the precision-recall curve (AUPR) score compared to the suboptimal method. On the other benchmark dataset, REDDA also obtains relative improvements of 2.48% on the AUC score and 4.93% on the AUPR score. Specifically, case studies also indicate that REDDA can give valid predictions for the discovery of -new indications for drugs and new therapies for diseases. The overall results provide an inspiring potential for REDDA in the in silico drug development. The proposed benchmark dataset and source code are available in https://github.com/gu-yaowen/REDDA.
Stroke patients tend to suffer from immobility, which increases the possibility of post-stroke complications. Urinary tract infections (UTIs) are one of the complications as an independent predictor of poor prognosis of stroke patients. However, the incidence of new UTIs onsets during hospitalization was rare in most datasets with a prevalence of 4%. This imbalanced data distribution sets obstacles to establishing an accurate prediction model. Our study aimed to develop an effective prediction model to identify UTIs risk in immobile stroke patients, and (2) to compare its prediction performance with traditional machine learning models. We tackled this problem by building a Siamese Network leveraging commonly used clinical features to identifying patients with UTIs risk. Model derivation and validation were based on a nationwide dataset including 3982 Chinese patients. Results showed that the Siamese Network performed better than traditional machine learning models in imbalanced datasets (Sensitivity: 0.810; AUC: 0.828).
Computational methods have been widely applied to resolve various core issues in drug discovery, such as molecular property prediction. In recent years, a data-driven computational method-deep learning had achieved a number of impressive successes in various domains. In drug discovery, graph neural networks (GNNs) take molecular graph data as input and learn graph-level representations in non-Euclidean space. An enormous amount of well-performed GNNs have been proposed for molecular graph learning. Meanwhile, efficient use of molecular data during training process, however, has not been paid enough attention. Curriculum learning (CL) is proposed as a training strategy by rearranging training queue based on calculated samples' difficulties, yet the effectiveness of CL method has not been determined in molecular graph learning. In this study, inspired by chemical domain knowledge and task prior information, we proposed a novel CL-based training strategy to improve the training efficiency of molecular graph learning, called CurrMG. Consisting of a difficulty measurer and a training scheduler, CurrMG is designed as a plug-and-play module, which is model-independent and easy-to-use on molecular data. Extensive experiments demonstrated that molecular graph learning models could benefit from CurrMG and gain noticeable improvement on five GNN models and eight molecular property prediction tasks (overall improvement is 4.08%). We further observed CurrMG's encouraging potential in resource-constrained molecular property prediction. These results indicate that CurrMG can be used as a reliable and efficient training strategy for molecular graph learning. Availability: The source code is available in https://github.com/gu-yaowen/CurrMG.
[目的]构建和比较抗结核杆菌药物虚拟筛选模型,助力抗结核药物的研发.[方法]提出一种基于课程式学习优化的图神经网络模型GNN-MTB,用于抗结核杆菌抑制剂的虚拟筛选.进一步,从开放数据库中收集整理抗结核杆菌药物筛选相关基准数据集,将GNN-MTB模型与4种常规机器学习模型和两种图神经网络模型在基准数据集上进行性能比较.[结果]对10 789条抗结核杆菌药物虚拟筛选实验数据的分析结果显示,GNN-MTB模型的预测性能(AUC为0.912,AUPR为0.679)优于传统的机器学习模型和图神经网络模型的性能表现(平均AUC为0.878~0.900,平均AUPR为0.600~0.673),平均AUC和AUPR的最大提升幅度达3.872%和13.167%.同时,开源GNN-MTB模型并构建抗结核杆菌药物虚拟筛选预测工具以供广大抗结核杆菌药物研究者使用.[局限]未纳入药物敏感性和菌株耐药性相关分析.[结论]GNN-MTB模型取得良好性能,可探索将其应用于抗结核病药物研发.同时,研究框架也可为其他疾病药物的虚拟筛选提供参考.
介绍自编码器、生成式对抗网络、BERT等无监督深度学习方法,阐述其在电子健康档案数据挖掘中的应用以及存在的挑战,指出无监督深度学习技术能够加速医疗知识发现和临床决策支持,促进个性化医学发展.