Artificial intelligence (AI) is an effective tool to accelerate drug discovery and cut costs in discovery processes. Many successful AI applications are reported in the early stages of small molecule drug discovery. However, most of those applications require a deep understanding of software and hardware, and focus on a single field that implies data normalization and transfer between those applications is still a challenge for normal users. It usually limits the application of AI in drug discovery. Here, based on a series of robust models, we formed a one-stop, general purpose, and AI-based drug discovery platform, MolProphet, to provide complete functionalities in the early stages of small molecule drug discovery, including AI-based target pocket prediction, hit discovery and lead optimization, and compound targeting, as well as abundant analyzing tools to check the results. MolProphet is an accessible and user-friendly web-based platform that is fully designed according to the practices in the drug discovery industry. The molecule screened, generated, or optimized by the MolProphet is purchasable and synthesizable at low cost but with good drug-likeness. More than 400 users from industry and academia have used MolProphet in their work. We hope this platform can provide a powerful solution to assist each normal researcher in drug design and related research areas. It is available for everyone at https://www.molprophet.com/.
The global health implications of fine particulate matter (PM2.5) underscore the imperative need for research into its toxicity and chemical composition. In this study, zebrafish embryos exposed to the water-soluble components of PM2.5 from two cities (Harbin and Hangzhou) with differences in air quality, underwent microscopic examination to identify primary target organs. The Harbin PM2.5 induced dose-dependent organ malformation in zebrafish, indicating a higher level of toxicity than that of the Hangzhou sample. Harbin PM2.5 led to severe deformities such as pericardial edema and a high mortality rate, while the Hangzhou sample exhibited hepatotoxicity, causing delayed yolk sac absorption. The experimental determination of PM2.5 constituents was followed by the application of four algorithms for predictive toxicological assessment. The random forest algorithm correctly predicted each of the effect classes and showed the best performance, suggesting that zebrafish malformation rates were strongly correlated with water-soluble components of PM2.5. Feature selection identified the water-soluble ions F- and Cl- and metallic elements Al, K, Mn, and Be as potential key components affecting zebrafish development. This study provides new insights into the developmental toxicity of PM2.5 and offers a new approach for predicting and exploring the health effects of PM2.5.
目的 应用随机森林模型和Logistic回归模型分析新型冠状病毒肺炎(COVID-19)发病的影响因素,为COVID-19防治提供依据.方法 选择2020年1月17日—2月17日浙江省6家医院收治的COVID-19疑似病例为研究对象,通过问卷收集人口学资料、既往基础疾病、流行病学史、临床表现、实验室检测指标和肺部影像学表现等,分别建立随机森林模型和Logistic回归模型分析COVID-19发病的影响因素.结果 共纳入786例COVID-19疑似病例,其中确诊病例336例,占42.75%.随机森林模型分析结果显示,COVID-19发病的影响因素重要性排名前十位依次是白细胞计数正常或减少、胸部影像学表现、淋巴细胞计数减少、聚集性发病、发病前14天内接触来自疫区的发热或呼吸道症状患者、乏力、发病前14天内有疫区旅居史、呼吸困难、鼻塞流涕和肌肉酸痛.多因素Logistic回归模型分析结果显示,发病前14天内有疫区旅居史(OR=8.440,95%CI:4.204~16.944)、发病前14天内接触来自疫区的发热或呼吸道症状患者(OR=2.967,95%CI:1.630~5.402)、聚集性发病(OR=25.164,95%CI:11.833~53.516)、乏力(OR=2.710,95%CI:1.490~4.930)、呼吸困难(OR=5.276,95%CI:2.076~13.410)、肌肉酸痛(OR=14.187,95%CI:1.998~100.730)、白细胞计数正常或减少(OR=1.750,95%CI:1.659~1.852)和胸部影像学表现(OR=6.291,95%CI:4.315~9.171)均与COVID-19存在统计学关联.结论 两种模型的分析结果相似,随机森林模型显示各个因素在COVID-19发病中的重要程度,Logistic回归模型可直观解释不同因素的风险度.
Early determination of coronavirus disease 2019 (COVID-19) pneumonia from numerous suspected cases is critical for the early isolation and treatment of patients. The purpose of the study was to develop and validate a rapid screening model to predict early COVID-19 pneumonia from suspected cases using a random forest algorithm in China. A total of 914 initially suspected COVID-19 pneumonia in multiple centers were prospectively included. The computer-assisted embedding method was used to screen the variables. The random forest algorithm was adopted to build a rapid screening model based on the training set. The screening model was evaluated by the confusion matrix and receiver operating characteristic (ROC) analysis in the validation. The rapid screening model was set up based on 4 epidemiological features, 3 clinical manifestations, decreased white blood cell count and lymphocytes, and imaging changes on chest X-ray or computed tomography. The area under the ROC curve was 0.956, and the model had a sensitivity of 83.82% and a specificity of 89.57%. The confusion matrix revealed that the prospective screening model had an accuracy of 87.0% for predicting early COVID-19 pneumonia. Here, we developed and validated a rapid screening model that could predict early COVID-19 pneumonia with high sensitivity and specificity. The use of this model to screen for COVID-19 pneumonia have epidemiological and clinical significance.
Objective: To improve the timeliness for the early COVID-19 infection diagnosis, it is essential to develop a decision- making tool to assist early diagnosis of COVID-19 patients in fever clinics. Materials and methods: This paper aims at extracting risk factors from clinical data of 912 early COVID-19 infected patients and utilizing four types of traditional machine learning approaches including Logistic Regression (LR), Support Vector Machine (SVM), Decision Tree (DT), Random Forest (RF) and a deep learning-based method for diagnosis of early COVID-19. Results: The results show that the LR predictive model presents a higher specificity rate of 0.95, an Area Under the receiver operating Curve (AUC) of 0.971 and an improved sensitivity rate of 0.82, which makes it optimal for the screening of early COVID-19 infection. We also perform the verification for generality of the best model (LR predictive model) among Zhejiang population, and analyse the contribution of the factors to the predictive models. Discussions: Under the background of COVID-19 pandemic, the early diagnosis of COVID-19 still face severe challenges, a decision-making tool assisting early diagnosis of COVID-19 patients is vital for fever clinics. Conclusions: Our manuscript describes and highlights the ability of machine learning methods for improving the accuracy and timeliness of early COVID-19 infection diagnosis. The higher AUC of our LR-base predictive model makes it a more conducive method for assisting COVID-19 diagnosis. The optimal model has been encapsulated as a mobile application (APP) and implemented in some hospitals in Zhejiang Province.
Novel coronavirus pneumonia (NCP) has been widely spread in China and several other countries. Early finding of this pneumonia from huge numbers of suspects gives clinicians a big challenge. The aim of the study was to develop a rapid screening model for early predicting NCP in a Zhejiang population, as well as its utility in other areas. A total of 880 participants who were initially suspected of NCP from January 17 to February 19 were included. Potential predictors were selected via stepwise logistic regression analysis. The model was established based on epidemiological features, clinical manifestations, white blood cell count, and pulmonary imaging changes, with the area under receiver operating characteristic (AUROC) curve of 0.920. At a cut-off value of 1.0, the model could determine NCP with a sensitivity of 85% and a specificity of 82.3%. We further developed a simplified model by combining the geographical regions and rounding the coefficients, with the AUROC of 0.909, as well as a model without epidemiological factors with the AUROC of 0.859. The study demonstrated that the screening model was a helpful and cost-effective tool for early predicting NCP and had great clinical significance given the high activity of NCP.
COVID-19 is a newly emerging infectious disease, which is generally susceptible to human beings and has caused huge losses to people's health. Acute respiratory distress syndrome (ARDS) is one of the common clinical manifestations of severe COVID-19 and it is also responsible for the current shortage of ventilators worldwide. This study aims to analyze the clinical characteristics of COVID-19 ARDS patients and establish a diagnostic system based on artificial intelligence (AI) method to predict the probability of ARDS in COVID-19 patients. We collected clinical data of 659 COVID-19 patients from 11 regions in China. The clinical characteristics of the ARDS group and no-ARDS group of COVID-19 patients were elaborately compared and both traditional machine learning algorithms and deep learning-based method were used to build the prediction models. Results indicated that the median age of ARDS patients was 56.5 years old, which was significantly older than those with non-ARDS by 7.5 years. Male and patients with BMI > 25 were more likely to develop ARDS. The clinical features of ARDS patients included cough (80.3%), polypnea (59.2%), lung consolidation (53.9%), secondary bacterial infection (30.3%), and comorbidities such as hypertension (48.7%). Abnormal biochemical indicators such as lymphocyte count, CK, NLR, AST, LDH, and CRP were all strongly related to the aggravation of ARDS. Furthermore, through various AI methods for modeling and prediction effect evaluation based on the above risk factors, decision tree achieved the best AUC, accuracy, sensitivity and specificity in identifying the mild patients who were easy to develop ARDS, which undoubtedly helped to deliver proper care and optimize use of limited resources.
With the dramatically fast spread of COVID-9, real-time reverse transcription polymerase chain reaction (RT-PCR) test has become the gold standard method for confirmation of COVID-19 infection. However, RT-PCR tests are complicated in operation andIt usually takes 5-6 hours or even longer to get the result. Additionally, due to the low virus loads in early COVID-19 patients, RT-PCR tests display false negative results in a number of cases. Analyzing complex medical datasets based on machine learning provides health care workers excellent opportunities for developing a simple and efficient COVID-19 diagnostic system. This paper aims at extracting risk factors from clinical data of early COVID-19 infected patients and utilizing four types of traditional machine learning approaches including logistic regression(LR), support vector machine(SVM), decision tree(DT), random forest(RF) and a deep learning-based method for diagnosis of early COVID-19. The results show that the LR predictive model presents a higher specificity rate of 0.95, an area under the receiver operating curve (AUC) of 0.971 and an improved sensitivity rate of 0.82, which makes it optimal for the screening of early COVID-19 infection. We also perform the verification for generality of the best model (LR predictive model) among Zhejiang population, and analyze the contribution of the factors to the predictive models. Our manuscript describes and highlights the ability of machine learning methods for improving the accuracy and timeliness of early COVID-19 infection diagnosis. The higher AUC of our LR-base predictive model makes it a more conducive method for assisting COVID-19 diagnosis. The optimal model has been encapsulated as a mobile application (APP) and implemented in some hospitals in Zhejiang Province.
Due to the limitation of self-repairing capability for cartilage injury, the construction of tissue engineering in vitro has been an ideal treatment to repair tissue injury. In this paper, hydroxyapatite (Hap) and chitosan (Chi) were selected to fabricate the scaffold through low temperature deposition manufacturing (LDM) technique. The scaffold was characterized with interconnected structure and high porosity, as well as lower toxicity to cells (TDC-5-EGPE). Animal experiment was performed, Twelve white New Zealand rabbits were randomly divided into two groups, the side of the thyroid cartilage was removed, Chi-HAP composite scaffold was implanted into the cartilage defect as the experimental group A. Group B was treated for thyroid cartilage defects without any treatment. After 10 weeks, hematoxylin-eosin(HE) staining and S-O staining were carried out on the injured tissues. The result showed that newborn chondrocytes were found in repaired areas for group A, and there are no new cells found for group B. Therefore, Chi-HAP composite scaffolds formed by LDM possess biological activity for repairing injury cartilage.