Early diagnosis of cancer is crucial to improving the long‐term survival rate of patients. However, commonly used tumor markers lack sensitivity and specificity for screening purposes. Herein, 10 diagnostic models for 10 common types of cancer are developed by extreme gradient boosting, incorporating 66 laboratory parameters. The datasets consist of a retrospective cohort of 737 503 training and 184 012 validation cases, and a prospective cohort of 174 894 cases for model testing. The areas under the curve of the 10 diagnostic models range from 0.763 to 0.993. Notably, the different models have varying numbers of identical parameters among the 66 test features. Additionally, SHapley Additive exPlanation analysis reveals that 54 nontumor markers contributed significantly to the models. Cosine similarity analysis and clustering analysis demonstrate that some of the 10 cancers share common pathophysiological characteristics. Feature‐based inference graph models are thus performed and infer relationships between nontumor index parameters and cancers with strong correlations. In conclusion, a machine learning‐based pan‐cancer early warning system has been established in this study, which can guide doctors in selecting more accurate testing indicators and assessing the risk of 10 types of cancer with greater precision.
BACKGROUND:The impact of COVID-19 on public health has mandated an 'all hands on deck' scientific response. The current clinical study and basic research on COVID-19 are mainly based on existing publications or our knowledge of coronavirus. However, efficiently retrieval of accurate, relevant knowledge on COVID-19 can pose significant challenges for researchers.METHODS:To improve quality in accessing important literature findings, we developed a novel natural language processing (NLP) method to automatically recognize the associations among potential targeted host organ systems, associated clinical manifestations, and pathways. We further validated these associations through clinician experts' evaluations and prioritize candidate drug targets through bioinformatics network analysis.RESULTS:We found that the angiotensin-converting enzyme 2 (ACE2), a receptor that SARS-CoV-2 required for cell entry, is associated with cardiovascular and endocrine organ system and diseases. Furthermore, we found SARS-CoV-2 is associated with some important pathways such as IL-6, TNF-alpha, and IL-1 beta-induced dyslipidemia, which are related to inflammation, lipogenesis, and oxidative stress mechanisms, suggesting potential drug candidates.CONCLUSION:We prioritized the list of therapeutic targets involved in antiviral and immune modulating drugs for experimental validation, rendering it valuable during public health crises marked by stresses on clinical and research capacity. Our automatic intelligence pipeline also contributes to other novel and emerging disease management and treatments in the future.
Laboratory medicine plays an important role in clinical diagnosis. However, no laboratory‐based artificial intelligence (AI) diagnostic system has been applied in current clinical practice due to the lack of robustness and interpretability. Although many attempts have been made, it is still difficult for doctors to adopt the existing machine learning (ML) patterns in interpreting laboratory (lab) big data. Here, a knowledge‐and‐data‐driven laboratory diagnostic system is developed, termed AI‐based Lab tEst tO diagNosis (AI LEON), by integrating an innovative knowledge graph analysis framework and “mixed XGboost and Genetic Algorithm (MiXG)” technique to simulate the doctor's laboratory‐based diagnosis. To establish AI LEON, we included 89 116 949 laboratory data and 10 423 581 diagnosis data points from 730 113 participants. Among them, 686 626 participants were recruited for training and validating purposes with the remaining for testing purposes. AI LEON automatically identified and analyzed 2071 lab indexes, resulting in multiple disease recommendations that involved 441 common diseases in ten organ systems. AI LEON exhibited outstanding transparency and interpretability in three universal clinical application scenarios and outperformed human physicians in interpreting lab reports. AI LEON is an advanced intelligent system that enables a comprehensive interpretation of lab big data, which substantially improves the clinical diagnosis.
Early diagnosis and clear differentiation of pancreatic ductal adenocarcinoma (PDAC) from chronic pancreatitis (CP) is clinically challenging. A machine learning model is developed for the diagnosis of PDAC. The model is induced using a dataset of 13 987 participants, of which 12 402 are used for training the model and the remaining 1585 for testing purposes. One thousand sixty‐six laboratory variables are reduced to 18 measures using standard filtering and feature importance methods. Then, five machine learning classifiers are evaluated for the study. Hyperparameter optimization for each classifier is carried out, and the optimal algorithm is established using a tenfold cross validation on the training data. Finally, gradient boosting decision tree‐based ternary classifier composed of 18 routine laboratory variables (GBDT‐TC 18 ) is established. In the test cohort, GBDT‐TC 18 differentiates PDAC from CP and healthy control (HC) with an accuracy better than carbohydrate antigen 19‐9 (CA19‐9)‐based diagnosis. It also maintains a high diagnostic accuracy for stages I, IIA, and IIB PDAC, small‐sized PDAC, body and tail adenocarcinoma, CA19‐9‐negative PDAC, and nonjaundice PDAC. What's more, GBDT‐TC 18 shows a higher accuracy than CA19‐9 in distinguishing PDAC from CP. GBDT‐TC 18 can be used to augment the capability of doctors for early and differential diagnosis of PDAC.
Efficiently learning representations of clinical concepts (i. e., symptoms, lab test, etc.) from unstructured clinical notes of electronic health record (EHR) data remain significant challenges, since each patient may have multiple visits at different times and each visit may contain different sequential concepts. Therefore, learning distributed representations from temporal patterns of clinical notes is an essential step for downstream applications on EHR data. However, existing methods for EHR representation learning can not adequately capture either contextual information per-visit or temporal information at multiple visits. In this study, we developed a new vector embedding method called EHR2Vec that can learn semantically-meaningful representations of clinical concepts. EHR2Vec incorporated the self-attention structure and showed its utility in accurately identifying relevant clinical concept entities considering time sequence information from multiple visits. Using EHR data from systemic lupus erythematosus (SLE) patients as a case study, we showed EHR2Vec outperforms in identifying interpretable representations compared to other well-known methods including Word2Vec and Med2Vec, according to clinical experts' evaluations.
目的 提出一种基于注意力机制的药物词向量生成模型Drug2vec,对药物信息做向量化表示,并与Word2vec和Med2vec模型比较向量转化效果.方法 使用注意力机制捕获医疗实体对中心词的作用,提出Drug2vec模型,将非结构化电子病历中的医疗实体转化为向量.使用包含14219例系统性红斑狼疮(SLE)患者和963个药物实体的数据集测试Drug2vec模型生成词向量的效果,并且与广泛应用的语言概念空间向量转化模型Word2vec和Med2vec进行对比.结果 在SLE患者数据集中,Drug2vec模型产生的药物词向量准确度优于Word2vec和Med2vec模型.药物词向量相似度排序结果显示Drug2vec模型产生的向量结果符合临床医师的用药顺序.结论 Drug2vec模型可以更精确地利用周围实体修正中心药物实体,从而产生更准确的药物向量.
Drug-drug interactions (DDIs) are one of the indispensable factors leading to adverse event reactions. Considering the unique structure of AERS (Food and Drug Administration Adverse Event Reporting System (FDA AERS)) reports, we changed the scope of the window value in the original skip-gram algorithm, then propose a language concept representation model and extract features of drug name and reaction information from large-scale AERS reports. The validation of our scheme was tested and verified by comparing with vectors originated from the cooccurrence matrix in tenfold cross-validation. In the verification of description enrichment of the DrugBank DDI database, accuracy was calculated for measurement. The average area under the receiver operating characteristic curve of logistic regression classifiers based on the proposed language model is 6% higher than that of the cooccurrence matrix. At the same time, the average accuracy in five severe adverse event classes is 88%. These results indicate that our language model can be useful for extracting drug and reaction features from large-scale AERS reports.
Background As the prototype of autoimmune disease, systemic lupus erythematosus (SLE) has complex and diverse clinical manifestations which may be harmful even life-threaten. Its pitiful that we can just passively respond to these serious complications. It will be a great advantage if the high-risk groups could be predicated and prevented with pre-treatment. The raising of risk prediction models depends on the collection of patient phenotypes, which are scattered in various forms and very cumbersome. In this study, we collected the largest database of complete medical record of inpatients of lupus in China. The clinical phenotype database was generated by using natural language processing (NLP) techniques, then lupus nephritis (LN) prediction model was built. Methods A total of 14,439 SLE patients were collected from the rheumatology and immunology departments of 13 Chinese tertiary hospitals in this study, including 13 062 females (90.46%), with an average age of 33.4 years, and the time span of EMR (Electronic Medical Records) was from October 28, 2001 to March 31, 2017. It includes basic information about patients, physical examination, inspection and diagnostic information, etc. We designed a hybrid NLP system combined NLP technical and expert knowledge at the same time, which was named as Deep Phenotyping System (DPS), to extract all the phenotypic information recorded in EMR. Based on these standard formatted entities, the machine learning and deep learning prediction methods are used to predict the LN in SLE. Results The DPS efficiently processed EMR data, and its accuracy, precision, and recall were each greater than 93%. It extracted 73 794 entities from 14,439 SLE cases, each with time attributes, and produced 18,785,000,640 entities. Thus, a LN prediction model was raised, which the likelihood of lupus patients without nephritis will develop lupus nephritis within half and one year can be predicted.) More than 35 000 phenotypes was used in this model and it was verified with independent samples. The best accuracyACC and area under the curve (AUC) can be achieved 0.88 and 0.86 respectively. Conclusions The comprehensive SLE phenotype database constructed by NLP greatly improves the research efficiency of lupus clinical phenotype. We first proposed a predictive model of lupus nephritis, which is high applicability and efficiency. The experimental results of good close and open testing fully demonstrate the authenticity and practicality of this database. The research process and method based on real world data are also applicable to predict other important complications of lupus. Funding Source(s): None