Accurate diagnosis of lung nodules is essential for detection and assessment of lung cancer. The present contribution proposes a descriptive model for diagnostic classification of lung nodules by jointly using deep and spectral features from the 3D surface structure of nodules. To the best of our knowledge, this is the first work that utilizes a point cloud (PC)-based deep network for extracting nodule shape features. The PC-based deep network takes into account the 3D context of a nodule; meanwhile, it is extensively less computationally intensive. The spectral features prevent over-fitting, a common problem of deep networks trained by relatively small dataset in the medical imaging domain, and compensates for missing information of mesh connections. Experimental results reveal that our descriptive model demonstrates high sensitivity (87.23%) as well as high specificity (89.80%) with a total accuracy of 88.54% for reliable and accurate prediction of lung nodule malignancy.
OBJECTIVE:To develop machine learning models for classifying the severity of opioid overdose events from clinical data. MATERIALS AND METHODS:Opioid overdoses were identified by diagnoses codes from the Marshfield Clinic population and assigned a severity score via chart review to form a gold standard set of labels. Three primary feature sets were constructed from disparate data sources surrounding each event and used to train machine learning models for phenotyping. RESULTS:Random forest and penalized logistic regression models gave the best performance with cross-validated mean areas under the ROC curves (AUCs) for all severity classes of 0.893 and 0.882 respectively. Features derived from a common data model outperformed features collected from disparate data sources for the same cohort of patients (AUCs 0.893 versus 0.837, p value = 0.002). The addition of features extracted from free text to machine learning models also increased AUCs from 0.827 to 0.893 (p value < 0.0001). Key word features extracted using natural language processing (NLP) such as 'Narcan' and 'Endotracheal Tube' are important for classifying overdose event severity. CONCLUSION:Random forest models using features derived from a common data model and free text can be effective for classifying opioid overdose events.
Calciphylaxis is a disorder that results in necrotic cutaneous lesions with a high rate of mortality. Due to its rarity and complexity, the risk factors for and the disease mechanism of calciphylaxis are not fully understood. This work focuses on the use of machine learning to both predict disease risk and model the contributing factors learned from an electronic health record data set. We present the results of four modeling approaches on several subpopulations of patients with chronic kidney disease (CKD). We find that modeling calciphylaxis risk with random forests learned from binary feature data produces strong models, and in the case of predicting calciphylaxis development among stage 4 CKD patients, we achieve an AUC-ROC of 0.8718. This ability to successfully predict calciphylaxis may provide an excellent opportunity for clinical translation of the predictive models presented in this paper.
Blindness or vision impairment, one of the top ten disabilities among men and women, targets more than 7 million Americans of all ages. Accessible visual information is of paramount importance to improve independence and safety of blind and visually impaired people, and there is a pressing need to develop smart automated systems to assist their navigation, specifically in unfamiliar healthcare environments, such as clinics, hospitals, and urgent cares. This contribution focused on developing computer vision algorithms composed with a deep neural network to assist visually impaired individual's mobility in clinical environments by accurately detecting doors, stairs, and signages, the most remarkable landmarks. Quantitative experiments demonstrate that with enough number of training samples, the network recognizes the objects of interest with an accuracy of over 98% within a fraction of a second.
Entity matching (EM) nds disparate data instances that refer to the same real-world entity. EM is critical in health informatics, and will become even more so in the age of Big Data and data science. Many EM systems have been developed. In this paper, we rst discuss why it is still very dicult for domain scientists to use such EM systems. We then describe CloudMatcher, a cloud/crowd service for EM that we have been building. CloudMatcher aims to be a fast, easy-to-use, scalable, and highly available EM service on the Web. We motivate CloudMatcher then describe its design and implementation. Next, we describe its deployment in the past six months, providing a detailed analysis of its performance over four representative datasets. Finally, we discuss lessons learned.
Machine learning as an advanced computational technology has been around for several years in discovering patterns from diverse biomedical data sources and providing excellent capabilities ranging from gene annotation to predictive phenotyping. However, machine learning strategies remain underused in small and medium-scale biomedical research labs where they have been collaboratively providing a reasonable amount of scientific knowledge. While most machine learning algorithms are complicated in code, theses labs and individual researchers could accomplish iterative data analysis using different machine learning techniques if they had access to highly available machine learning components and powerful computational infrastructures. In this contribution, we provide a comparison of several state-of-the-art Machine Learning-as-a-Service platforms along with their capabilities in medical informatics. In addition, we performed several analyses to examine the qualitative and quantitative attributes of two Machine Learning-as-a-Service environments namely “BigML” and “Algorithmia”.
Every single day, a massive amount of text data is generated by different medical data sources, such as scientific literature, medical web pages, health-related social media, clinical notes, and drug reviews. Processing this wealth of data is indeed a daunting task, and it forces us to adopt smart and scalable computational strategies, including machine intelligence, big data analytics, and distributed architecture. In this contribution, we designed and developed an open-source big data neural network toolkit, namely bigNN which tackles the problem of large-scale biomedical text classification in an efficient fashion, facilitating fast prototyping and reproducible text analytics researches. bigNN scales up a word2vec-based neural network model over Apache Spark 2.10 and Hadoop Distributed File System (HDFS) 2.7.3, allowing for more efficient big data sentence classification. The toolkit supports big data computing, and simplifies rapid application development in sentence analysis by allowing users to configure and examine different internal parameters of both Apache Spark and the neural network model. bigNN is fully documented, and it is publicly and freely available at https://github.com/bircatmcri/bigNN.
Background: The study of adverse drug events (ADEs) is a tenured topic in medical literature. In recent years, increasing numbers of scientific articles and health-related social media posts have been generated and shared daily, albeit with very limited use for ADE study and with little known about the content with respect to ADEs.Objective: The aim of this study was to develop a big data analytics strategy that mines the content of scientific articles and health-related Web-based social media to detect and identify ADEs.Methods: We analyzed the following two data sources: (1) biomedical articles and (2) health-related social media blog posts. We developed an intelligent and scalable text mining solution on big data infrastructures composed of Apache Spark, natural language processing, and machine learning. This was combined with an Elasticsearch No-SQL distributed database to explore and visualize ADEs.Results: The accuracy, precision, recall, and area under receiver operating characteristic of the system were 92.7%, 93.6%, 93.0%, and 0.905, respectively, and showed better results in comparison with traditional approaches in the literature. This work not only detected and classified ADE sentences from big data biomedical literature but also scientifically visualized ADE interactions.Conclusions: To the best of our knowledge, this work is the first to investigate a big data machine learning strategy for ADE discovery on massive datasets downloaded from PubMed Central and social media. This contribution illustrates possible capacities in big data biomedical text analysis using advanced computational methods with real-time update from new data published on a daily basis.
Antibacterial fluoromicas were prepared by ion-exchanging fluoromicas with different antibacterial agents including various quaternary ammonium compounds, AgNO3, and norfloxacin. Antibacterial activities of the ion-exchanged fluoromicas were determined against Staphylococcus aureus and Escherichia coli. Minimum inhibitory concentration (MIC) and zone of inhibition (ZOI) tests were performed to determine both antibacterial effectiveness and mode of action associated with the fluoromicas. All treated fluoromicas showed excellent antibacterial activities against both types of bacteria. The antibacterial activities of treated fluoromicas were found to be either better than or the same as those of neat antibacterial agents. The repeated antibacterial activity tests demonstrated the extended activity of these systems.
A genome-wide scan in 60 bipolar affective disorder (BPAD) affected sib-pairs (ASPs) identified linkage on chromosome 21 at 21q22 (D21S1446, NPL = 1.42, P = 0.08), a BPAD susceptibility locus supported by multiple studies. Although this linkage only approaches significance, the peak marker is located 12 Kb upstream of S100B, a neurotrophic factor implicated in the pathology of psychiatric disorders, including BPAD and schizophrenia. We hypothesized that the linkage signal at 21q22 may result from pathogenic disease variants within S100B and performed an association analysis of this gene in a collection of 125 BPAD type I trios. S100B single nucleotide polymorphisms (SNPs) rs2839350 (P = 0.022) and rs3788266 (P = 0.031) were significantly associated with BPAD. Since variants within S100B have also been associated with schizophrenia susceptibility, we reanalyzed the data in trios with a history of psychosis, a phenotype in common between the two disorders. SNPs rs2339350 (P = 0.016) and rs3788266 (P = 0.009) were more significantly associated in the psychotic subset. Increased significance was also obtained at the haplotype level. Interestingly, SNP rs3788266 is located within a consensus-binding site for Six-family transcription factors suggesting that this variant may directly affect S100B gene expression. Fine-mapping analyses of 21q22 have previously identified transient receptor potential gene melastatin 2 (TRPM2), which is 2 Mb upstream of S100B, as a possible BPAD susceptibility gene at 21q22. We also performed a family-based association analysis of TRPM2 which did not reveal any evidence for association of this gene with BPAD. Overall, our findings suggest that variants within the S100B gene predispose to a psychotic subtype of BPAD, possibly via alteration of gene expression.
Bipolar disorder (BPD) is a complex genetic disorder with cycling symptoms of depression and mania. Despite the extreme complexity of this psychiatric disorder, attempts to localize genes which confer vulnerability to the disorder have had some success. Chromosomal regions including 4p16, 12q24, 18p11, 18q22, and 21q21 have been repeatedly linked to BPD in different populations. Here we present the results of a whole genome scan for linkage to BPD in an Irish population. Our most significant result was at 14q24 which yielded a non-parametric LOD (NPL) score of 3.27 at the D14S588 marker with a nominal P-value of 0.0006 under a narrow (bipolar type I only) model of affection. We previously reported linkage to 14q22-24 in a subset of the families tested in this analysis. We also obtained suggestive evidence for linkage at 4q21, 9p21, 12q24, and 16p13, chromosomal regions that have all been previously linked to BPD. Additionally, we report on a novel approach to linkage analysis, STRUCTURE-Guided Linkage Analysis (SGLA), which is designed to reduce genetic heterogeneity and increase the power to detect linkage. Application of this technique resulted in more highly significant evidence for linkage of BPD to three regions including 16p13, a locus that has been repeatedly linked to numerous psychiatric disorders.