The goal of this study was to assess trends in postpartum hemorrhage (PPH), its risk factors, and maternal comorbidity burden in the United States using aggregate data from the Evolve to Next-Gen Accrual to Clinical Trials (ENACT) network. This federated network employs interactive querying of electronic health record data repositories in academic medical centers nationwide. We conducted repeated annual cross-sectional analyses to evaluate PPH occurrence and comorbidities across various ethnoracial and sociodemographic groups, starting with a large cohort of 1,287,675 unique delivery hospitalizations collected from 22 ENACT sites between 2005 and 2022. During this time, there was a statistically significant increasing trend in the prevalence of PPH, rising from 5,634 to 10,504 PPH per 100,000 deliveries (Ptrend <0.001). Our findings revealed a continuous upward trend in PPH rates that remained consistent among women with ≥ 1 comorbid conditions (Ptrend <0.001) and those with ≥ 1 maternal risk factor (Ptrend <0.001). This result aligns with prior studies and extends beyond the time periods previously reported. Overall, Native Hawaiian or Other Pacific Islander women had the highest PPH prevalence (~ 13%), followed by Asian (9.8%), American Indian or Alaska Native (8.9%), multirace (8.6%), Black or African American (8.4%) and White (7.4%) women. The top PPH risk factor identified was placenta previa or accreta, while the top comorbidity was antepartum hemorrhage / placental abruption. The most common cause of PPH, namely uterine atony, was prevalent in ENACT data. Our analysis highlights significant ethnoracial disparities and underscores the need for targeted preventative interventions.
Infectious mononucleosis is mostly caused by the Epstein-Barr virus (EBV), and can spread through infected people sharing food and drinks with others. Once this virus gets into your system, it is there to stay. The virus can get activated when a person has low immunity and can cause major complications. Furthermore, if physicians miss the diagnosis of this disease, and prescribe penicillin-based antibiotics, it can cause severe rash and adverse reactions that compromise patient safety. This paper develops a simple Hidden Markov Model using which a Viterbi algorithm provides the maximum a posteriori probability estimate for the most likely hidden state path, given a sequence of symptoms arising as observations from a patient with hidden EBV positive or negative states. Apart from bringing awareness to help reduce missed diagnoses and subsequent adverse events, this work provides a tool for health care systems to better incorporate prompts during electronic medical record (EMR) interactions to help physicians catch potential missed diagnoses during a visit. This research demonstrates how statistical models can be used to assess likelihood of underlying conditions that require tests to be offered by physicians in order to make a definitive diagnosis. The model developed and applied herein for estimating likelihood of EBV infection from a series of observations has the potential to alter guidelines within healthcare systems to ensure that the safety of patients, particularly teens, is not compromised due to a lack of definitive diagnosis for Mono at point of care.
Purpose Chronic low back pain (cLBP) is a common health condition worldwide and a leading cause of disability with an estimated lifetime prevalence of 80-90% in industrialized countries. However, we have had limited success in treating cLBP likely due to its non-specific heterogeneous nature that goes beyond detectable anatomical changes. We propose that omics technologies as precision medicine tools are well suited to provide insight into its pathophysiology and provide diagnostic markers and therapeutic targets. Therefore, in this review, we explore the current state of omics technologies in the diagnosis and classification of cLBP. We identify factors that may serve as markers to differentiate between acute and chronic cases of low back pain (LBP). Finally, we also discuss some challenges that must be overcome to successfully apply precision medicine to the diagnosis and treatment of cLBP. Methods A literature search for the current applications of omics technologies to chronic low back pain was performed using the following search terms-"back pain, " "low back pain, " "proteomics, " "transcriptomics ", "epigenomics, " "genomics, " "omics. " We reviewed molecular markers identified from 35 studies which hold promise in providing information regarding molecular insights into cLBP. Results GWAS studies have found evidence for the role of single nucleotide polymorphisms (SNPs) associated with pain pathways in individuals with cLBP. Epigenomic modifications in patients with cLBP have been found to be enriched among genes involved in immune signaling and inflammation. Transcriptomics profiles of patients with cLBP show multiple lines of evidence for the role of inflammation in cLBP. The glycomics profiles of patients with cLBP are similar to those of patients with inflammatory conditions. Proteomics and microbiomics show promise but have limited studies currently. Conclusion Omics technologies have identified associations between inflammatory and pain pathways in the pathophysiology of cLBP. However, in order to integrate information across the range of studies, it is important for the field to identify and adopt standardized definitions of cLBP and control patients. Additionally, most papers have applied a single omics method to a sampling of cLBP patients which have yielded limited insight into the pathophysiology of cLBP. Therefore, we recommend a multi-omics approach applied to large global consortia for advancing subphenotyping and better management of cLBP, via improved identification of diagnostic markers and therapeutic targets.
Purpose Exposures related to beryllium (Be) are an enduring concern among workers in the nuclear weapons and other high-tech industries, calling for regular and rigorous biological monitoring. Conventional biomonitoring of Be in urine is not informative of cumulative exposure nor health outcomes. Biomarkers of exposure to Be based on non-invasive biomonitoring could help refine disease risk assessment. In a cohort of workers with Be exposure, we employed blood plasma extracellular vesicles (EVs) to discover novel biomarkers of exposure to Be. Methods EVs were isolated from plasma using size-exclusion chromatography and subjected to mass spectrometry-based proteomics. A protein-based classifier was developed using LASSO regression and validated by ELISA. Results We discovered a dual biomarker signature comprising zymogen granule protein 16B and putative protein FAM10A4 that differentiated between Be-exposed and -unexposed subjects. ELISA-based quantification of the biomarkers in an independent cohort of samples confirmed higher expression of the signature in the Be-exposed group, displaying high predictive accuracy (AUROC = 0.919). Furthermore, the biomarkers efficiently discriminated high- and low-exposure groups (AUROC = 0.749). Conclusions This is the first report of EV biomarkers associated with Be exposure and exposure levels. The biomarkers could be implemented in resource-limited settings for Be exposure assessment.
Modeling factors influencing disease phenotypes, from biomarker profiling study datasets, is a critical task in biomedicine. Such datasets are typically generated from high-throughput 'omic' technologies, which help examine disease mechanisms at an unprecedented resolution. These datasets are challenging because they are high-dimensional. The disease mechanisms they study are also complex because many diseases are multifactorial, resulting from the collective activity of several factors, each with a small effect. Bayesian rule learning (BRL) is a rule model inferred from learning Bayesian networks from data, and has been shown to be effective in modeling high-dimensional datasets. However, BRL is not efficient at modeling multifactorial diseases since it suffers from data fragmentation during learning. In this paper, we overcome this limitation by implementing and evaluating three types of ensemble model combination strategies with BRL— uniform combination (UC; same as Bagging), Bayesian model averaging (BMA), and Bayesian model combination (BMC)— collectively called Ensemble Bayesian Rule Learning (EBRL). We also introduce a novel method to visualize EBRL models, called the Bayesian Rule Ensemble Visualizing tool (BREVity), which helps extract interpret the most important rule patterns guiding the predictions made by the ensemble model. Our results using twenty-five public, high-dimensional, gene expression datasets of multifactorial diseases, suggest that, both EBRL models using UC and BMC achieve better predictive performance than BMA and other classic machine learning methods. Furthermore, BMC is found to be more reliable than UC, when the ensemble includes sub-optimal models resulting from the stochasticity of the model search process. Together, EBRL and BREVity provides researchers a promising and novel tool for modeling multifactorial diseases from high-dimensional datasets that leverages strengths of ensemble methods for predictive performance, while also providing interpretable explanations for its predictions.
Background: Ongoing molecular profiling studies enabled by advances in biomedical technologies are producing vast amounts of ‘omic’ data for early detection, monitoring, and prognosis of diverse diseases. A major common limitation is the scarcity of biological samples, necessitating integrative modeling frameworks that can make optimal use of available data for disease classification tasks. Related data sets are often available from different studies, but may have been generated using different technology platforms. Thus, there is a critical need for flexible modeling methods that can handle data from diverse sources to facilitate the discovery of robust biomarkers that underlie disease regulatory processes. Results: In this paper, we introduce a novel framework called Knowledge Augmented Rule Learning (KARL), which incorporates two sources of knowledge, domain, and data, for pattern discovery from small and high-dimensional datasets, such as transcriptomic data. We propose KARL as a transfer rule learning framework in which knowledge of the domain is transferred to the learning process on data in order to 1) improve the reliability of the discovered patterns, and 2) study the knowledge of the domain when used along with data for modeling. In this work, we generated KARL models on gene expression datasets for five types of cancer, including brain, breast, colon, lung, and prostate. As our knowledge of the domain, we used the Ingenuity Knowledge Base (IKB) to extract genes related to hallmarks of cancer and annotated these prior relationships before learning classifiers from these datasets. Conclusions: Our results show that KARL produces, on average, rule models that are more robust classifiers than the baseline without such background knowledge, for our tasks of cancer prediction using 25 publicly available gene expression datasets. Moreover, KARL helped us learn insights about previously known relationships in these gene expression datasets, along with new relationships not input as known, to enable informed biomarker discovery for cancer prediction tasks. KARL can be applied to modeling similar data from any other domain and classification task. Future work would involve extensions to KARL to handle hierarchical knowledge to derive more general hypotheses to drive biomedicine.
Biomarker discovery is critical for both biomedical research and for clinical diagnostic, prognostic, and therapeutic decision-making. They help improve our understanding of the underlying physiological processes within an individual. Discovery of biomarkers from complex biomedical datasets is done using data mining algorithms. Hundreds of thousands of biomarkers have been discovered and reported in literature but only a few dozen have been found to be clinically useful. This discrepancy is because statistical significance is not clinical relevance. Statistical significance only accounts for the correctness of the learned associations. Clinical relevance, in addition to statistical significance, also accounts for clinical utility such as cost-effectiveness, non-invasiveness, efficacy, and safety of the proposed biomarkers. We need models that are statistically significant and clinically relevant, all the while keeping it interpretable. Interpretable classifiers are more actionable in medicine because they offer human-readable explanations for their predictions. Traditional data mining methods cannot account for clinical relevance. We formulate this as a knowledge discovery problem. In computer science, knowledge discovery in databases is “a non-trivial process of the extraction of valid, novel, potentially useful, and ultimately understandable patterns in data”. Bayesian Rule Learning (BRL) finds an optimal Bayesian network to explain the training data and translates that into an interpretable rule model. In this paper, we extend BRL for knowledge discovery (BRL-KD) to enable BRL to incorporate a clinical utility function to learn models that are clinically more relevant. We demonstrate this using a real-world dataset to predict cardiovascular disease outcome. We evaluate predictive performance with the area under the receiver operating characteristic curve (AUROC) and clinical utility with the cost of the model. We show that BRL-KD successfully generates a set of models offering different trade-offs between AUROC and cost. Based on the clinical standard, a model with an acceptable trade-off can then be chosen.
Finding optimal blood pressure (BP) target and BP treatment after acute ischemic or hemorrhagic strokes is an area of controversy and a significant unmet need in the critical care of stroke victims. Numerous large prospective clinical trials have been done to address this question but have generated neutral or conflicting results. One major limitation that may have contributed to so many neutral or conflicting clinical trial results is the “one-size fit all” approach to BP targets, while the optimal BP target likely varies between individuals. We address this problem with the Acute Intervention Model of Blood Pressure (AIM-BP) framework: an individualized, human interpretable model of BP and its control in the acute care setting. The framework consists of two components: one, a model of BP homeostasis and the various effects that perturb it; and two, a parameter estimator that can learn clinically important model parameters on a patient by patient basis. By estimating the parameters of the AIM-BP model for a given patient, the effectiveness of antihypertensive medication can be quantified separately from the patient’s spontaneous BP trends. We hypothesize that the AIM-BP is a sufficient framework for estimating parameters of a homeostasis perturbation model of a stroke patient’s BP time course and the AIM-BP parameter estimator can do so as accurately and consistently as a state-of-the-art maximum likelihood estimation method. We demonstrate that this is the case in a proof of concept of the AIM-BP framework, using simulated clinical scenarios modeled on stroke patients from real world intensive care datasets.
AIM To develop a framework to incorporate background domain knowledge into classification rule learning for knowledge discovery in biomedicine. METHODS Bayesian rule learning (BRL) is a rule-based classifier that uses a greedy best-first search over a space of Bayesian belief-networks (BN) to find the optimal BN to explain the input dataset, and then infers classification rules from this BN. BRL uses a Bayesian score to evaluate the quality of BNs. In this paper, we extended the Bayesian score to include informative structure priors, which encodes our prior domain knowledge about the dataset. We call this extension of BRL as BRLp. The structure prior has a λ hyperparameter that allows the user to tune the degree of incorporation of the prior knowledge in the model learning process. We studied the effect of λ on model learning using a simulated dataset and a real-world lung cancer prognostic biomarker dataset, by measuring the degree of incorporation of our specified prior knowledge. We also monitored its effect on the model predictive performance. Finally, we compared BRLp to other state-of-the-art classifiers commonly used in biomedicine. RESULTS We evaluated the degree of incorporation of prior knowledge into BRLp, with simulated data by measuring the Graph Edit Distance between the true data-generating model and the model learned by BRLp. We specified the true model using informative structure priors. We observed that by increasing the value of λ we were able to increase the influence of the specified structure priors on model learning. A large value of λ of BRLp caused it to return the true model. This also led to a gain in predictive performance measured by area under the receiver operator characteristic curve (AUC). We then obtained a publicly available real-world lung cancer prognostic biomarker dataset and specified a known biomarker from literature [the epidermal growth factor receptor (EGFR) gene]. We again observed that larger values of λ led to an increased incorporation of EGFR into the final BRLp model. This relevant background knowledge also led to a gain in AUC. CONCLUSION BRLp enables tunable structure priors to be incorporated during Bayesian classification rule learning that integrates data and knowledge as demonstrated using lung cancer biomarker data.
Deep neural networks are increasingly being used in both supervised learning for classification tasks and unsupervised learning to derive complex patterns from the input data. However, the successful implementation of deep neural networks using neuroimaging datasets requires adequate sample size for training and well-defined signal intensity based structural differentiation. There is a lack of effective automated diagnostic tools for the reliable detection of brain dysmaturation in the neonatal period, related to small sample size and complex undifferentiated brain structures, despite both translational research and clinical importance. Volumetric information alone is insufficient for diagnosis. In this study, we developed a computational framework for the automated classification of brain dysmaturation from neonatal MRI, by combining a specific deep neural network implementation with neonatal structural brain segmentation as a method for both clinical pattern recognition and data-driven inference into the underlying structural morphology. We implemented three-dimensional convolution neural networks (3D-CNNs) to specifically classify dysplastic cerebelli, a subset of surface-based subcortical brain dysmaturation, in term infants born with congenital heart disease. We obtained a 0.985 ± 0. 0241-classification accuracy of subtle cerebellar dysplasia in CHD using 10-fold cross-validation. Furthermore, the hidden layer activations and class activation maps depicted regional vulnerability of the superior surface of the cerebellum, (composed of mostly the posterior lobe and the midline vermis), in regards to differentiating the dysplastic process from normal tissue. The posterior lobe and the midline vermis provide regional differentiation that is relevant to not only to the clinical diagnosis of cerebellar dysplasia, but also genetic mechanisms and neurodevelopmental outcome correlates. These findings not only contribute to the detection and classification of a subset of neonatal brain dysmaturation, but also provide insight to the pathogenesis of cerebellar dysplasia in CHD. In addition, this is one of the first examples of the application of deep learning to a neuroimaging dataset, in which the hidden layer activation revealed diagnostically and biologically relevant features about the clinical pathogenesis. The code developed for this project is open source, published under the BSD License, and designed to be generalizable to applications both within and beyond neonatal brain imaging.
Many clinical research datasets have a large percentage of missing values that directly impacts their usefulness in yielding high accuracy classifiers when used for training in supervised machine learning. While missing value imputation methods have been shown to work well with smaller percentages of missing values, their ability to impute sparse clinical research data can be problem specific. We previously attempted to learn quantitative guidelines for ordering cardiac magnetic resonance imaging during the evaluation for pediatric cardiomyopathy, but missing data significantly reduced our usable sample size. In this work, we sought to determine if increasing the usable sample size through imputation would allow us to learn better guidelines. We first review several machine learning methods for estimating missing data. Then, we apply four popular methods (mean imputation, decision tree, k-nearest neighbors, and self-organizing maps) to a clinical research dataset of pediatric patients undergoing evaluation for cardiomyopathy. Using Bayesian Rule Learning (BRL) to learn ruleset models, we compared the performance of imputation-augmented models versus unaugmented models. We found that all four imputation-augmented models performed similarly to unaugmented models. While imputation did not improve performance, it did provide evidence for the robustness of our learned models.
The comprehensibility of good predictive models learned from high-dimensional gene expression data is attractive because it can lead to biomarker discovery. Several good classifiers provide comparable predictive performance but differ in their abilities to summarize the observed data. We extend a Bayesian Rule Learning (BRL-GSS) algorithm, previously shown to be a significantly better predictor than other classical approaches in this domain. It searches a space of Bayesian networks using a decision tree representation of its parameters with global constraints, and infers a set of IF-THEN rules. The number of parameters and therefore the number of rules are combinatorial in the number of predictor variables in the model. We relax these global constraints to learn a more expressive local structure with BRL-LSS. BRL-LSS entails a more parsimonious set of rules because it does not have to generate all combinatorial rules. The search space of local structures is much richer than the space of global structures. We design the BRL-LSS with the same worst-case time-complexity as BRL-GSS while exploring a richer and more complex model space. We measure predictive performance using Area Under the ROC curve (AUC) and Accuracy. We measure model parsimony performance by noting the average number of rules and variables needed to describe the observed data. We evaluate the predictive and parsimony performance of BRL-GSS, BRL-LSS and the state-of-the-art C4.5 decision tree algorithm, across 10-fold cross-validation using ten microarray gene-expression diagnostic datasets. In these experiments, we observe that BRL-LSS is similar to BRL-GSS in terms of predictive performance, while generating a much more parsimonious set of rules to explain the same observed data. BRL-LSS also needs fewer variables than C4.5 to explain the data with similar predictive performance. We also conduct a feasibility study to demonstrate the general applicability of our BRL methods on the newer RNA sequencing gene-expression data.
1) Introduction: Brain parcellation is an important processing step in the analysis of structural brain MRI. Existing software implementations are optimized for fully developed adult brains, and provide inadequate results when applied to neonatal brain imaging. 2) Methods: We developed a semi-automated pipeline, NeBSS, for extracting 50 discrete brain structures from neonatal brain MRI, using an atlas registration method that leverages the existing ALBERT neonatal atlas 3) Results: We demonstrate a simple linear workflow for neonatal brain parcellation. NeBSS is robust to variation in imaging acquisition protocol and magnet field strength. 4) Conclusion: NeBSS is a robust pipeline capable of parcellating neonatal brain MRIs using a simple processing workflow. NeBSS fills a need in clinical translational research in neonatal imaging, where existing automated or semi-automated implementations are too rigid to be successfully applied to multi-center neuroprotection studies and clinically heterogeneous cohorts. The software is open source and freely available.
BACKGROUND:Adenocarcinoma (ADC) and squamous cell carcinoma (SCC) are the most prevalent histological types among lung cancers. Distinguishing between these subtypes is critically important because they have different implications for prognosis and treatment. Normally, histopathological analyses are used to distinguish between the two, where the tissue samples are collected based on small endoscopic samples or needle aspirations. However, the lack of cell architecture in these small tissue samples hampers the process of distinguishing between the two subtypes. Molecular profiling can also be used to discriminate between the two lung cancer subtypes, on condition that the biopsy is composed of at least 50 % of tumor cells. However, for some cases, the tissue composition of a biopsy might be a mix of tumor and tumor-adjacent histologically normal tissue (TAHN). When this happens, a new biopsy is required, with associated cost, risks and discomfort to the patient. To avoid this problem, we hypothesize that a computational method can distinguish between lung cancer subtypes given tumor and TAHN tissue.METHODS:Using publicly available datasets for gene expression and DNA methylation, we applied four classification tasks, depending on the possible combinations of tumor and TAHN tissue. First, we used a feature selector (ReliefF/Limma) to select relevant variables, which were then used to build a simple naïve Bayes classification model. Then, we evaluated the classification performance of our models by measuring the area under the receiver operating characteristic curve (AUC). Finally, we analyzed the relevance of the selected genes using hierarchical clustering and IPA® software for gene functional analysis.RESULTS:All Bayesian models achieved high classification performance (AUC > 0.94), which were confirmed by hierarchical cluster analysis. From the genes selected, 25 (93 %) were found to be related to cancer (19 were associated with ADC or SCC), confirming the biological relevance of our method.CONCLUSIONS:The results from this study confirm that computational methods using tumor and TAHN tissue can serve as a prognostic tool for lung cancer subtype classification. Our study complements results from other studies where TAHN tissue has been used as prognostic tool for prostate cancer. The clinical implications of this finding could greatly benefit lung cancer patients.
Human microbiome data from genomic sequencing technologies is fast accumulating, giving us insights into bacterial taxa that contribute to health and disease. The predictive modeling of such microbiota count data for the classification of human infection from parasitic worms, such as helminths, can help in the detection and management across global populations. Real-world datasets of microbiome experiments are typically sparse, containing hundreds of measurements for bacterial species, of which only a few are detected in the bio-specimens that are analyzed. This feature of microbiome data produces the challenge of needing more observations for accurate predictive modeling and has been dealt with previously, using different methods of feature reduction. To our knowledge, integrative methods, such as transfer learning, have not yet been explored in the microbiome domain as a way to deal with data sparsity by incorporating knowledge of different but related datasets. One way of incorporating this knowledge is by using a meaningful mapping among features of these datasets. In this paper, we claim that this mapping would exist among members of each individual cluster, grouped based on phylogenetic dependency among taxa and their association to the phenotype. We validate our claim by showing that models incorporating associations in such a grouped feature space result in no performance deterioration for the given classification task. In this paper, we test our hypothesis by using classification models that detect helminth infection in microbiota of human fecal samples obtained from Indonesia and Liberia countries. In our experiments, we first learn binary classifiers for helminth infection detection by using Naive Bayes, Support Vector Machines, Multilayer Perceptrons, and Random Forest methods. In the next step, we add taxonomic modeling by using the SMART-scan module to group the data, and learn classifiers using the same four methods, to test the validity of the achieved groupings. We observed a 6% to 23% and 7% to 26% performance improvement based on the Area Under the receiver operating characteristic (ROC) Curve (AUC) and Balanced Accuracy (Bacc) measures, respectively, over 10 runs of 10-fold cross-validation. These results show that using phylogenetic dependency for grouping our microbiota data actually results in a noticeable improvement in classification performance for helminth infection detection. These promising results from this feasibility study demonstrate that methods such as SMART-scan can be utilized in the future for knowledge transfer from different but related microbiome datasets by phylogenetically-related functional mapping, to enable novel integrative biomarker discovery.
In this era of precision medicine, understanding the epigenetic differences in lung cancer subtypes could lead to personalized therapies by possibly reversing these alterations. Traditional methods for analyzing microarray data rely on the use of known pathways. We propose a novel workflow, called Junction trees to Knowledge (J2K) framework, for creating interpretable graphical representations that can be derived directly from in silico analysis of microarray data. Our workflow has three steps, preprocessing (discretization and feature selection), construction of a Bayesian network and, its subsequent transformation into a Junction tree. We used data from the Cancer Genome Atlas to perform preliminary analyses of this J2K framework. We found relevant cliques of methylated sites that are junctions of the network along with potential methylation biomarkers in the lung cancer pathogenesis.
A major challenge in the diagnosis and treatment of brain tumors is tissue heterogeneity leading to mixed treatment response. Additionally, they are often difficult or at very high risk for biopsy, further hindering the clinical management process. To overcome this, novel advanced imaging methods are increasingly being adapted clinically to identify useful noninvasive biomarkers capable of disease stage characterization and treatment response prediction. One promising technique is called functional diffusion mapping (fDM), which uses diffusion-weighted imaging (DWI) to generate parametric maps between two imaging time points in order to identify significant voxel-wise changes in water diffusion within the tumor tissue. Here we introduce serial functional diffusion mapping (sfDM), an extension of existing fDM methods, to analyze the entire tumor diffusion profile along the temporal course of the disease. sfDM provides the tools necessary to analyze a tumor data set in the context of spatiotemporal parametric mapping: the image registration pipeline, biomarker extraction, and visualization tools. We present the general workflow of the pipeline, along with a typical use case for the software. sfDM is written in Python and is freely available as an open-source package under the Berkley Software Distribution (BSD) license to promote transparency and reproducibility.
Background: Most 'transcriptomic' data from microarrays are generated from small sample sizes compared to the large number of measured biomarkers, making it very difficult to build accurate and generalizable disease state classification models. Integrating information from different, but related, 'transcriptomic' data may help build better classification models. However, most proposed methods for integrative analysis of 'transcriptomic' data cannot incorporate domain knowledge, which can improve model performance. To this end, we have developed a methodology that leverages transfer rule learning and functional modules, which we call TRL-FM, to capture and abstract domain knowledge in the form of classification rules to facilitate integrative modeling of multiple gene expression data. TRL-FM is an extension of the transfer rule learner (TRL) that we developed previously. The goal of this study was to test our hypothesis that "an integrative model obtained via the TRL-FM approach outperforms traditional models based on single gene expression data sources".Results: To evaluate the feasibility of the TRL-FM framework, we compared the area under the ROC curve (AUC) of models developed with TRL-FM and other traditional methods, using 21 microarray datasets generated from three studies on brain cancer, prostate cancer, and lung disease, respectively. The results show that TRL-FM statistically significantly outperforms TRL as well as traditional models based on single source data. In addition, TRL-FM performed better than other integrative models driven by meta-analysis and cross-platform data merging.Conclusions: The capability of utilizing transferred abstract knowledge derived from source data using feature mapping enables the TRL-FM framework to mimic the human process of learning and adaptation when performing related tasks. The novel TRL-FM methodology for integrative modeling for multiple 'transcriptomic' datasets is able to intelligently incorporate domain knowledge that traditional methods might disregard, to boost predictive power and generalization performance. In this study, TRL-FM's abstraction of knowledge is achieved in the form of functional modules, but the overall framework is generalizable in that different approaches of acquiring abstract knowledge can be integrated into this framework.
Pediatric cardiomyopathies are a heterogeneous group of disorders in which early detection and treatment can drastically alter the course of the disease. Cardiac MRIs have become an increasingly popular option in evaluating for cardiomyopathies, though they remain expensive and time-consuming compared to echocardiography. Because guidelines for the use of cardiac MRI are vague, we investigated the quantitative characteristics of cardiac echo that predict subsequent positive cardiac MRI. Measurements were extracted from echo reports, processed, and fed into a Bayesian rule learning system. We discovered that ejection fraction and interventricular septum thickness were particularly important predictors of positive cardiac MRI. These features may help justify obtaining a cardiac MRI when echo results are inconclusive.
BACKGROUND:Pediatric cardiomyopathies are a rare, yet heterogeneous group of pathologies of the myocardium that are routinely examined clinically using Cardiovascular Magnetic Resonance Imaging (cMRI). This gold standard powerful non-invasive tool yields high resolution temporal images that characterize myocardial tissue. The complexities associated with the annotation of images and extraction of markers, necessitate the development of efficient workflows to acquire, manage and transform this data into actionable knowledge for patient care to reduce mortality and morbidity.METHODS:We develop and test a novel informatics framework called cMRI-BED for biomarker extraction and discovery from such complex pediatric cMRI data that includes the use of a suite of tools for image processing, marker extraction and predictive modeling. We applied our workflow to obtain and analyze a dataset of 83 de-identified cases and controls containing cMRI-derived biomarkers for classifying positive versus negative findings of cardiomyopathy in children. Bayesian rule learning (BRL) methods were applied to derive understandable models in the form of propositional rules with posterior probabilities pertaining to their validity. Popular machine learning methods in the WEKA data mining toolkit were applied using default parameters to assess cross-validation performance of this dataset using accuracy and percentage area under ROC curve (AUC) measures.RESULTS:The best 10-fold cross validation predictive performance obtained on this cMRI-derived biomarker dataset was 80.72% accuracy and 79.6% AUC by a BRL decision tree model, which is promising from this type of rare data. Moreover, we were able to verify that mycocardial delayed enhancement (MDE) status, which is known to be an important qualitative factor in the classification of cardiomyopathies, is picked up by our rule models as an important variable for prediction.CONCLUSIONS:Preliminary results show the feasibility of our framework for processing such data while also yielding actionable predictive classification rules that can augment knowledge conveyed in cardiac radiology outcome reports. Interactions between MDE status and other cMRI parameters that are depicted in our rules warrant further investigation and validation. Predictive rules learned from cMRI data to classify positive and negative findings of cardiomyopathy can enhance scientific understanding of the underlying interactions among imaging-derived parameters.