Conventional research methodologies and data analytic approaches in psychiatric research are unable to reliably infer causal relations without experimental designs, or to make inferences about the functional properties of the complex systems in which psychiatric disorders are embedded. This article describes a series of studies to validate a novel hybrid computational approach--the Complex Systems-Causal Network (CS-CN) method-designed to integrate causal discovery within a complex systems framework for psychiatric research. The CS-CN method was first applied to an existing dataset on psychopathology in 163 children hospitalized with injuries (validation study). Next, it was applied to a much larger dataset of traumatized children (replication study). Finally, the CS-CN method was applied in a controlled experiment using a 'gold standard' dataset for causal discovery and compared with other methods for accurately detecting causal variables (resimulation controlled experiment). The CS-CN method successfully detected a causal network of 111 variables and 167 bivariate relations in the initial validation study. This causal network had well-defined adaptive properties and a set of variables was found that disproportionally contributed to these properties. Modeling the removal of these variables resulted in significant loss of adaptive properties. The CS-CN method was successfully applied in the replication study and performed better than traditional statistical methods, and similarly to state-of-the-art causal discovery algorithms in the causal detection experiment. The CS-CN method was validated, replicated, and yielded both novel and previously validated findings related to risk factors and potential treatments of psychiatric disorders. The novel approach yields both fine-grain (micro) and high-level (macro) insights and thus represents a promising approach for complex systems-oriented research in psychiatry.
BACKGROUND:Predicting Posttraumatic Stress Disorder (PTSD) is a pre-requisite for targeted prevention. Current research has identified group-level risk-indicators, many of which (e.g., head trauma, receiving opiates) concern but a subset of survivors. Identifying interchangeable sets of risk indicators may increase the efficiency of early risk assessment. The study goal is to use supervised machine learning (ML) to uncover interchangeable, maximally predictive combinations of early risk indicators.METHODS:Data variables (features) reflecting event characteristics, emergency department (ED) records and early symptoms were collected in 957 trauma survivors within ten days of ED admission, and used to predict PTSD symptom trajectories during the following fifteen months. A Target Information Equivalence Algorithm (TIE*) identified all minimal sets of features (Markov Boundaries; MBs) that maximized the prediction of a non-remitting PTSD symptom trajectory when integrated in a support vector machine (SVM). The predictive accuracy of each set of predictors was evaluated in a repeated 10-fold cross-validation and expressed as average area under the Receiver Operating Characteristics curve (AUC) for all validation trials.RESULTS:The average number of MBs per cross validation was 800. MBs' mean AUC was 0.75 (95% range: 0.67-0.80). The average number of features per MB was 18 (range: 12-32) with 13 features present in over 75% of the sets.CONCLUSIONS:Our findings support the hypothesized existence of multiple and interchangeable sets of risk indicators that equally and exhaustively predict non-remitting PTSD. ML's ability to increase prediction versatility is a promising step towards developing algorithmic, knowledge-based, personalized prediction of post-traumatic psychopathology.
Field of cancerization in the airway epithelium has been increasingly examined to understand early pathogenesis of non-small cell lung cancer. However, the extent of field of cancerization throughout the lung airways is unclear. Here we sought to determine the differential gene and microRNA expressions associated with field of cancerization in the peripheral airway epithelial cells of patients with lung adenocarcinoma. We obtained peripheral airway brushings from smoker controls (n=13) and from the lung contralateral to the tumor in cancer patients (n=17). We performed gene and microRNA expression profiling on these peripheral airway epithelial cells using Affymetrix GeneChip and TaqMan Array. Integrated gene and microRNA analysis was performed to identify significant molecular pathways. We identified 26 mRNAs and 5 miRNAs that were significantly (FDR <0.1) up-regulated and 38 mRNAs and 12 miRNAs that were significantly down-regulated in the cancer patients when compared to smoker controls. Functional analysis identified differential transcriptomic expressions related to tumorigenesis. Integration of miRNA-mRNA data into interaction network analysis showed modulation of the extracellular signal-regulated kinase/mitogen-activated protein kinase (ERK/MAPK) pathway in the contralateral lung field of cancerization. In conclusion, patients with lung adenocarcinoma have tumor related molecules and pathways in histologically normal appearing peripheral airway epithelial cells, a substantial distance from the tumor itself. This finding can potentially provide new biomarkers for early detection of lung cancer and novel therapeutic targets.
ObjectiveInflammatory mediators, such as prostaglandin E2 (PGE2) and interleukin‐1β (IL‐1β), are produced by osteoarthritic (OA) joint tissue, where they may contribute to disease pathogenesis. We undertook the present study to examine whether inflammation, evidenced in plasma and peripheral blood leukocytes (PBLs), reflects the presence, progression, or specific symptoms of symptomatic knee OA.MethodsPatients with symptomatic knee OA were enrolled in a 24‐month prospective study of radiographic progression. Standardized knee radiographs were obtained at baseline and 24 months. At baseline, levels of the plasma lipids PGE2 and 15‐hydroxyeicosatetraenoic acid (15‐HETE) were measured, and transcriptome analysis of PBLs was performed by microarray and quantitative polymerase chain reaction.ResultsBaseline PGE2 synthase (PGES) levels determined by PBL microarray gene expression and plasma PGE2 levels distinguished patients with symptomatic knee OA from non‐OA controls (area under the receiver operating characteristic curve [AUC] 0.87 and 0.89, respectively, P < 0.0001). Baseline plasma 15‐HETE levels were significantly elevated in patients with symptomatic knee OA versus non‐OA controls (P < 0.0195). In the 146 patients who completed the 24‐month study, elevated baseline expression of IL‐1β, tumor necrosis factor α, and cyclooxygenase 2 (COX‐2) messenger RNA in PBLs predicted higher risk of radiographic progression as evidenced by joint space narrowing (JSN). In a multivariate model, AUC point estimates of models containing COX‐2 in combination with demographic traits overlapped the confidence interval of the base model in 2 of the 3 JSN outcome measures (JSN >0.0 mm, JSN >0.2 mm, and JSN >0.5 mm; AUC 0.62–0.67).ConclusionThe inflammatory plasma lipid biomarkers PGE2 and 15‐HETE identify patients with symptomatic knee OA, and the PBL inflammatory transcriptome identifies a subset of patients with symptomatic knee OA who are at increased risk of radiographic progression. These findings may reflect low‐grade inflammation in OA and may be useful as diagnostic and prognostic biomarkers in clinical development of disease‐modifying OA drugs.
Objective: Pro-and anti-inflammatory mediators, such as IL-1 beta and IL1Ra, are produced by joint tissues in osteoarthritis (OA), where they may contribute to pathogenesis. We examined whether inflammatory events occurring within joints are reflected in plasma of patients with symptomatic knee osteoarthritis (SKOA).Design: 111 SKOA subjects with medial disease completed a 24-month prospective study of clinical and radiographic progression, with clinical assessment and specimen collection at 6-month intervals. The plasma biochemical marker IL1Ra was assessed at baseline and 18 months; other plasma biochemical markers were assessed only at 18 months, including IL-1 beta, TNF alpha, VEGF, IL-6, IL-6R alpha, IL-17A, IL-17A/F, IL17F, CRP, sTNF-RII, and MMP-2.Results: In cross-sectional studies, WOMAC (total, pain, function) and plasma IL1Ra were modestly associated with radiographic severity after adjustment for age, gender and body mass index (BMI). In addition, elevation of plasma IL1Ra predicted joint space narrowing (JSN) at 24 months. BMI did associate with progression in some but not all analyses. Causal graph analysis indicated a positive association of IL1Ra with JSN; an interaction between IL1Ra and BMI suggested either that BMI influences IL1Ra or that a hidden confounder influences both BMI and IL1Ra. Other protein biomarkers examined in this study did not associate with radiographic progression or severity.Conclusions: Plasma levels of IL1Ra were modestly associated with the severity and progression of SKOA in a causal fashion, independent of other risk factors. The findings may be useful in the search for prognostic biomarkers and development of disease-modifying OA drugs. (C) 2015 Osteoarthritis Research Society International. Published by Elsevier Ltd. All rights reserved.
An important aspect to performing text categorization is selecting appropriate supervised classification and feature selection methods. A comprehensive benchmark is needed to inform best practices in this broad application field. Previous benchmarks have evaluated performance for a few supervised classification and feature selection methods and limited ways to optimize them. The present work updates prior benchmarks by increasing the number of classifiers and feature selection methods order of magnitude, including adding recently developed, state-of-the-art methods. Specifically, this study used 229 text categorization data sets/tasks, and evaluated 28 classification methods (both well-established and proprietary/commercial) and 19 feature selection methods according to 4 classification performance metrics. We report several key findings that will be helpful in establishing best methodological practices for text categorization.
An important aspect to performing text categorization is selecting appropriate supervised classification and feature selection methods. A comprehensive benchmark is needed to inform best practices in this broad application field. Previous benchmarks have evaluated performance for a few supervised classification and feature selection methods and limited ways to optimize them. The present work updates prior benchmarks by increasing the number of classifiers and feature selection methods order of magnitude, including adding recently developed, state‐of‐the‐art methods. Specifically, this study used 229 text categorization data sets/tasks, and evaluated 28 classification methods (both well‐established and proprietary/commercial) and 19 feature selection methods according to 4 classification performance metrics. We report several key findings that will be helpful in establishing best methodological practices for text categorization.
Psoriasis is a common chronic inflammatory disease of the skin. We sought to use bacterial community abundance data to assess the feasibility of developing multivariate molecular signatures for differentiation of cutaneous psoriatic lesions, clinically unaffected contralateral skin from psoriatic patients, and similar cutaneous loci in matched healthy control subjects. Using 16S rRNA high-throughput DNA sequencing, we assayed the cutaneous microbiome for 51 such matched specimen triplets including subjects of both genders, different age groups, ethnicities and multiple body sites. None of the subjects had recently received relevant treatments or antibiotics. We found that molecular signatures for the diagnosis of psoriasis result in significant accuracy ranging from 0.75 to 0.89 AUC, depending on the classification task. We also found a significant effect of DNA sequencing and downstream analysis protocols on the accuracy of molecular signatures. Our results demonstrate that it is feasible to develop accurate molecular signatures for the diagnosis of psoriasis from microbiomic data.
BACKGROUND:Recent advances in next-generation DNA sequencing enable rapid high-throughput quantitation of microbial community composition in human samples, opening up a new field of microbiomics. One of the promises of this field is linking abundances of microbial taxa to phenotypic and physiological states, which can inform development of new diagnostic, personalized medicine, and forensic modalities. Prior research has demonstrated the feasibility of applying machine learning methods to perform body site and subject classification with microbiomic data. However, it is currently unknown which classifiers perform best among the many available alternatives for classification with microbiomic data.RESULTS:In this work, we performed a systematic comparison of 18 major classification methods, 5 feature selection methods, and 2 accuracy metrics using 8 datasets spanning 1,802 human samples and various classification tasks: body site and subject classification and diagnosis.CONCLUSIONS:We found that random forests, support vector machines, kernel ridge regression, and Bayesian logistic regression with Laplace priors are the most effective machine learning techniques for performing accurate classification from these microbiomic data.
BACKGROUND:Oncogenic mechanisms in small-cell lung cancer remain poorly understood leaving this tumor with the worst prognosis among all lung cancers. Unlike other cancer types, sequencing genomic approaches have been of limited success in small-cell lung cancer, i.e., no mutated oncogenes with potential driver characteristics have emerged, as it is the case for activating mutations of epidermal growth factor receptor in non-small-cell lung cancer. Differential gene expression analysis has also produced SCLC signatures with limited application, since they are generally not robust across datasets. Nonetheless, additional genomic approaches are warranted, due to the increasing availability of suitable small-cell lung cancer datasets. Gene co-expression network approaches are a recent and promising avenue, since they have been successful in identifying gene modules that drive phenotypic traits in several biological systems, including other cancer types.RESULTS:We derived an SCLC-specific classifier from weighted gene co-expression network analysis (WGCNA) of a lung cancer dataset. The classifier, termed SCLC-specific hub network (SSHN), robustly separates SCLC from other lung cancer types across multiple datasets and multiple platforms, including RNA-seq and shotgun proteomics. The classifier was also conserved in SCLC cell lines. SSHN is enriched for co-expressed signaling network hubs strongly associated with the SCLC phenotype. Twenty of these hubs are actionable kinases with oncogenic potential, among which spleen tyrosine kinase (SYK) exhibits one of the highest overall statistical associations to SCLC. In patient tissue microarrays and cell lines, SCLC can be separated into SYK-positive and -negative. SYK siRNA decreases proliferation rate and increases cell death of SYK-positive SCLC cell lines, suggesting a role for SYK as an oncogenic driver in a subset of SCLC.CONCLUSIONS:SCLC treatment has thus far been limited to chemotherapy and radiation. Our WGCNA analysis identifies SYK both as a candidate biomarker to stratify SCLC patients and as a potential therapeutic target. In summary, WGCNA represents an alternative strategy to large scale sequencing for the identification of potential oncogenic drivers, based on a systems view of signaling networks. This strategy is especially useful in cancer types where no actionable mutations have emerged.
A common type of reliability data is the right censored time‐to‐failure data. In this article, we developed a control chart to monitor the time‐to‐failure data in the presence of right censoring using weighted rank tests. On the basis of the asymptotic properties of the rank statistics, we derived the generic formulae for the operating characteristic functions of the control chart to show the relationship between type I error probability, type II error probability, sample size, and hazard rate change. We presented case studies to illustrate the design procedure and the effectiveness of the proposed control chart system. We also investigated and compared the performance of the proposed monitoring procedure with some available monitoring techniques for nonconformities. Copyright © 2011 John Wiley & Sons, Ltd.
The Proportional Hazards (PH) model is an important type of failure time regression model which relates the occurrence probability of critical failures to influential factors. However, little research work has been done on detecting changes in the PH models fitted based on different sets of reliability data. This paper develops the methods for change detection in the Cox PH models, also known as Semiparametric PH model, for reliability prediction and/or assessment of the time‐to‐failure data collected from different subjects. The effectiveness of the developed methods is illustrated through numerical studies and real‐world data analysis. The developed technique possesses wide applicability to the systems and processes where the Cox PH model fits the reliability data well. Copyright © 2010 John Wiley & Sons, Ltd.
Recognizing facial expressions from facial video sequences is an important and unsolved problem. Among many factors that contribute to the challenges of this task are: non-frontal facial poses, poorly aligned face images, large variations in the temporal scale of facial expressions, and the subtle differences between different subjects for the same facial expression etc. A successful video-based facial expression analysis system should be able to handle at least the following problems: robust face tracking, or spatial alignment of the faces, video segmentation, effective feature representation and selection schemes which are robust to face mis-alignment and temporal normalization by sequential classifier. In this work we report several advances we made in building various components of a system for classifying facial expressions from video inputs. Particularly, my work focus on robust face tracking, facial feature representation and selection under different face alignment conditions, sequential modeling for facial expression recognition. We performed extensive experiments using the proposed algorithms on publicly available dataset and achieved state of the art performances.
In this paper, we systematically study the effect of poorly registered faces on the training and inferring stages of traditional face recognition algorithms. We then propose a novel multiple-instance based subspace learning scheme for face recognition. In this approach, we iteratively update the subspace training instances according to diverse densities, using class-balanced supervised clustering. We test our multiple instance subspace learning algorithm with Fisherface for the application of face recognition. Experimental results show that the proposed learning algorithm can improve the robustness of current methods with poorly aligned training and testing data.
Reliable 3D tracking is still a difficult task. Most parametrized 3D deformable models rely on the accurate extraction of image features for updating their parameters, and are prone to failures when the underlying feature distribution assumptions are invalid. Active Shape Models (ASMs), on the other hand, are based on learning, and thus require fewer reliable local image features than parametrized 3D models, but fail easily when they encounter a situation for which they were not trained. In this paper, we develop an integrated framework that combines the strengths of both 3D deformable models and ASMs. The 3D model governs the overall shape, orientation and location, and provides the basis for statistical inference on both the image features and the parameters. The ASMs, in contrast, provide the majority of reliable 2D image features over time, and aid in recovering from drift and total occlusions. The framework dynamically selects among different ASMs to compensate for large viewpoint changes due to head rotations. This integration allows the robust tracking effaces and the estimation of both their rigid and non- rigid motions. We demonstrate the strength of the framework in experiments that include automated 3D model fitting and facial expression tracking for a variety of applications, including sign language.
We present a Dynamic Data Driven Application System (DDDAS) to track 2D shapes across large pose variations by learning non-linear shape manifold as overlapping, piecewise linear subspaces. The learned subspaces adaptively adjust to the subject by tracking the shapes independently using Kanade Lucas Tomasi(KLT) point tracker. The novelty of our approach is that the tracking of feature points is used to generate independent training examples for updating the learned shape manifold and the appearance model. We use landmark based shape analysis to train a Gaussian mixture model over the aligned shapes and learn a Point Distribution Model(PDM) for each of the mixture components. The target 2D shape is searched by first maximizing the mixture probability density for the local feature intensity profiles along the normal followed by constraining the global shape using the most probable PDM cluster. The feature shapes are robustly tracked across multiple frames by dynamically switching between the PDMs. The tracked 2D facial features are used deform the 3D face mask.The main advantage of the 3D deformable face models is the reduced dimensionality. The smaller number of degree of freedom makes the system more robust and enables capturing subtle facial expressions as change of only a few parameters. We demonstrate the results on tracking facial features and provide several empirical results to validate our approach. Our framework runs close to real time at 25 frames per second.
Chan-Su Lee合作论文数Department of Computer Science
Rutgers University2
Ahmed Elgammal合作论文数Art and Artificial Intelligence Laboratory, Computational Biomedicine Imaging and Modeling Center, Rutgers University;Department of Computer Science, Rutgers University2