This study investigated group differences and longitudinal changes in brain volume before and after trauma-focused cognitive behavioral therapy (TF-CBT) in 20 unmedicated youth with maltreatment-related posttraumatic stress disorder (PTSD) and 20 non-trauma-exposed healthy control (HC) participants. We collected MRI scans of brain anatomy before and after 5 months of TF-CBT or the same time interval for the HC group. FreeSurfer software was used to segment brain images into 95 cortical and subcortical volumes, which were submitted to optimal scaling regression with lasso variable selection. The resulting model of group differences at baseline included larger right medial orbital frontal and left posterior cingulate corticies and smaller right midcingulate and right precuneus corticies in the PTSD relative to the HC group, R2 = .67. The model of group differences in pre- to posttreatment change included greater longitudinal changes in right rostral middle frontal, left pars triangularis, right entorhinal, and left cuneus corticies in the PTSD relative to the HC group, R2 = .69. Within the PTSD group, pre- to posttreatment symptom improvement was modeled by longitudinal decreases in the left posterior cingulate cortex, R2 = .45, and predicted by baseline measures of a smaller right isthmus (retrosplenial) cingulate and larger left caudate, R2 = .77. In sum, treatment was associated with longitudinal changes in brain regions that support executive functioning but not those that discriminated PTSD from HC participants at baseline. Additionally, results confirm a role for the posterior/retrosplenial cingulate as a correlate of PTSD symptom improvement and predictor of treatment outcome.
In this paper we combine two important extensions of ordinary least squares regression: regularization and optimal scaling. Optimal scaling (sometimes also called optimal scoring) has originally been developed for categorical data, and the process finds quantifications for the categories that are optimal for the regression model in the sense that they maximize the multiple correlation. Although the optimal scaling method was developed initially for variables with a limited number of categories, optimal transformations of continuous variables are a special case. We will consider a variety of transformation types; typically we use step functions for categorical variables, and smooth (spline) functions for continuous variables. Both types of functions can be restricted to be monotonic, preserving the ordinal information in the data. In addition to optimal scaling, three regularization methods will be considered: Ridge regression, the Lasso, and the Elastic Net. The resulting method will be called ROS Regression (Regularized Optimal Scaling Regression. We will show that the basic OS algorithm provides straightforward and efficient estimation of the regularized regression coefficients, automatically gives the Group Lasso and Blockwise Sparse Regression, and extends them with monotonicity properties. We will show that Optimal Scaling linearizes nonlinear relationships between predictors and outcome, and improves upon the condition of the predictor correlation matrix, increasing (on average) the conditional independence of the predictors. Alternative options for regularization of either regression coefficients or category quantifications are mentioned. Extended examples are provided. Keywords: Categorical Data, Optimal Scaling, Conditional Independence, Step Functions, Splines, Monotonic Transformations, Regularization, Lasso, Elastic Net, Group Lasso, Blockwise Sparse Regression.
Introduction: Biological therapies have greatly improved the treatment efficacy in rheumatoid arthritis (RA). However, in clinical practice a significant proportion of patients experience an inadequate response to treatment. The aim of this study is to classify responding and non-responding rheumatoid arthritis patients treated with biological therapies, based on clinical parameters and symptoms used in Western and Chinese medicine. Methods: Cold and Heat symptoms assessed by a Chinese medicine (CM) questionnaire and Western clinical data were collected as baseline data, before initiating biological therapy. Categorical principal components analysis with forced classification (CATPCA-FC) approach was applied to the baseline data set to classify responders and non-responders. Results: In this study, 61 RA patients were characterized using a CM questionnaire and clinical measurements. The combination of baseline symptoms ('preference for warm food', 'weak tendon severity') and clinical parameters (positive rheumatoid factor/anti-cyclic citrullinated peptide antibody, C-reactive protein, creatinine) were able to differentiate responders from non-responders to biological therapies with a positive predictive value of 82.35% and a misclassification rate of 24.59%. Adding CM symptom variables in addition to clinical data did not improve the classification of responders, but it did show 8.3% improvement in classifying non-responders. Conclusions: No significant differences were found between the three classification models. Adding CM symptoms to the clinical parameters in the combined model improved the classification of non-responders. Although this improvement is not significant in the current study, we consider it worthwhile to further investigate the potential of adding symptom variables for improving treatment efficacy.
Objective The aim is to characterize subgroups or phenotypes of rheumatoid arthritis (RA) patients using a systems biology approach. The discovery of subtypes of rheumatoid arthritis patients is an essential research area for the improvement of response to therapy and the development of personalized medicine strategies. Methods In this study, 39 RA patients are phenotyped using clinical chemistry measurements, urine and plasma metabolomics analysis and symptom profiles. In addition, a Chinese medicine expert classified each RA patient as a Cold or Heat type according to Chinese medicine theory. Multivariate data analysis techniques are employed to detect and validate biochemical and symptom relationships with the classification. Results The questionnaire items ‘Red joints’, ‘Swollen joints’, ‘Warm joints’ suggest differences in the level of inflammation between the groups although c-reactive protein (CRP) and rheumatoid factor (RHF) levels were equal. Multivariate analysis of the urine metabolomics data revealed that the levels of 11 acylcarnitines were lower in the Cold RA than in the Heat RA patients, suggesting differences in muscle breakdown. Additionally, higher dehydroepiandrosterone sulfate (DHEAS) levels in Heat patients compared to Cold patients were found suggesting that the Cold RA group has a more suppressed hypothalamic-pituitary-adrenal (HPA) axis function. Conclusion Significant and relevant biochemical differences are found between Cold and Heat RA patients. Differences in immune function, HPA axis involvement and muscle breakdown point towards opportunities to tailor disease management strategies to each of the subgroups RA patient.
This article is set up as a tutorial for nonlinear principal components analysis (NLPCA), systematically guiding the reader through the process of analyzing actual data on personality assessment by the Rorschach Inkblot Test. NLPCA is a more flexible alternative to linear PCA that can handle the analysis of possibly nonlinearly related variables with different types of measurement level. The method is particularly suited to analyze nominal (qualitative) and ordinal (e.g., Likert-type) data, possibly combined with numeric data. The program CATPCA from the Categories module in SPSS is used in the analyses, but the method description can easily be generalized to other software packages.
BACKGROUND:The future of personalized medicine depends on advanced diagnostic tools to characterize responders and non-responders to treatment. Systems diagnosis is a new approach which aims to capture a large amount of symptom information from patients to characterize relevant sub-groups.METHODOLOGY:49 patients with a rheumatic disease were characterized using a systems diagnosis questionnaire containing 106 questions based on Chinese and Western medicine symptoms. Categorical principal component analysis (CATPCA) was used to discover differences in symptom patterns between the patients. Two Chinese medicine experts where subsequently asked to rank the Cold and Heat status of all the patients based on the questionnaires. These rankings were used to study the Cold and Heat symptoms used by these practitioners.FINDINGS:The CATPCA analysis results in three dimensions. The first dimension is a general factor (40.2% explained variance). In the second dimension (12.5% explained variance) 'anxious', 'worrying', 'uneasy feeling' and 'distressed' were interpreted as the Internal disease stage, and 'aggravate in wind', 'fear of wind' and 'aversion to cold' as the External disease stage. In the third dimension (10.4% explained variance) 'panting s', 'superficial breathing', 'shortness of breath s', 'shortness of breath f' and 'aversion to cold' were interpreted as Cold and 'restless', 'nervous', 'warm feeling', 'dry mouth s' and 'thirst' as Heat related. 'Aversion to cold', 'fear of wind' and 'pain aggravates with cold' are most related to the experts Cold rankings and 'aversion to heat', 'fullness of chest' and 'dry mouth' to the Heat rankings.CONCLUSIONS:This study shows that the presented systems diagnosis questionnaire is able to identify groups of symptoms that are relevant for sub-typing patients with a rheumatic disease.
The component structure of 14 Likert-type items measuring different aspects of job satisfaction was investigated using nonlinear Principal Components Analysis (NLPCA). NLPCA allows for analyzing these items at an ordinal or interval level. The participants were 2066 workers from five types of social service organizations. Our results suggest that taking into account the ordinal nature of the items was most appropriate. On the basis of a stability study, a two-component structure was found, from which we extracted two subscales ("Motivation" and "Hygiene") with reliabilities of .81 and .77. A Multiple Group analysis confirmed this structure. We also investigated whether workers in the five types of organizations differed with respect to the component structure, employing a feature of the program CATPCA. We found that the organizations did not differ much with respect to the job satisfaction components.
In explorative regression studies, linear models are often applied without questioning the linearity of the relations between the predictor variables and the dependent variable, or linear relations are taken as an approximation. In this study, the method of regression with optimal scaling transformations is demonstrated. This method does not require predefined nonlinear functions and results in easy-to-interpret transformations that will show the form of the relations. The method is illustrated using data from a German multicenter project on the indication criteria for inpatient or day clinic psychotherapy treatment. The indication criteria to include in the regression model were selected with the Lasso, which is a tool for predictor selection that overcomes the disadvantages of stepwise regression methods. The resulting prediction model indicates that treatment status is (approximately) linearly related to some criteria and nonlinearly related to others.
Principal components analysis (PCA) is used to explore the structure of data sets containing linearly related numeric variables. Alternatively, nonlinear PCA can handle possibly nonlinearly related numeric as well as nonnumeric variables. For linear PCA, the stability of its solution can be established under the assumption of multivariate normality. For nonlinear PCA, however, standard options for establishing stability are not provided. The authors use the nonparametric bootstrap procedure to assess the stability of nonlinear PCA results, applied to empirical data. They use confidence intervals for the variable transformations and confidence ellipses for the eigenvalues, the component loadings, and the person scores. They discuss the balanced version of the bootstrap, bias estimation, and Procrustes rotation. To provide a benchmark, the same bootstrap procedure is applied to linear PCA on the same data. On the basis of the results, the authors advise using at least 1,000 bootstrap samples, using Procrustes rotation on the bootstrap results, examining the bootstrap distributions along with the confidence regions, and merging categories with small marginal frequencies to reduce the variance of the bootstrap results.
The authors provide a didactic treatment of nonlinear (categorical) principal components analysis (PCA). This method is the nonlinear equivalent of standard PCA and reduces the observed variables to a number of uncorrelated principal components. The most important advantages of nonlinear over linear PCA are that it incorporates nominal and ordinal variables and that it can handle and discover nonlinear relationships between variables. Also, nonlinear PCA can deal with variables at their appropriate measurement level; for example, it can treat Likert-type scales ordinally instead of numerically. Every observed value of a variable can be referred to as a category. While performing PCA, nonlinear PCA converts every category to a numeric value, in accordance with the variable's analysis level, using optimal quantification. The authors discuss how optimal quantification is carried out, what analysis levels are, which decisions have to be made when applying nonlinear PCA, and how the results can be interpreted. The strengths and limitations of the method are discussed. An example applying nonlinear PCA to empirical data using the program CATPCA (J. J. Meulman, W. J. Heiser, & SPSS, 2004) is provided.
The component structure of a 14-item scale measuring different aspects of job satisfaction was investigated, and its stability among different types of organizations. The participants were 2066 workers from 220 organizations in the Italian social service sector. The job satisfaction items were measured at an ordinal scaling level and analyzed by Categorical Principal Components analysis (CATPCA) using monotonic (spline) transformations. The sample was randomly divided into a training set and a test set. CATPCA was applied to the training set and resulted in a two-component solution. Results of a Multiple Group analysis for the test set confirmed the validity of the training set solution. From the two-component solution, we extracted two subscales with reliabilities (Cronbach's α) of .81 and .77. The subscales reflect motivator and hygiene aspects of job satisfaction, and appeared in line with Herzberg's theory. The different types of organizations have the same two-component structure of job satisfaction.
This chapter focuses on the analysis of ordinal and nominal multivariate data, using a special variety of principal components analysis that includes nonlinear optimal scaling transformation of the variables. Since the early 1930s, classical statistical methods have been adapted in various ways to suit the particular characteristics of social and behavioral science research. Research in these areas often results in data that are nonnumerical, with measurements recorded on scales having an uncertain unit of measurement. Data would typically consist of qualitative or categorical variables that describe the persons in a limited number of categories. The zero point of these scales is uncertain, the relationships among the different categories is often unknown, and although frequently it can be assumed that the categories are ordered, their mutual distances might still be unknown. The uncertainty in the unit of measurement is not just a matter of measurement error because its variability may have a systematic component. For example, in the data set that will be used throughout this chapter as an illustration, concerning feelings of national identity and involving 25,000 respondents in 23 different countries all over the world (International Social Survey Programme [ISSP], 1995), there are variables indicating how close the respondents feel toward their neighborhood, town, and country, measured on a 5-point scale with labels ranging from not close at all to very close. This response format is typical for a lot of behavioral research and definitely is not numerical (even though the categories are ordered and can be coded numerically).
The central topic of this thesis is the CATREG approach to nonlinear regression. This approach finds optimal quantifications for categorical variables and/or nonlinear transformations for numerical variables in regression analysis. (CATREG is implemented in SPSS Categories by the author of the thesis; the relevant parts of the Categories manual are included in the appendix.) The first chapter of the thesis provides a non-technical introduction to the CATREG approach, illustrated with graphs. The more technical part of the thesis includes (1) a solution to the local minima problem for monotone transformations, as well as a study of the effect of several data conditions on the incidence and severeness of local minima, (2) the incorporation into CATREG of a particular resampling method (the .632 bootstrap) for assessing prediction accuracy, and (3) the incorporation into CATREG of several regularization methods (Ridge Regression, the Lasso, and the Elastic Net) for stabilizing the estimates of the regression coefficients and transformations. The technical part is followed by a chapter describing a bulimia nervosa study in which the CATREG-Lasso and the .632 bootstrap are applied.