Studies of resemblance for disorders and other traits measured at the binary (yes/no) level between relatives frequently contain individuals who are currently in the negative category but who will become positive in future. For example, a 10-year-old may develop depression in the future, but is as yet unaffected. Such censoring can substantially bias estimates of correlation between relatives. To overcome this problem we develop a model for the association between liability to a disorder, and its age at onset. The model is designed for data from pairs of relatives to enable estimation of the correlation between an individuals' liability to disorder and their age at onset. Usually, such information is not available at the individual level, because age at onset is uniquely available when onset has occurred. Lacking variation in disorder status, data from non-related persons cannot estimate the covariance between liability and age at onset. Data from relatives can resolve this issue when there is a correlation in liability between the relatives, because different age at onset distributions would be expected in concordant vs. discordant pairs of relatives. Greater severity and worse outcomes are often observed among those with earlier onset, so a correlation between disorder liability and age at onset seems likely in many cases. In this article we present the basic theory of the model, implemented as a mixture distribution, and an application to cannabis use in a Virginia Twin Study of Adolescent Behavioral Development. A negative association of (-.212) between age at onset an liability was found, with confidence intervals of -.263 to -.152, which do not cross zero. The method contrasts with Cox Proportional Hazards, in which disorder liability and onset timing are treated as a single dimension.
With models and research designs ever increasing in complexity, the foundational question of model identification is more important than ever. The determination of whether or not a model can be fit at all or fit to some particular data set is the essence of model identification. In this article, we pull from previously published work on data-independent model identification applicable to a broad set of structural equation models, and extend it further to include extremely flexible exogenous covariate effects and also to include data-dependent empirical model identification. For illustrative purposes, we apply this model identification solution to several small examples for which the answer is already known, including a real data example from the National Longitudinal Survey of Youth; however, the method applies similarly to models that are far from simple to comprehend. The solution is implemented in the open-source OpenMx package in R.
Introduction: The availability of large-scale biobanks linking genetic data, rich phenotypes, and biological measures is a powerful opportunity for scientific discovery. However, real-world collections frequently have extensive missingness. While missing data prediction is possible, performance is significantly impaired by block-wise missingness inherent to many biobanks. Methods: To address this, we developed Missingness Adapted Group-wise Informed Clustered (MAGIC)-LASSO which performs hierarchical clustering of variables based on missingness followed by sequential Group LASSO within clusters. Variables are pre-filtered for missingness and balance between training and target sets with final models built using stepwise inclusion of features ranked by completeness. This research has been conducted using the UK Biobank (n > 500 k) to predict unmeasured Alcohol Use Disorders Identification Test (AUDIT) scores. Results: The phenotypic correlation between measured and predicted total score was 0.67 while genetic correlations between independent subjects was high >0.86. Discussion: Phenotypic and genetic correlations in real data application, as well as simulations, demonstrate the method has significant accuracy and utility for increasing power for genetic loci discovery.
Weighted least squares (WLS) estimation has tremendous utility for structural equation models (SEMs) that include ordered categorical data. However, some common statistical programs for fitting these models lack full flexibility in specification. Moreover, the popular approach toward model identification of ordinal variables arbitrarily constrains researchers and limits their ability to construct theories involving ordinal data. We develop a novel approach for identifying SEMs with ordered categorical variables that allows a wider variety of model specifications and theories, furthermore implementing this approach in OpenMx. We review WLS and ordinal variables both in general and as implemented in common software programs. Then we discuss the novel approach taken by the OpenMx implementation and derive analytic criteria for ordinal data identification. We conclude by posing new research questions that may now be addressed.
Background Psychotic disorders and schizotypal traits aggregate in the relatives of probands with schizophrenia. It is currently unclear how variability in symptom dimensions in schizophrenia probands and their relatives is associated with polygenic liability to psychiatric disorders. Aims To investigate whether polygenic risk scores (PRSs) can predict symptom dimensions in members of multiplex families with schizophrenia. Method The largest genome-wide data-sets for schizophrenia, bipolar disorder and major depressive disorder were used to construct PRSs in 861 participants from the Irish Study of High-Density Multiplex Schizophrenia Families. Symptom dimensions were derived using the Operational Criteria Checklist for Psychotic Disorders in participants with a history of a psychotic episode, and the Structured Interview for Schizotypy in participants without a history of a psychotic episode. Mixed-effects linear regression models were used to assess the relationship between PRS and symptom dimensions across the psychosis spectrum. Results Schizophrenia PRS is significantly associated with the negative/disorganised symptom dimension in participants with a history of a psychotic episode (P = 2.31 × 10−4) and negative dimension in participants without a history of a psychotic episode (P = 1.42 × 10−3). Bipolar disorder PRS is significantly associated with the manic symptom dimension in participants with a history of a psychotic episode (P = 3.70 × 10−4). No association with major depressive disorder PRS was observed. Conclusions Polygenic liability to schizophrenia is associated with higher negative/disorganised symptoms in participants with a history of a psychotic episode and negative symptoms in participants without a history of a psychotic episode in multiplex families with schizophrenia. These results provide genetic evidence in support of the spectrum model of schizophrenia, and support the view that negative and disorganised symptoms may have greater genetic basis than positive symptoms, making them better indices of familial liability to schizophrenia.
Alcohol use (i.e., quantity, frequency) and alcohol use disorder (AUD) are common, associated with adverse outcomes, and genetically-influenced. Genome-wide association studies (GWAS) identified genetic loci associated with both. AUD is positively genetically associated with psychopathology, while alcohol use (e.g., drinks per week) is negatively associated or NS related to psychopathology. We wanted to test if these genetic associations extended to life satisfaction, as there is an interest in understanding the associations between psychopathology-related traits and constructs that are not just the absence of psychopathology, but positive outcomes (e.g., well-being variables). Thus, we used Genomic Structural Equation Modeling (gSEM) to analyze summary-level genomic data (i.e., effects of genetic variants on constructs of interest) from large-scale GWAS of European ancestry individuals. Results suggest that the best-fitting model is a Bifactor Model, in which unique alcohol use, unique AUD, and common alcohol factors are extracted. The genetic correlation (r(g)) between life satisfaction-AUD specific factor was near zero, the r(g) with the alcohol use specific factor was positive and significant, and the r(g) with the common alcohol factor was negative and significant. Findings indicate that life satisfaction shares genetic etiology with typical alcohol use and life dissatisfaction shares genetic etiology with heavy alcohol use.
SummaryThere is a moderate association between poor sleep and psychological distress. There are marked sex differences in the prevalence of both variables, with females outnumbering males. However, the origin of these sex differences remains unclear. The objectives of this study were to: (1) study genetic and environmental influences on the relationship between poor sleep quality and psychological distress; and (2) test possible sex differences in this relationship. The sample comprised 3544 participants from the Murcia Twin Registry. Univariate and multivariate twin models were fitted to estimate the magnitude of genetic and environmental influences on both individual variance and covariance between poor sleep quality and psychological distress. Sleep quality and psychological distress were measured using the Pittsburgh Sleep Quality Index and the EuroQol five‐dimensions questionnaire, respectively. The results reveal a strong genetic association between poor sleep quality and psychological distress, which accounts for 44% (95%CI: 27%–61%) of the association between these two variables. Substantial genetic (rA = 0.50; 95%CI: 0.32, 0.67) and non‐shared environmental (rE = 0.41; 95%CI: 0.30, 0.52) correlations were also found, indicating a moderate overlap between genetic (and non‐shared environmental) factors influencing both phenotypes. Equating sexes in sex‐limitation models did not result in significant decreases in model fit. Despite the remarkable sex differences in the prevalence of both poor sleep quality and psychological distress, there were no sex differences in the genetic and environmental influences on these variables. This suggests that genetic factors play a similar role for men and women in explaining individual differences in both phenotypes and their relationship.
Psychotic and affective disorders often aggregate in the relatives of probands with schizophrenia, and genetic studies show substantial genetic correlation among schizophrenia, bipolar disorder, and major depressive disorder. In this study, we examined the polygenic risk burden of bipolar disorder and major depressive disorder in 257 multiplex schizophrenia families ( N = 1005) from the Irish Study of High-Density Multiplex Schizophrenia Families versus 2205 ancestry-matched controls. Our results indicate that members of multiplex schizophrenia families have an increased polygenic risk for bipolar disorder and major depressive disorder compared to population controls. However, this observation is largely attributable to the part of the genetic risk that bipolar disorder or major depressive disorder share with schizophrenia due to genetic correlation, rather than the affective portion of the genetic risk unique to them. These findings suggest that a complete interpretation of cross-disorder polygenic risks in multiplex families requires an assessment of the relative contribution of shared versus unique genetic factors to account for genetic correlations across psychiatric disorders.
Genome-wide association studies (GWAS) have successfully identified common variants associated with BMI. However, the stability of aggregate genetic variation influencing BMI from midlife and beyond is unknown. By analysing 165,717 men and 193,073 women from the UKBiobank, we performed BMI GWAS on six independent five-year age intervals between 40 and 72 years. We then applied genomic structural equation modeling to test competing hypotheses regarding the stability of genetic effects for BMI. LDSR genetic correlations between BMI assessed between ages 40 to 73 were all very high and ranged 0.89 to 1.00. Genomic structural equation modeling revealed that molecular genetic variance in BMI at each age interval could not be explained by the accumulation of any age-specific genetic influences or autore-gressive processes. Instead, a common set of stable genetic influences appears to underpin genome-wide variation in BMI from middle to early old age in men and women alike.
readmission (p African American, and Hispanic patients showed in (p in
Purpose: Posttraumatic Stress Disorder (PTSD) is associated with increased alcohol use and alcohol use disorder (AUD), which are all moderately heritable. Studies suggest the genetic association between PTSD and alcohol use differs from that of PTSD and AUD, but further analysis is needed. Basic procedures: We used genomic Structural Equation Modeling (genomicSEM) to analyze summary statistics from large-scale genome-wide association studies (GWAS) of European Ancestry participants to investigate the genetic relationships between PTSD (both diagnosis and re-experiencing symptom severity) and a range of alcohol use and AUD phenotypes. Main findings: When we differentiated genetic factors for alcohol use and AUD we observed improved model fit relative to models with all alcohol-related indicators loading onto a single factor. The genetic correlations (rG) of PTSD were quite discrepant for the alcohol use and AUD factors. This was true when modeled as a threecorrelated-factor model (PTSD-AUD rG:.36, p <.001; PTSD-alcohol use rG: 0.17, p <.001) and as a Bifactor model, in which the common and unique portions of alcohol phenotypes were pulled out into an AUD-specific factor (rG with PTSD:.40, p <.001), AU-specific factor (rG with PTSD: 0.57, p <.001), and a common alcohol factor (rG with PTSD:.16, NS). Principal conclusions: These results indicate the genetic architecture of alcohol use and AUD are differentially associated with PTSD. When the portions of variance unique to alcohol use and AUD are extracted, their genetic associations with PTSD vary substantially, suggesting different genetic architectures of alcohol phenotypes in people with PTSD.
This study proposes transformation functions and matrices between coefficients in the original and reparameterized parameter spaces for an existing linear-linear piecewise model to derive the interpretable coefficients directly related to the underlying change pattern. Additionally, the study extends the existing model to allow individual measurement occasions and investigates predictors for individual differences in change patterns. We present the proposed methods with simulation studies and a real-world data analysis. Our simulation study demonstrates that the method can generally provide an unbiased and accurate point estimate and appropriate confidence interval coverage for each parameter. The empirical analysis shows that the model can estimate the growth factor coefficients and path coefficients directly related to the underlying developmental process, thereby providing meaningful interpretation.
Multiplex families have higher recurrence risk of schizophrenia compared to the families of sporadic cases, but the source of this increased recurrence risk is unknown. We used schizophrenia genome-wide association study data (N = 156,509) to construct polygenic risk scores (PRS) in 1005 individuals from 257 multiplex schizophrenia families, 2114 ancestry-matched sporadic cases, and 2205 population controls, to evaluate whether increased PRS can explain the higher recurrence risk of schizophrenia in multiplex families compared to ancestry-matched sporadic cases. Using mixed-effects logistic regression with family structure modeled as a random effect, we show that SCZ PRS in familial cases does not differ significantly from sporadic cases either with, or without family history (FH) of psychotic disorders (All sporadic cases p = 0.90, FH+ cases p = 0.88, FH- cases p = 0.82). These results indicate that increased burden of common schizophrenia risk variation as indexed by current SCZ PRS, is unlikely to account for the higher recurrence risk of schizophrenia in multiplex families. In the absence of elevated PRS, segregation of rare risk variation or environmental influences unique to the families may explain the increased familial recurrence risk. These findings also further validate a genetically influenced psychosis spectrum, as shown by a continuous increase of common SCZ risk variation burden from unaffected relatives to schizophrenia cases in multiplex families. Finally, these results suggest that common risk variation loading are unlikely to be predictive of schizophrenia recurrence risk in the families of index probands, and additional components of genetic risk must be identified and included in order to improve recurrence risk prediction.
Externalizing behavior is substantially affected by genetic effects, which are moderated by environmental exposures. However, little is known about whether these moderation effects differ depending on individual characteristics, and whether moderation of environmental effects generalizes across different environmental domains. With a large sample (N = 1,441 individuals) of early adolescent twins (ages 11 and 13), using a longitudinal multi-informant design, we tested interaction effects between negative emotionality and both positive and negative aspects of three key social domains: parents, peers, and schools, on the phenotypic variance as well as the etiology of externalizing. Negative emotionality moderated some of the environmental effects on the phenotypic, genetic, and environmental variance in externalizing, with adolescents at both ends of the negative emotionality distribution showing different patterns of sensitivity to the tested environmental influences. This is the first use of gene-environment interaction twin models to test individual differences in environmental sensitivity, offering a new approach to study such effects.
A serrated polyposis syndrome was diagnosed in a 26-year-old female presenting with gastrointestinal symptoms. Screening for other lesions of the gastrointestinal tract showed a serpiginous looking papilla, described as possibly dysplastic. Histological analysis of biopsies showed a serrated lesion. This case describes the first known association between a duodenal serrated lesion and serrated polyposis syndrome. Upper GI screening is probably of little interest in this setting. In patients with upper GI serrated lesions, we recommend screening colonoscopy.
Empirical researchers are usually interested in investigating the impacts that baseline covariates have when uncovering sample heterogeneity and separating samples into more homogeneous groups. However, a considerable number of studies in the structural equation modeling (SEM) framework usually start with vague hypotheses in terms of heterogeneity and possible causes. It suggests that (1) the determination and specification of a proper model with covariates is not straightforward, and (2) the exploration process may be computationally intensive given that a model in the SEM framework is usually complicated and the pool of candidate covariates is usually huge in the psychological and educational domain where the SEM framework is widely employed. Following Bakk and Kuha (2017), this article presents a two-step growth mixture model (GMM) that examines the relationship between latent classes of nonlinear trajectories and baseline characteristics. Our simulation studies demonstrate that the proposed model is capable of clustering the nonlinear change patterns, and estimating the parameters of interest unbiasedly, precisely, as well as exhibiting appropriate confidence interval coverage. Considering the pool of candidate covariates is usually huge and highly correlated, this study also proposes implementing exploratory factor analysis (EFA) to reduce the dimension of covariate space. We illustrate how to use the hybrid method, the two-step GMM and EFA, to efficiently explore the heterogeneity of nonlinear trajectories of longitudinal mathematics achievement data.
The availability of large-scale biobanks linking rich phenotypes and biological measures is a powerful opportunity for scientific discovery. However, real-world collections frequently have extensive non-random missingness. While missing data prediction is possible, performance is significantly impaired by block-wise missingness inherent to many biobanks. To address this, we developed Missingness Adapted Group-wise Informed Clustered (MAGIC)-LASSO which performs hierarchical clustering of variables based on missingness followed by sequential Group LASSO within clusters. Variables are pre-filtered for missingness and balance between training and target sets with final models built using stepwise inclusion of features ranked by completeness. This research has been conducted using the UK Biobank (n>500k) to predict unmeasured Alcohol Use Disorders Identification Test (AUDIT) scores. The phenotypic correlation between measured and predicted total score was 0.67 while genetic correlations between independent subjects was high >0.86, demonstrating the method has significant accuracy and utility.