
A common goal of longitudinal research in psychology and related disciplines is the study of causal effects over time. An often overlooked assumption in this context is temporal alignment, referring to a match between the timescale at which a longitudinal process operates, the timescale at which data were collected, and the time frame in measurements. We present multiple scenarios, representing both the analysis of intensive longitudinal data and panel data, in which this assumption is violated due to time aggregation (i.e., aggregating scores over multiple occasions) and/or systematic undersampling (i.e., measurement intervals that are too long in relation to the longitudinal process). Our simulations demonstrate the large biases that can occur in estimates of lag-0 and lag-1 effects using various modeling approaches, even when there is no effect in the data. We discuss implications of these results and briefly consider future lines of methodological research to address this problem.
Mental health issues, particularly depression, arise from complex symptom interactions that traditional psychiatric models often fail to capture. We address this issue by introducing a continuous-time mechanistic network model of depressive symptoms, inspired by empirical findings. The model captures bistability between healthy and depressed states, illustrates transitions and tipping points, and shows post-shock persistence. It also allows the study of resilience by simulating responses to external shocks. Importantly, we demonstrate how the same underlying dynamics can give rise to different statistical networks depending on the timing of observation, thereby showing how statistical networks estimated from simulated observations can differ in density across healthy, depressed, and shock phases, even when the underlying mechanistic couplings remain unchanged. Using synthetic data across resilience levels, we compare model-derived networks to those estimated from a large-scale dataset (n = 23,283, HELIUS study), showing strong agreement and revealing that in-degree in the mechanistic network maps onto statistical centrality. Our aim is not to propose a definitive model of depression. Rather, the model provides both a plausible representation of symptom dynamics and an illustrative framework for linking mechanistic processes to statistical symptom networks, thereby clarifying how dynamic mechanisms can underpin diverse empirical findings across individuals and populations.
Growth curve modeling (GCM) has been widely used in social and behavioral sciences to analyze longitudinal data. However, it remains a significant challenge for GCM to handle missing data in longitudinal research, especially when data are nonnormally distributed. Although the robust median-based Bayesian approach for GCM developed by Tong effectively addresses both ignorable and nonignorable missing data in various data scenarios, particularly for nonnormally distributed data, no dedicated software was previously available to implement this advanced method, posing a barrier for researchers without extensive statistical or programming backgrounds. This article introduces a newly developed R package, Romeb, which streamlines the application of the robust median-based Bayesian linear GCM approach. An empirical example is provided to demonstrate the functionality of Romeb, accompanied by multiple figures and diagnostic tests.
For decades, researchers have debated whether the magnitude of the Flynn Effect-intergenerational increases in mean IQ scores-varies across cognitive domains and subdomains, and whether the Flynn Effect reflects gains in general cognitive ability (g). We conducted a cross-domain, subtest-level investigation of the Flynn Effect across middle childhood and early adolescence (ages 7-15 years, N = 1187, 89% White, 9% Black, 52% female) in longitudinal cognitive ability data collected prospectively between 1957 and 1999 using three versions of the Wechsler Intelligence Scale for Children. Results provided clear evidence of the Flynn Effect as both increases in mean IQ score across generational cohort and decreases in mean IQ across test versions. Flynn Effect magnitude differed substantially across domains and subtests. Performance IQ gains were larger than full-scale and verbal IQ gains in test version-based analyses, but cohort-based estimates showed modest differences in an opposite pattern (VIQ > FSIQ > PIQ). Variance in Flynn Effect magnitude across subtests did not follow a discernible pattern. The strength of the Flynn Effect on individual subtests was not strictly proportional to each subtest's g-loading, with performance IQ subtests showing larger increases than would be expected given their g-loadings and verbal subtests showing smaller-than-expected gains.
Substance use disorder data are often collected by asking individuals to endorse (or not endorse) a set of items pertaining to various diagnostic criteria. The result is a bipartite network, which can be represented by a two-mode binary matrix with rows corresponding to the individuals and columns to the items. Two-mode blockmodeling is an exploratory data analysis approach for bipartite networks that establishes partitions of both the individuals and items. Some two-mode blockmodeling methods are deterministic, whereas the latent blockmodel is stochastic and grounded by an underlying statistical model. A simulation study comparing the latent blockmodel and two deterministic blockmodeling methods revealed that the methods often perform comparably with respect to recovery of the true (known) cluster memberships when the number of clusters for both individuals and items is prespecified. However, the results also showed that one of the deterministic methods is unsuitable for sparse bipartite networks. A key advantage of the latent blockmodel method is a principled approach to selection of the number of clusters for individuals and items. We also demonstrate the effectiveness of the latent blockmodel via comparison to deterministic blockmodeling for a multiple-substance use disorder data set from the literature.
In this study we advance causal mediation analysis for nominal mediators within a potential outcome framework. Through Monte Carlo simulations, we compared three inference methods for testing total natural indirect effects (TNIE) and pure natural indirect effects (PNIE): non-parametric bootstrapping, parametric resampling, and Bayesian estimation. Results showed that nominal mediation models yielded accurate estimates across conditions, with both maximum likelihood and Bayesian approaches performing well. All three inference methods maintained acceptable Type I error control and achieved adequate statistical power in larger samples, with comparable performance across approaches. We provide an empirical illustration using survey data on healthcare payment types and mental health help-seeking behavior to illustrate the model's utility. Findings suggest that researchers can reliably estimate and test nominal mediation effects using any of the three inference approaches, providing a robust methodological framework for investigating causal pathways involving categorical mediating variables in social, behavioral, and health sciences.
Alcohol use disorder (AUD) research faces significant challenges in capturing individual heterogeneity and complex temporal patterns in drinking behaviors. Standard statistical methods fail to account for within-person variability and between-person differences, while existing machine learning algorithms are not designed for hierarchical, longitudinal data structures common in AUD research. We developed a comprehensive R package implementing 30 Bayesian machine learning functions specifically designed for alcohol use research, spanning interpretable linear and logistic regression to flexible Bayesian additive regression trees (BART), all with mixed-effects and time-trend extensions. We demonstrate the package capabilities using alcohol use data from two studies: the ABQDrinQ longitudinal cohort (n = 190) and the COMBINE clinical trial (n = 1,383). Key findings include strong associations between concurrent substance use and alcohol consumption (nicotine use associated with 13.5% increase in drinks and 14 times higher odds of drinking), discovery of nonlinear age effects on drinking variability (peak at ages 25-30), and high-accuracy daily predictions (median correlation 0.82, median absolute error 1.0 drinks). The Bayesian framework provides uncertainty quantification essential for both research and clinical applications, while the range of algorithms allows researchers to navigate the complexity-interpretability tradeoff. While developed for alcohol research, the methodological framework addresses statistical challenges common across substance use research.
Current models for assessing response accuracy and response times in testing environments typically overlook variations in speed and ability within individuals. Instead, they often treat these as residual variances, missing the dynamic changes in individual performance throughout a test. Additionally, the influence of item position and the ordinal nature of response accuracy remain underexplored. This paper introduces a comprehensive modelling framework that integrates item responses, response times, and item position to better understand skill acquisition and latent speed changes. Our approach uses the Bivariate Generalised Linear Item Response Theory (B-GLIRT) model, capturing the dual impact of ability on response accuracy and the interplay between ability and speed on response times. We extend this model by incorporating random effects for item and individual-specific variations. The proposed model further explores how item positioning affects test performance and provides diagnostic insights into individual differences. The paper also discusses parameter estimation, model identification, and applications to real-world data, illustrating the practical implications of our findings in computer-based learning assessments.
Measuring daily romantic relationship quality is important for understanding conflict, support, and satisfaction processes in near real-time. Although there now exists a great deal of research on determinants and outcomes of daily relationship quality, these investigations often rely on several untested assumptions regarding sources of, and consistency in, variability. In this study, we tested the variance components for daily relationship quality, how consistent these measurements are, and whether they are impacted by individual differences in attachment security. Six daily, nonexperimental self-reports of relationship satisfaction and functioning (e.g., conflict) from 101 couples were analyzed using generalizability theory. Results demonstrate a large amount of residual in daily reports of relationship quality, yet significant portions of variance attributable to person- and day-level processes. Variance proportions also changed depending on the index of relationship functioning examined, and individual differences in attachment security moderated these results. Lastly, daily reports of relationship satisfaction are relatively stable when submitted to models using persons as the object of measurement. These findings have implications for planning and interpreting future diary and ecological momentary assessment/experience sampling method studies.
Recognizing that complex networks of skills typically exhibit hierarchical and modular organization, this article presents a Modularized Higher-Order Diagnostic Classification Model (MHO-DCM) designed to capture hierarchical relationships among attributes organized into clustered subdomains. Central to the proposed method is a representation of attribute hierarchies in which attributes are grouped into cognitively coherent subgraphs nested within a single higher-order ability continuum. We adopt a nominal response model framework in item response theory and leverage standard maximum likelihood estimation (MLE). In parallel, we demonstrate that sequential higher-order latent structural models can likewise be implemented in a modularized fashion within an MLE framework. The performance of the proposed models is examined through simulation studies assessing parameter recovery, classification accuracy, and null rejection rates of goodness-of-fit measures. An empirical demonstration showcases how the framework can be applied in practice, highlighting its advantages in flexibility, interpretability, and the richer diagnostic insights it affords.