
An understanding of the body size of individuals and the relationships between different dimensions is critical for monitoring the status and the health of a wildlife population. Morphometric data have traditionally been collected by physically handling and measuring individual animals, but recent technological advancements allow researchers to deploy sophisticated but affordable instruments, like drones and camera traps, to take photos of individual animals from which morphometric measurements are extracted. However, morphometric data obtained from photographs can be less accurate than those from direct measurement. In this paper we propose a new linear mixed-effects model approach, involving multivariate normal distributions, to describe the latent true measurements and to accommodate measurement error. Our method can be used to estimate relationships between dimensions, to predict the measurement of any one dimension from any subset of other dimensions, and to make inference on allometry, including testing for allometric growth. We demonstrate the use of our method with an application to morphometric data of the reef manta ray Mobula alfredi collected by drones.
Genetic interactions are essential for understanding the risk and progression of complex diseases. However, signals from individual genes and their pairwise interactions are often weak; most phenotypes are driven by alterations in a limited number of pathways and interactions between them. Identifying such pathways and their interactions is critical in biomedical research. Although traditional analyses have extended beyond main effects to include gene-gene interactions, most existing methods remain at the gene level and fail to capture higher-level pathway interactions. In this paper we propose a novel group-level model that jointly identifies key pathways and their interactions associated with clinical outcomes such as disease status or survival. The model involves estimating a high-dimensional binary matrix, which presents significant computational challenges. To overcome this, we reformulate the problem as a standard high-dimensional estimation task with hierarchical and exclusivity constraints and develop a two-stage estimation procedure. Theoretical analysis, simulation studies, and applications to TCGA breast cancer and Michigan lung cancer datasets demonstrate the superior performance of our method. In particular, our approach yields biologically meaningful insights, reveals novel gene-pathway mechanisms, and achieves substantially improved prediction accuracy and sensitivity, with comparable specificity to competing methods, including those modeling gene-gene interactions or employing two-step procedures that separately estimate pathways and their interactions.
Given the complex interactions between the microbiome, the host, and external factors, causal mediation analysis is essential for unraveling how dysbiosis or microbial imbalance mediates the effects of interventions or environmental exposures on health outcomes. However, zero inflation in microbiome count data complicates high-dimensional mediation analysis, as frequently employed zero-inflated models often struggle to distinguish true zero inflation from the underlying count distribution. To address this, we employ a zero-inflated negative binomial distribution for each mediator within a Bayesian framework, incorporating spike-and-slab priors to enable sparsity in estimating natural indirect effects (NIE) and an informative prior based on nonzero counts to improve dispersion estimation and account for zero sources. Recognizing the distinct biological mechanisms underlying microbial presence vs. abundance, we developed a dual mediation model, ZIMMA, to separate the NIE into pathways for mediator abundance and prevalence. Extensive simulations demonstrate ZIMMA’s superior performance in capturing distinct mediation mechanisms for both rare and abundant species compared to existing methods. Its application to human microbiome studies examining the effects of dietary intake and metabolic syndrome underscores its efficacy in identifying key microbial mediators, offering valuable insights and biological interpretation into the role of the microbiome in disease physiology and health sciences.
Somatic mutations, which accumulate in cells after conception, drive cancer development. Their characteristic patterns, namely mutational signatures, reflect the underlying mutational processes and have provided valuable insights into cancer etiology, evolution and therapeutic strategies. While non-negative matrix factorization (NMF) is commonly used to infer de novo mutational signatures and their activities, it requires large datasets for reliable estimation. When the sample size is limited, signature refitting is typically used, which estimates the signature activities using a set of reference signatures derived from external studies. However, current signature refitting methods often use the full list of reference signatures, leading to overfitting and compromised interpretability and accuracy. Despite its importance, the problem of selecting an appropriate subset of reference signatures received little attention. We proposed BayesSigRefitting, a Bayesian model selection framework to select an optimal subset of reference signatures for accurate refitting. Our approach employs a Bayesian hierarchical model with a sparsity-inducing Laplace prior, and the Shotgun Stochastic Search (SSS) algorithm to efficiently explore possible signature subsets and identify the optimal one. We established the model selection consistency of BayesSigRefitting and demonstrated, through simulation and real data studies across seven cancer types, that it outperformed existing methods in both signature selection and signature activity estimation. These findings highlight the potential of BayesSigRefitting to enhance the accuracy and reliability of mutational signature analysis, especially in settings with limited sample sizes.
The rapid growth of distributed energy resources (DERs) presents both opportunities and operational challenges for electric grid management. Accurately predicting DER adoption is critical for proactive infrastructure planning, but the inherent uncertainty and spatial disparity of DER growth complicate traditional forecasting approaches. Moreover, the hierarchical structure of distribution grids demands that predictions satisfy statistical guarantees at both the circuit and substation levels, a nontrivial requirement for reliable decision-making. In this paper we propose a novel uncertainty quantification framework for DER adoption predictions that ensures validity across hierarchical grid structures. Leveraging a multivariate Hawkes process to model DER adoption dynamics and a tailored split conformal prediction algorithm, we introduce a new nonconformity score that preserves statistical guarantees under aggregation while maintaining prediction efficiency. We establish theoretical validity under mild conditions and, through empirical evaluation on customer-level solar panel installation data from Indianapolis, Indiana, demonstrate that our method consistently outperforms existing baselines in both predictive accuracy and uncertainty calibration.
This paper is motivated by neuroscience studies aimed at understanding causal or predictive interactions between nodes in a brain network using multimodal brain activity data. To assess Granger causality, we introduce a flexible framework through a general class of models that accommodate mixed types of data (binary, count, continuous and positive components) formulated in a generalized linear model (GLM) fashion. To conduct statistical inference for causality, we propose a Bayesian mixed time series model that incorporates spike-and-slab priors on selected parameters. This framework enables effective selection of causal ordering and offers robust uncertainty quantification. The proposed methods are then applied to brain activity data, including both spike train and local field potential (LFP) activity, recorded as animals (rats) performed a complex sequence memory task. The proposed methodology provides critical insights into the causal relationship between band-specific spectral power in the LFP and subsequent spiking activity. Specifically, power in the LFP beta band is predictive of spiking activity 300 milliseconds later, providing a novel analytical tool for this area of emerging interest in neuroscience and demonstrating its usefulness and flexibility in the study of causality in general.
Time-dependent regionalization, or spatially restricted grouping, is a significant area of research focused on understanding the evolution of spatial clusters over time. In this study we adopt a probabilistic approach to regionalization, conceptualizing it as a random partition of geographic space at each time point, with the sequence of spatial partitions exhibiting time dependency. This methodology facilitates inference regarding the temporal dynamics of clusters. We employ a product partition prior for the random partitions at each time point, introducing temporal correlation among partitions through the temporal structure associated with prior cohesions. To explore partition search space effectively and ensure spatially constrained clustering, we utilize random spanning trees. This research is motivated by a pertinent applied problem: the identification of spatial and temporal patterns associated with mosquito-borne diseases. Given the overdispersion inherent in this type of data, we propose a spatiotemporal Poisson mixture model in which both mean and dispersion parameters vary according to spatiotemporal covariates. We apply the proposed model to analyze weekly reported cases of dengue from 2018 to 2023 in the Southeast region of Brazil. Additionally, we assess modeling performance using simulated data. Results indicate that our model is competitive in analyzing the temporal evolution of spatial clustering.
The discontinuous moulting process in crustaceans poses fundamental challenges for growth modelling and can lead to biologically implausible estimates of asymptotic size under traditional continuous-growth frameworks such as the von Bertalanffy curve. We develop a stochastic growth model that jointly characterises moult increment (MI) and intermoult period (IP) through a unified convolution-based likelihood. Individual growth is represented within a Lévy-inspired jump framework that enforces monotone but discontinuous size trajectories by modelling MI through a beta-type jump process with biologically realistic support, while IP is described by a gamma generalised linear model. This construction yields a joint likelihood for MI and IP and permits maximum likelihood estimation together with profile-likelihood inference for parameters governing both the size and timing of moults. In an application to tank data on Panulirus ornatus, we compare several plausible jump and waiting-time distributions and find that the beta-based specification provides a substantially better fit than gamma or inverse Gaussian alternatives, while yielding von Bertalanffy-type population summaries compatible with current stock-assessment practice. A simulation study shows that the proposed estimators recover key growth characteristics and reproduce the observed stepwise trajectories under realistic sample sizes. By embedding biologically constrained increments and intermoult timing within a stochastic jump-process framework, the model provides a flexible tool for analysing discontinuous growth in crustaceans and related biological systems, and illustrates how modern jump-process methods can inform fisheries management.
In the high-dimensional landscape, addressing the challenges of covariance regression with high-dimensional predictors has posed difficulties for conventional methodologies. This paper addresses these hurdles by presenting a novel approach for high-dimensional inference with covariance matrix outcomes. The proposed methodology is demonstrated through its application in identifying patterns of brain co-activation observed in functional magnetic resonance imaging (fMRI) experiments and in revealing the predictive role of brain structural connectivity mapped through diffusion tensor imaging (DTI). In the pursuit of dependable statistical inference, we introduce an integrative approach based on penalized estimation. This approach combines data splitting, variable selection, aggregation of low-dimensional estimators, and robust variance estimation. It enables the construction of reliable confidence intervals for covariate coefficients, supported by theoretical confidence levels under specified conditions, where asymptotic distributions are provided. Through various simulation studies, the proposed approach performs well for covariance regression in the presence of high-dimensional predictors. This innovative approach is applied to the Lifespan Human Connectome Project (HCP) Aging Study. Brain networks and corresponding regions are identified, where regional DTI metrics predict within network resting-state functional connectivity. The findings are in line with established knowledge of the human brain.
The extent to which the American public is politically polarized is of great interest in the lay and academic communities. To study opinion polarization, political scientists and public opinion researchers examine the distribution of respondents on survey items, using visual comparison of histograms and/or measures such as variances and bimodality coefficients. We prove these measures fail to align with prevailing conceptualizations of polarization put forth in the literature. To remedy this situation, we specify several properties a measure of polarization consistent with these conceptualizations should possess: in particular, it should increase as a distribution spreads away from a center toward the poles and/or as clustering below or above this center increases. We then propose a p-Wasserstein bipolarization index that satisfies these properties and measures the distance between the distribution of an item and a most polarized distribution with all mass concentrated on the lower and upper endpoints of the scale, using the index to examine bipolarization in attitudes toward governmental COVID-19 vaccine mandates across 11 countries: the U.S. and U.K. are most polarized, China, France, and India the least polarized, with Spain, Colombia, Italy, Brazil, Australia, and Canada occupying an intermediate position.
Our society is under constant threat from outbreaks of various infectious diseases, such as COVID-19, Zika, and others. The recent COVID-19 pandemic has claimed millions of lives and caused devastating disruption to our daily routines. Prompt detection of outbreaks and effective disease surveillance are critical yet challenging, due to the complex spatiotemporal dynamics of infectious disease spread. Existing analytical tools often rely on restrictive assumptions, such as data independence and specific parametric distributions, that are rarely valid in practical scenarios. Additionally, effective disease surveillance demands sequential decision-making since decisions should be made or updated whenever new data become available. But many existing methods were designed for retrospective data observed in a prespecified time interval and would not be effective for disease surveillance. To address these limitations, we develop a cumulative sum (CUSUM) control chart for sequentially monitoring the evolution of spatiotemporal disease incidence rates. The new chart is constructed under the sequential learning framework that continuously incorporates new data in updating the estimated baseline model. Unlike traditional methods, the new method can capture complex data structure including spatiotemporal data variation and correlation. Numerical studies indicate that it achieves faster and more reliable detection of disease outbreaks compared to some traditional methods.
In a diagnostic test using multiplex assay, each individual biomarker is often expected to have monotonic association with the disease outcome, and, therefore, the underlying disease classification rule is partially ordered with respect to the biomarkers. Nonparametric estimation of the classification rule can be accomplished by projecting an unconstrained Bayes estimator onto the partial ordering subspace. However, computing the projection is challenging as it involves performing maximization over a constrained parameter space whose size grows exponentially with the sample size. We introduce a novel sequential update method for projection-based nonparametric estimation of the disease classification rule and propose new recursive algorithms to implement the method. The proposed algorithms yields the exact Bayes solution that maximizes the posterior gain with respect to a classification-type gain function. When compared to an existing algorithm that gives approximate Bayes solution, our algorithms accomplish the same tasks with much reduced elapsed time in simulation study. We apply the sequential update method to evaluate a human papillomavirus test for cervical cancer precursor lesions and derive diagnostic rule that improves accuracy on existing estimation methods including parametric logistic regression and monotone generalized additive models.
Use of standard random forests may not guarantee reliable small area estimates unless a rich source of predictors explains the between-area heterogeneity. We propose mixed effects random forests with area random effects for small area estimation of general parameters. A new fitting algorithm with an embedded bootstrap-bias correction for the random forest residual variance is presented. Point estimators of small area parameters are constructed using a smearing estimator of the area-specific distribution function. Non-parametric block bootstrap is used for MSE estimation. The methodology is evaluated using household consumption data from Mozambique to derive district estimates of head count ratio and poverty gap. Comparisons to the empirical best predictor under a linear mixed model and to a synthetic estimator under the random forest are presented. Estimates are further contrasted to 2023 World Bank estimates and to design-unbiased direct estimates. The results show: (a) the advantages from including random effects in random forests, (b) the importance of data transformations for machine learning methods, (c) robustness properties of random forest-type methods, and (d) the importance of bias correcting the naive estimator of the random forest residual variance. Our conclusions demonstrate that a black-box approach to using machine learning methods should be avoided.
Human disease network (HDN) analysis, which jointly considers a large number of diseases and focuses on their interconnections, is getting increasingly popular and can shed important insight not possessed by individual-disease-based analysis. Multiple network analysis techniques have been developed for HDNs, although new developments are still strongly needed. In this article we adopt latent space modeling, which has proven powerful in other network analysis contexts and offers unique, insightful interpretations, but has been limitedly applied in HDN analysis. Different from some other types of network analysis and some other HDN analyses (such as gene-centric ones), in this article we pay unique attention to modeling temporal variations. For this purpose, a penalization approach is developed, which can identify time regions with constant network structures (that correspond to ignorable changes) as well as those with smooth variations. The statistical and computational properties are rigorously established. With Medicare data-one of the most powerful medical claims databases-we analyze the admission records of 133 million hospital inpatient treatments from January 2008 to December 2019. Sensible findings are made on disease interconnections and clustering structures. Additionally, the temporal variations, which have not been revealed in the literature, are found to be interpretable. The analysis can provide a new way for connecting and grouping diseases and assist in understanding and planning medical resources.
In this paper our objective is to identify the risk factors associated with adolescent marijuana use in Washington State, utilizing data from the 2018 and 2021 Healthy Youth Survey (HYS). Despite the survey's assurance of anonymity, the possibility of over-or underreporting exists due to various reasons, such as fear of being exposed, social stigma, and peer pressure. We are also interested in identifying factors that are associated with the occurrence of misreport. To achieve these goals, we develop a full Bayesian framework with a two-level latent linear regression model. The top level is for the true marijuana use response, and the second level is for the occurrence of misreporting. An informative prior is designed to seamlessly incorporate the domain knowledge or prior information while minimizing the risk of prior misspecification. We propose a partially collapsed Gibbs sampling algorithm with a Metropolis-Hastings step to sample the regression coefficients. Simulation studies have been conducted to demonstrate the superior performance of the proposed method over alternative approaches. Our analysis of HYS data discovers multiple factors for identifying at-risk adolescents and informing future prevention efforts.
High-dimensional measurements are often correlated, which motivates their approximation by factor models. This holds also true when features are engineered via low-dimensional interactions or kernel tricks. This often results in overparametrization and requires a fast dimensionality reduction. We propose a simple technique to enhance the performance of supervised learning algorithms by augmenting features with factors extracted from design matrices and their transformations. This is implemented by using the factors and idiosyncratic residuals which significantly weaken the correlations between input variables and hence increase the interpretability of learning algorithms and numerical stability. Extensive experiments on various algorithms and real-world data in diverse fields are carried out, among which we put special emphasis on the stock return prediction problem with Chinese financial news data due to the increasing interest in NLP problems in financial studies. We verify the capability of the proposed feature augmentation approach to boost overall prediction performance with the same algorithm. The approach bridges a gap in research that has been overlooked in previous studies, which focus either on collecting additional data or constructing more powerful algorithms, whereas our method lies in between these two directions using a simple PCA augmentation.
In survival analysis, accurate identification of latent classes is essential to effectively account for potential hidden population heterogeneity. In response to this challenge, we introduce the latent class discrete survival (LaCDS) model. LaCDS employs a finite-mixture model structure within the context of the discrete failure time model and implements the expectation-maximization algorithm for efficient optimization. Through extensive simulation studies, we evaluate the performance of LaCDS in comparison to other methods. Our results demonstrate LaCDS's superior ability to identify population heterogeneities, both in terms of baseline hazards and coefficients. Additionally, it is robust under both discrete and continuous simulation mechanisms. We apply LaCDS and other methods to identify subgroups among kidney transplant patients within the Organ Procurement and Transplantation Network (OPTN) study. Our findings underscore the superior accuracy of LaCDS in subgrouping homogeneous patients compared to existing methods.
We propose a Multilevel Dynamic Factor Model (ML-DFM) to capture the common global and region-specific stochastic trends in monthly centre and log-range temperatures observed at 68 locations across the Iberian Peninsula from January 1930 to December 2020. The specification of common trends is based on the analysis of temperatures at each location using unobserved component models, which decompose temperatures into trend, seasonal, and transitory components. First, we show that the centre and log-range temperatures evolve independently. Second, we remove the seasonal component before analysing common trends. Third, we find that centre temperature trends are well approximated by a smooth, integrated random walk with a time-varying slope. In contrast, a stochastic level better captures the dynamics of the log-range. The ML-DFM is estimated using an EM algorithm extended here to accommodate nonstationary factors. We show that, although the commonality in centre-temperature trends is considerable, the regional components remain relevant, particularly at the log-range.
Brain-computer interfaces (BCIs) enable direct communication between the brain and computers, providing critical tools for people with disabilities to communicate with the world. The performance of BCIs is often evaluated using BCI-utility, a comprehensive metric that balances both accuracy and speed in communication. This paper introduces a Bayesian reinforcement learning framework to optimize the BCI-utility of the P300 BCI, a BCI system that identifies a user's intended character on a virtual keyboard by analyzing EEG responses to stimuli. We construct confidence scores for each character based on EEG responses and then propose a unified learning framework that explicitly maximizes BCI utility. It integrates two key components: an early stopping policy and a dynamic stimulus selection policy. The early stopping policy is optimized using an actor-critic algorithm, while a Gaussian process-based Bayesian model is developed to learn transition dynamics to guide the selection of the next stimulus. The proposed framework effectively addresses critical implementation challenges, including pauses between characters, double-target issues, and delays caused by the time required for EEG responses. Extensive simulations under varying signalto-noise ratios (SNRs) and evaluations on recorded human EEG data demonstrate that our method significantly improves BCI-utility compared to existing approaches. This work highlights the potential of reinforcement learning to improve the performance and usability of P300 BCI systems.
This work is motivated by randomized clinical trial NCT03281876 (November 2017-March 2020), whose secondary aim was to evaluate the efficacy of the NTHi-Mcat vaccine vs. placebo in preventing recurrent severe exacerbations among patients with acute exacerbations of chronic obstructive pulmonary disease (AECOPD). The published analysis (Vaccine 40 (2022) 5924-5932; Lancet Respir. Med. 10 (2022) 435-446) aimed to estimate the ratio of the expected number of exacerbations one experienced by the end of study in the vaccinated vs. the placebo arm. One, therefore, regressed the number of exacerbations one experienced by the last point in time one is uncensored on treatment and baseline covariates with offset the logarithm of the observation time to account for different follow-up times. In this paper we demonstrate that this approach is prone to selection bias due to: (i) selective withdrawal and (ii) selective timing of the outcome measurements. We show that inverse probability of censoring weigting (IPCW), a common approach to adjust for dependent censoring under the assumption that censoring is non-informative given the observed covariate history, does not suffice to restore the unbiasedness of the treatment effect estimator under the above-mentioned type of analysis. To address this, we propose hazard inverse probability of censoring weighting (HIPCW). This novel weighting technique preserves the simplicity of IPCW but extracts more efficiency by using each individual's last recorded outcome. We validate the proposed approach through extensive simulations and compare with: (i) IPCW at a single time, (ii) a variant of the existing IPCW-GEE routines (J. Amer. Statist. Assoc. 90 (1995) 106-121) and (iii) IPCW-based estimators derived from the popular Andersen-Gill model (J. R. Stat. Soc. Ser. B. Stat. Methodol. 66 (2004) 239-257). We illustrate the routines through reanalysing clinical trial NCT03281876.