
Evidence syntheses are often updated as new trials become available. A cumulative meta-analysis repeats a meta-analysis in chronological order, and trial sequential analysis applies group-sequential principles to cumulative meta-analysis to control the risk of spurious findings due to repeated testing. Despite the growing popularity of trial sequential analysis, existing methods rely heavily on normal approximations derived from interim analyses of randomized controlled trials, where participants are typically more homogeneous than in meta-analyses. In random-effects meta-analyses, the conventional assumption that the synthesized effect estimate follows a normal distribution can perform poorly when the number of studies is small. In such settings, the Hartung-Knapp-Sidik-Jonkman method, which is based on a t distribution, has been recommended for more reliable inference. This article introduces refined trial sequential procedures based on cumulative meta-analytic t statistics. The proposed methods are designed to reduce the risk of premature or spurious conclusions in updating evidence syntheses, particularly when between-study heterogeneity is present. Numerical studies demonstrate that the proposed methods provide improved control of type I error compared with existing methods, although the degree of improvement depends on the magnitude of heterogeneity and the true effect size.
Abstract Multi-parametric magnetic resonance imaging (mpMRI) plays a critical role in the detection and localization of prostate cancer (PCa). 3D mpMRI represents prostate anatomy as a volume assembled from multiple 2D slices, providing richer information for lesion detection and localization than analysis based on a single slice alone. Existing statistical 3D approaches often struggle to ensure spatial contiguity of detected regions and may fail to capture the complex geometry of cancer boundaries. In addition, our motivating application uses 3D prostate mpMRI data from a single-site study, where the limited sample size can make deep learning approaches less practical for reliable training of high-dimensional 3D models. To address these challenges, we develop BFSP-3D, a Bayesian functional spatial partitioning method for automatic lesion detection in 3D imaging data. The proposed method employs a transformation from Cartesian to spherical coordinates to flexibly represent smooth 3D lesion boundaries. An adaptive Metropolis–Hastings algorithm is used to jointly estimate the boundary surface and the parameters of the underlying Gaussian process. Simulation studies demonstrate that BFSP-3D is robust to varying spatial configurations and lesion shapes. We illustrate the method using 3D mpMRI data of the prostate.
Abstract The human body harbours diverse microbiomes that play a vital role in health, with microbial interactions providing key insights into many disease mechanisms. Uncovering the complex microbiome community structure is essential to understanding the role of the microbiome in disease progression and susceptibility. This paper introduces a novel weighted stochastic infinite block model for analyzing fully connected microbial interaction networks. Prior to model fitting, microbial interaction networks are constructed using a recently developed semi-parametric rank-based correlation method that accounts for zero-inflation and latent confounding during network construction. By adopting a Bayesian nonparametric approach, our model clusters taxa into distinct communities and estimates the number of communities without the need for ad hoc selection or post-processing. Through simulations and real data from postmenopausal women with recurrent urinary tract infections, we show our model achieves superior clustering accuracy over alternatives and uncovers several novel biological insights. Notably, this is the first study to explore urinary microbiome interaction networks while preserving all information from taxonomic abundance data.
Abstract Randomized response (RR) techniques (RRTs) protect privacy but yield latent sensitive variables. Existing models primarily focus on predicting the sensitive attribute rather than characterizing its association with non-sensitive covariates. We propose a likelihood-based joint modelling framework for a latent binary sensitive variable and an observed non-sensitive variable, conditional on covariates, under an unrelated-question RR design. By parameterizing their joint distribution through a logistic regression model for the sensitive variable and a conditional model for the non-sensitive variable, the framework enables assessment of conditional association between the two variables given covariates. Conditional independence is formally evaluated using likelihood-based tests. Parameter estimation is implemented via an expectation–maximization algorithm that accounts for the RRT mechanism, with established large-sample properties. Simulation studies and an application to the 2012 Taiwan Social Change Survey demonstrate the finite-sample performance and practical utility of the proposed method. All findings are interpreted as conditional associations rather than causal effects.
Abstract We introduce a novel approach that integrates compositional data analysis (CoDA) in the forecasting of hierarchical and grouped time series and apply it to improve mortality forecasts. By leveraging the inherent decomposition of mortality data by cause of death and sex, our framework addresses key challenges in mortality forecasting, ensuring internal coherence across different levels of aggregation while maintaining high predictive accuracy. Our empirical results on Italian cause-specific mortality data demonstrate that the proposed approach significantly improves the reliability of forecasts compared to the traditional CoDA, especially for total mortality forecasts. Consistent improvements in forecasting accuracy are also found for total deaths by cause and by sex.
Cancer remains the second most prevalent cause of death in the United States, claiming 605,213 lives in 2021, surpassing COVID-19 deaths. The cancer mortality rate continued to decline between 2019 and 2020, dropping by 1.5%, marking a significant 33% decrease since 1991. This ongoing improvement primarily mirrors advances in treatment, allowing patients to achieve clinical remission and recovery. Now, a cancer patient is simultaneously exposed to the risk of primary cancer as well as other risks, such as other cancer(s) or other diseases, leading to a competing risks scenario. Analysis of survival data under competing risks and the presence of cured patients have been extensively studied individually, but there is limited work in the current literature that models the possibility of cure from one risk in the presence of competing risks. Moreover, such a model should allow for the possibility of cure from the cause-specific risk of the primary cancer; however, the overall survival probability should eventually approach zero, thereby incorporating the prevalent belief of eventual failure with certainty. We propose a novel unified competing risks cure model, based on the cause-specific hazard approach, that satisfies the aforementioned desired properties. The conditions required to establish model identifiability are studied in detail. To find the maximum likelihood estimates of the model parameters, a computationally efficient expectation maximization algorithm is developed. An extensive simulation study is carried out to demonstrate the performance of the proposed model and estimation method under different parameter settings and in the presence of multiple competing risks. Finally, an application is illustrated using breast cancer data from the SEER cancer database.
We propose in this paper a framework for performing overlapping clustering in the presence of time dependence, where the main goal is to identify overlapping coarse parcelations of the brain based on resting state fMRI time series. Our procedure is based on the Latent OVErlapping (LOVE) clustering method of Bing et. al (2020) which we extend to weakly dependent time series. Although the method is developed in an fMRI context, it is general enough to be directly applicable to other contexts, such as gene expression data, where one has at disposal multiple time series and is interested in identifying overlapping groups of similar elements.
Social science researchers are making increasing efforts to replicate important experiments, compare effect estimates and analyse their discrepancies. This article is motivated by a replication of a psychological experiment in which a decline in the effect of an eye movement intervention on false memory has been attributed to a distributional shift in participants' depression scores. Understanding how such shifts contribute to effect discrepancy is crucial for result reporting, scientific understanding, and treatment deployment. We propose a framework using generalizability methods to decompose effect size discrepancy between two studies into contributions of various sources of distribution shifts. We address several unique challenges in the application, including incorporating commonly cited sources of shift within a unified framework, providing interpretable summaries of contributions, dealing with limited sample sizes and selection bias. In the motivating study, our analysis reveals that, despite the notable shift in observed depression scores, due to limited effect heterogeneity, their contribution to effect discrepancy is less compelling than that of unobserved factors. In two additional pairs of experiments, our method offers context-dependent insights on the contributions of different factors in effect generalization. The proposed tools are especially useful early in the scientific process when there are not enough studies for a meta-analysis.
We propose a regression model for cylindrical response variables that arise in the analysis of events occurring randomly over time. Each response consists of two components, one related to the timing of the event (circular), and one related to the intensity or the consequences of the event (linear). We use the multivariate generalized Laplace distribution whose parameters make this model more flexible than-as well as a generalization of-the ordinary multivariate Gaussian model. For inference, we propose standard maximum likelihood procedures. We investigate about 134,000 road traffic accidents from the Fatality Analysis Reporting System of the US National Highway Traffic Safety Administration. We define the circular component as the time of the day at which the accident occurs, and the linear component as the gravity of the event. The latter comprises an overall score for the severity of the injuries experienced by the people involved in the accidents, along with the median age of the victims. Our analysis shows that the timing and severity of the accidents are influenced by temporal and geographical factors.
Traditional methods for analysing composite survival endpoints, such as the proportional hazards model, proportional mean, proportional win-fraction models (win ratio), and multi-state models, are limited by the subjective nature of event ranking, or by their inability to integrate multiple events into a unified measure. We propose a statistical method that transforms clinical observations into 'pseudo-death' by rescaling nonfatal events on the scale of death (considering death as the gold standard for survival analysis), generating scaling factors representing the conditional probability of death given the observed data. We then utilize pseudo-death to conduct the Z test and permutation test for evaluating treatment efficacy at a single time point, and the Wald test for multiple time points. We apply pseudo-death analysis to data from the V325 trial, which evaluates the addition of docetaxel to a cisplatin and fluorouracil regimen compared to cisplatin and fluorouracil alone in advanced gastric cancer. The pseudo-death framework detects significant treatment effects missed by traditional methods. In addition, simulation studies indicate that the proposed pseudo-death approach outperforms existing composite endpoint methods in various scenarios. Due to its simplicity, interpretability, and efficiency, the pseudo-death approach serves as a powerful tool for analysing composite endpoints.
Violence Against Women (VAW) is a pervasive and often underreported phenomenon, posing significant challenges for measurement and policy intervention. In Italy, the lack of recent and regularly updated ad hoc survey data has motivated our interest in official registers as an alternative source of geographically referenced information for estimating violence rates; however, these data are subject to significant underreporting. This study analyses VAW across Italian provinces using police registry data from 2020 within a hierarchical Bayesian Poisson regression framework. We adopt a Pogit model that accounts for the reporting mechanism, enabling joint inference on both the incidence of violence and the probability of reporting at the provincial level, incorporating socioeconomic and spatial covariates. To evaluate the robustness of the proposed approach, we conduct a simulation study assessing the sensitivity of posterior inferences to prior specification and covariate choice. The results indicate that the model accurately recovers both model parameters and true counts, even when only proxy covariates are available to model underreporting. Overall, the results underscore the importance of covariate quality and prior specification in improving inference and informing policies, particularly in settings where administrative data are biased or incomplete.
Small area estimation methods are generally based on models which assume normal errors, but many types of data do not follow a normal distribution. Several approaches have been suggested to deal with skewed data, including transformations (with and without bias correction), robust models which are less affected by the tails of the distributions and building models directly with skewed error distributions. We investigate the properties of models for transformed data with a real dataset which mimics a structural business survey. This contributes to the understanding of which tools are best for small area estimation with skewed data. We investigate the sensitivity of results to different shift parameters (used to make methods practical when data contain zeroes) and transformation parameters. The empirical best predictor (EBP) approach is found to be a flexible way to fit transformation-based models without the need for development of bias adjustments in back transformation. We prefer the EBP log-shift and EBP dual power which have good performance in our example (noting that the variables affecting the weighting are included in the model) because of their adaptability to new datasets. The bias-corrected empirical best estimator has similar performance in our example but is tailored to the log transformation.
Causal inference from a randomized trial becomes challenging when interest focuses on a specific subpopulation and when the causal effect involves both the randomized binary treatment and a non-randomized continuous exposure. This study is motivated by a clinical trial conducted among women of childbearing potential to evaluate the effectiveness of a malaria vaccine during pregnancy within the subpopulation of women who would have become pregnant under placebo conditions. Although the primary policy-relevant contrast concerns the binary randomized treatment-vaccine versus placebo-there also exists a continuous post-randomization exposure-exposure to pregnancy-characterized by the timing of conception. Given the seasonality of malaria transmission and waning of vaccine-induced immunity, vaccine effectiveness during pregnancy may vary with conception timing, motivating causal evaluations that consider both exposures. We propose methodological approaches for causal inference in this setting, study their large-sample and finite-sample properties, and apply them to the malaria vaccine trial to provide a comprehensive assessment of vaccine effectiveness against pregnancy malaria in the target subpopulation while accounting for temporal variation in disease transmission.
A multi-site city air-quality dataset should be considered distributed data as it is generated from multiple geographically dispersed sources, such as air quality sensors or monitoring stations. In various fields, distributed systems are increasingly employed to handle data collected from diverse sources, often resulting in datasets that are heavy-tailed, asymmetric, or heterogeneous. Robust expectile regression combines the computational efficiency of expectile regression with its robustness in handling heavy-tailed response distributions and outliers. This paper extends robust expectile regression to communication-efficient distributed systems and applies it to the analysis of multi-site air-quality datasets. The proposed distributed estimators achieve both computational and communication efficiency while delivering statistical performance comparable to global estimators, as demonstrated through both theoretical analysis and numerical experiments.
Establishing the causal relationship between treatment and outcome is a primary objective in many biomedical studies. In case-control studies, individuals are selected based on their outcome status, complicating causal inference due to the retrospective design. Further complications arise when the observed outcomes are subject to misclassifications. As a motivating example, the Global Enteric Multicenter Study utilizes a case-control design to investigate the causal effect of Cryptosporidium infection on diarrhoea in African children. The classification of diarrhoea status (outcome of interest) is based on caregiver-reported symptoms, which can differ from the true status. In fact, caregiver-reported classification has been reported to have a low sensitivity of approximately 16.8%. The presence of both case-control sampling and outcome misclassification poses a great challenge in accurately estimating causal effects, and we aim to resolve this issue in this paper. We establish nonparametric identifiability of the average treatment effect and conditional average treatment effect under both nondifferential and differential outcome misclassification scenarios by leveraging external information on disease prevalence and misclassification rates, and propose two novel estimation methods for the nondifferential scenarios. Extensive simulation studies and two real-data examples are provided to evaluate the finite-sample performance of the proposed estimators.
Heat stress is a growing public health concern in cities as urban dwellers are exposed to combined effect of global warming and urban heat islands. We develop a cutting-edge Bayesian Hierarchical Model that incorporates data from connected vehicles and citizen weather stations to draw hourly urban air temperature maps at hectometric resolution. To overpass the uncertainty of opportunistic observations, we set priors on the measurement error into two separate likelihoods, one per each data source. The model with joint-likelihoods is compared to two single models, one for the on-board thermometers and the other for citizen weather stations. All models, inferred with the INLA-SPDE approach, are evaluated against an independent professional network in the French city of Dijon. The maps are consistent with the reference network. The performance of the joint-likelihood model with both data sources exceeds the other two. Its Root Mean Square Error is less than 1 degrees C for more than 75% of the 714 hourly maps. These results open up new perspectives in urban climatology and for the post-processing of numerical weather forecasts on cities. They will also support research on urban heat exposure and all actors involved in sustainable urban planning.
Conventional modelling of networks evolving in time focuses on capturing variations in the network structure. However, the network might be static from the origin or experience only deterministic, regulated changes in its structure, providing either a physical infrastructure or a specified connection arrangement for some other processes. Thus, to detect the change in network use, we need to focus on the processes happening on the network. In this work, we present the concept of monitoring random temporal edge network processes that take place on the edges of a graph with a fixed structure. Our framework is based on the generalized network autoregressive statistical models with time-dependent exogenous variables (GNARX models) and cumulative sum control charts. To demonstrate its effective detection of various types of changes, we conduct a simulation study and monitor cross-border physical electricity flows in Europe.
Ecological surveys often generate count data to assess biodiversity and ecosystem dynamics, capturing species abundances or behavioural interactions such as pollination and resource competition. In this study, we analyse flower-invertebrate interactions in the Brazilian Atlantic Forest using a dataset compiled from multiple regional surveys. We examine 5,150 recorded interactions spanning four plant orders and 28 invertebrate orders, including effective pollination, general floral visits, and contact-based reproductive behaviours. Our main objectives are to model the distribution of interaction frequencies within the resulting network and to cluster flower-invertebrate interactions. These tasks are complicated by the heavy-tailed nature of the data. To address this, we develop a toolkit for sampling and fitting the Polylogarithm distribution, a discrete model capable of capturing heavy-tailed behaviour. We show that the Polylogarithm distribution unifies several well-known discrete heavy-tailed distributions, but its probability function is intractable. We overcome this limitation by designing a rejection sampler to generate exact draws, which forms the basis of a Bayesian inferential framework using Approximate Bayesian Computation. Finally, we extend the model to a finite mixture formulation to identify ecological interaction patterns. This clustering approach reveals three main plant groups and improves the fit to the empirical degree sequence.
Motivated by the analysis of complex single-cell gene expression data we propose a Bayesian class of generalized factor models for high dimensional count data. The developed methodology allows us to incorporate external knowledge, such as biological pathways, into the model's prior distribution. This approach promotes sparsity in the factor loadings facilitating their interpretation and that of the corresponding latent factors. We demonstrate the effectiveness of our model on single-cell RNA sequencing data obtained from cord blood mononuclear cells, revealing promising insights into the role of pathways in characterizing gene relationships and extracting valuable information about unobserved cell traits.
Genetic profiles of cancer patients are expected to vary with tumour progression. Therefore, it may be desirable to incorporate interactions between tumour stage and high-dimensional omics variables in prognostic models. These interactions may be confounded with other clinical risk factors. We present a novel interaction model for colorectal cancer prognosis based on 20,000+ omics variables, the tumour stage, and a small set of clinical risk factors. The model consists of a regression tree, fitted with only tumour stage and clinical risk factors, and omics-based regressions in the leaf nodes. To stabilize estimation of the node-specific omics effects, we develop a fusion-type penalized likelihood estimator, for which we derive shrinkage limits and computationally efficient tuning of hyperparameters. We show the benefit of the fused estimator in simulations. The colorectal cancer application reveals that FusedTree obtains good model fit compared to competitors and hence benefits from the incorporation of interaction effects. Furthermore, we develop a post hoc test suggesting that the overall omics effect does not further improve prognosis for subgroups of colorectal cancer patients.