
Social science researchers are making increasing efforts to replicate important experiments, compare effect estimates and analyse their discrepancies. This article is motivated by a replication of a psychological experiment in which a decline in the effect of an eye movement intervention on false memory has been attributed to a distributional shift in participants' depression scores. Understanding how such shifts contribute to effect discrepancy is crucial for result reporting, scientific understanding, and treatment deployment. We propose a framework using generalizability methods to decompose effect size discrepancy between two studies into contributions of various sources of distribution shifts. We address several unique challenges in the application, including incorporating commonly cited sources of shift within a unified framework, providing interpretable summaries of contributions, dealing with limited sample sizes and selection bias. In the motivating study, our analysis reveals that, despite the notable shift in observed depression scores, due to limited effect heterogeneity, their contribution to effect discrepancy is less compelling than that of unobserved factors. In two additional pairs of experiments, our method offers context-dependent insights on the contributions of different factors in effect generalization. The proposed tools are especially useful early in the scientific process when there are not enough studies for a meta-analysis.
We propose a regression model for cylindrical response variables that arise in the analysis of events occurring randomly over time. Each response consists of two components, one related to the timing of the event (circular), and one related to the intensity or the consequences of the event (linear). We use the multivariate generalized Laplace distribution whose parameters make this model more flexible than-as well as a generalization of-the ordinary multivariate Gaussian model. For inference, we propose standard maximum likelihood procedures. We investigate about 134,000 road traffic accidents from the Fatality Analysis Reporting System of the US National Highway Traffic Safety Administration. We define the circular component as the time of the day at which the accident occurs, and the linear component as the gravity of the event. The latter comprises an overall score for the severity of the injuries experienced by the people involved in the accidents, along with the median age of the victims. Our analysis shows that the timing and severity of the accidents are influenced by temporal and geographical factors.
Traditional methods for analysing composite survival endpoints, such as the proportional hazards model, proportional mean, proportional win-fraction models (win ratio), and multi-state models, are limited by the subjective nature of event ranking, or by their inability to integrate multiple events into a unified measure. We propose a statistical method that transforms clinical observations into 'pseudo-death' by rescaling nonfatal events on the scale of death (considering death as the gold standard for survival analysis), generating scaling factors representing the conditional probability of death given the observed data. We then utilize pseudo-death to conduct the Z test and permutation test for evaluating treatment efficacy at a single time point, and the Wald test for multiple time points. We apply pseudo-death analysis to data from the V325 trial, which evaluates the addition of docetaxel to a cisplatin and fluorouracil regimen compared to cisplatin and fluorouracil alone in advanced gastric cancer. The pseudo-death framework detects significant treatment effects missed by traditional methods. In addition, simulation studies indicate that the proposed pseudo-death approach outperforms existing composite endpoint methods in various scenarios. Due to its simplicity, interpretability, and efficiency, the pseudo-death approach serves as a powerful tool for analysing composite endpoints.
Violence Against Women (VAW) is a pervasive and often underreported phenomenon, posing significant challenges for measurement and policy intervention. In Italy, the lack of recent and regularly updated ad hoc survey data has motivated our interest in official registers as an alternative source of geographically referenced information for estimating violence rates; however, these data are subject to significant underreporting. This study analyses VAW across Italian provinces using police registry data from 2020 within a hierarchical Bayesian Poisson regression framework. We adopt a Pogit model that accounts for the reporting mechanism, enabling joint inference on both the incidence of violence and the probability of reporting at the provincial level, incorporating socioeconomic and spatial covariates. To evaluate the robustness of the proposed approach, we conduct a simulation study assessing the sensitivity of posterior inferences to prior specification and covariate choice. The results indicate that the model accurately recovers both model parameters and true counts, even when only proxy covariates are available to model underreporting. Overall, the results underscore the importance of covariate quality and prior specification in improving inference and informing policies, particularly in settings where administrative data are biased or incomplete.
Small area estimation methods are generally based on models which assume normal errors, but many types of data do not follow a normal distribution. Several approaches have been suggested to deal with skewed data, including transformations (with and without bias correction), robust models which are less affected by the tails of the distributions and building models directly with skewed error distributions. We investigate the properties of models for transformed data with a real dataset which mimics a structural business survey. This contributes to the understanding of which tools are best for small area estimation with skewed data. We investigate the sensitivity of results to different shift parameters (used to make methods practical when data contain zeroes) and transformation parameters. The empirical best predictor (EBP) approach is found to be a flexible way to fit transformation-based models without the need for development of bias adjustments in back transformation. We prefer the EBP log-shift and EBP dual power which have good performance in our example (noting that the variables affecting the weighting are included in the model) because of their adaptability to new datasets. The bias-corrected empirical best estimator has similar performance in our example but is tailored to the log transformation.
Causal inference from a randomized trial becomes challenging when interest focuses on a specific subpopulation and when the causal effect involves both the randomized binary treatment and a non-randomized continuous exposure. This study is motivated by a clinical trial conducted among women of childbearing potential to evaluate the effectiveness of a malaria vaccine during pregnancy within the subpopulation of women who would have become pregnant under placebo conditions. Although the primary policy-relevant contrast concerns the binary randomized treatment-vaccine versus placebo-there also exists a continuous post-randomization exposure-exposure to pregnancy-characterized by the timing of conception. Given the seasonality of malaria transmission and waning of vaccine-induced immunity, vaccine effectiveness during pregnancy may vary with conception timing, motivating causal evaluations that consider both exposures. We propose methodological approaches for causal inference in this setting, study their large-sample and finite-sample properties, and apply them to the malaria vaccine trial to provide a comprehensive assessment of vaccine effectiveness against pregnancy malaria in the target subpopulation while accounting for temporal variation in disease transmission.
A multi-site city air-quality dataset should be considered distributed data as it is generated from multiple geographically dispersed sources, such as air quality sensors or monitoring stations. In various fields, distributed systems are increasingly employed to handle data collected from diverse sources, often resulting in datasets that are heavy-tailed, asymmetric, or heterogeneous. Robust expectile regression combines the computational efficiency of expectile regression with its robustness in handling heavy-tailed response distributions and outliers. This paper extends robust expectile regression to communication-efficient distributed systems and applies it to the analysis of multi-site air-quality datasets. The proposed distributed estimators achieve both computational and communication efficiency while delivering statistical performance comparable to global estimators, as demonstrated through both theoretical analysis and numerical experiments.
Establishing the causal relationship between treatment and outcome is a primary objective in many biomedical studies. In case-control studies, individuals are selected based on their outcome status, complicating causal inference due to the retrospective design. Further complications arise when the observed outcomes are subject to misclassifications. As a motivating example, the Global Enteric Multicenter Study utilizes a case-control design to investigate the causal effect of Cryptosporidium infection on diarrhoea in African children. The classification of diarrhoea status (outcome of interest) is based on caregiver-reported symptoms, which can differ from the true status. In fact, caregiver-reported classification has been reported to have a low sensitivity of approximately 16.8%. The presence of both case-control sampling and outcome misclassification poses a great challenge in accurately estimating causal effects, and we aim to resolve this issue in this paper. We establish nonparametric identifiability of the average treatment effect and conditional average treatment effect under both nondifferential and differential outcome misclassification scenarios by leveraging external information on disease prevalence and misclassification rates, and propose two novel estimation methods for the nondifferential scenarios. Extensive simulation studies and two real-data examples are provided to evaluate the finite-sample performance of the proposed estimators.
Heat stress is a growing public health concern in cities as urban dwellers are exposed to combined effect of global warming and urban heat islands. We develop a cutting-edge Bayesian Hierarchical Model that incorporates data from connected vehicles and citizen weather stations to draw hourly urban air temperature maps at hectometric resolution. To overpass the uncertainty of opportunistic observations, we set priors on the measurement error into two separate likelihoods, one per each data source. The model with joint-likelihoods is compared to two single models, one for the on-board thermometers and the other for citizen weather stations. All models, inferred with the INLA-SPDE approach, are evaluated against an independent professional network in the French city of Dijon. The maps are consistent with the reference network. The performance of the joint-likelihood model with both data sources exceeds the other two. Its Root Mean Square Error is less than 1 degrees C for more than 75% of the 714 hourly maps. These results open up new perspectives in urban climatology and for the post-processing of numerical weather forecasts on cities. They will also support research on urban heat exposure and all actors involved in sustainable urban planning.
Conventional modelling of networks evolving in time focuses on capturing variations in the network structure. However, the network might be static from the origin or experience only deterministic, regulated changes in its structure, providing either a physical infrastructure or a specified connection arrangement for some other processes. Thus, to detect the change in network use, we need to focus on the processes happening on the network. In this work, we present the concept of monitoring random temporal edge network processes that take place on the edges of a graph with a fixed structure. Our framework is based on the generalized network autoregressive statistical models with time-dependent exogenous variables (GNARX models) and cumulative sum control charts. To demonstrate its effective detection of various types of changes, we conduct a simulation study and monitor cross-border physical electricity flows in Europe.
Ecological surveys often generate count data to assess biodiversity and ecosystem dynamics, capturing species abundances or behavioural interactions such as pollination and resource competition. In this study, we analyse flower-invertebrate interactions in the Brazilian Atlantic Forest using a dataset compiled from multiple regional surveys. We examine 5,150 recorded interactions spanning four plant orders and 28 invertebrate orders, including effective pollination, general floral visits, and contact-based reproductive behaviours. Our main objectives are to model the distribution of interaction frequencies within the resulting network and to cluster flower-invertebrate interactions. These tasks are complicated by the heavy-tailed nature of the data. To address this, we develop a toolkit for sampling and fitting the Polylogarithm distribution, a discrete model capable of capturing heavy-tailed behaviour. We show that the Polylogarithm distribution unifies several well-known discrete heavy-tailed distributions, but its probability function is intractable. We overcome this limitation by designing a rejection sampler to generate exact draws, which forms the basis of a Bayesian inferential framework using Approximate Bayesian Computation. Finally, we extend the model to a finite mixture formulation to identify ecological interaction patterns. This clustering approach reveals three main plant groups and improves the fit to the empirical degree sequence.
Motivated by the analysis of complex single-cell gene expression data we propose a Bayesian class of generalized factor models for high dimensional count data. The developed methodology allows us to incorporate external knowledge, such as biological pathways, into the model's prior distribution. This approach promotes sparsity in the factor loadings facilitating their interpretation and that of the corresponding latent factors. We demonstrate the effectiveness of our model on single-cell RNA sequencing data obtained from cord blood mononuclear cells, revealing promising insights into the role of pathways in characterizing gene relationships and extracting valuable information about unobserved cell traits.
Genetic profiles of cancer patients are expected to vary with tumour progression. Therefore, it may be desirable to incorporate interactions between tumour stage and high-dimensional omics variables in prognostic models. These interactions may be confounded with other clinical risk factors. We present a novel interaction model for colorectal cancer prognosis based on 20,000+ omics variables, the tumour stage, and a small set of clinical risk factors. The model consists of a regression tree, fitted with only tumour stage and clinical risk factors, and omics-based regressions in the leaf nodes. To stabilize estimation of the node-specific omics effects, we develop a fusion-type penalized likelihood estimator, for which we derive shrinkage limits and computationally efficient tuning of hyperparameters. We show the benefit of the fused estimator in simulations. The colorectal cancer application reveals that FusedTree obtains good model fit compared to competitors and hence benefits from the incorporation of interaction effects. Furthermore, we develop a post hoc test suggesting that the overall omics effect does not further improve prognosis for subgroups of colorectal cancer patients.
Administrative health data contain rich information for investigating public health issues; however, many restrictions and regulations apply to their use. Such data are usually not in the conventional format for statistical analysis since the databases are created and maintained to serve non-research purposes and only information for people who seek health services is recorded. Analysis of administrative health data is thus challenging in general. We aim to develop a tool for understanding how the mental health of youth aged younger than 18 years evolves over time through administrative records of mental health related emergency department (MHED) visits. The MHED records, a set of zero-truncated recurrent events data, are integrated with relevant population census information, and framed into a set of doubly censored recurrent events data. We present innovative strategies for overcoming the zero-truncation induced by the data collection and for processing doubly censored recurrent events data. We compare the dynamic patterns and exposure impacts of the MHED visits in two decades with a loosely structured model. The findings are verified empirically via simulation. The asymptotic properties of the proposed estimator are established. Through exploring the paediatric MHED visit records, we provide new insights into children/youths mental health changes over time.
The National Crime Victimization Survey (NCVS) gathers information on criminal victimizations for individuals in a representative sample of United States households. The NCVS provides authoritative data on the rates of many types of violent crimes, including simple assault, robbery, and aggravated assault. Estimates are of interest for small domains defined by the intersection of sex with detailed age divisions. Standard survey estimators for these domains suffer from instability due to small sample sizes. Model-based small area procedures are needed to obtain more reliable estimates. We employ a multivariate Bayesian model to obtain small area estimates for domains defined by intersections of sex with specific age categories. We construct estimates for four types of violent crimes in each of two time periods. We compare a model with a log transformation to a model fit to the data in the original scale. We compare small area predictors based on a selected model to the direct estimators.
We present a novel approach to ecological risk assessment by recasting the species sensitivity distribution (SSD) method within a Bayesian nonparametric (BNP) framework. Widely mandated by environmental regulatory bodies globally, SSD has faced criticism due to its historical reliance on parametric assumptions when modelling species variability. By adopting nonparametric mixture models, we address this limitation, establishing a statistically robust foundation for SSD. Our BNP approach offers several advantages, including its efficacy in handling small datasets or censored data, which are common in ecological risk assessment, and its ability to provide principled uncertainty quantification alongside simultaneous density estimation and clustering. We utilize a specific nonparametric prior as the mixing measure, chosen for its robust clustering properties, a crucial consideration given the lack of strong prior beliefs about the number of components. Through simulation studies and analysis of real datasets, we demonstrate the superiority of our BNP-SSD over classical SSD methods. We also provide a BNP-SSD Shiny application, making our methodology available to the Ecotoxicology community. Moreover, we exploit the inherent clustering structure of the mixture model to explore patterns in species sensitivity. Our findings underscore the effectiveness of the proposed approach in improving ecological risk assessment methodologies.
Crash risk affects every driver in some capacity and quantifying this risk is a key outcome in evaluating transportation safety. This study develops a statistical model for estimating continuous crash risk surfaces from discrete, multivariate crash count data. We focus on Interstate 15 (I-15) in Utah, USA, using aggregated crash counts released by the Utah Department of Transportation (UDOT). We model segment-level crash counts as a realization of an aggregated multivariate point pattern with continuous intensity surfaces; thus enabling continuous spatial risk estimation. We link roadway characteristics to crash risk using multivariate random forests, capturing nonlinear relationships and correlations between crash types. This approach supports the identification of high-risk areas and contributing roadway features, ultimately aiming to support decisions related to targeted safety interventions.
The Carbon Footprint (CFP), derived from household consumption survey data, is a crucial indicator for assessing the impact of human consumption on greenhouse gas emissions. The definition and measurement of personal CFP rely on data that require the comparability between classifications of firm production and household consumption. In this article, we address three key aspects related to the definition and estimation of personal CFP. First, to compute the CFP, we build upon the methodology developed by Pang et al. (2020, Urban carbon footprints: A consumption-based approach for Swiss households. Environmental Research Communications, 2(1), 011003.) by incorporating a conversion factor matrix into the formulas, using official Eurostat data. This matrix serves as a bridge between macroeconomic data across different statistical classifications of production and consumption. Second, aiming to conduct inferential procedures on CFP, we select the probability distribution that best fits CFP empirical distribution. The Generalized Beta Distribution of the Second Kind (GB2) provides the best fit. Third, in order to map local CFP through reliable estimates we propose a Small Area Estimation (SAE) model based on Generalized Additive Models for Location, Scale, and Shape (SAE-GAMLSS), assuming CFP follows a GB2 distribution. Finally, we emphasize the significance of mapping per-capita CFP based on reliable local estimates to support the implementation of effective place-based policies.
Bike sharing is an increasingly popular mobility choice as it is a sustainable, healthy, and economically viable transportation mode. By interpreting rides between bike stations over time as temporal events connecting two bike stations, relational event models can provide important insights into this phenomenon. The focus of relational event models, as a typical event history model, is normally on dyadic or node-specific covariates, as global covariates are considered nuisance parameters in a partial likelihood approach. As full likelihood approaches are infeasible given the sheer size of the relational process, we propose an innovative sampling approach of temporally shifted non-events to recover important global drivers of the relational process. The method combines nested case-control sampling on a time-shifted version of the event process. This leads to a partial likelihood of the relational event process that is identical to that of a degenerate logistic additive model, enabling efficient estimation of both global and non-global covariate effects. The computational effectiveness of the method is demonstrated through a simulation study. The analysis of around 350,000 bike rides in the Washington D.C. area reveals significant influences of weather and time of day on bike sharing dynamics, besides a number of traditional node-specific and dyadic covariates.
In this paper, we propose a new flexible family of distributions for data that consist of three angles, two angles and one linear component, or one angle and two linear components. To achieve this, we equip the recently proposed trivariate wrapped Cauchy copula with non-uniform marginals and develop a parameter estimation procedure. We compare our model to its main competitors for analysing trivariate data and provide some evidence of its advantages. We illustrate our new model using toroidal data from protein bioinformatics of conformational angles, and cylindrical data from climate science related to a buoy in the Adriatic Sea. The paper is motivated by these real trivariate datasets.