
This paper presents a Bayesian hierarchical model designed to address spatial location perturbation in Demographic and Health Survey (DHS) data. This perturbation can matter significantly in analyses that rely on fine-scale geographic proximity and precise distance calculations. The method builds on measurement error modelling by incorporating prior information on the perturbation process reported by the DHS programme. To enhance estimation, we further adjust for error using a Moran's I operator matrix, which links cluster mid-points with areal-level spatial units such as municipal districts or upazilas. Simulation studies and empirical applications with DHS data from Bangladesh and Ghana demonstrate that the proposed method produces more concise and reliable posterior estimates at the cluster level. While it does not recover the true cluster locations, the approach effectively mitigates bias introduced by spatial perturbation and provides a robust adjustment for measurement error. This framework offers a balance between protecting confidentiality and ensuring analytical validity, which supports more robust spatial inference in population health research.
In-hospital mortality and length of stay are fundamental metrics for evaluating healthcare quality, patient outcomes, and resource utilization. While length of stay reflects hospital efficiency and capacity management, mortality provides insights into patient safety and the effectiveness of clinical interventions. These outcomes are interdependent, and demographic, clinical and laboratory factors simultaneously influence both hospitalization duration and mortality. To address this, a copula additive distributional regression framework is employed, enabling the joint modelling of these hospital metrics as functions of covariate effects. Application to COVID-19 data demonstrates that key predictors, including age, oxygenation and inflammation markers, modulate the dependence between mortality and hospitalization duration. The joint modelling approach provides a probabilistic, patient-level characterization of the interplay between these indicators, supporting risk stratification, resource planning and actionable clinical decision-making.
Behaviour by individuals or organizations is often interdependent. Social contagion posits that behaviour spreads from unit to unit due to the presence of network or equivalence relations as transmission pathways. Contagion of a single behaviour has been modelled in cross-sectional and temporal data contexts. But existing statistical approaches have not been able to identify multiple contagion pathways in temporal processes where multiple actors can display or adopt multiple behaviours. This data structure and problem setting is common, for example in health behaviours by peers, treaty ratification by states, the spread of wildfire incidents in forests, or the diffusion of policies or political beliefs. We explore the application of bipartite relational event models of actors and behaviours and find that temporally backward-looking specifications confound social contagion with prior similarity, the tendency of similar units to adopt the same behaviour independently. We construct a set of sufficient statistics parsing information bidirectionally along the event sequence to establish an atemporal prior similarity null distribution against which contagion hypotheses for multiple pathways can be tested. Using simulations and four empirical cases, we show the efficacy of this parametric approach for disentangling contagion from prior similarity, contributing to causal inference for temporal networks.
Corporations, unions, and other interest groups have become key sponsors of television advertising in US elections after the Supreme Court's decision in Citizens United v. FEC that eliminated restrictions on such spending. This paper estimates the partisan effects of ads sponsored by these groups to obtain a more complete picture of voter behaviour and electoral politics. Advertising strategies vary over the course of the campaign, making marginal structural models a natural tool for this setting. Unfortunately, this approach requires an assumption of no unobserved confounders between the treatment and outcome, which may not be plausible with observational electoral data. We propose a novel weighting estimator with propensity-score fixed effects to adjust for time-constant unmeasured confounding in marginal structural models of fixed-length treatment histories. This estimator is consistent and asymptotically normal when the number of units and time periods grow at a similar rate. Unlike traditional fixed effect models, this approach works even when the outcome is only measured at a single point in time as in our setting. Against conventional wisdom, we find interest group ads are only effective when run by Democratic groups, and these effects are most prominent after Donald Trump became a presidential candidate in 2015.
Selection bias correction is generally crucial to have a reliable inference for a nonprobability sample. Often the correction is applied with the two-sample setup. That is, along with the nonprobability sample we are interested in, a probability sample sharing some common auxiliary variables is used for constructing correction weights for the nonprobability sample. The two-sample setup allows one to calculate weighted estimates for population parameters of interest based on the nonprobability sample. Since the nonprobability sample is usually easy to collect, we often end up with a large nonprobability sample and a small probability sample. The imbalance of the two samples may cause difficulties in modelling the propensity of units to be included in the nonprobability sample. This paper discusses some often-seen solutions for imbalanced samples in machine learning literature, i.e. undersampling, Synthetic Minority Oversampling Technique (SMOTE), and a mixture of both. A selection bias correction framework is adjusted to incorporate the imbalance solutions. Three evaluation studies across different types of data sets are shown. The results indicate that SMOTE has the potential to deal with imbalances in selection bias correction, while further study is needed to identify in which scenarios SMOTE works well.
Due to the potential association between the longitudinal and time-to-event data, these two types of data are often jointly analysed to obtain less biased and more efficient inferences. A regular joint model (RJM) normally assumes there exist subject-specific latent random effects or classes shared by the longitudinal and time-to-event processes and the two processes are conditionally independent given these latent variables. Under this assumption, the joint likelihood of the two processes is straightforward to derive and their association, as well as the heterogeneity among the population, are naturally introduced by the unobservable latent variables. However, because of the unobservable nature of these latent variables, the conditional independence assumption is difficult to verify. Therefore, in addition to the time-invariant random effects, a time-varying bivariate copula is introduced to account for the extra time-dependent association between the two processes. The proposed time-varying bivariate copula joint model includes a RJM as a special case under specific copulas. Our study indicates the proposed model is robust to copula misspecification in parameter estimation and superior in predicting survival probabilities compared to a RJM. A real data application on the primary biliary cirrhosis data is used to illustrate the merits of this method.
This paper introduces an area-level Dirichlet mixed model for predicting compositional indicators of small areas. Direct estimators of the domain category proportions of a classification variable are the target variables of the new model. Once the model has been selected and fitted to the data, predictors of proportions, totals and rates of small areas are obtained and their mean square errors are estimated by parametric bootstrap. Several simulation experiments, designed to analyse the behaviour of the fitting algorithm, the small area predictors and the bootstrap procedure, are carried out. An application to real data from the Spanish Labour Force Survey, in the last quarter of 2022, is given. The target is the estimation of proportions of employed, unemployed and inactive people and unemployment rates by province, sex and age group.
The survey quality predictor (SQP) is a web-based tool designed to predict the measurement quality of survey questions based on up to 72 manually coded formal and linguistic characteristics (e.g. domain, response scale properties, linguistic complexity). Users must input these features following a detailed coding manual, after which a trained random forest (RF) model predicts measurement quality. Here, we evaluate whether measurement quality can instead be predicted directly from the natural language text of a survey question, eliminating the need for manual coding. We find a fine-tuned language model can predict survey item quality based solely on the question and answer options, achieving performance comparable to the RF model currently implemented in SQP, which is trained on manually coded features. Specifically, we fine-tuned xlm-RoBERTa, a multilingual transformer-based model trained on multiple text corpora in over 100 languages, using the SQP dataset. Our findings suggest the current SQP web interface (https://sqp.gesis.org), which requires users to manually code the 72 features, can be simplified. A redesigned, more user-friendly web interface could allow users to only enter the survey question and answer options, with the model automatically predicting measurement quality. This would enhance accessibility and lower the barrier for researchers and practitioners using SQP.
We statistically analyse the link between extreme weather events, urban expansion, and tree cover loss in a data set of regions in 163 countries from 2001 to 2018. While previous research has focused on these phenomena individually, we provide quantitative evidence on the chain of events in a whole cycle. A simultaneous equation model allows us to capture the interrelations and to quantify their strength. We show how droughts and floods can drive urban expansion, triggering deforestation, which in turn increases vulnerability to disasters. In a heterogeneity analysis across continents and levels of development, we uncover differences in the strength of the cycle and the statistical significance of its links. These findings emphasize the need for targeted, region-specific adaptation policies. We stress the importance of multilevel governance and land use planning to balance urban growth, environmental sustainability, and disaster impact reduction.
When researchers recruit participants for an online survey via Facebook advertisements, a targeting algorithm determines which individuals receive the survey invitation. Because the algorithm maximizes the number of clicks on the survey link, predominantly easy-to-reach population groups in a country or region might be reached (simple demographic targeting: SDT). To reduce representation bias, researchers can target selected demographic population groups separately (complex demographic targeting: CDT). In this study, we evaluate whether CDT effectively leads to less bias in univariate, bivariate, and multivariate estimates than SDT. This is done by comparing socio-demographic and health-related measures of health surveys recruited via Facebook with and without CDT in six African countries against probability-based survey benchmarks. Independent of the targeting strategy, our results show that many estimates were strongly biased, especially the univariate estimates of education and Internet use, whereas other estimates (e.g. HIV knowledge) were less affected. Although we found minor evidence that the CDT method reduced bias in univariate estimates, bias in relationship estimates stayed similar. Moreover, the study shows that univariate estimates are generally more often biased than relationship estimates and that bias is partially related to coverage issues. The success of weighting in bias reduction was mixed.
This paper introduces a novel multivariate composite estimator for the Labour Force Survey (LFS). The estimator improves upon traditional methods by simultaneously modelling all labour market categories, such as employment, unemployment, and nonparticipation. This multivariate approach avoids the asymmetrical treatment of a residual category and leverages the correlation structure between the different labour market categories. The paper also presents a framework for explicitly incorporating and estimating wave-specific biases, which can vary over time. We derive analytical formulas for the estimator's variance and demonstrate its properties using data from the Norwegian LFS. The empirical results show that the estimator provides substantial precision gains for estimates of change over time compared to the direct estimator. By not assuming a smooth trend for the underlying population values, the proposed method also offers a robust alternative to state-space models, particularly during periods of high economic volatility.
Mortality patterns in closely related subpopulations often exhibit similarities, suggesting that mortality forecasts for individual subpopulations could be enhanced by borrowing strength from larger related groups. In this article, we focus on multipopulation mortality modelling, in which the data form a multiway mortality array comprising mortality rates of populations disaggregated by various sociodemographic attributes, such as gender, age, smoking/nonsmoking, and country or region. Each dimension of the array corresponds to one attribute. First, we propose a tensor autoregressive (TAR) model to efficiently model and forecast such multiway mortality arrays. Unlike existing vector autoregressive models, the TAR model preserves the multiway structure and more effectively incorporates patterns across groups and attributes. The proposed low-rank TAR models capture underlying low-dimensional tensor dynamics by utilizing the CANDECOMP/PARAFAC (CP) and Tucker decompositions. This yields a significant dimensionality reduction and a flexible model transformation. The CP decomposition addresses the overparameterization problem, while the Tucker decomposition enables demographic interpretations across multiple attributes. Finally, an empirical analysis using three-way mortality data (age, population, and gender) demonstrates that the proposed models achieve strong in-sample fit and satisfactory out-of-sample forecasting performance. Furthermore, we demonstrate that the proposed low-rank TAR models ensure coherence and nondivergence.
Opportunities for statisticians in the age of blockchains are now beginning to unfold, with researchers pursuing bold solutions to heretofore unsolved problems with this technology. While the notion of an 'immutable record' is not novel, given the advent of modern computing, new methods to transact and maintain robust system of records arise out of blockchains. Analytical challenges from blockchain ecosystems can substantially benefit from statistical methodologies. Inferential methods can be deployed to gain insights about aberrations in blockchains, thus making this technology more resilient to fraud or other criminal interference. This special issue of Statistics in Society underscores the critical need for statistical expertise in blockchains, particularly as it concerns how our global society embraces this new and accessible method of enabling transactions. Many 'use cases' for blockchains involve cryptocurrencies, but other applications also employ blockchain technologies. The nature of real estate transactions, the maintenance of health records, or even the manner in which works of art change hands could all leverage efficiencies afforded by blockchains. The hope of this special issue is to encourage more statisticians to get involved in tackling important problems in blockchain data analysis to advance how society can benefit from this technology.
Understanding how relationships among global financial markets change over time is crucial for effective risk management, portfolio diversification, and risk assessment. Motivated by the recent episodes of market turmoil, in this paper we analyse daily returns for 21 major indices covering cryptocurrencies, equities, energy commodities, and exchange rates from 2017 to 2025 by developing a novel time-varying graphical model for detecting the evolution of conditional dependency structures in financial markets. To identify temporal shifts in market regimes and account for the characteristics of returns, we exploit nonparanormal distributions with state-dependent parameters that evolve according to a latent finite-state semi-Markov chain. Our methodology results in regime-specific graphs that capture dynamic network connectivity while preserving the tractability of Gaussian methods for identifying conditional dependencies. Model estimation is carried out with a penalized Expectation-Maximization algorithm to induce sparsity in the state-specific precision matrices, without parametric assumptions about the states' sojourn distributions.
This article leverages Wasserstein Propagation in Social Network to propose a novel distributional framework for the inference of cyber risk across interconnected economic systems. Cyber attacks represent an increasing threat to global security and economic stability, making the assessment of cyber risk particularly challenging, especially in countries with limited data availability. Using a comprehensive dataset on worldwide cyber attacks along with a set of macroeconomic indicators, we estimate risk profiles for all countries, including those with sparse information. Our approach reveals critical interdependencies and vulnerabilities among countries, highlighting the interconnected nature of cyber risks. The findings demonstrate the value of social network analysis in modelling cyber risk and the related uncertainty within cyberspace. Furthermore, the results offer actionable insights to strengthen global cybersecurity policies and improve resilience against cyber threats.
The paper develops a dynamic panel stochastic frontier model that incorporates firms' intertemporal decision behaviour and short-run stagnant adjustments to the production process. Its dynamic specification recognizes short-run output adjustment costs, where final output may be only partially adjusted to the optimum level. In nesting previous panel stochastic frontier models, our new approach delivers a flexible framework that accommodates heterogeneous technologies and latent time-varying inefficiency effects. In addition, our model handles endogeneity issues related to flexible inputs. Model inference is based on a Bayesian framework, where Markov Chain Monte Carlo (MCMC) techniques are utilized. Through extensive simulations, we demonstrate the robustness of the model in small and moderate samples. Last, we present our model in an empirical example, analysing publicly listed UK companies operating in the manufacturing and construction sector over the period 2004-2022. A general finding is that most firms exhibit stagnant production processes, with the half-life for adjusting supply to be as high as 6 quarters. The estimated average technical efficiency is 89%. Our findings underscore the importance of accounting for dynamic frictions and heterogeneity when evaluating firm performance and designing productivity-enhancing policies.
The paper introduces a novel framework for small area estimation based on spatio-temporal M-quantile regression. The proposed approach extends the Geographically Weighted Regression by incorporating both spatial and temporal weighting schemes, and integrates them with the M-quantile modelling to effectively capture local distributional features across space and time. The resulting predictors are specifically designed for out-of-sample prediction in small domains and are accompanied by analytical estimators of their mean squared error. The methodology is evaluated through extensive simulation studies, demonstrating strong robustness to spatio-temporal dependence and the presence of outliers at both unit and area levels. An application to county-level air quality data in the United States (2016-2023) highlights the predictive performance and practical relevance of the proposed methods.
In meta-analysis, the conventional random-effects model (REM) is commonly used when heterogeneity is thought probable a priori. However, there are no standard tests of the adequacy of this model. To address this issue, we propose a new third-order Cochran statistic. Our new J index, and associated test, measures departures from the REM, analogous to the way that Cochran's heterogeneity statistic and the I2 index are used to quantify departures from the fixed effect model. We also introduce a conceptually simple plot that is useful in showing the presence of unusual features of outcome data, that may cast doubt on conventional modelling approaches. A measure of the diversity of study sizes is also introduced. The performance of the proposed methodology is explored in a simulation study and illustrated by analysing some well-travelled datasets. Our proposals can be used as descriptive statistics, to convey the adequacy of conventional statistical methodologies, or help determine model choice.
Scrutiny of referees in football is ever-increasing and likely at an all-time high. Is this fair? How consistent are modern day referees? In this work, the number of yellow cards given to the home and away teams over four seasons of data from 2018 to 2022 in each of the 'Big 5' European leagues in men's football is modelled. This allows us to make an assessment of the heterogeneity amongst both referees, primarily, and teams, secondarily. The underdispersed nature of the data from the small counts, and likely dependence of the cards issued within a game, leads to a bivariate mean-parameterized Conway-Maxwell-Poisson copula model to analyse the data. We also model home advantage, and examine whether this was diluted during COVID-19, and league effects, allowing for an assessment of how consistent referees are across leagues and which league has the most heterogeneous 'men in the middle'. We find that teams and referees are, perhaps surprisingly, similarly heterogeneous and that referees in Serie A appear to be the least variable. The underlying model has wider use in other fields when the researcher is analysing small multivariate counts with possible bidispersion.
Recent failures of major cryptocurrency exchanges and widespread crypto scams highlight the urgent need for stronger regulatory oversight. Meanwhile, tensor-based frameworks are increasingly used to model complex interdependencies across crypto-trading platforms. A key challenge is reliably estimating precision matrices for tensor-valued processes, which can uncover hidden connections and risk propagation between exchanges, including undisclosed trading relationships and cross-platform exposures and systemic financial risks overall. However, most existing tensor precision matrix estimations assume independence, failing to account for the temporal dependencies that characterize crypto-markets. This limitation proved particularly damaging in events such as the FTX collapse, where interconnected trading activities and hidden leverage generated rapid systemic risk across platforms. To address this gap, we propose WeDTLasso, a novel graphical model for estimating precision matrices of temporally dependent tensor-valued data, able to capture how risks propagate through crypto-exchange networks over time. We derive theoretical guarantees in the form of non-asymptotic near-oracle error bounds and demonstrate WeDTLasso's value for regulatory oversight by analysing cross-platform risk transmission in cryptocurrency markets. Our results show that accounting for temporal dependencies is essential for identifying systemic risks and market manipulation and can deliver up to 50% improvements in Sharpe ratios for portfolio construction compared to existing methods.