
ABSTRACT Statisticians commonly find themselves working in isolation, with some managed exclusively by nonstatisticians. Mentoring is critical to help statisticians build their careers and network. We created a national mentoring program to support statisticians across Australia through the national professional statistical organisation: the Statistical Society of Australia. This national mentoring program is now in its sixth year, including the pilot, and is well established in the Society. In this article, we describe why and how this program was created, outlining some of our challenges and how the program has evolved. We provide advice for others interested in creating a mentoring program for statisticians.
ABSTRACT Autocorrelated bivariate count data frequently arise in criminal, environmental and financial studies, where capturing both serial dependence and cross‐series interaction is essential for statistical modelling and inference. In many applications, such dynamics are further influenced by exogenous covariates such as policy interventions or environmental factors, leading to time‐varying dependence structures that are not adequately captured by standard models. Existing bivariate integer‐valued autoregressive (BINAR) models mainly rely on constant or observation‐driven coefficients and rarely incorporate covariate information, which restricts their ability to represent evolving dependence in multivariate count processes. To address this limitation, we propose a covariate‐driven doubly stochastic bivariate integer‐valued autoregressive process, in which the thinning mechanism evolves jointly with past observations and exogenous covariates. This formulation extends the classical BINAR framework by allowing the dependence structure to vary dynamically under both internal and external driving mechanisms. The basic statistical properties of the proposed process are derived, and two estimation methods are developed, including an EM‐based algorithm. Monte Carlo simulations and a real data application are conducted to assess finite‐sample performance and robustness under different settings.
ABSTRACT This paper studies variance estimation in difference‐in‐differences (DID) designs under covariate matching. While matching is widely used to improve covariate balance, it induces a dependence structure that complicates variance estimation. In particular, covariate‐based matching generates positive correlation within matched pairs, which reduces the variance of the DID estimator. Standard variance estimators fail to account for this design‐induced covariance and therefore yield systematically conservative standard errors. We characterize the asymptotic variance implied by the matched design and propose a projection‐based variance estimator to remove variation attributable to the matching covariates. Simulation results show that the proposed estimator achieves accurate coverage, whereas standard methods substantially overestimate uncertainty. An empirical application illustrates the practical implications for inference.
ABSTRACT Basis precision matrix plays fundamental roles in the study of compositional data. Constrained and unconstrained optimization methods are usually used to estimate basis precision matrices. Zhang and He (2019) discuss that problem by an unconstrained optimization estimator. However, the authors do not give any convergence rates. Zhang et al. (2025) define their constrained optimization estimator and provide a convergence rate. In this paper, we extend Zhang and He's work. It turns out that our unconstrained ‐regularized estimator attains the same convergence rate as the one of Zhang et al. Moreover, our estimator outperforms Zhang et al.'s for numerical experiments.
Sufficient dimension reduction (SDR) in survival regression aims to identify low-dimensional structures that preserve the relationship between survival time and predictors. Classical SDR methods, such as sliced inverse regression (SIR), rely on strong assumptions such as linearity, constant variance and coverage conditions, which are often violated in practical data. To address this issue, we propose an SDR framework for survival regression under non-elliptical predictor distributions. The proposed fused Cox proportional hazards with transformation-mean methods estimate the survival informative predictor subspace by integrating three components: the Cox proportional hazards direction, single response transformations and kernel fusion over multiple slicing schemes. Numerical studies and real data analysis show that our proposed methods outperform existing SIR-based approaches, particularly when the linearity condition fails. They also remain robust across various sample sizes, predictor dimensions and censoring levels. Overall, the proposed fused methods offer a reliable tool for dimension reduction in survival regression.
ABSTRACT Prior work established upper bounds for spiked covariance matrices under Gaussian and sub‐Gaussian distributions in the Schatten‐q norm, a specific class of unitarily invariant norms. In contrast, our work investigates differentially private estimation for general covariance matrices under sub‐Gaussian distributions, and we analyse the estimation error under any unitarily invariant norm. Furthermore, we establish the conditions for the invertibility of the covariance matrix and provide an upper bound on its spectral norm under differential privacy. To demonstrate the optimality of the upper bound, we derive provide a lower bound estimation. We also use numerical experiments to illustrate the effectiveness and tightness of the upper bound estimation.
Individuals who are connected to each other in a network are more likely to exhibit similar behaviours. This implies that the formation of a dyadic link could be inferred from common factors shared by two nodes. In this study, we propose a varying coefficient dyadic interaction model with a random effect to understand user relationships in social networks. Specifically, we assume that the generation of a dyadic link is driven by an observed common factor. Owing to the existence of heterogeneity, we assume that the influence of the common factor is a linear combination of several moderate variables. Obtaining a maximum likelihood estimator in this situation is infeasible. To address this issue, we propose a pseudo-maximum likelihood estimator, which is feasible and can substantially alleviate the computational cost. Although the proposed estimator is not as accurate as a maximum likelihood estimator, it can be used when heterogeneity exists or the sample size is extremely large. We then conduct numerical studies to assess the finite sample performance of the proposed method. Finally, to empirically examine the usefulness and practicality of the proposed method, we apply the model to two different real-world datasets. The empirical results demonstrate the influence of common factors on dyadic link generation.
Burg entropy, developed by John P. Burg in maximum entropy spectrum analysis, has inspired significant advancements in information theory. Expanding upon the cumulative residual Burg entropy, we present the concept of cumulative past Burg entropy and establish four characterisations of symmetric distributions. A characterisation result is used to provide a goodness-of-fit test for symmetry, demonstrating robust power and competitiveness with established methods available in the literature. We additionally offer a comprehensive Burg entropy characterisation of the uniform distribution and formulate a general theory that delineates distributions like the truncated exponential and truncated normal under appropriate conditions. We demonstrate that the dual of the modified Burg entropy aligns with positively scaled extropy and establish that, within the skew-normal family, it reaches its maximum at the standard normal distribution.
The COVID-19 pandemic has underscored the critical need for accurate mortality forecasts to shape effective public-health strategies. Our study introduces an advanced prediction framework that integrates online change point detection with a novel training scheme to enhance the forecasting of COVID-19 mortality. Empirical analyses across national and state datasets in the United States demonstrate that our methodology not only boosts prediction accuracy but also significantly reduces model training time. The proposed hybrid models, which combine the strengths of various model candidates tailored to specific intervals identified by change points, consistently outperform baseline models. They achieve the lowest sMAPE scores and realize execution times that are over 99% faster than traditional models. This work highlights our approach's adaptability to the dynamic nature of the pandemic, marking a significant advancement in real-time infectious disease monitoring and supporting public-health decision-making.
Subgroup analysis is important in practice because real-world data typically come from heterogeneous populations, where meaningful patterns can differ substantially across subpopulations. Correctly identifying these subgroups can improve prediction accuracy, prevent biased or misleading conclusions, and support more effective, targeted decision-making. While most existing subgroup analysis methods are developed for complete data, in this paper we propose a novel and robust approach for censored data under heterogeneous accelerated failure time (AFT) models. Specifically, we combine inverse probability weighting, M-estimation, and concave pairwise fusion penalization to simultaneously identify subgroups and estimate covariate effects for heterogeneous censored data, without requiring prior knowledge of individual subgroup memberships. We further develop an efficient RISA-ADMM algorithm to implement the method and establish its convergence. Furthermore, we derive the theoretical properties of the proposed estimators under mild regularity conditions. Extensive simulations and an application to the German credit dataset demonstrate the robustness and effectiveness of our approach.
This paper is concerned with the testing on regression coefficients of high-dimensional partially linear regression models. An improved power enhancement test method is proposed based on the -statistic type test via random integration technique, and the asymptotic distribution of the proposed test statistic is investigated under a more general distribution assumption of the explanatory random vector. It is shown that the proposed test method is more universal and it outperforms the existing method in terms of power under some situations by the asymptotic relative efficiency. The finite-sample performance of the proposed test method is evaluated through numerical studies, which demonstrate that the proposed test method is more powerful than the existing method in a wide range of alternative hypothesis settings.
Crime studies have garnered significant interest in recent times due to the rising rate of criminal activities. Understanding the spatial and temporal clustering of crime events such as burglary, robbery and theft has become essential. In this study, we propose a non-parametric self-exciting point process model with Gaussian mixture kernels (GMKs), along with a truncation technique for the intensity function. The proposed GMK can give accurate estimation while simultaneously reducing computational time. We first evaluate the performance of the proposed model using simulated data. Subsequently, we demonstrate its applicability in analysing the clustering patterns of theft in Philadelphia and robbery in Chicago. Residual analysis is employed as a diagnostic tool to assess the effectiveness of the model. To further evaluate its predictive capability, we generate 3-day forecasts, illustrating the potential of the self-exciting point process in capturing the clustering behaviour of these crime types.
Conformal prediction is a powerful framework for constructing distribution-free prediction regions with guaranteed coverage. Existing conformal prediction methods are often developed in a centralized setting, assuming that all data are accessible at a single location. However, data are often distributed across multiple machines due to communication constraints or data access restrictions, making conventional conformal prediction methods infeasible. In this paper, we propose a distributed conformal prediction method that leverages random forests without centralizing raw data. To construct prediction sets, we design a communication-efficient distributed quantile estimation algorithm using binary search. Additionally, we investigate a distributed conformalized quantile regression approach that adapts the lengths of prediction sets to heteroscedasticity. To improve predictive efficiency, especially under complex error distributions such as skewed or multimodal structures, we further introduce a distributed conformal prediction method based on highest density sets, yielding narrower and more informative prediction regions. We establish upper and lower bounds on the coverage of the proposed methods. Numerical simulations and an illustrative airline data example demonstrate the effectiveness of our methods.
This paper is a follow-up to an earlier study, which identified a new class of generalized Bayes estimators with a particularly simple form for estimating a normal variance under entropy loss. Although the previous work established the Bayesianity of these estimators, it did not provide a closed-form result for their minimaxity. In this paper, we revisit the problem and establish a definitive closed-form minimaxity result for this class of simple Bayes estimators.
Basis precision matrix plays fundamental roles in the study of compositional data. Constrained and unconstrained optimization methods are usually used to estimate basis precision matrices. Zhang and He (2019) discuss that problem by an unconstrained optimization estimator. However, the authors do not give any convergence rates. Zhang et al. (2025) define their constrained optimization estimator and provide a convergence rate. In this paper, we extend Zhang and He's work. It turns out that our unconstrained -regularized estimator attains the same convergence rate as the one of Zhang et al. Moreover, our estimator outperforms Zhang et al.'s for numerical experiments.
In computer experiments, space-filling designs with favourable low-dimensional projection properties are crucial for the efficient exploration of the design space, especially when many factors are involved but only a few are active. Among existing space-filling designs, uniform projection designs stand out for their guaranteed uniformity across all two-dimensional projections. This paper presents a series of novel and efficient algebraic constructions for uniform projection designs. By using orthogonal arrays, permuted good lattice point designs and -equidistant designs, we generate a rich class of uniform projection designs with flexible sizes. Additionally, we construct column-orthogonal designs that are also nearly optimal under the uniform projection criterion. Numerical comparisons demonstrate the superiority of our proposed constructions compared to existing methods.
Abscission, the shedding of organismal parts, depends on physiological events whose optimal timing is crucial for species survival. Environmental variations impact species development, particularly abscission processes across multiple developmental stages, identifying which environmental factors modulate abscission and when is essential in the context of climate change. Treating environmental variables as time series-groups of temporally correlated variables raises statistical challenges for selecting relevant groups (environmental variables) and their correlated components (time periods). We address these objectives by introducing the Bayesian fused and fusion priors through a general parameterization. We highlight a trade-off between priors used on differences versus coefficients, demonstrating that horseshoe-type priors on both differences and coefficients, with appropriate parameterizations, achieve effective selection, estimation and algorithmic stability regardless of group number or size. Our study focuses on fruit abscission in oil palms, which affects bunch harvest timing. Abscission disruption can impact oil yield and quality, consequently affecting economic returns. This application, based on experimental data from Benin, illustrates how our proposed priors successfully select both environmental variables and developmental stages involved in bunch harvest timing.
In practical applications, time series data are often affected by overreporting. To better capture these data features, this paper proposes an integer-valued process with overreporting. We discuss the bias of naive estimation that ignores overreporting and provides the corrected estimating equations. Besides, two estimation methodologies are explored that operate without requiring the error probability information. The Yule-Walker estimation is employed to estimate the parameters of the integer-valued autoregressive process under the influence of overreporting. In particular, to address the unobservability of the true process caused by spurious counts, a pseudo-maximum likelihood method based on the expectation-maximization algorithm and Gibbs sampling is presented. The asymptotic properties of the proposed estimators are studied. Simulation results demonstrate satisfactory estimation performance, and the model is further applied to a real dataset for validation.
This paper investigates the parameter estimation problem for panel data with a multilayer network structure. Considering the heterogeneous behaviours or attributes exhibited by individuals, we propose a multilayer network autoregressive model with group structure. We allow individuals to have a latent group structure, whereby individuals within the same group share common slope. We assume that time-invariant fixed effects are individually heterogeneous, which endows the model with a wide range of applications. To identify the latent group structure, we use the initial MLE for clustering purposes. Our theoretical analysis demonstrates that this two-stage least squares (2SLS) estimator possesses desirable asymptotic normality properties. We conduct a detailed analysis of the factors influencing estimation errors and employ a jackknife method to reduce the bias of the momentum effect. Extensive simulation experiments and real data analyses validate the effectiveness of our proposed method.
Double-censored data can occur in many areas like clinical trials, epidemiology, cancer research and survival analysis. Such data usual refers failure time of interest can be observed only if it is within an observation interval or a window, encompassing instances of left censoring (events occurring before observation begins) and right censoring (events occurring after observation ends). Many methods have been proposed for their analyses and most of the existing methods assume that or apply only to the situation where the censoring is independent. However, it is well-known this may not be true or one may face informative censoring in many cases. To address this, we proposed a generalized accelerated hazards frailty model that allows for dependent censoring among other advantages and includes the proportional hazards, accelerated failure time and traditional accelerated hazards frailty models as special cases. For inference, we propose a joint model-based sieve maximum likelihood approach and develop an EM-based algorithm for its implementation. Also the profile approach is adapted for variance estimation and the proposed estimator of regression parameters is shown to be consistent and asymptotically normal. Furthermore an extensive simulation study is performed and suggests that the proposed method works well in practical situations and it is applied to a set of real data from an AIDS clinical trial that motivated this study.