
Patients affected by Parkinson’s disease (PD) usually experience a range of symptoms, including movement disorder, sleep disturbance, and brain structural changes, and may start treatment at certain time points throughout clinical course. Current research focusing on a single modality (e.g., movement disorder) fails to display the full picture of the disease progression and its connection with timing of symptomatic therapy. In this paper, we integrate longitudinal multi-type measurements from multiple modalities (e.g., binary/ordinal clinical measures, continuous neuroimaging biomarkers) and uncover their relationship with the initiation of PD treatment using a latent Gaussian process joint model. The dependence structure between observed multi-type biomarkers is characterized by several underlying unobserved Gaussian processes through a generalized factor model. Instead of assuming a certain parametric form a priori, the Gaussian processes, with their philosophy to let data speak for themselves, are expected to capture the possibly nonlinear and complex progression profiles of PD-related measurements. The obtained Gaussian processes are also incorporated into a varying coefficient Cox model so as to jointly monitor the longitudinal PD-related measurements and their time-varying effects on time to initiation of PD treatment. The application of our proposed method to the Parkinson’s Progression Markers Initiative dataset yields a comprehensive understanding of PD progression by synthesizing information from clinical, biological, and neuroimaging perspectives.
Correlated proportions appear in many real-world applications and present a unique challenge in finding an appropriate probabilistic model and performing inference due to their constrained nature. The bivariate beta is a natural extension of the well-known beta distribution to the space of correlated quantities on [0, 1]^2 , but its construction is not unique. Over the years, many bivariate beta distributions have been proposed, ranging from three to eight or more parameters, and for which the joint density and distribution moments vary in terms of mathematical tractability. As a consequence, the literature on statistical inference for correlated proportion data is scattered. In this paper, we address this gap by thoroughly studying two rather different bivariate models from their construction to inference and diagnostics. We investigate two bivariate constructions namely the four-parameter bivariate beta proposed by Olkin and Trikalinos (2015) (OT) and a bivariate logistic normal distribution, which has one extra parameter. We provide a thorough account of inference, exploring both classical (frequentist) and Bayesian approaches to estimation, utilising the method of moments and latent variable augmentation coupled with Hamiltonian Monte Carlo, respectively. The elicitation of a bivariate beta distribution as a prior is also discussed. Further, we develop diagnostics for checking model fit and adequacy and test their performance with Monte Carlo experiments under well-specified and misspecified data-generating settings. We illustrate these methods using data on childhood vaccination coverage and on sensitivity/specificity from COVID-19 tests, comparing the OT bivariate beta model with a bivariate logit-normal model.
To overcome computational difficulties in Bayesian nonparametric regression with massive data, this paper develops a new Bayesian nonparametric technique, called Bayesian reconstruction regression. Bayesian reconstruction regression combines Bayes’ rule and the reconstruction parameterization approach, which can accurately approximate a multivariate function with relatively few basis functions. It is shown that, with the number of basis functions being the sample size, Bayesian reconstruction regression can be approximately equivalent to Bayesian Gaussian process regression. We discuss the implementation of Bayesian reconstruction regression within both the empirical and full Bayesian frameworks, with informative and non-informative priors. Numerical results including an analysis of a metro computer simulation indicate that the proposed methods perform better than popular simplifications of Bayesian Gaussian process regression.
Research on dissimilarity indexes has received limited attention in the Small Area Estimation literature. Motivated by the high variability of direct estimators, this paper proposes a new statistical methodology for estimating Duncan Segregation Indexes in small areas. Empirical best, simplified, and plug-in predictors are developed based on a unit-level logit mixed model. The mean squared error of the predictors is estimated using parametric bootstrap procedures, with and without bias correction. Simulation experiments are conducted to evaluate the additional variability arising from the estimation of auxiliary information and to assess the performance of the proposed predictors and their mean squared error estimators. The results are compared with those obtained from direct estimators and from models fitted to area-level data. The methodology is applied to the Spanish Labour Force Survey to assess sex segregation in the workplace.
In repeated measure designs with multiple groups, the primary purpose is to compare different groups in various aspects. For several reasons, the number of repetitions can depend on the group, making the usage of existing approaches impossible. We want to develop an approach which can be used not only for a possibly increasing number of groups but also for group-depending dimension d_i , which is allowed to go to infinity. This is a unique high-dimensional asymptotic framework impressing through its variety and doing without usual conditions on the relation between sample size and dimension. New and innovative estimators are developed to find a statistical test in this setting, and the asymptotic distribution of a quadratic-form-based test statistic is investigated. Finally, an extensive simulation study is done to investigate the role of the single group’s dimensions.
In this paper, the rates of uniformly strong consistency for some estimators of the cumulative hazard function, the cumulative distribution function, the truncation probability, and the hazard rate are investigated under random left truncated and negatively dependent data. The results improve and extend the corresponding ones in the literature to general negatively dependent and random left truncated data. Some simulation studies and a real data example are also provided to support the theoretical results.
Scheike et al. propose a dynamic regression augmentation method to improve estimation of treatment effects on recurrent events using auxiliary covariates in randomized clinical trials. This discussion considers the causal interpretation of such effects within the potential outcomes framework. Marginal effects for the full randomized population are distinguished from effects among survivors, and conditioning on survival is noted to potentially introduce posttreatment selection when survival is influenced by treatment. These considerations complement the authors' approach by clarifying the interpretation of survivor-based analyses when survival may be treatment-dependent.
We treat the problem of measuring and testing for the equality of multivariate distributions. We construct a new measure by combining the concept of projection averaging with the theory of optimal transport. Our measure possesses the necessary and intuitive property as a metric index. Briefly, it is zero almost surely if and only if two multivariate random variables have the same distribution. The sample counterparts of this measure can be expressed elegantly and enjoy desirable theoretical properties. Based on the corresponding U- and V-statistics, we propose two nonparametric tests for assessing the equality of two multivariate probability distributions. The proposed tests are free of underlying population distributions and are consistent against all fixed alternatives. Moreover, our procedures avoid the need to estimate any nuisance parameters or impose moment assumptions, which broadens the scope of the proposed testing applications and reduces the computational burden. We also demonstrate the favorable finite-sample performance of the proposed tests through extensive simulations and real data examples, particularly under certain models involving location differences.
In this study, we first propose the functional quadratic spatial autoregressive model, which can effectively capture spatial dependencies, linear relationships and nonlinear interactions in functional data. Then, we develop a generalized statistical inference framework for this model, where the functional component is treated by principal component analysis, followed by parameter estimation employing the generalized method of moments. Subsequently, we construct two residual-based test statistics to assess the model’s goodness of fit. Under some regularity conditions, we derive the asymptotic properties, encompassing the asymptotic normality of estimators for the finite parametric vector, the optimal convergence rate for nonparametric functions, and the asymptotic distributions of the proposed test statistics under null hypothesis and local or global alternative hypothesis. To determine the critical values for model checking statistics, we introduce a wild bootstrap procedure, and the asymptotic validity of this bootstrap-based testing approach is also discussed. The finite-sample performance of our methodology is evaluated through Monte Carlo simulations, and its practical usefulness is illustrated through two real-data applications to Spanish meteorological data and economic growth data. The results substantiate the efficacy of our approach in the real-world scenarios.
This paper studies hypothesis testing for structured scale matrices in the multivariate t distribution. We develop likelihood ratio and Rao score tests for general structural constraints on the scale matrix and derive the Wald test for hypotheses in which the scale matrix is restricted to a quadratic subspace. The latter class includes, in particular, diagonality, sphericity, compound symmetry, circular Toeplitz structures, and their block extensions with structured subblocks. Since the construction of the test statistics requires maximum likelihood estimation, we derive the likelihood equations under quadratic subspace restrictions. Although likelihood ratio and score tests have been investigated for covariance structure in the multivariate normal model, the Wald test has not been developed in that setting. Because the multivariate normal distribution arises as a limiting case of the multivariate t model, our results also provide a new contribution to the normal framework. Finite sample performance is examined via simulation, and a real data example illustrates the proposed methodology.
We in this paper propose a robust estimation technique for a first-order autoregressive process characterized by the root ρ _n=1+c/k_n , employing a kernel mode-based objective function construed on the mode value. Compared to traditional least squares or maximum likelihood approaches, the newly developed method exhibits enhanced robustness against outliers and non-normal errors. We suggest a computationally efficient mode expectation-maximization algorithm leveraging a Gaussian kernel to numerically estimate the coefficient. Under mild assumptions, we derive the asymptotic distributions of the resulting kernel mode-based estimator, assuming that k_n increases to infinity at a slower rate than n. Specifically, for c<0 , we establish a convergence rate of √(nk_n) with a normal limit distribution, while for c>0 , the convergence rate is k_n ρ ^n_n with a Cauchy limit distribution. Monte Carlo simulations are presented to illustrate the favorable finite sample performance of the proposed estimation procedure. Furthermore, we extend these results to the general autoregressive process with a coefficient satisfying n |1-ρ _n |→∞ under weaker initial conditions. The convergence rates are demonstrated to be [n(1-ρ ^2_n)^-1]^1/2 and ρ ^n_n/(ρ ^2_n-1) for nearly stationary and mildly explosive cases, respectively.
Detecting the absence of the mediation effect is a major focus of mediation analysis. When treatment does not influence mediators that do not affect the outcome either, testing the natural indirect effect is an interesting issue, since the asymptotic distributions of existing test statistics vary under different sub-null hypotheses. This paper introduces a novel statistical inference procedure tailored for high-dimensional mediation structures to address the issue of test conservativeness under some nontrivial sub-null hypotheses. We first suggest a procedure using a partial penalized least squares estimation and compute the inner product of the treatment-mediator and the mediator-outcome coefficients. Based on the product, we develop a Wald-type test to handle the case where the mediators affect the outcome. When the mediators do not affect the outcome, the Wald-type test statistic fails to maintain the significance level. We then construct another test for the significance level maintenance. The final test is an adaptive-to-sub-null hybrid of the two tests, which can flexibly accommodate different sub-null hypotheses and ensures that the limiting null distributions converge to a common Chi-square distribution uniformly across all sub-null hypotheses. Numerical studies are conducted to assess the finite-sample performances of the proposed test and make comparisons with existing methodologies.