The Bermuda Department of Statistics, subject to the Bermuda Statistics Act of 2002 as further amended, reports to the head of Cabinet Office. It was created mainly for the collection, compilation, analysis and publication of statistical information and so facilitate the development of a statistical system for Bermuda which is also a component of the regional statistical systems of CARICOM, together with other member countries of this regional institution..
We study mean change point testing problems for high-dimensional data, with exponentially- or polynomially-decaying tails. In each case, depending on the ℓ0-norm of the mean change vector, we separately consider dense and sparse regimes. We characterise the boundary between the dense and sparse regimes under the above two tail conditions for the first time in the change point literature and propose novel testing procedures that attain optimal rates in each of the four regimes up to a poly-iterated logarithmic factor. To be specific, when the error distributions possess exponentially-decaying tails, a near-optimal CUSUM-type statistic is considered. As for polynomially-decaying tails, admitting bounded α-th moments for some α ≥ 4, we introduce a median-of-means-type test statistic that achieves a near-optimal testing rate in both dense and sparse regimes. Our investigation in the even more challenging case of 2 ≤ α < 4, unveils a new phenomenon that the minimax testing rate has no sparse regime, i.e. testing sparse changes is information-theoretically as hard as testing dense changes. Finally, we consider various extensions where we also obtain near-optimal performances, including testing against multiple change points, allowing temporal dependence as well as fewer than two finite moments in the data generating mechanisms. We also show how sub-Gaussian rates can be achieved when an additional minimal spacing condition is imposed under the alternative hypothesis.
We propose an autoregressive framework for modelling dynamic networks with dependent edges. It encompasses the models which accommodate, for example, transitivity, density-dependent and other stylized features often observed in real network data. By assuming the edges of network at each time are independent conditionally on their lagged values, the models, which exhibit a close connection with temporal ERGMs, facilitate both simulation and the maximum likelihood estimation in the straightforward manner. Due to the possible large number of parameters in the models, the initial MLEs may suffer from slow convergence rates. An improved estimator for each component parameter is proposed based on an iteration based on the projection which mitigates the impact of the other parameters (Chang et al., 2021, 2023). Based on a martingale difference structure, the asymptotic distribution of the improved estimator is derived without the stationarity assumption. The limiting distribution is not normal in general, and it reduces to normal when the underlying process satisfies some mixing conditions. Illustration with a transitivity model was carried out in both simulation and a real network data set.
Multivariate histograms are difficult to construct due to the curse of dimensionality. Motivated by $k$-d trees in computer science, we show how to construct an efficient data-adaptive partition of Euclidean space that possesses the following two properties: With high confidence the distribution from which the data are generated is close to uniform on each rectangle of the partition; and despite the data-dependent construction we can give guaranteed finite sample simultaneous confidence intervals for the probabilities (and hence for the average densities) of each rectangle in the partition. This partition will automatically adapt to the sizes of the regions where the distribution is close to uniform. The methodology produces confidence intervals whose widths depend only on the probability content of the rectangles and not on the dimensionality of the space, thus avoiding the curse of dimensionality. Moreover, the widths essentially match the optimal widths in the univariate setting. The simultaneous validity of the confidence intervals allows to use this construction, which we call {\sl Beta-trees}, for various data-analytic purposes. We illustrate this by using Beta-trees for visualizing data and for multivariate mode-hunting.
Although the significance of big data and data science in predicting health outcomes and identifying causal factors is widely recognized, their application in health disparities research remains limited. Understanding health disparities in the visually impaired population requires examining their health behavior patterns and health literacy levels, which can longitudinally impact their health and well-being. In prior research, one of the authors of the study conducted an online survey with 2718 participants using 5 validated self-reported questionnaires, such as Health Promoting Lifestyle Profile II, Health Literacy Questionnaire, eHealth Literacy Questionnaire, General Self-Efficacy, and Center for Epidemiological Studies Depression. Analysis of the online survey data demonstrated that individuals with blindness exhibited significantly higher levels of health-promoting behaviors, health literacy, and eHealth literacy compared to those with moderate and severe low vision. In this study, the research team developed an R Shiny web application as a follow-up to the online survey to disseminate its findings reproducibly and interactively. The R Shiny web application is expected to facilitate reproducible as well as interactive data analysis and sharing more efficiently than traditional methods, such as appendices or Supplementary Materials, Supplemental Digital Content 2, http://links.lww.com/CIN/A463 in academic journals. Extending the research cycle with open datasets and reproducible data analysis can deepen our understanding of health disparities and foster greater collaboration among researchers with similar interests.
Non-cognitive, neuropsychiatric symptoms (NPS) are nearly universal in Alzheimer’s disease (AD), but investigation of their underlying biology is complicated by comparative medicine approaches that incompletely capture spontaneous disease, primarily using transgenic rodent models. The aged companion dog, which spontaneously develops an AD-like disease called canine cognitive dysfunction (CCD), may help fill this translational gap. Using data from the Dog Aging Project with > 10,000 aged dogs (> 8 years old), we identify numerous behaviors in dogs “at-risk” for and with CCD that mirror NPS in humans. Compared to dogs without CCD, our analysis shows that dogs with CCD are less physically active, exhibit fewer previously trained behaviors, demonstrate fewer motivated behaviors, have more daytime sleep, demonstrate more separation anxiety, have altered anxious responses to novelty, have changes in aggressive behaviors, and exhibit lower appetite. Using k-means clustering, we did not find evidence for behavioral sub-phenotypes. Overall, our analysis of a large number of aged dogs suggests clinically significant NPS are associated with CCD and that the companion dog may serve as an important comparative medicine approach to understand these debilitating symptoms across species.