We evaluated the integration of the Rapid Interactive screening Test for Autism in Toddlers (RITA-T) model in a community, comparing autism spectrum disorder (ASD) toddlers' demographic and socioeconomic characteristics. Of 394 ASD toddlers, 323 were screened with RITA-T. Those screened were from more deprived areas, traveled farther and were diagnosed earlier. The model improved the diagnosis of ASD in underserved areas.
MobileVirtual Reality (VR) and panoramic video streaming rely on interactive panoramic scene delivery to provide desirable user experiences. However, it is pretty challenging to support multiple users via the wireless network since a panoramic scene typically consumes 4 similar to 6x bandwidth compared with a regular video with the same resolution. Motivated by the fact that users only perceive the Field-of-View (FoV), we employ the autoregressive process to predict the user's motion and stream only part of the panoramic content. Notably, we analytically characterize the effect of the delivered portion on the user's successful viewing probability. Then, we formulate an optimization problem to maximize the application-level throughput (which measures the average rate for successful viewing the desired content instead of raw network throughput) while providing a regular service. In addition, we impose three main constraints to our problem: minimum required service rate, maximum allowable energy consumption, and wireless interference. We then propose a novel scheduling algorithm that incorporates users' successful viewing probabilities and asymptotically maximizes application-level throughput while providing service regularity guarantees. We conduct real-trace simulations to evaluate the efficiency of our algorithm.
Objective: To evaluate improved identification and the generalization of the RITA-T (Rapid interactive Screening Test for Autism in Toddlers) model through partnerships with Primary Care (PC), Early Intervention (EI), and Autism Diagnosticians. Methods: Over 3 years (2018-2021), 15 EI and 9 PC (MD and NP) centers participated in this project. We trained providers on the RITA-T and established screening models. We reviewed charts of all toddlers referred through this model and compared wait times, and diagnoses, to those evaluated through regular referral in a tertiary-based autism clinic. We also examined the RITA-T psychometrics. Results: 377 toddlers met our inclusion criteria. Wait time for diagnosis was an average of 2.8 months and led to further collaboration between community providers. RITA-T cut-off scores stayed consistent. Providers reported improved confidence and easy integration of this model. Conclusions: This model is generalizable and improves the Early Identification of ASD.
Large-scale multiple perturbation experiments have the potential to reveal a more detailed understanding of the molecular pathways that respond to genetic and environmental changes. A key question in these studies is which gene expression changes are important for the response to the perturbation. This problem is challenging because (i) the functional form of the nonlinear relationship between gene expression and the perturbation is unknown and (ii) identification of the most important genes is a high-dimensional variable selection problem. To deal with these challenges, we present here a method based on the model-X knockoffs framework and Deep Neural Networks to identify significant gene expression changes in multiple perturbation experiments. This approach makes no assumptions on the functional form of the dependence between the responses and the perturbations and it enjoys finite sample false discovery rate control for the selected set of important gene expression responses. We apply this approach to the Library of Integrated Network-Based Cellular Signature data sets which is a National Institutes of Health Common Fund program that catalogs how human cells globally respond to chemical, genetic and disease perturbations. We identified important genes whose expression is directly modulated in response to perturbation with anthracycline, vorinostat, trichostatin-a, geldanamycin and sirolimus. We compare the set of important genes that respond to these small molecules to identify co-responsive pathways. Identification of which genes respond to specific perturbation stressors can provide better understanding of the underlying mechanisms of disease and advance the identification of new drug targets.
We consider the problem of sequential multiple hypothesis testing with nontrivial data collection costs. This problem appears, for example, when conducting biological experiments to identify differentially expressed genes of a disease process. This work builds on the generalized α-investing framework which enables control of the marginal false discovery rate in a sequential testing setting. We make a theoretical analysis of the long term asymptotic behavior of α-wealth which motivates a consideration of sample size in the α-investing decision rule. Posing the testing process as a game with nature, we construct a decision rule that optimizes the expected α-wealth reward (ERO) and provides an optimal sample size for each test. Empirical results show that a cost-aware ERO decision rule correctly rejects more false null hypotheses than other methods for $n=1$ where n is the sample size. When the sample size is not fixed cost-aware ERO uses a prior on the null hypothesis to adaptively allocate of the sample budget to each test. We extend cost-aware ERO investing to finite-horizon testing which enables the decision rule to allocate samples in a non-myopic manner. Finally, empirical tests on real data sets from biological experiments show that cost-aware ERO balances the allocation of samples to an individual test against the allocation of samples across multiple tests.
ObjectiveCervical cancer, the fourth leading cancer diagnosed in women, has brought great attention to cervical cancer screening to eliminate cervical cancer. In this study, we analyzed two waves of provincially representative data from northeastern China's National Health Services Survey (NHSS) in 2013 and 2018, to investigate the temporal changes and socioeconomic inequalities in the cervical cancer screening rate in northeastern China.MethodsData from two waves (2013 and 2018) of the NHSS deployed in Jilin Province were analyzed. We included women aged 15–64 years old and considered the occurrence of any cervical screening in the past 12 months to measure the cervical cancer screening rate in correlation with the annual per-capita household income, educational attainment, health insurance, and other socioeconomic characteristics.ResultsA total of 11,616 women aged 15–64 years were eligible for inclusion. Among all participants, 7,069 participants (61.11%) were from rural areas. The rate of cervical cancer screening increased from 2013 to 2018 [odds ratio (OR): 1.06; 95% confidence interval (CI): 1.04–1.09, p < 0.001]. In total, the cervical cancer screening rate was higher among participants who lived in urban areas than rural areas (OR: 1.20; 95% CI: 1.03–1.39, p = 0.020). The rate was also higher among those with the highest household income per capita (OR: 1.30; 95% CI: 1.07–1.56, p = 0.007), with higher educational attainment (p < 0.001), and with health insurance (p < 0.05), respectively. The rate of cervical cancer screening was also significantly associated with parity (OR: 1.62; 95% CI: 1.23–2.41, p = 0.001) and marital status (OR: 1.45; 95% CI: 1.15–1.81, p = 0.001) but not ethnicity (OR: 1.41; 95% CI: 0.95–1.36, p = 0.164).ConclusionCervical cancer screening coverage improved from 2013 to 2018 in northeastern China but remains far below the target 70% screening rate proposed by the World Health Organization. Although rural-urban inequality disappeared over time, other socioeconomic inequalities remained.
Infantile spasms are a serious epilepsy syndrome with a poor prognosis. Electroencephalography (EEG) has been a key component in the prognosis and treatment of infantile spasms. This multi-center study protocol is developed to investigate interrater and intrarater agreement of an electroencephalographic grading scale—the Burden of Amplitudes and Epileptiform Discharges (BASED) score among electroencephalographers. Thirty children, aged 0–2 years, with infantile spasms who were hospitalized in the Chinese PLA General Hospital will be recruited into this study by stratified sampling. Seven electroencephalographers from different Class A tertiary hospitals will select a 5-min epoch with the most severe epileptiform discharge, score the EEG reports, and provide the basis for the scoring. The 420 (30 × 7 × 2) scoring results provided by electroencephalographers in two rounds can be analyzed statistically using weighted kappa (weighted $$\kappa$$ ) statstic, Fless’ kappa (Fless’ $$\kappa$$ ) statistic, and intraclass correlation coefficient (ICC) to calculate the interrater and intrarater agreement. We will recruit more electroencephalographers than were included in previous studies to assess the interrater and intrarater agreement in the selection of 5-min EEG epochs, the BASED scores, and the basis for scoring. If the BASED score has an adequate interrater and intrarater agreement, the score will have more significance for guiding the clinical management and for predicting the prognosis of patients with infantile spasms.
Population quantiles are important parameters in many applications. Enthusiasm for the development of effective statistical inference procedures for quantiles and their functions has been high for the past decade. In this article, we study inference methods for quantiles when multiple samples from linked populations are available. The research problems we consider have a wide range of applications. For example, to study the evolution of the economic status of a country, economists monitor changes in the quantiles of annual household incomes, based on multiple survey datasets collected annually. Even with multiple samples, a routine approach would estimate the quantiles of different populations separately. Such approaches ignore the fact that these populations are linked and share some intrinsic latent structure. Recently, many researchers have advocated the use of the density ratio model (DRM) to account for this latent structure and have developed more efficient procedures based on pooled data. The nonparametric empirical likelihood (EL) is subsequently employed. Interestingly, there has been no discussion in this context of the EL-based likelihood ratio test (ELRT) for population quantiles. We explore the use of the ELRT for hypotheses concerning quantiles and confidence regions under the DRM. We show that the ELRT statistic has a chi-square limiting distribution under the null hypothesis. Simulation experiments show that the chi-square distributions approximate the finite-sample distributions well and lead to accurate tests and confidence regions. The DRM helps to improve statistical efficiency. We also give a real-data example to illustrate the efficiency of the proposed method.
Multi-user panoramic video streaming demands 4~6× bandwidth of a regular video with the same resolution, which poses a significant challenge on the wireless scheduling design to achieve desired performance. On the other hand, recent studies reveal that one can effectively predict the user's Field-of-View (FoV) and thus simply deliver the corresponding portion instead of the entire scenes. Motivated by this important fact, we aim to employ autoregressive process for motion prediction and analytically characterize the user's successful viewing probability as a function of the delivered portion. Then, we consider the problem of wireless scheduling design with the goal of maximizing application-level throughput (i.e., average rate for successfully viewing the desired content) and service regularity performance (i.e., how often each user gets successful views) subject to the minimum required service rate and wireless interference constraints. As such, we incorporate users' successful viewing probabilities into our scheduling design and develop a scheduling algorithm that not only asymptotically achieves the optimal application-level throughput but also provides service regularity guarantees. Finally, we perform simulations to demonstrate the efficiency of our proposed algorithm using a real dataset of users' head motion.
Marine food webs are structured through a combination of top-down and bottom-up processes. In coral reef ecosystems, fish size is related to life-history characteristics and size-based indicators can represent the distribution and flow of energy through the food web. Thus, size spectra can be a useful tool for investigating the impacts of both fishing and habitat condition on the health and productivity of coral reef fisheries. In addition, coral reef fisheries are often data-limited and size spectra analysis can be a relatively cost-effective and simple method for assessing fish populations. Abundance size spectra are widely used and quantify the relationship between organism size and relative abundance. Previous studies that have investigated the impacts of fishing and habitat condition together on the size distribution of coral reef fishes, however, have aggregated all fishes regardless of taxonomic identity. This leads to a poor understanding of how fishes with different feeding strategies, body size-abundance relationships, or catchability might be influenced by top-down and bottom-up drivers. To address this gap, we quantified size spectra slopes of carnivorous and herbivorous coral reef fishes across three regions of Indonesia representing a gradient in fishing pressure and habitat conditions. We show that fishing pressure was the dominant driver of size spectra slopes such that they became steeper as fishing pressure increased, which was due to the removal of large-bodied fishes. When considering fish functional groups separately, however, carnivore size spectra slopes were more heavily impacted by fishing than herbivores. Also, structural complexity, which can mediate predator-prey interactions and provisioning of resources, was a relatively important driver of herbivore size spectra slopes such that slopes were shallower in more complex habitats. Our results show that size spectra slopes can be used as indicators of fishing pressure on coral reef fishes, but aggregating fish regardless of trophic identity or functional role overlooks differential impacts of fishing pressure and habitat condition on carnivore and herbivore size distributions.
Feature selection is central to contemporary high-dimensional data analysis. Group structure among features arises naturally in various scientific problems. Many methods have been proposed to incorporate the group structure information into feature selection. However, these methods are normally restricted to a linear regression setting. To relax the linear constraint, we design a new Deep Neural Network (DNN) architecture and integrating it with the recently proposed knockoff technique to perform nonlinear group-feature selection with controlled group-wise False Discovery Rate (gFDR). Experimental results on high-dimensional synthetic data demonstrate that our method achieves the highest power and accurate gFDR control compared with state-of-the-art methods. The performance of Deep-gKnock is especially superior in the following five situations: (1) nonlinearity relationship; (2) dimension p greater than sample size n; (3) high between-group correlation; (4) high within-group correlation; (5) large number of associated groups. And Deep-gKnock is also demonstrated to be robust to the misspecification of the feature distribution and the change of network architecture. Moreover, Deep-gKnock achieves scientifically meaningful group-feature selection results for cutting-edge real world datasets.
Sparse partial least squares (SPLS) is widely used in applied sciences as a method that performs dimension reduction and variable selection simultaneously in linear regression. Several implementations of SPLS have been derived, among which the SPLS proposed in Chun and Keles (J. R. Stat. Soc. Ser. B. Stat. Methodol. 72 (2010) 3-25) is very popular and highly cited. However, for all of these implementations, the theoretical properties of SPLS are largely unknown. In this paper, we propose a new version of SPLS, called the envelope-based SPLS, using a connection between envelope models and partial least squares (PLS). We establish the consistency, oracle property and asymptotic normality of the envelope-based SPLS estimator. The large-sample scenario and high-dimensional scenario are both considered. We also develop the envelope-based SPLS estimators under the context of generalized linear models, and discuss its theoretical properties including consistency, oracle property and asymptotic distribution. Numerical experiments and examples show that the envelope-based SPLS estimator has better variable selection and prediction performance over the SPLS estimator (J. R. Stat. Soc. Ser. B. Stat. Methodol. 72 (2010) 3-25).
The quantile regression method is a valuable complement to the classical mean regression, helping to ensure robust and comprehensive data analyses in a variety of applications. We propose a novel envelope quantile regression (EQR) method that adapts a nascent technique called enveloping to improve the efficiency of the standard quantile regression. The proposed method aims to identify the material and immaterial information in a quantile regression model, and then use only the material information for estimation. By excluding the immaterial information, the EQR method has the potential to substantially reduce estimation variability. Unlike existing envelope model approaches, which rely mainly on the likelihood framework, our proposed estimator is defined through a set of nonsmooth estimating equations. We facilitate the estimation via the generalized method of moments, and derive the asymptotic normality of the proposed estimator by applying empirical process techniques. Furthermore, we establish that the EQR is asymptotically more efficient than (or at least as asymptotically efficient as) the standard quantile regression estimators, without imposing stringent conditions. Hence, our work advances the envelope model theory to general distribution-free settings. We demonstrate the effectiveness of the proposed method via Monte Carlo simulations and real data examples.
The quantile regression method is a valuable complement to the classical mean regression, helping to ensure robust and comprehensive data analyses in a variety of applications. We propose a novel envelope quantile regression (EQR) method that adapts a nascent technique called enveloping to improve the efficiency of the standard quantile regression. The proposed method aims to identify the material and immaterial information in a quantile regression model, and then use only the material information for estimation. By excluding the immaterial information, the EQR method has the potential to substantially reduce estimation variability. Unlike existing envelope model approaches, which rely mainly on the likelihood framework, our proposed estimator is defined through a set of nonsmooth estimating equations. We facilitate the estimation via the generalized method of moments, and derive the asymptotic normality of the proposed estimator by applying empirical process techniques. Furthermore, we establish that the EQR is asymptotically more efficient than (or at least as asymptotically efficient as) the standard quantile regression estimators, without imposing stringent conditions. Hence, our work advances the envelope model theory to general distribution-free settings. We demonstrate the effectiveness of the proposed method via Monte Carlo simulations and real data examples. ∗Shanshan Ding and Zhihua Su are co-first authors. Shanshan Ding is Assistant Professor, Department of Applied Economics and Statistics, University of Delaware. Zhihua Su is Associate Professor, Department of Statistics, University of Florida. Guangyu Zhu is Assistant Professor, Department of Computer Science and Statistics, University of Rhode Island. Lan Wang is Professor, School of Statistics, University of Minnesota. The research of Shanshan Ding is partially supported by National Science Foundation grant DMS-1916376. The research of Zhihua Su is supported by National Science Foundation grant DMS1407460. The research of Lan Wang is supported by National Science Foundation grant DMS-1512267. 1 Statistica Sinica: Preprint doi:10.5705/ss.202018.0060
Multi-parameter one-sided hypothesis test problems arise naturally in many applications. We are particularly interested in effective tests for monitoring multiple quality indices in forestry products. Our search reveals that there are many effective statistical methods in the literature for normal data, and that they can easily be adapted for non-normal data. We find that the beautiful likelihood ratio test is unsatisfactory, because in order to control the size, it must cope with the least favorable distributions at the cost of power. In this paper, we find a novel way to slightly ease the size control, obtaining a much more powerful test. Simulation confirms that the new test retains good control of the type I error and is markedly more powerful than the likelihood ratio test as well as many competitors based on normal data. The new method performs well in the context of monitoring multiple quality indices.
The envelope model allows efficient estimation in multivariate linear regression. In this paper, we propose the sparse envelope model, which is motivated by applications where some response variables are invariant with respect to changes of the predictors and have zero regression coefficients. The envelope estimator is consistent but not sparse, and in many situations it is important to identify the response variables for which the regression coefficients are zero. The sparse envelope model performs variable selection on the responses and preserves the efficiency gains offered by the envelope model. Response variable selection arises naturally in many applications, but has not been studied as thoroughly as predictor variable selection. In this paper, we discuss response variable selection in both the standard multivariate linear regression and the envelope contexts. In response variable selection, even if a response has zero coefficients, it should still be retained to improve the estimation efficiency of the nonzero coefficients. This is different from the practice in predictor variable selection. We establish consistency and the oracle property and obtain the asymptotic distribution of the sparse envelope estimator.