![Sankhya. Series B. [Methodological.]](https://originalfileserver.aminer.cn/sys/aminer/magazine.png)
It is well recognized that relationships between variables are not always linear or even monotonic. For example, the expressions of cell-cycle, or circadian clock genes, or the abundance of microbes in a dynamic ecology are not expected to be linear. Furthermore, unknown to the researcher, there may be heterogeneous subgroups or clusters in the data. Researchers may be interested in discovering those clusters and derive an overall measure of association between variables of interest accounting for the different clusters as well as deriving associations within each cluster. Although standard concepts of correlations, such as the Pearson or Spearman, are widely used to describe overall associations, they can be misleading in such situations. As researchers continue to generate complex high dimensional data with hidden substructures or clusters, there is an urgent need for a measure that correctly quantifies associations between variables while agnostically accounting for hidden clusters in the data. Using clustering algorithms which are able to detect hidden clusters and association measures which are suitable for quantifying arbitrary relationships within each clusters, we develop a novel association procedure called CLuster based Association Measures (CLAM) to describe association between pairs of univariate as well as multivariate variables. The method is not limited to any specific form of association and is well-suited for heterogeneous data with hidden clusters, which are common in biomedical research. Performance of CLAM is evaluated using a synthetic data as well as real data from diverse applications, such as fission yeast (S. pombe) cell-cycle genes data, intestinal microbiome data from IBD patients, and three well-known imaging data sets, namely DrivFace data, Landsat data, and COIL data.
Often linear regression is used to estimate mediation effects. In many instances the underlying relationships may not be linear. Although, the exact functional form of the relationship may be unknown, based on the underlying science, one may hypothesize the shape of the relationship. For these reasons, we develop a novel shape-restricted inference-based methodology for conducting mediation analysis. This work is motivated by an application in fetal endocrinology where researchers are interested in understanding the effects of pesticide application on birth weight, with human chorionic gonadotropin (hCG) as the mediator. Using the proposed methodology on a population-level prenatal screening program data, with hCG as the mediator, we discovered that while the natural direct effects suggest a positive association between pesticide application and birth weight, the natural indirect effects were negative.
Word embeddings are a fundamental tool in natural language processing. Currently, word embedding methods are evaluated on the basis of empirical performance on benchmark data sets, and there is a lack of rigorous understanding of their theoretical properties. This paper studies word embeddings from a statistical theoretical perspective, which is essential for formal inference and uncertainty quantification. We propose a copula-based statistical model for text data and show that under this model, the now-classical Word2Vec method can be interpreted as a statistical estimation method for estimating the theoretical pointwise mutual information (PMI). Next, by building on the work of Levy and Goldberg (2014), we develop a missing value-based estimator as a statistically tractable and interpretable alternative to the Word2Vec approach. The estimation error of this estimator is comparable to Word2Vec and improves upon the truncation-based method proposed by Levy and Goldberg (2014). The proposed estimator also performs comparably to Word2Vec in a benchmark sentiment analysis task on the IMDb Movie Reviews data set.
Tree-based methods have become one of the most flexible, intuitive, and powerful analytic tools for exploring complex data structures. The best documented, and arguably most popular uses of tree-based methods are in biomedical research, where multivariate outcomes occur commonly (e.g. diastolic and systolic blood pressure and nerve conduction measures in studies of neuropathy). Existing tree-based methods for multivariate outcomes do not appropriately take into account the correlation that exists in such data. In this paper, we develop goodness-of-split measures for building multivariate regression trees for continuous multivariate outcomes. We propose two general approaches: minimizing within-node homogeneity and maximizing between-node separation. Within-node homogeneity is measured using the average Mahalanobis distance and the determinant of the variance-covariance matrix. Between-node separation is measured using the Mahalanobis distance, Euclidean distance and standardized Euclidean distance. To enhance prediction accuracy we extend the single multivariate regression tree to an ensemble of multivariate trees. Extensive simulations are presented to examine the properties of our goodness-of-split measures. Finally, the proposed methods are illustrated using two clinical datasets of neuropathy and pediatric cardiac surgery.
The drastic improvement in data collection and acquisition technologies has enabled scientists to collect a great amount of data. With the growing dataset size, typically comes a growing complexity of data structures and of complex models to account for the data structures. How to estimate the parameters of complex models has put a great challenge on current statistical methods. This paper proposes a blockwise consistency approach as a potential solution to the problem, which works by iteratively finding consistent estimates for each block of parameters conditional on the current estimates of the parameters in other blocks. The blockwise consistency approach decomposes the high-dimensional parameter estimation problem into a series of lower-dimensional parameter estimation problems, which often have much simpler structures than the original problem and thus can be easily solved. Moreover, under the framework provided by the blockwise consistency approach, a variety of methods, such as Bayesian and frequentist methods, can be jointly used to achieve a consistent estimator for the original high-dimensional complex model. The blockwise consistency approach is illustrated using high-dimensional linear regression with both univariate and multivariate responses. The results of both problems show that the blockwise consistency approach can provide drastic improvements over the existing methods. Extension of the blockwise consistency approach to many other complex models is straightforward.
We consider the problem of estimating the trend for a spatial random process model expressed as Z(x) = μ(x) + ε(x) + δ(x), where the trend μ is a smooth random function, ε(x) is a mean zero, stationary random process, and {δ(x)} are assumed to be i.i.d. noise with zero mean. We propose a new model for stochastic trend in ℝ d by generalizing the notion of a structural model for trend in time series. We estimate the stochastic trend nonparametrically using a local linear regression method and derive the asymptotic mean squared error of the trend estimate under the proposed model for trend. Our results show that the asymptotic mean squared error for the stochastic trend is of the same order of magnitude as that of a deterministic trend of comparable complexity. This result suggests from the point of view of estimation under stationary noise, it is immaterial whether the trend is treated as deterministic or stochastic. Moreover, we show that the rate of convergence of the estimator is determined by the degree of decay of the correlation function of the stationary process ε(x) and this rate can be different from the usual rate of convergence found in the literature on nonparametric function estimation. We also propose a data dependent selection procedure for the bandwidth parameter which is based on a generalization of Mallow's C p criterion. We illustrate the methodology by simulation studies and by analyzing a data on surface temperature anomalies.
We consider the finite sample performance of a new nonparametric method for bioassay and benchmark analysis in risk assessment, which averages isotonic MLEs based on disjoint subgroups of dosages, and whose asymptotic behavior is essentially optimal (Bhattacharya and Lin, Stat Probab Lett 80:1947–1953, 2010). It is compared with three other methods, including the leading kernel-based method, called DNP, due to Dette et al. (J Am Stat Assoc 100:503–510, 2005) and Dette and Scheder (J Stat Comput Simul 80(5):527–544, 2010). In simulation studies, the present method, termed NAM, outperforms the DNP in the majority of cases considered, although both methods generally do well. In small samples, NAM and DNP both outperform the MLE.
In this article we introduce a new procedure for estimating population parameters under inequality constraints (known as order restrictions) when the unrestricted maximum liklelihood estimator (UMLE) is multivariate normally distributed with a known covariance matrix. Furthermore, a Dunnett-type test procedure along with the corresponding simultaneous confidence intervals are proposed for drawing inferences on elementary contrasts of population parameters under order restrictions. The proposed methodology is motivated by estimation and testing problems encountered in the analysis of covariance models. It is well-known that the restricted maximum likelihood estimator (RMLE) may perform poorly under certain conditions in terms of quadratic loss. For example, when the UMLE is distributed according to multivariate normal distribution with means satisfying simple tree order restriction and the dimension of the population mean vector is large. We investigate the performance of the proposed estimator analytically as well as using computer simulations and discover that the proposed method does not fail in the situations where RMLE fails. We illustrate the proposed methodology by re-analyzing a recently published rat uterotrophic bioassay data.
The number of scientific publications on semiparametric methods per year has been steadily increasing since the early 1980s. This increased interest has happened in spite of the fact that the novelty of semiparametrics for its own sake has run its course, and semiparametric methods are by now considered classical. The underlying reasons for this continued interest include the genuine scientific utility of semiparametric models combined with the breadth and depth of the many theoretical questions that remain to be answered. Empirical process techniques are an essential research tool for many of these questions. Moreover, both semiparametric methods and empirical processes are playing an increasingly valuable role in high dimensional data analysis and in other emerging areas in statistics. The topics are very fruitful and intriguing for new researchers to engage in. Graduate programs in statistics, biostatistics and econometrics can and should include more empirical processes and semiparametrics in their teaching in order to ensure a sufficient supply of suitably qualified researchers.
I first want to thank Professors Moulinath Banerjee (MB hereafter) and Jon Wellner (JW hereafter) for their very illuminating and helpful discussions on my review paper. I am particularly appreciative of the very up-to-date material provided on research topics that I did not discuss in the review as well as their intriguing comments on graduate education. I will now comment briefly on these two items and then conclude with some final words.
"A class of analytical models to study the distribution of maternal age at different births from the data on age-specific fertility rates has been presented. Deriving the distributions and means of maternal age at birth of any specific order, final parity and at next-to-last birth, we have extended the approach to estimate parity progression ratios and the ultimate parity distribution of women in the population.... We illustrate computations of various components of the model expressions with the current fertility experiences of the United States for 1970."
"In this paper, a set of two probability models have been derived to describe the variation in the length of open birth interval of women having given birth to a child during the last 'T' years of their current reproductive age. The first model is derived by assuming the reproduction process as steady-state, the second is obtained by varying the fecundability parameter involved in the first model after the last birth. These models are applied to the three sets of data, one collected from [the Indian] Varanasi-survey, 1969-70 and the other two generated from the data on age-specific fertility rates using the life table technique. The biological parameters such as fecundability and secondary sterility have been estimated using some simple procedure of estimation."
The paper presents a proof of certain conjecture about the LRT in a special case. The conjecture says that if we impose some restrictions on the alternative space then this leads to some gain in the power function. In this paper we consider the bivariate normal distribution with alternatives restricted to some nested cones. We show that the smaller the cone the more powerful is the LRT. The proof is based on writing the difference between the two power functions in the form E(DELTA)g(X) where X:N(DELTA, 1) and g(x) is a function that has one sign change. Using some results of Brown et al. (1981), we conclude that E(DELTA)g(X) has also one sign change. By utilizing Stein's (1956) result, the proof of the conjecture follows immediately.
Recent developments in statistical quality control include the study of the process capability. Much attention has been given to the estimation and the tests Of hypotheses in this regard. This paper brings together the recent work that has been done iu this Of particular interest is the distribution theory of the statistics used for this purpose. Various violations of the assumptions are examined. Extensions to the bivariate normal situation are alw reviewed.
In this paper, based on a simple procedure on the lines of Menken (1979), the estimates of fecundability of migrated couples in rural areas of Eastern Uttar Pradesh have been obtained for females of different age-groups. The estimates are significantly larger than the level of fecundability estimated for non-migrated couples of the area. Some possible explanations for observing such a large gap are also presented.
In this paper we consider the problem of optimally comparing v test treatments with a control using block designs where b1 blocks axe of size k1 and b2 blocks are of size k2, k1 > k2. Following Majumdar and Notz (1983), some sufficient conditions are derived for balanced treatment unequal block designs to be A and MV-optimal in these situations and methods of constructing optimal designs which satisfy the sufficient conditions obtained are discussed.
Correlation studies on various agronomic and physiological traits in plants throw useful light for constructing selection indices in crop improvement programs. Large sample bias and mean squared error of genotypic correlation between two traits from a set of fixed number of genotypes grown in randomized complete block design have been evaluated. These results are obtained for general bivariate distributions of the two characters on genotypes and also on plot errors. Simplication has been indicated for nine combinations of normal and non-normal bivariate populations consisting of bivariate normal, bivariate t and bivariate chisquare distributions. Several transformations for genotypic correlations have been considered and approximations for their biases and mean squared errors obtained. A simulation study was undertaken to examine the closeness of approximations of bias and mean squared error and departure of the distribution of various transformation from normality. We found close agreement of approximated values of bias and mean squared error of the estimate of genotypic co"elation coefficient with their corresponding simulated values for a number of populations and parameter sets. Samiuddin's and tanh inverse transformations showed high skewness and kurtosis values. Arcsine transformation was close to normal. Data from 40 or more genotypes evaluated with three or more replications generally provide valid estimate of genotypic correlation. Arcsine transformations was found robust over the distributions considered and the approximations made were satisfactorily close for this transformation.
A generalised definition of outliers is given. The serial numbers of individual items in a sample collected for detecting outliers among them using group tests we recoded by the combinations of a suitable asymmetrical factorial. Using these factorial encoders, a method of forming group samples is provided. After the test results in the form of positive or negative results are known, a method of unique detection of outliers is given when the number of outliers is not known in advance. The present technique leads to a reduction of group sizes which is necessary to avoid risk of dilution effect.
We show that the estimator Q(sr) of the population total Y, introduced in Srivastava (1985), contains as special cases the well known Mickey and Hartley-Ross estimators of Y.