
The joint adaptive progressive censoring scheme proposed by Sultana et al. (2021) reduces experimental time and cost in life-testing experiments. In this paper, we extend this scheme by developing statistical inference for two independent Chen populations under jointly adaptive progressive Type-II censoring. Unlike earlier works that mainly focused on exponential models, we consider the flexible Chen distribution, which can model both increasing and bathtub-shaped hazard rates. We obtain maximum likelihood estimates and construct asymptotic and bootstrap confidence intervals. We also apply Bayesian methods to derive credible intervals under squared error loss and LINEX loss functions. The methods are illustrated using real data. Likelihood ratio tests are used to examine whether the two populations have similar shape parameters under both single-sample and joint-sample approaches. An extensive simulation study is carried out to evaluate the performance of the methods and to suggest suitable censoring schemes. The results show that the joint censoring scheme performs better than the single-sample censoring scheme.
Francis Galton, in his seminal 1886 work titled “Regression Towards Mediocrity in Hereditary Stature", scrutinized the heights of 928 adult offspring along with 205 pairs of parents to elucidate the new concept of “regression toward the mean”. Despite the recurrent use of Galton’s data to exemplify this concept, several studies have contested its conformity to his chosen straight-line linear regression model, attributing a discrepancy to an apparent non-linearity in the data, now recognised as “Galton’s bend". In our paper, we provide a possible explanation for the presence of Galton’s bend by advancing the theory of errors-in-variables regression where we show that E[Y | x] is only linear for specific cases of errors-in-variables models.
In clinical trials, multiple correlated continuous and binary variables are often employed as primary endpoints, and evaluation using multiple endpoints establishes evidence for the efficacy of the test treatment. When dealing with multiple endpoints, the familywise error rate of statistical tests must be kept below the nominal significance level. In several studies, procedures have been developed that can be applied to multiple primary endpoints with only a superiority test. However, a procedure that simultaneously incorporates non-inferiority and superiority tests and includes multiple continuous as well as binary variables has not yet been discussed. In this study, we propose a testing procedure that recognizes the efficacy of test treatment only when the superiority of at least one endpoint and the non-inferiority of the remaining endpoints are achieved. The type I error rates and powers of the proposed testing procedure are evaluated through simulations and are compared with those of the closed testing procedure. Regardless of the correlation between endpoints, sample size, or the magnitude of the difference between endpoints in the two groups, the type I error rate was shown not to be inflated. The proposed testing procedure showed higher power than the closed testing procedure. When any of the endpoints are within the non-inferiority margin, the power is drastically reduced depending on the correlation coefficient, indicating the importance of obtaining reliable information a priori.
This article presents two novel goodness-of-fit tests for the Rayleigh distribution, specifically designed for analyzing Type-II censored data. The development of these tests is based on the Kullback–Leibler information criterion. Both tests exhibit consistency, and one of the test statistics possesses a nonnegative property akin to the Kullback–Leibler information. To assess the performance of the proposed tests, a comprehensive simulation study is conducted. The simulation results provide percentile points and power values, offering valuable insights into the effectiveness of the tests in detecting departures from the Rayleigh distribution. Furthermore, a real-life data analysis is included to demonstrate the practical application of the proposed tests. This application showcases how these tests can be utilized to assess the goodness-of-fit of the Rayleigh distribution when analyzing real-world data.
In this paper a discrete version of the Exponentiated Weibull-Geometric distribution called Discrete Exponentiated Weibull-Geometric(DEWG) distribution is developed and its properties are studied. The parameters are estimated using both classical and Bayesian approaches. The goodness of fit of the newly developed model is evaluated using different simulated datasets and a real dataset. The new model is compared with existing models such as the Poisson, Negative Binomial, and Discrete Weibull distributions using AIC (Akaike Information Criterion), BIC (Bayesian Information Criterion), and the Kolmogorov–Smirnov test. Furthermore, we develop a novel regression model for count data based on the DEWG distribution, termed the DEWG regression model. The estimation of its parameters is carried out via maximum likelihood (ML) and Bayesian methods. The performance of the newly developed regression model is evaluated using several simulated datasets, and its practical usefulness is demonstrated with a real dataset. The new model is further compared with the existing models using the Vuong test.
This paper investigates the role of wavelet based models in analyzing periodic time series. The proposed method leads to a parsimonious periodic autoregressive moving average (PARMA) model with a reduced number of parameters compared to the commonly used methods based on Fourier analysis for such series. Our simulation studies and the data analysis illustrate the parsimonious nature and the forecasting efficiency of the proposed models.
In this paper, we present the testing of many hypotheses on multiple streams of observations that are driven by Lévy processes. This is applicable for sequential decision-making on the state of multi-sensor systems. In one case, each sensor receives or does not receive a signal obstructed by noise. In another, each sensor receives data driven by Lévy processes with large or small jumps. In either case, these give rise to 2^n possibilities. Infinitesimal generators are presented and analyzed. Bounds for infinitesimal generators in terms of super-solutions and sub-solutions are computed. An application of this procedure for the stochastic model is also presented in relation to the financial market.
In this paper, we introduce a method for selecting among k experimental treatments, each having two Bernoulli endpoints (e.g., efficacy and safety), the subset that contains the treatments whose efficacy and safety rates are superior to those of a control treatment. We then identify the treatment within the selected subset that has the highest efficacy rate. If no treatment meets the criteria of surpassing the control in both efficacy and safety, the control treatment is selected instead. Throughout, we consider the case in which the association between the two binary endpoints is characterized by a common, known odds ratio ϕ shared by the control and all experimental arms. We employ a quadrinomial distribution to perform the exact calculations and derive the formulas for the proposed procedure. All designs use the exact counts of outcomes rather than the typical normal approximation, allowing for more accurate sample size determination to meet the required probability guarantees.
The study of human genetic variation among populations has attracted much interest for almost 100 years, since the first realization of the differences in allele frequencies that are the result of human history, geography, and demography. Measures of genetic differentiation were developed not only to summarize this diversity but also to make inferences about the underlying forces that shape this variation. Likewise, the study of genetic similarities and differences among related individuals is also about 100 years old. These individual patterns of variation also derive from the events of randomness in the transmission of DNA from generation to generation, but on a much shorter time-scale. In this centenary volume in honor of Professor C. R. Rao, a brief review of approaches to the analysis of human genetic variation is appropriate: Professor Rao also, among his many major contributions, developed multivariate measures of population differentiation and statistical approaches to the analysis of genetic diversity Rao (Theor. Popul. Biol., 21, 24–43, 1982). As genetic and genomic technologies evolve, the available data provide new insights into the structure of human genetic variation and new measures and methods have been developed to address questions of ancestry and admixture in our genomes. The advent of widely available SNP data activated a wave of analyses and many interesting results. These data led to a merging of individual-based and population-based analyses with the ability to analyze genotypic information on many individuals who may represent a number of different human populations or whose genome may be a mixture of recognized separate populations. However, the inheritance of DNA is segmental in nature, resulting in an additional dimension of haplotypic variation, which can reveal more fine-structured patterns of genetic variation. The ability to study human genetic variation at the individual level offers hope for personalized medicine, and is a stated motivation for many human genetic studies. However, within many populations, and more so between them, the high dimensionality of variation means that current methods to adjust for population structure for GWAS or risk analyses may fail in this regard. An understanding of the complexities of human genetic variation is important to a recognition of the limitations of current genetic assessments of individual health risks.
Tuberculosis is one of the leading challenges for global health, it occurs throughout the world at a varying incidence rate. India, among Asian countries, faces the most significant burden of tuberculosis, with it being a major cause of mortality and morbidity. Evaluating its burden in relevance to vulnerable population can help in developing appropriate control measures. Hence, this study aims to assess the influence of socio-economic vulnerability and other contributing factors on the risk of tuberculosis. A multilevel logistic regression model is used to assess the risk and its variance partition at multiple levels such as individual, district and state. Data used for the study are obtained from the household file of National Family Health Survey – 5. Bihar, Kerala and a set of north-eastern states are having the highest burden of tuberculosis. Living Standard, Household Composition, Health and Resource Accessibility, and Infrastructure and Sanitation are identified as the four latent variables of socio-economic vulnerability and they have significant odds ratio (OR: 0.73, 0.97, 0.99, 1.04) for the risk of tuberculosis respectively. In the final model, it is observed that district level variation (10
Fertility is one of the most significant determinants of population dynamics and growth. It is also necessary to obtain estimates of fertility rate, trends, and patterns of a country/region in order to plan and monitor the socioeconomic development. The Parity Progression Ratio (PPR) has acquired dominant place in the study of fertility which gives the probability that a woman giving birth of a given order at a time will ever proceed to the next birth. In this study, we have proposed a methodology to get estimates of PPR values for different parities by utilising data on only the open birth interval (OBI) of different birth cohorts. This methodology also provides a procedure to get estimates of mean closed birth intervals (CBIs). We have applied this methodology to the fifth round of National Family Health Survey (NFHS-V) data to get estimates of PPR values and mean closed birth intervals (CBIs) up to parity three. Further, a comparison between the estimates obtained using the proposed methodology and those derived from earlier approaches demonstrates that our method yields consistent results across different parities. Additionally, simulation studies have also been done to assess the efficiency and reliability of the proposed methodology.
Rao’s spacing test is a widely used nonparametric method for assessing uniformity on the circle. However, its broader applicability in practical settings has been limited because the null distribution is not easily calculated. As a result, practitioners have traditionally depend on pre-tabulated critical values computed for a limited set of sample sizes, which restricts the flexibility and generality of the method. In this paper, we address this limitation by recursively computing higher-order moments of the Rao’s spacing test statistic and employing the Gram-Charlier expansion to derive an accurate approximation to its null distribution. This approach allows for the efficient and direct computation of P-values for arbitrary sample sizes, thereby eliminating the dependency on existing critical value tables. Moreover, we confirm that our method remains accurate and effective even for large sample sizes that are not represented in current tables, thus overcoming a significant practical limitation. Comparative evaluations with published critical values and saddlepoint approximations demonstrate that our method achieves a high degree of accuracy across a wide range of sample sizes. These findings greatly improve the practicality and usability of Rao’s spacing test in both theoretical investigations and applied statistical analyses.
In this paper we consider a stratified clusters based finite population and develop a design cum model unbiased (DCMU) predictor for the finite population total using the two-stage cluster samples based survey data. The same design cum model based principle is applied to develop a valid variance function for the predictor and its unbiased estimate for the purpose of confidence intervals construction when needed. As far as the existing studies using a model unbiased prediction approach are concerned, unlike their claims, we demonstrate in this paper that the ordinary and/or generalized least square estimates computed conditionally on the given sample can never be model unbiased (MU) for the regression parameters involved in the prediction function which makes the existing MU prediction a flawed approach. Turning back to the design cum model unbiased predictions, while the aforementioned ordinary least square estimators in an independence setup can be used for DCMU prediction, the generalized least square estimators under a correlation setup are DCM biased and they are useless for any DCMU predictions. We, instead, use a doubly weighted (accommodating both sampling and inverse correlation weights) regression estimator which is DCMU for the parameter leading to DCMU predictions. The variance of the doubly weighted estimator based predictor and its unbiased estimates are provided by exploiting the DCM based approach. As the clusters based survey data are frequently encountered in general by statistical agencies, the proposed methodological study should be highly beneficial to those practitioners among other applied statistics researchers.
The descriptive analysis of real data can be used to determine the levels and trends of human fertility. However, some essential and profound properties of the phenomenon, such as fecundability and sterility that are inherent in nature and cannot be observed directly. They can only be estimated using appropriate probability models. This paper attempts to develop some probability models to describe the distribution of waiting time from marriage to first conception based on the data from two different periods i.e. 1969-70 and 2015-16 for Varanasi District. The proposed models effectively address the challenge of variability in fecundability within heterogeneous groups of married women, which is often overlooked by traditional mathematical models. The uniqueness of this approach lies in assuming an appropriate hazard rate instead of the constant hazard rate for fecundability. Further, the proposed model based on appropriate hazard, generalized considering the parameter as a random variable. The generalized distribution provides a better understanding of the data on waiting time to first conception. Utilizing the above two data sets of waiting time to first conception, proposed probability models are used to estimate the parameters.
We present a partial collection of tributes, biographies, celebrations of life and obituaries of Prof. C. R. Rao. We have limited our search to only those articles – in print or on social media, which were published in English. No claim is made for the exhaustiveness of this collection.
This paper analyses fertility transition in India during 1985–2020 based on the data from the official sample registration system. The analysis reveals that fertility transition in the country is contingent upon the way age-specific fertility rates are aggregated into a single composite indicator of fertility. When the simple arithmetic mean of age-specific fertility rates is used as a composite indicator of fertility, fertility in India has decreased almost linearly. However, when the geometric mean of age-specific fertility rates is used as the composite indicator of fertility, fertility transition in India appears to have stalled during the period 2011–2013. The analysis also reveals that the change in marital fertility accounted for only about 35 per cent of the change in the simple arithmetic mean of age-specific fertility rates but more than half of the change in the geometric mean of age-specific fertility rates. The paper suggests that fertility transition should not be analysed in terms of the trend in the simple arithmetic mean of age-specific fertility rates or, equivalently, total fertility rate but should be analysed in terms of the trend in the geometric mean of age-specific fertility rates.
When multitudes of features can plausibly be associated with a response, both privacy considerations and model parsimony suggest grouping them to increase the predictive power of a regression model. Specifically, the identification of groups of predictors significantly associated with the response variable eases further downstream analysis and decision-making. This paper proposes a new data analysis methodology that utilizes the high-dimensional predictor space to construct an implicit network with weighted edges to identify significant associations between the response and the predictors. Using a population model for groups of predictors defined via network-wide metrics, a new supervised grouping algorithm is proposed to determine the correct group, with probability tending to one as the sample size diverges to infinity. For this reason, we establish several theoretical properties of the estimates of network-wide metrics. A novel model-assisted bootstrap procedure that substantially decreases computational complexity is developed, facilitating the assessment of uncertainty in the estimates of network-wide metrics. The proposed methods account for several challenges that arise in the high-dimensional data setting, including (i) a large number of predictors, (ii) uncertainty regarding the true statistical model, and (iii) model selection variability. The performance of the proposed methods is demonstrated through numerical experiments, data from sports analytics, and breast cancer data.
This paper develops a new prediction framework for the extreme episodes of cross-border capital flows in a mixed-frequency binary choice panel setting. The four well-established episodes are redefined and are modelled not only individually, as routinely assessed, but also jointly, in pairs. The model predicts quarterly event probabilities using daily and monthly macro-financial predictors. The time series of out-of-sample predictions is assessed using formal forecast verification techniques in full information and pseudo-real-time settings. The panel model predicts with significant forecast skill with respect to a random classifier and an intercept-only benchmark – more prominently so for three of the six episodes. The pseudo-real-time analysis also shows that the predictions can beat the benchmarks and thus generate meaningful early warnings a quarter in advance of the target quarter. Accuracy remains stable as time elapses and more information becomes available.
I provide a brief nontechnical review of the major lifetime work of Prof. C R Rao from a personal point of view. Emphasis is on the areas where his work has left a lasting impact.
A continuous-time Markov model is developed to account for transitions between discrete states of the infected and uninfected disease status of individuals in a population. Home-based quarantining of the infected individual is considered as one of the states of the model. The mean survival time, the mean hospitalisation time and the mean home isolation time have been derived using matrix calculus. Theoretical relationships between transition probabilities, mean times in state, and transition intensities have been obtained. Extended continuous time Markov model is also discussed by taking into account the age of the infected individual that could prove useful in analysing the impact of different age groups on the transition probabilities. Analysis is then conducted to understand elasticity and sensitivity of parameters. Computations have been shown keeping in mind a general hospital data, and some individuals could be under quarantine.