
A longitudinal data set where the system is characterized by its states in place of the values of the underlying random variables taken over time, can be modeled using a Markov model. In case of psychological data, the Markov models are beneficial as problems may progress or regress over time thus exhibiting the shift in the states of the system. These models, when applied to cohort studies may indicate at a general shift in the psychological health of the cohort under study. In a study regarding the psychological health of young adults in higher education in India during COVID-19 pandemic period, three independent surveys were conducted using the Strength and Difficulty Questionnaire (SDQ). 162 respondents were found to have been participated in all three surveys. A Markov chain model was used to study the transition of the respondents’ psychological health over different phases of the pandemic duration in respect of the observed scores of the two components of SDQ viz, the ‘Difficulty’ score and the ‘Impact’ score; and the estimated ‘Impact’ scores obtained from the observed ‘Difficulty’ scores on application of the Quantile Regression and Quantile Regression Neural Network. For all the three data sets, the Markov model indicated at the prominent shift from a ‘Normal’ state to the ‘Borderline’ and the ‘Abnormal’ states of SDQ. Moreover, the stationary distributions showed significantly higher probabilities of being in the ‘Borderline’ and the ‘Abnormal’ states during the pandemic period than what is suggested by the psychological manuals in standard times.
Count data represent the number of occurrences of an event within a fixed time or space and arise frequently in areas such as epidemiology, insurance, demography, and reliability analysis. In many practical situations, zero counts are structurally absent or unobserved, resulting in zero-truncated count data. Standard zero-truncated models may be inadequate when such data exhibit substantial over-dispersion. In this paper, we introduce and study a zero-truncated hyper-negative binomial distribution (ZTHNBD) to addressthis limitation. This work constitutes the first systematic investigation of the ZTHNBD. Several important statistical properties of the proposed distribution are derived, including the probability mass function, cumulative distribution function, mode, log-concavity, survival function, and hazard function.Recurrence relations for probabilities, raw moments, and factorial moments are also obtained. Parameter estimation is carried out using the method of maximum likelihood, and a generalized likelihood ratio test is developed to assess the significance of the additional parameter. The practical usefulness of the proposed model is demonstrated using multiple real-life data sets, where it provides an improved fit compared to existing zero-truncated models based on goodness of fit measures and information criteria. A brief simulation study is conducted to examine the finite sample performance of the maximum likelihood estimators.
This study describes the profile of low-income mothers who already had children and received prenatal care in the municipality of Inimutaba, Minas Gerais, between January 2020 and December 2023. Sociodemographic variables were analyzed to identify the factors most strongly associated with the number of children per woman. The sample included 134 women with an average age of 29 years, the majority of whom self-identified as 86.56%, living on a minimum wage 94%, and having completed high school 94%. To investigate the determinants of fertility within this population, we applied an Ordinary Least Squares (OLS) regression and a Classification and Regression Tree (CART) algorithm. Although the in-sample OLS model initially exhibited a high coefficient of determination approx 0.91%, internal validation through 10-fold cross-validation demonstrated substantially lower generalization performance approx 0.37%, indicating overfitting. The most relevant predictors after model refinement were family size, household income, marital status (single), and age. The CART model showed similar patterns. The fully grown tree overfitted the training data, whereas cost-complexity pruning with an optimal parameter of 0.113% improved test accuracy to 74%. Despite this improvement, the model exhibited limited discriminative ability (AUC = 0.55), reinforcing the importance of careful interpretation in small epidemiological datasets. Overall, the findings indicate that fertility patterns among low-income mothers in Inimutaba are closely linked to socioeconomic vulnerability. Machine learning and statistical models, when combined with rigorous internal validation, can help identify high-risk subgroups and support targeted interventions for maternal health.
This paper introduces C-SOMGOAM, a self-organizing migrating grasshopper optimization algorithm designed for constrained optimization (CO). The algorithm integrates features from the Grasshopper Optimization Algorithm (GOA), the Self-Organizing Migrating Algorithm (SOMA), and a non-uniform mutation operator. A key contribution of this work is the incorporation of a penalty-free constraint-handling mechanism into GOA to effectively address CO problems. Unlike some traditional methods, the proposed GOA variant explicitly operates on a population of solutions. Starting with a randomly selected population, the algorithm iteratively modifies these solutions to improve the approximation of the global optimum. C-SOMGOAM is not only simple to implement but also capable of generating feasible and high-quality solutions. To assess its performance, the algorithm is evaluated on ten constrained benchmark problems and three engineering design problems from the literature. The effectiveness of C-SOMGOAM is demonstrated through comprehensive result analysis and comparative evaluation against existing variants.
Gas flaring is the most common source of global warming, causing environmental pollution and ecological disturbances. This study evaluated the environmental impact of gas flaring and possible bacterial isolates that can be employed in the remediation of petroleum hydrocarbon polluted soil. Soil samples collected from Ologbo Settlement in Edo State, Nigeria were analyzed for their physicochemical and petroleum hydrocarbon parameters. The Shake flask biodegradation test was carried with screening of hydrocarbon degrading bacterial isolates which were characterized using the 16S RNA analysis technique. Across the six locations (40, 80, and 120, 160, 200 and 1000 m) sampled from 40m to 1000m and the control, it was observed that there was a progressive increase in the soil pH, moisture content and electricity conductivity. In the other hand, there was a gradual decrease in the soil temperature, total hydrocarbon content and total organic carbon. The total petroleum hydrocarbon (TPH), oil and grease and polycyclic aromatic hydrocarbon (PAH) content showed statistical significance (p<0.05) compare to the control which implies that distances from the flare sites significantly influence the hydrocarbon parameters of the soil. From the molecular characterization of the bacterial isolates, the four isolates were Acinetobacter tandoii strain BASG143, Bacillus cereus strain Ou9, Bacillus subtilis strain BS3902 and Pseudomonas aeruginosa strain KAVKOI. These results show that these bacteria strains were able to degrade hydrocarbon contaminants.
Mass media refers to any communication platform that reaches large audiences and is used to inform, entertain, and educate people. Ideally, access to mass media should be distributed equally across all sections of society, regardless of gender, caste, economic status, religion, place of residence, or level of education. Yet, only a limited number of studies—particularly those using recent data—have examined this issue. The present study specifically investigates exposure to mass media among individuals who reported using at least one of the following—radio, television, newspapers, or magazines at least once a week. The data for this study was obtained from the National Family Health Survey (2019–21) among 15–49 years age group. Logistic regression model was performed to assess the gender differences for men and women. In hierarchical regression model, background (age, caste, religion and marital status) were entered on the first step, place of residence (urban/rural) on the second, education on the third and wealth index on the fourth step. Analysis reveals that only 41.7 percent women and 52.1 percent men aged 15–49 years were regularly exposed to the mass media in Uttar Pradesh, India. Majority of the respondents (39.1 percent of women and 42 percent of men) were exposed to television in the age group 15-49. The finding indicated that regular media exposure is significantly less among women as compared to men (AOR = 0.71, 95% CI:0.68–0.74). Women who received twelve or more years of education were 2.8 times (AOR = 2.80, 95% CI:2.67–2.93) more likely to be exposed to mass media as compared to the category of no education whereas men were about 5.4 times more likely to be exposed towards mass media (AOR = 5.43, 95% C.I:4.69–6.30). Furthermore, women in the rich wealth index were 5.4 times more likely (AOR = 5.38, 95% C.I:5.2–5.6) to be exposed but men were only 3.4 times (AOR = 3.31, 95% CI:2.96–3.70) exposed to mass media compared to the poor wealth index. However, rural women had lower odds of mass media exposure than urban (AOR=0.66, 95% CI:0.63–0.69), rural men had lower likelihood of mass media exposure than urban men (AOR=0.67, 95% CI:0.60–0.76). The study demonstrates that men are more likely than women to be exposed towards mass media in Uttar Pradesh, India. The findings of this paper provide evidence that education level and wealth index were the main significant predictor variables for gender differences towards overall mass media exposure. It suggests that greater focus must be placed on the marginalized population, especially women.
The family decision process is a complex phenomenon. At some point, a couple may decide to interfere with the childbirth process in some way. The result of this decision is a sudden event that limits family size and gender composition. An important consideration in a study dealing with the number of children families wish to have is whether these desires include preference as to the child's sex. The purpose of this paper is to propose five hypothetical rules that reflect the current preferences of parents regarding the size and gender composition of their children. Various combinations of marital durations, levels of fecundity, and rest periods (gestational period plus amenorrhea period) have been used to calculate the expected waiting time for each case to reach the desired family size. Additionally, for each case, an estimate of the truncation bias has been obtained. The lower limit of the range in each case is constructed by that value of the expected waiting time which corresponds to the pair of parameters andh = 0.92 years. The upper limit is the value of the expected waiting time derived by assuming and h = 1.25 years for an infinite duration of the marriage. As a result of this information, couples may be able to plan their families in such a way that all their financial and social obligations have been met by the time they plan to retire, which means that it will help make decisions about family planning at the micro level.
In this paper we have suggested a family of estimators for estimating the population mean of the study variable with the dual use of auxiliary information in sample surveys. In addition to Haq et al. (2017) many other estimators are members of the proposed family of estimators. The bias and mean squared error of the suggested family of estimators are obtained up to the first order of approximation. We have compared the proposed family of estimators with some existing estimators and derived the conditions under which the suggested family of estimators is more efficient than the existing estimators. An empirical study is carried out to demonstrate the performance of the proposed family of estimators over existing estimators.
Modeling lifetime and reliability data requires flexible probability distributions capable of capturing diverse hazard rate behaviors and skewness patterns. Classical models often fail to adequately represent left-skewed data with bounded ranges. To address this limitation, the Reflected-Shifted-Truncated Maxwell (RSTM) distribution is introduced as an extension of the classical Maxwell model through reflection, shifting, and truncation. Key statistical properties including moments, hazard rate behavior, and stress–strength reliability are derived. Parameters are estimated using the maximum likelihood method for both complete and right-censored data, and estimator performance is assessed via simulation studies. The effectiveness of the RSTM distribution is illustrated through two fiberglass strength datasets, representative of left-skewed lifetime data. Comparative analysis based on information-theoretic measures demonstrates that the RSTM distribution consistently outperforms competing models, underscoring its potential as a robust tool for modeling leftskewed lifetime and reliability data.
Sport has increasingly evolved into an interdisciplinary research field, where statistical modelling plays a central role in performance analysis, injury prediction, tactical optimisation, and audience behaviour. Although traditional methods like linear regression and generalised linear models are frequently employed, their inability to handle complex data structures has led to the expanding usage of more flexible approaches. In this context, generalised additive models for location, scale and shape (GAMLSS) are one of the most flexible statistical frameworks currently available, enabling the simultaneous modelling of multiple distributional parameters. Hence, in this paper, we perform a detailed systematic review exploring the application of GAMLSS in sports science, drawing on peer-reviewed articles. The results show a variety of applications, such as the development of reference growth curves, athlete performance modelling, match-fixing detection, and forecasting of sport-related consumer behaviour. GAMLSS have proven especially useful in contexts where traditional models are inadequate, offering enhanced flexibility in capturing distributional nuances. Nonetheless, opportunities remain to integrate GAMLSS with machine learning techniques and to extend their use across underexplored domains in sport. This review contributes to the field by outlining current trends, highlighting methodological strengths, and identifying promising directions for future research.
Sickle Cell Anaemia (SCA) is a genetic blood disorder caused by a mutation in the haemoglobin gene, leading to the production of abnormal haemoglobin known as haemoglobin S. This abnormal haemoglobin causes red blood cells to become rigid, sticky, and shaped like a crescent or sickle, which obstructs blood flow and leads to various complications such as pain, infections, and potential damage to nerves and organs (kidneys, liver and spleen). This research utilizes a two-level factorial experiment to evaluate the impact of four major factors (Age, Sex, Genotype, and Rhythm) on six distinct blood pressure (BP) indices: Systolic Blood Pressure (SBP), Diastolic Blood Pressure (DBP), Pulse Rate (PR), Pulse Pressure (PP), Mean Arterial Pressure (MAP), and Rate Pressure Product (RPP). The experimental units consist of young adults with Sickle Cell Anaemia (SCA) and Haemoglobin AA (HbAA). The results of the analysis indicate that Age and Genotype are the major significant factors affecting blood pressure (BP) indices. Meanwhile, Pulse Pressure (PP) appears to be more sensitive to the aforementioned factors when compared to SBP or DBP. Also, the interaction effects between Age and Genotype, and between Age and Sex demonstrate clinical relevance. Importantly, these results highlight the importance of early detection of abnormal cardiovascular symptoms and open ways for further heart disease diagnostic tests and treatments in young adults. It is also worthwhile to note that Pulse Pressure (PP) provides a more comprehensive measure for abnormal cardiovascular detection within young adults.
Problems involving comparisons of treatment effectiveness for multivariate responses are common in various fields of knowledge. Typically, methods for comparing vectors of means use the Bonferroni inequality to construct conservative tests, avoiding the complexities of the exact distribution of the maximum test statistic T2 max, the maximum of Hotelling’s T2. In high-dimensional scenarios, traditional methods are not viable as they depend on the inverse of the sample covariance matrix, which becomes singular. To address this problem, Dempster’s trace criterion can be used, and a second alternative is Ahmad’s test statistic Tig, however, in both cases, the Bonferroni inequality is used. Another issue is that both in the estimation process and in the exact distribution of statistics for multiple comparison tests, there is a need to deal with sophisticated and complex numerical methods. These facts make these approximations not readily usable. To try to overcome these challenges, this work proposes multivariate multiple comparison tests with the control treatment using the nonparametric bootstrap method. The performance of the tests was evaluated through experimentwise type I error rates (EER) and power in different scenarios using Monte Carlo simulation and the R program. The results showed that for homoscedastic scenarios, the proposed bootstrap test ATB showed more effective control of EER, in addition to having higher power, regardless of whether the distribution was normal or not, in both low and high-dimensional contexts. Thus, the ATB test proved to be the most recommended alternative in these situations. For heteroscedastic scenarios, it was not possible to identify a clearly superior test, but in several circumstances, the proposed bootstrap tests demonstrated superior performance compared to their respective asymptotic versions.
Tropical forests play a crucial role in regulating the climate and maintaining the global carbon cycle. Carbon is primarily stored in plant biomass. In Africa, these forests still present significant uncertainty in the carbon balance. This is due to scarce inventories and a lack of continuous monitoring. Mopane forests (Colophospermum mopane) are typical of dry tropical regions and have important ecological and socioeconomic roles. They are widely used by local communities. In Mozambique, charcoal production mainly drives degradation of these forests and loss of aboveground carbon (AGC). The aim of this study is to analyze aboveground carbon stock in Mopane forests in the Chicualacuala District, Southern Mozambique. The data used refer to georeferenced tree locations in a fragment of native forest. Aboveground biomass was individually estimated through a pantropical equation, using as predictor variables: diameter at breast height (DBH), total height, and basic wood density. From the biomass estimates, AGC was calculated and subsequently analyzed using geostatistical methods such as variograms and ordinary kriging. The results indicated that the ordinary kriging model partially captured the spatial pattern of AGC and, as evidenced by cross-validation, showed a moderate positive correlation (54.66%) between observed and predicted values. This level of correlation suggests that while the model provides a reasonable prediction of AGC, there is still considerable unexplained variability, which highlights areas for further model refinement or additional data collection.
Sensory evaluation studies often rely on hedonic scales that yield ordinal and bounded data. Traditional statistical techniques, although widely applied, often fail to account to the boundedness and interdependencies characteristic of sensory responses. This study proposes a Bayesian beta regression framework specifically designed for sensory data, aiming to extend existing methodologies by addressing the joint behavior of multiple attributes. The approach models scores rescaled to the (0,1) interval via a beta distribution, effectively capturing formulation effects and correlations among sensory attributes. To this end, we assumed a hierarchical structure for the regression coefficients. We chose weak prior distributions to avoid strong or subjective assumptions that might distort the results thereby, making the analysis more robust. Without strong restrictions on the prior model, the posterior inference was conducted via Markov Chain Monte Carlo (MCMC) methods. By jointly analyzing all attributes within a hierarchical structure, the method enables direct estimation of inter-attribute associations and offers a more integrated interpretation of product performance. Applied to a grape juice acceptance study, the model not only identified the most preferred formulations but also unveiled meaningful patterns across sensory dimensions. With this Bayesian construction, we present a singular contribution that provides a reliable alternative for researchers seeking to analyze bounded sensory responses while simultaneously exploring the multivariate nature of consumer perception.
The global rapid spread of COVID-19, partially driven by asymptomatic undetected carriers, needs appropriate modeling to inform effective intervention strategies. This research develops and describes a deterministic compartmental model with six epidemiological classes: Susceptible (S), Exposed (E), Asymptomatic Infected (IA), Symptomatic Infected (IS), Quarantined (Q), and Recovered (R), all encompassing the SEIAISQR framework. The model incorporates the progression of asymptomatic infections to symptomatic phases prior to progressing into quarantine, which captures realistic disease state progressions. By means of analytical methods, the basic reproduction number (R0) is derived from the nextgeneration matrix, and the local and global stability of the disease-free steady state are established under the R0 < 1 assumption. Sensitivity analysis reveals that transmission rates, progression rates, and quarantine policies drive R0, and transmission through asymptomatic carriers is dominant. Euler’s method-based numerical simulation shows that asymptomatic undetected carriers have a high contribution to persistent disease transmission, while effective quarantine and extensive recovery severely inhibit epidemic persistence. This paper stresses the critical necessity for robust public health measures involving mass testing, successful contact tracing, and isolation of asymptomatic patients in a bid to combat COVID-19.
This paper proposes new 16-run foldover designs aimed at minimizing the number of level changes. The method involves selecting n - p independent columns from the main effects and interaction effects of a 24 full factorial experiment while avoiding run duplication to construct a 2n-p fractional factorial design. The remaining p columns are then generated using these selected independent columns to maintain cost efficiency. The number of level changes and trend-free factors are computed for all possible fold-over plans for both standard and new designs. The performance of the new designs is compared to standard designs proposed by Li and Lin (2003), Cheng and Steinberg (1991) and Coster (1993) using the criteria of minimum level change, maximum trend-free factors and uniformity exhibiting improved performance and cost-effectiveness.
Carbon emission has become a major challenge everywhere, such as in the production process, holding items, deteriorating items, waste disposal, and transporting items. Two warehouses are a realistic approach in inventory modeling. Inflation and lead time play a key role in making this study closer to reality. Uncertainty is also a very realistic approach for any organization. In this study, we developed a supply chain model in which there is one retailer and one supplier. The retailer holds its inventory in two warehouses. Our objective is to find the optimal total cost, cycle length, and supplier's lead time. In this study, we calculate the total cost in three different ways: first, for the crisp model; second, for the fuzzy model using the signed distance method; and third, using the graded mean integration method. We carried out the numerical solution using the software MATHEMATICA 12.0. From numerical illustration, we find that the total cost is minimum for the fuzzy model using the graded mean integration method. Sensitivity analysis is carried out to see the behavior of different parameters on the total cost.
This paper introduces four new types of multiple comparison procedures. Their performance was evaluated in large experimental scenarios using Monte Carlo simulations, with a focus on comparing experimentwise error rates and statistical power against those of the traditional Tukey, Student-Newman-Keuls (SNK), and Scott-Knott tests. To facilitate the practical application of these methods, we developed the R package midrangeMCP. The proposed procedures are: the Midrange Tukey test (TM), the Midrange Student-Newman-Keuls test (SNKM), the Mean Grouping based on the Midrange (MGM), and the Mean Grouping based on the Range (MGR). The TM, MGM, and MGR tests outperformed their classical counterparts (Tukey and Scott-Knott), except under scenarios with partially true null hypotheses. The SNKM test was the only one that consistently underperformed compared to its original version (the SNK test) across all evaluated scenarios. Among the proposed methods, the MGM test stood out for its superior performance and its additional advantage of avoiding ambiguity in group assignments.
Emotion detection plays a vital role in understanding human sentiments and behaviors across various applications, including customer feedback analysis and mental health monitoring. This research assesses the efficiency of different algorithms for machine learning in detecting emotions in text data. A meticulously curated dataset is utilized for the study. The research compares conventional models like Logistic Regression (LR), Random Forest (RF), Support Vector Machines (SVM), and Naive Bayes (NB) with deep learning models like Long Short-Term Memory (LSTM), Convolutional Neural Networks (CNN), and Bidirectional Encoder Representations from Transformers (BERT). The performance of each algorithm is assessed using accuracy, precision, recall, and F1 score. BERT exhibits superiority over other models, achieving the maximum accuracy of 0.8867 and F1 Score of 0.8871. CNN and SVM also display commendable performance. While the traditional models perform adequately, they are surpassed by deep learning models, with Naive Bayes showing the lowest metrics. This study underscores the significance of selecting models based on specific application requirements, taking into account factors like interpretability and efficiency. Future research endeavors may explore multimodal approaches, model interpretability, bias reduction, and real-time applications, thereby contributing to the advancement of emotion detection in text.
In sampling theory, the researchers are often dependent on estimators that use only current sample data to estimate population parameters. However, the hybrid exponentially weighted moving average (HEWMA) approach incorporates both current and past sample information and helps increasing the efficiency of the estimators. This enables us to develop an improved estimation procedure for temporal surveys based on HEWMA. We develop memory-type log estimator of population mean based on HEWMA under simple random sampling (SRS). We derive the bias and mean square error (MSE) of the developed estimator to the first-order approximation. The efficiency conditions are established by comparing the MSE of the proposed estimator with the MSE of the available traditional and memory-type estimators. To validate our theoretical findings, we conduct a simulation study utilizing hypothetically drawn population. A real data illustration of the developed methods is also presented. The findings demonstrate that our approach integrates past and present sample information and enhances the estimators’ efficacy.