Big data classification is a basic task in data mining in identifying the class labels for instances based on a set of features. The Naive Bayes classifier is one of the most commonly used methods for classification. Although the strong feature independence assumption in the Naive Bayes classifier makes it a tractable method for learning, this assumption may not hold in real-world applications. The correlated Naive Bayes classifier is a generalization of the Naive Bayes classification model considering the dependencies between features. In this paper, we propose a novel non-parametric Bayesian classification model called feature weighted correlative naive Bayes. In the first place, a kernel density method augments the weight of each training instance iteratively based on the estimated posterior probability. Next, we incorporate this probability into the conditional log-likelihood formula, and finally we optimize the weight of each feature value for each class by maximizing the conditional log-likelihood. Experiments have been conducted on the dataset posed learner has been compared with the other existing state-of-the-art competitors. The experimental results have demonstrated the effectiveness and efficiency of our proposed learning algorithm.
The finite mixture models provide a mathematical approach in the form of combining multiple probability distributions and creating a parametric framework for modeling unknown distributions. In some situations, the population is divided into several heterogeneous groups, and in each group, individuals may be somewhat homogeneous, so in this case, the fitting of the linear mixed model will not be appropriate. A suitable solution to solve this problem is to fitting a mixture of linear mixed models that take into account the heterogeneity in the population. In this paper, we propose mixture-mixed models and shrinkage methods to estimate its parameters. To do this, we use the penalized likelihood approach and propose an iterative shrinkage estimation approach. We call this used iterative algorithm iterative shrinkage weighted least square based on the least square. Iterative weighted least square algorithm has been used to estimate the parameters in the mixture model before. We conduct monte carlo simulations to study the performance of the estimators in terms of their simulated relative efficiencies.
Technical advances have enabled the conversion of agricultural waste into high-value-added products. However, the lack of proper recycling in developing countries leads to environmental pollution, public health risks, and lost economic opportunities. This survey examined farmers’ willingness to recycle agricultural waste in Kamyaran County, western Iran (n = 375 , analyzed using structural equation modeling). The theory of planned behavior (TPB) was extended by incorporating “moral norms” and “knowledge” as additional explanatory constructs. The results indicated that while farmers’ knowledge regarding waste recycling strategies and value-added products was limited, they exhibited a high willingness to recycle waste properly rather than discarding or burning it. The extended TPB model explained a substantial proportion of the variance in willingness to recycle (R 2 = 0.70 ). Attitude, subjective norms, and moral norms were identified as significant predictors of willingness. Furthermore, moral norms mediated the effect of subjective norms, whereas attitude mediated the relationship between knowledge and willingness. Based on these findings, we recommend leveraging educational programs and mass media to enhance farmers’ knowledge of the practical and economic benefits of recycling. Additionally, agricultural extension services should be utilized to cultivate positive attitudes and strengthen ethical and social norms toward sustainable waste management.
This article proposes a novel enhancement to semi-parametric models by integrating shrinkage to improve the predictive accuracy and performance of Ridge and ecpc models. By employing the EM algorithm, updated parameter estimates are derived, demonstrating significant error reductions in metrics such as mean squared error (MSE), sum of squared errors (SSE), and geometric mean squared error (GMSE). Extensive simulation studies reveal consistent improvements in model performance, with shrinkage enhancing accuracy and robustness across various scenarios. The findings highlight the potential of shrinkage in refining semi-parametric models, offering a more accurate framework for statistical predictions. This study paves the way for future research into applying shrinkage techniques to diverse data types and model structures, setting a benchmark for model optimization in statistical analysis.
In recent years, the diagnosis of diseases using artificial intelligence and machine learning algorithms has gained significant importance. Using data from relevant medical studies enables the extraction of valuable insights that can reduce the occurrence of numerous fatalities. One of the rapidly growing chronic diseases is diabetes, which has shown a growing prevalence due to urbanization and reduced physical activity. Hence, early detection of diabetes in individuals has immense significance. This paper utilizes a dataset comprising information from individuals who underwent diabetes diagnostic tests and employs classification techniques to determine whether their test results were positive or negative for diabetes. The novelty of this work lies in the comparative analysis of Bayesian classifiers and boosting methods, which have not been extensively explored in the literature. The utilized classification methods include Bayesian classifiers such as Bayesian Support Vector Machine, Bayesian k-nearest neighbor, Bayesian decision tree and boosting methods like Catboost, Adaboost, and XGboost. Performance evaluation metrics, including accuracy, precision, recall, F1-score, and ROC curve analysis, are employed to compare the efficacy of these methods in analyzing the data. The findings of this study will contribute to the advancement of accurate and efficient diabetes diagnosis using machine learning techniques, potentially helping in early intervention and management of the disease.
Poverty, an intricate global challenge influenced by economic, political, and social elements, is characterized by a deficiency in crucial resources, necessitating collective efforts towards its mitigation as embodied in the United Nations' Sustainable Development Goals. The Gini coefficient is a statistical instrument used by nations to measure income inequality, economic status, and social disparity, as escalated income inequality often parallels high poverty rates. Despite its standard annual computation, impeded by logistical hurdles and the gradual transformation of income inequality, we suggest that short-term forecasting of the Gini coefficient could offer instantaneous comprehension of shifts in income inequality during swift transitions, such as variances due to seasonal employment patterns in the expanding gig economy. System Identification (SI), a methodology utilized in domains like engineering and mathematical modeling to construct or refine dynamic system models from captured data, relies significantly on the Nonlinear Auto-Regressive (NAR) model due to its reliability and capability of integrating nonlinear functions, complemented by contemporary machine learning strategies and computational algorithms to approximate complex system dynamics to address these limitations. In this study, we introduce a NAR Multi-Layer Perceptron (MLP) approach for brief term estimation of the Gini coefficient. Several parameters were tested to discover the optimal model for Malaysia's Gini coefficient within 1987–2015, namely the output lag space, hidden units, and initial random seeds. The One-Step-Ahead (OSA), residual correlation, and residual histograms were used to test the validity of the model. The results demonstrate the model's efficacy over a 28-year period with superior model fit (MSE: 1.14 × 10−7) and uncorrelated residuals, thereby substantiating the model's validity and usefulness for predicting short-term variations in much smaller time steps compared to traditional manual approaches.
This paper provides a comprehensive overview of working memory, a crucial cognitive construct for learning, reasoning, and intellectual abilities. It introduces Baddeley and Hitch’s multi-component model, highlighting the construct’s essential role in facilitating learning, comprehension, problem-solving, and attention. The paper then analyzes four key models: Baddeley’s multicomponent framework, Cowan’s integrated memory network model, Engle’s attentional control-based model, and Oberauer’s tripartite structure, exploring their shared and differing perspectives. Recent studies on working memory using electroencephalogram are reviewed, identifying core frequency bands associated with cognitive states and predicting individual working memory capacity from electroencephalogram. Conventional working memory assessment methods are also discussed, emphasizing their relative advantages and limitations based on specific goals and practical considerations. Overall, this paper integrates theoretical models, neural correlates, and practical applications to provide a comprehensive overview of working memory research, highlighting its interdisciplinary nature for future studies.
This study proposes the use of Generative Adversarial Networks (GANs), specifically Lightweight GANs (LGANs), as a novel approach to revitalize the batik industry in Malaysia and Indonesia, which is currently experiencing a decline in interest among young artists. By automating the generation of innovative batik designs, this technology aims to bridge the gap between traditional craftsmanship and modern innovation, offering a significant opportunity for economic upliftment and skill development for the economically underprivileged B40 community. LGANs are chosen for their efficiency in training and their capability to produce high-quality outputs, making them particularly suited for creating intricate batik patterns. The research evaluates LGANs' effectiveness in generating novel batik designs, comparing the results with those of traditional manual methods. Findings suggest that LGANs are not only capable of producing distinctive and complex designs but also do so with greater efficiency and accuracy, demonstrating the potential of this technology to attract young artists and provide sustainable income opportunities for the B40 community. This study highlights the synergy between artificial intelligence and traditional artistry as a promising direction for revitalizing the batik industry, expanding its global reach, and preserving cultural heritage while fostering innovation and inclusivity.
This paper explores an audio-based on-road vehicle classification method that utilizes visual representations of sound through spectrograms, scalograms, and their fusion as features, classified using a modified VGG16 Convolutional Neural Network (CNN) architecture. The proposed method offers a non-intrusive, potentially less costly, and environmentally adaptable alternative to traditional sensor-based and computer vision techniques. Our results indicate that the fusion of scalogram and spectrogram features provides enhanced accuracy and reliability in distinguishing between vehicle types. Performance metrics such as training and loss, alongside precision and recall of classes, support the efficacy of a richer feature set in improving classification outcomes. The fusion features demonstrate a marked improvement in distinguishing closely related vehicle classes like 'Cars' and 'Trucks'. These findings underline the potential of our approach in refining and expanding vehicle classification systems for intelligent traffic monitoring and management.
The freshness of roasted coffee is a critical factor in determining its quality and flavor. Freshly roasted coffee has a distinct aroma and taste, with bright and lively flavors that are often lost in stale or old beans. However, roasted coffee beans are best taken until 14 days after it is roasted. This research attempts to establish whether Microwave Non-Destructive Testing (MNDT) can be used as a tool to determine the freshness of coffee. To achieve this, an intelligent method for determining the freshness of roasted coffee was based on MNDT data. The MNDT method collects s-parameter readings from roasted coffee every 5 days of interval by passing microwaves through them. The s-parameter readings will be fed to an Error-Correcting Output Coding Support Vector Machine (ECOC-SVM) to assess the degree of oxidation level in roasted coffee after several days it is being freshly roasted.
Graphical mixture models provide a powerful tool to visually depict conditional independencies or dependencies between heterogeneous high-dimensional data. When the random variables corresponding to the vertices are continuous, mixture component densities are assumed to be multivariate normal with different covariance matrices, leading to introduction of the Gaussian graphical mixture model (GGMM). The nonparanormal graphical mixture model (NGMM) replaces the restrictive normal assumption with a semiparametric Gaussian copula, which extends the nonparanormal graphical model and mixture models. Such an extension allows us to simultaneously estimate cluster assignments and cluster-specific graphical model structure. To enable such analyses, we propose a regularized estimation scheme with two forms of l1 penalty function (conventional and unconventional) via the expectation-maximization algorithm to learn a finite mixture of nonparanormal graphical models. We illustrate the performance of our method through a simulation study under both ideal and noisy settings. We also apply the proposed methodology to a breast cancer data set to diagnose malignant or benign tumors in patients. The results showed that the combination of NGMM together with unconventional penalty, which was named as NGMM1 during this study, provides the most efficient approach in clustering and estimating the graphical model structure.
Gaussian graphical model, which assumes that the variables of interest jointly follow a multivariate normal distribution with a sparse precision matrix, has been widely used to study the intrinsic dependence between the variables. However, we often encounter non-normal data, and we use variable transformation to achieve normality. Motivated by sparse additive models, the nonparanormal model extends the Gaussian graphical model to semiparametric Gaussian copula models, in which it is assumed that the variables follow a joint multivariate normal distribution after a set of unknown smooth monotone transformations. The basic methods of estimating undirected graphs in high-dimensional settings apply the simple random sampling (SRS) method to estimate the parameters of the model. We consider the ranked set sampling (RSS) approach for estimating the nonparanormal graphical model. Computationally, we show that using RSS leads to a better inference about the unknown features of the nonparanormal graphical model, and the proposed rank-based estimator performs as well as its counterparts estimators defined based on SRS. We study the numerical performance of the proposed method through a simulation study under both ideal and noisy settings and apply it to a genomic data set.
OBJECTIVE:Metabolic syndrome (MetS) is a complex multifactorial disorder that considerably burdens healthcare systems. We aim to classify MetS using regularized machine learning models in the presence of the risk variants of GCKR, BUD13 and APOA5, and environmental risk factors.MATERIALS AND METHODS:A cohort study was conducted on 2,346 cases and 2,203 controls from eligible Tehran Cardiometabolic Genetic Study (TCGS) participants whose data were collected from 1999 to 2017. We used different regularization approaches [least absolute shrinkage and selection operator (LASSO), ridge regression (RR), elasticnet (ENET), adaptive LASSO (aLASSO), and adaptive ENET (aENET)] and a classical logistic regression (LR) model to classify MetS and select influential variables that predict MetS. Demographics, clinical features, and common polymorphisms in the GCKR, BUD13 and APOA5 genes of eligible participants were assessed to classify TCGS participant status in MetS development. The models' performance was evaluated by 10-repeated 10-fold crossvalidation. Various assessment measures of sensitivity, specificity, classification accuracy, and area under the receiver operating characteristic curve (AUC-ROC) and AUC-precision-recall (AUC-PR) curves were used to compare the models.RESULTS:During the follow-up period, 50.38% of participants developed MetS. The groups were not similar in terms of baseline characteristics and risk variants. MetS was significantly associated with age, gender, schooling years, body mass index (BMI), and alternate alleles in all the risk variants, as indicated by LR. A comparison of accuracy, AUCROC, and AUC-PR metrics indicated that the regularization models outperformed LR. Regularized machine learning models provided comparable classification performances, whereas the aLASSO model was more parsimonious and selected fewer predictors.CONCLUSION:Regularized machine learning models provided more accurate and parsimonious MetS classifying models. These high-performing diagnostic models can lay the foundation for clinical decision support tools that use genetic and demographical variables to locate individuals at high risk for MetS.
BACKGROUND:The Naive Bayes (NB) classifier is a powerful supervised algorithm widely used in Machine Learning (ML). However, its effectiveness relies on a strict assumption of conditional independence, which is often violated in real-world scenarios. To address this limitation, various studies have explored extensions of NB that tackle the issue of non-conditional independence in the data. These approaches can be broadly categorized into two main categories: feature selection and structure expansion. In this particular study, we propose a novel approach to enhancing NB by introducing a latent variable as the parent of the attributes. We define this latent variable using a flexible technique called Bayesian Latent Class Analysis (BLCA). As a result, our final model combines the strengths of NB and BLCA, giving rise to what we refer to as NB-BLCA. By incorporating the latent variable, we aim to capture complex dependencies among the attributes and improve the overall performance of the classifier.METHODS:Both Expectation-Maximization (EM) algorithm and the Gibbs sampling approach were offered for parameter learning. A simulation study was conducted to evaluate the classification of the model in comparison with the ordinary NB model. In addition, real-world data related to 976 Gastric Cancer (GC) and 1189 Non-ulcer dyspepsia (NUD) patients was used to show the model's performance in an actual application. The validity of models was evaluated using the 10-fold cross-validation.RESULTS:The presented model was superior to ordinary NB in all the simulation scenarios according to higher classification sensitivity and specificity in test data. The NB-BLCA model using Gibbs sampling accuracy was 87.77 (95% CI: 84.87-90.29). This index was estimated at 77.22 (95% CI: 73.64-80.53) and 74.71 (95% CI: 71.02-78.15) for the NB-BLCA model using the EM algorithm and ordinary NB classifier, respectively.CONCLUSIONS:When considering the modification of the NB classifier, incorporating a latent component into the model offers numerous advantages, particularly within medical and health-related contexts. By doing so, the researchers can bypass the extensive search algorithm and structure learning required in the local learning and structure extension approach. The inclusion of latent class variables allows for the integration of all attributes during model construction. Consequently, the NB-BLCA model serves as a suitable alternative to conventional NB classifiers when the assumption of independence is violated, especially in domains pertaining to health and medicine.
In the literature on modeling heterogeneous data via mixture models, it is generally assumed that the samples are drawn from the underlying population using the simple random sampling (SRS) technique. This study exploits the bivariate ranked set sampling (BVRSS) technique to learn finite mixture models. We generalize the expectation-maximization (EM) algorithm under univariate RSS to the bivariate case. Computationally, through a simulation study under a noisy setting, we compare the performance of the proposed rank-based estimators with that of the SRS-based competitors in estimating unknown parameters and cluster assignments. The proposed methodology is applied to a breast cancer data set to diagnose malignant or benign tumors in patients. The results showed that the extra rank information in BVRSS samples leads to a better inference about the unknown features of mixture models.
This paper is concerned with estimating the risk measure, Value-at-Risk (VaR), without considering the usual hypothesis used in parametric methods. A non-parametric method is used to fit severity and frequency loss distributions in collective risk models. In addition, an optimum bandwidth is estimated. The model is then applied to insurance claims data from a particular insurance company. As a result of the new model, the outcomes show better accuracy, for both light-tailed and heavy-tailed distributions
The Hawkes process models have been recently become a popular tool for modeling and analysisof neural spike trains. In this article, motivated by neuronal spike trains study, we propose a novelmultivariate generalized linear Hawkes process model, where covariates are included in the intensityfunction. We consider the problem of simultaneous variable selection and estimation for the multivariategeneralized linear Hawkes process in the high-dimensional regime. Estimation of the intensity function ofthe high-dimensional point process is considered within a nonparametric framework, applying B-splinesand the SCAD penalty for matters of sparsity. We apply the Doob-Kolmogorov inequality and themartingale central limit theory to establish the consistency and asymptotic normality of the resultingestimators. Finally, we illustrate the performance of our proposal through simulation and demonstrateits utility by applying it to the neuron spike train data set.
This research describes an intelligent method for differentiating coffee roasting levels based on Microwave Non- Destructive Testing (MNDT) data. The MNDT method collects s-parameter readings from several types of coffee (dark, medium, and light roast) by passing microwaves through them. Error-Correcting Output Coding Support Vector Machine (ECOC-SVM) was fed a multi-layer perceptron neural network to assess the degree of different coffee roasts. With a small number of hidden units, the ECOC-SVM could identify between the various roasts (with 6,400 data points per sample).
One of the most widely used tools for quality engineers in quickly detecting assignable causes and monitoring processes is control charts. Duncan (in J Am Stat Assoc 51(274):228–242, 1956), in order to improve product quality and reduce the economic costs of the quality cycle, presented the first economic design of the $$ \bar{X} $$ control chart in the presence of multiple assignable causes. In his model and all the economic designs derived from it, it is assumed that after the occurrence of an assignable cause, no other assignable cause occurs until the correct alarm is issued. This assumption is unrealistic and impractical in production and service processes. Therefore, in this paper, we present a realistic and practical economic design in the presence of multiple assignable causes for the $$\bar{X}$$ control chart under the Weibull shock model in industry. The numerical results of our model show well that in the previous models, the average cost per unit time of the quality cycle is severely underestimated compared to the actual value. Therefore, it is suggested that in order to eliminate the shortcomings of the previous economic design in the presence of multiple assignable causes, in future research, they should be redesigned based on our proposed model.
Robust high dimensional estimation is one of the most important problems in statistics. In a high dimensional structure with a small number of non-zero observations, the dimension of the parameters is larger than the sample size. For modeling the sparsity of outlier response vector, we randomly selected a small number of observations and corrupted them arbitrarily. There are two distinct ways to overcome sparsity in the generalized linear model (GLM): in the parameter space, or in the space output. According to several studies in corrupted observation modeling, there is a relationship between robustness and sparsity. In this paper for obtaining robust high dimensional estimation, we proposed a finite mixture of the generalized linear models (FMGLMs). By using simulation with the expectation-maximization (EM) algorithm, we show improved modeling performance.