Missing data imputation is an essential preprocessing step in clinical survey data mining applications. Rough set imputation is one way to handle missing data. A major advantage of using rough sets is that only the information presented in the dataset itself is sufficient to perform the analysis. Hence, no additional information, external parameters, models, functions, grades, or subjective interpretations are necessary. While there are several studies on rough set data imputation, none has been conducted to measure the effect of such imputation on prediction. In this paper, we generate several simulation datasets based on an existing epidemiological dataset (MESA) to perform such study. To measure how well each dataset lends itself to the prediction model, we have used p-values from the Wald test. To evaluate the accuracy of the prediction, we have considered the width of 95% confidence interval for the probability of incontinence. Both imputed and non-imputed simulation datasets were fit to the prediction model and they both turned out to be significant (p-value <; 0.05). In addition, the Wald score shows a better fit for the imputed compared to non-imputed datasets. The average confidence interval width was decreased by 10.4% when the imputed dataset was used, i.e. higher precision was achieved. The results show that using the rough set method for missing data imputation on MESA data improves the predictive capability. Further studies are required to generalize this conclusion to other clinical survey datasets.
It is common for clinical data in survey trials to be incomplete and inconsistent for several reasons. Inconsistent data occur when more than one set of exclusive alternative questions are answered. One objective of this study was to identify and eliminate inconsistent data as an important data mining preprocessing step. We define three types of incomplete data: missing data due to skip pattern (SPMD), undetermined missing data (UMD), and genuine missing data (GMD). Identifying the type of missing data is another important objective as all missing data types cannot be treated the same. This goal cannot be achieved manually on large data of complex surveys since each subject should be processed individually. The analyses are accomplished in a mathematical framework by exploiting graph theoretic structure inherent in the questionnaire. An undirected graph is built using mutually inconsistent responses as well as its complement. The responses not in the largest maximal clique of complement graph are considered inconsistent. This guarantees removing as few responses as possible so that remaining ones are mutually consistent. Further, all potential paths in questionnaire’s graph are considered, based on the responses of subjects, to identify each type of incomplete data. Experiments are conducted on MESA data. Results show 15.4 % GMD , 9.8 % SPMD , 12.9 % UMD , and 0.021 % inconsistent data. Further utility of the approach is using a) the SPMD for data stratification, and b) inconsistent data for noise estimation. Proposed method is a preprocessing prerequisite for any data mining of clinical survey data.
Purpose Urinary incontinence (UI) is a chronic, costly condition that impairs quality of life. To identify older women most at risk, the Medical Epidemiologic and Social Aspects of Aging (MESA) datasets were mined to create a set of questions that can reliably predict future UI. Methods MESA data were collected during four household interviews at approximately 1 year intervals. Factors associated with becoming incontinent at the second interview (HH2) were identified using logistic regression (construction datasets). Based on p values and odds ratios, eight potential predictive factors with their 256 combinations and corresponding prediction probabilities formed the Continence Index. Its predictive and discriminatory capability was tested against the same cohort’s outcome in the fourth survey (HH4 validation datasets). Sensitivity analysis, area under receiver operating characteristic (ROC) curve, predicted probabilities and confidence intervals were used to statistically validate the Continence Index. Results Body mass index, sneezing, post-partum UI, urinary frequency, mild UI, belief of developing UI in the future, difficulty stopping urinary stream and remembering names emerged as the strongest predictors of UI. The confidence intervals for prediction probabilities strongly agreed between construction and validation datasets. Calculated sensitivity, specificity, false-positive and false-negative values revealed that the areas under the ROCs (0.802 and 0.799) for the construction and validation datasets, respectively, indicated good discriminatory capabilities of the index as a predictor. Conclusion The Continence Index will help identify older women most at risk of UI in order to apply targeted prevention strategies in women that are most likely to benefit.
Data stratification is the process of partitioning the data into distinct and non-overlapping groups since the study population consists of subpopulations that are of particular interest.In clinical data, once the data is stratified into sub populations based on a significant stratifying factor, different risk factors can be determined from each subpopulation.In this paper, the Fisher's Exact Test is used to determine the significant stratifying factors.The experiments are conducted on a simulated study and the Medical, Epidemiological and Social Aspects of Aging (MESA) data constructed for prediction of urinary incontinence.Results show that, smoking is the most significant stratifying factor of MESA data, showing that the smokers and non-smokers indicates different risk factors towards urinary incontinence and should be treated differently.
In this study, a Bayesian predictor of urinary incontinence (UI) is devised for screening older women. Risk factors identified from an epidemiological survey data as significant for UI, are utilized. The proposed Bayesian method combines an experimental design template with relevant information to construct a predictive index in terms of posterior probabilities. The computations are carried out on a longitudinal data called the Medical, Epidemiological and Social Aspects of Aging (MESA). The index is applied to the baseline and follow-up portions of the MESA data. The results show that, the percentage of the absolute relative change between the prior and posterior probabilities can be used as a decision tool to make conclusions on credibility of the class labels on continence and incontinence. The proposed index can be applied for immediate screening and for predicting future urinary incontinence in older women of comparable demographics as those presented in the MESA data.
Urinary Incontinence (UI) is a costly condition that decreases the quality of a patient's life and social engagement. Identification of UI risk factors may help early prevention and treatment of the condition. In this study we revisited the Medical, Epidemiological and Social Aspects of Aging (MESA) data collected in 1983. The experiments are conducted on a longitudinal dataset pertaining to the female-only population. A methodology that identifies skip patterns in order to facilitate MESA risk factor analysis is presented. The identified skip patterns are used to stratify MESA data. Based on the stratification performed, the important risk factors are then analyzed for each group of subjects. JRip rule extraction technique is utilized to determine the UI risk factors. Consequently, taking female hormones was determined as the most important stratifying feature. The dataset is then stratified to two subsets based on this stratifying feature. Education level, hearing problems, urine loss while coughing or sneezing, physical activity, stress and cancer are risk factors specific to taking female hormones. The common risk factors among both of the stratified groups were: stress, frequent sneezing, and low physical activity. Although there were common risk factors among both of the stratified groups these preliminary results show that different group of subjects have different risk factors, and therefore they should be provided with different UI predictive indices, diagnoses and possibly treatment plans.
A common problem in clinical survey trials is missing data. Skip patterns are one type of missing data in medical datasets, skipping a respondent over a group of questions that is not relevant to them. Applying any imputation technique to missing values caused by skip patterns may add misinformation. Moreover, skip pattern analysis provides detection of non-applicable data along with undetermined and inconsistent data. The Medical, Epidemiological and Social Aspects of Aging (MESA) questionnaire is responded by a large number of subjects which entails the need of an automated method. Manual methods may not provide reliable results and they are costly. A directed, acyclic graph is generated based on the questionnaire. A graph theory method is proposed to detect each missing data type. The method finds a minimal deletion set of nodes, that are the nodes once deleted, leaves a connected graph behind. The deleted nodes can be considered as noise. The experiments are conducted on a subset of the MESA data and the results show that there are 16.04% of non-applicable data, 7.09% of genuine missing data, 0.61% of undetermined data and 0.015% of inconsistent data. This method can be used for preprocessing the dataset and estimating the noise.
Mathematical analysis of existing data mining methods is not straightforward and in many cases it is not possible. Therefore, simulated data plays a central role in validation of data mining results in a given situation, i.e., noise, missing value and multicollinearity levels. This paper proposes a longitudinal binary data simulation focusing on presentation of the major challenge of infusing user-defined rules. Results of applying Apriori, PRAT, Prism, and JRip rule extraction methods on these simulated data in several missing value levels are presented in this paper. This simulation proved to be essential in verifying data mining results that we have generated on Medical Epidemiological and Social Aspects of Aging (MESA) data set.
Longitudinal data for studying urinary incontinence (UI) risk factors are rare. Data from one study, the hallmark Medical, Epidemiological, and Social Aspects of Aging (MESA), have been analyzed in the past; however, repeated measures analyses that are crucial for analyzing longitudinal data have not been applied. We tested a novel application of statistical methods to identify UI risk factors in older women. MESA data were collected at baseline and yearly from a sample of 1955 men and women in the community. Only women responding to the 762 baseline and 559 follow-up questions at one year in each respective survey were examined. To test their utility in mining large data sets, and as a preliminary step to creating a predictive index for developing UI, logistic regression, generalized estimating equations (GEEs), and proportional hazard regression (PHREG) methods were used on the existing MESA data. The GEE and PHREG combination identified 15 significant risk factors associated with developing UI out of which six of them, namely, urinary frequency, urgency, any urine loss, urine loss after emptying, subject's anticipation, and doctor's proactivity, are found most highly significant by both methods. These six factors are potential candidates for constructing a future UI predictive index.
Data mining is the discipline of systematically reviewing datasets to determine what patterns, concurrences and/or rule sets can be discovered. In an effort to understand the effectiveness of a proposed mining methodology a well defined dataset is required for model verification. To this end, a fully-controlled simulation was created by which several different feature selection algorithms were evaluated. In this paper, we present a comprehensive comparison between attribute selection methods when noise, missing values and multicollinearity are in question. Our results show that "Relief" and "information gain" have outperformed other feature selection methods available in Weka when considering both sensitivity and specificity measures. We have evaluated the following features selection methods: J48, Relief, information gain, consistency based feature selection and correlation based feature selection to see which one handles additive noise better. The sensitivity of consistency based feature selection was 11% higher than the average sensitivity of other methods. However, it's specificity was 37% lower than that of the average. It is important to note that sensitivity or specificity alone does not give enough support to a method to say that it is the best way to handle the data. The best method, when both sensitivity and specificity are considered, was information gain. This method outperformed the average of other methods by 1% and 20.5% when we considered its sensitivity and specificity, respectively. Also in this regard, a goal of our study was to see which feature selection method outperforms within the missing set of values. In this case, when looking again at sensitivity and specificities, Relief and information gain proved to outperform the other methods of our study by 7.2% and 12.4%, respectively. Our studies also show that when multicollinearity is embedded into the fully controlled dataset without any noise and missing values, the correlation based feature selection outperforms other methods. In summary Relief and information gain performed the best in all three situations in terms of its sensitivity and specificities.
Minimal square designs are proposed and compared. All treatment contrasts in both designs are estimable under the existence of two-way heterogeneity. That is, all designs are treatment-connected. Extended treatment-connected designs are generated by adding one column to minimal treatment-connected square designs. The extended designs not only have lower variances in paired comparisons of unreplicated treatments but also provide necessary degrees of freedom to estimate the process error. (M,S)-optimal extended designs are constructed systematically. Both square designs and their extensions have large numbers of unreplicated treatments.
A new class of row–column designs is proposed. These designs are saturated in terms of eliminating two-way heterogeneity with an additive model. The (m,s)-criterion is used to select optimal designs. It turns out that all (m,s)-optimal designs are binary. Square (m,s)-optimal designs are constructed and they are treatment-connected. Thus, all treatment contrasts are estimable regardless of the row and column effects.
In this paper, we determine a nonbinary s-optimal design over a class of minimally connected binary row-column designs.
Use of the (M,S) criterion to select and classify factorial designs is proposed and studied. The criterion is easy to deal with computationally and it is independent of the choice of treatment contrasts. It can be applied to two-level designs as well as multi-level symmetrical and asymmetrical designs. An important connection between the (M,S) and minimum aberration criteria is derived for regular fractional factorial designs. Relations between the (M,S) criterion and generalized minimum aberration criteria on nonregular designs are also discussed. The (M,S) criterion is then applied to study the projective properties of some nonregular designs.
A new class of binary, square designs is proposed. The proposed designs are saturated in terms of eliminating two-way heterogeneity. It is proved that the proposed binary, square designs are not only row-treatment and column-treatment connected but also treatment-connected.
In this note, we determine s- and (m,s)-optimal designs in the class of all minimal incomplete block designs under the mixed effects model.
Claudia D'Amato合作论文数Dipartimento di Informatica;Universita degli Studi di Bari1
Paul Bradley合作论文数ZirMed1