Predictive accuracy, as an estimation of a classifier’s future performance, has been studied for at least seventy years. With the advent of the modern computer era, techniques that may have been previously impractical are now calculable within a reasonable time frame. Within this chapter, three techniques of resampling, namely, leave-one-out, k-fold cross validation and bootstrapping; are investigated as methods of error rate estimation with application to variable precision rough set theory (VPRS). A prototype expert system is utilised to explore the nature of each resampling technique when VPRS is applied to an example dataset. The software produces a series of graphs and descriptive statistics, which are used to illustrate the characteristics of each technique with regards to VPRS, and comparisons are drawn between the results.
This chapter considers, and elucidates, the general methodology of rough set theory (RST), a nascent approach to rule based classification associated with soft computing. There are two parts of the elucidation undertaken in this chapter, firstly the levels of possible pre-processing necessary when undertaking an RST based analysis, and secondly the presentation of an analysis using variable precision rough sets (VPRS), a development on the original RST that allows for misclassification to exist in the constructed “if … then …” decision rules. Throughout the chapter, bespoke software underpins the pre-processing and VPRS analysis undertaken, including screenshots of its output. The problem of US bank credit ratings allows the pertinent demonstration of the soft computing approaches described throughout.
Rough Set Theory (RST), since its introduction in Pawlak (1982), continues to develop as an effective tool in data mining. Within a set theoretical structure, its remit is closely concerned with the classification of objects to decision attribute values, based on their description by a number of condition attributes. With regards to RST, this classification is through the construction of ‘if .. then ..’ decision rules. The development of RST has been in many directions, amongst the earliest was with the allowance for miss-classification in the constructed decision rules, namely the Variable Precision Rough Sets model (VPRS) (Ziarko, 1993), the recent references for this include; Beynon (2001), Mi et al. (2004), and Slezak and Ziarko (2005). Further developments of RST have included; its operation within a fuzzy environment (Greco et al., 2006), and using a dominance relation based approach (Greco et al., 2004). The regular major international conferences of ‘International Conference on Rough Sets and Current Trends in Computing’ (RSCTC, 2004) and ‘International Conference on Rough Sets, Fuzzy Sets, Data Mining and Granular Computing’ (RSFDGrC, 2005) continue to include RST research covering the varying directions of its development. This is true also for the associated book series entitled ‘Transactions on Rough Sets’ (Peters and Skowron, 2005), which further includes doctoral theses on this subject. What is true, is that RST is still evolving, with the eclectic attitude to its development meaning that the definitive concomitant RST data mining techniques are still to be realised. Grzymala-Busse and Ziarko (2000), in a defence of RST, discussed a number of points relevant to data mining, and also made comparisons between RST and other techniques. Within the area of data mining and the desire to identify relationships between condition attributes, the effectiveness of RST is particularly pertinent due to the inherent intent within RST type methodologies for data reduction and feature selection (Jensen and Shen, 2005). That is, subsets of condition attributes identified that perform the same role as all the condition attributes in a considered data set (termed ß-reducts in VPRS, see later). Chen (2001) addresses this, when discussing the original RST, they state it follows a reductionist approach and is lenient to inconsistent data (contradicting condition attributes - one aspect of underlying uncertainty). This encyclopaedia article describes and demonstrates the practical application of a RST type methodology in data mining, namely VPRS, using nascent software initially described in Griffiths and Beynon (2005). The use of VPRS, through its relative simplistic structure, outlines many of the rudiments of RST based methodologies. The software utilised is oriented towards ‘hands on’ data mining, with graphs presented that clearly elucidate ‘veins’ of possible information identified from ß-reducts, over different allowed levels of missclassification associated with the constructed decision rules (Beynon and Griffiths, 2004). Further findings are briefly reported when undertaking VPRS in a resampling environment, with leave-one-out and bootstrapping approaches adopted (Wisnowski et al., 2003). The importance of these results is in the identification of the more influential condition attributes, pertinent to accruing the most effective data mining results.
This dissertation considers, the Variable Precision Rough Sets (VPRS) model, and its development within a comprehensive software package (decision support system), incorporating methods of re sampling and classifier aggregation. The concept of /-reduct aggregation is introduced, as a novel approach to classifier aggregation within the VPRS framework. The software is applied to the credit rating prediction problem, in particularly, a full exposition of the prediction and classification of Fitch's Individual Bank Strength Ratings (FIBRs), to a number of banks from around the world is presented. The ethos of the developed software was to rely heavily on a simple 'point and click' interface, designed to make a VPRS analysis accessible to an analyst, who is not necessarily an expert in the field of VPRS or decision rule based systems. The development of the software has also benefited from consultations with managers from one of Europe's leading hedge funds, who gave valuable insight, advice and recommendations on what they considered as pertinent issues with regards to data mining, and what they would like to see from a modern data mining system. The elements within the developed software reflect each stage of the knowledge discovery process, namely, pre-processing, feature selection, data mining, interpretation and evaluation. The developed software encompasses three software packages, a pre-processing package incorporating some of the latest pre-processing and feature selection methods a VPRS data mining package, based on a novel interface, which presents the analyst with selectable /-reducts over the domain of / and a third more advanced VPRS data mining package, which essentially automates the vein graph interface for incorporation into a re-sampling environment, and also implements the introduced aggregated /-reduct, developed to optimise and stabilise the predictive accuracy of a set of decision rules induced from the aggregated /-reduct.
The variable precision rough sets model (VPRS) along with many derivatives of rough set theory (RST) necessitates a number of stages towards the final classification of objects. These include, (i) the identification of subsets of condition attributes (@b-reducts in VPRS) which have the same quality of classification as the whole set, (ii) the construction of sets of decision rules associated with the reducts and (iii) the classification of the individual objects by the decision rules. The expert system exposited here offers a decision maker (DM) the opportunity to fully view each of these stages, subsequently empowering an analyst to make choices during the analysis. Its particular innovation is the ability to visually present available @b-reducts, from which the DM can make their selection, a consequence of their own reasons or expectations of the analysis undertaken. The practical analysis considered here is applied on a real world application, the credit ratings of large banks and investment companies in Europe and North America. The snapshots of the expert system presented illustrate the variation in results from the 'asymmetric' consequences of the choice of @b-reducts considered.
The variable precision rough sets model (VPRS) is a development of the original rough set theory (RST) and allows for the partial (probabilistic) classification of objects. This paper introduces a prototype VPRS expert system. A number of processes for the identification of the VPRS related ß-reducts and their respective ß intervals over the domain of the ß parameter are included. Three data sets are utilised in the exposition of the expert system.
This paper concerns the study of destination choice modelling, more specifically identifying within some area (e.g. a city) the region where a particular store is the most favourable to be visited by individuals. An influence measure is constructed for each individual, which incorporates the modern technique known as Dempster–Shafer theory. Based on the evidence of the shopping destinations of individuals, geographical regions are found for levels of largest belief and plausibility (within Dempster–Shafer theory) for specific stores being the most favourable to visit. Additionally, this method may be used to identify the possible position of new stores, based on regions of most uncertainty or conflict in store choice. A prototype choice modelling system is introduced to enable the series of associated results to be easily visualized and analysed.