
This work indicates application of the notion of fuzzy $$\alpha $$ -cut in rough set theory and studies the properties of Gödel-like arrow in details.
Biomedical named entity recognition is becoming increasingly important to biomedical research due to a proliferation of articles and also due to the current pandemic disease. This paper addresses the task of automatically finding and recognizing biomedical entity types related to COVID (e.g., virus, cell, therapeutic) with tolerance rough sets. The task includes i) extracting nouns and their co-occurring contextual patterns from a large BioNER dataset related to COVID-19 and, ii) annotating unlabelled data with a semi-supervised learning algorithm using co-occurence statistics. 465,250 noun phrases and 6,222,196 contextual patterns were extracted from 29,500 articles using natural language text processing methods. Three categories were successfully classified at this time: virus, cell and therapeutic. Early precision@N results demonstrate that our proposed tolerant pattern learner (TPL) is able to constrain concept drift in all 3 categories during the iterative learning process.
The theory of rough sets has been studied extensively, both from foundation and application points of view, since its introduction by Pawlak in 1982. On the foundations side, a substantial part of work on rough set theory involves the study of its algebraic aspects and logics. The present work is in this direction, initiated through the study of categories of rough sets. Starting from two categories RSC and ROUGH of rough sets, it is shown that they are equivalent. Moreover, RSC, and thus ROUGH, are found to be a quasitopos, a structure slightly weaker than topos. The construction is then lifted to a more general set-up to give the category RSC( $$\mathscr {C}$$ ) with an arbitrary non-degenerate topos $$\mathscr {C}$$ serving as a ‘base’, just as sets constitute a base for defining rough sets. The category-theoretic study gives rise to two directions of work. In one direction of work, a particular example of RSC( $$\mathscr {C}$$ ) when $$\mathscr {C}$$ is the topos of monoid actions on sets is considered. It yields the monoid actions on rough sets and that of transformation semigroups (ts) for rough sets, leading to decomposition results. A semiautomaton for rough sets is also defined. In the other direction, we incorporate Iwinski’s notion of ‘relative rough complementation’ in the internal algebra of the quasitopos RSC( $$\mathscr {C}$$ ). This results in the introduction of two new classes of algebraic structures with two negations, namely contrapositionally complemented pseudo-Boolean algebra (ccpBa) and contrapositionally $${\vee }$$ complemented pseudo-Boolean algebra (c $$\vee $$ cpBa). Examples of ccpBas and c $$\vee $$ cpBas are developed, comparison with existing algebras is done and representation theorems are established. The logics ILM and ILM- $$\vee $$ corresponding to ccpBas and c $${\vee }$$ cpBas respectively are defined, and different relational semantics are obtained. It is shown that ILM is a proper extension of a variant JP $$'$$ of Peirce’s logic, defined by Segerberg in 1968. The inter-relationship between relational semantics and the algebraic semantics of ILM and ILM- $${\vee }$$ are investigated. Lastly, in the line of Dunn’s study of logics, the two negations are expressed without the help of the connective of implication, and the resulting logical and algebraic structures are also studied.
In this article an attempt is made to introduce and study the notion of single-valued neutrosophic rough continuous mapping, single-valued neutrosophic rough compactness via single-valued neutrosophic rough topological spaces (SVNRTS). By defining the concept of single-valued neutrosophic rough continuous function, single-valued neutrosophic rough compactness, we formulate and discuss several interesting results on SVNRTSs.
This paper provides an overview of the Rough Set Database System (the RSDS for short) for creating bibliographies on rough sets and related fields, as well as sharing and analysis. The current version of the RSDS includes a number of modifications, extensions and functional improvements compared to the previous versions of this system. The system was made in the client-server technology. Currently, the RSDS contains over 38 540 entries from nearly 42 860 authors. This system works on any computer connected to the Internet and is available at http://rsds.ur.edu.pl .
The concept of rational discourse is typically determined by subjective, normative, and rule based constraints in the context under consideration. It is typically determined by related ontologies, and coherence between associated concepts employed in the discourse. Classical rough approximations, and variants of variable precision rough sets (VPRS) including graded rough sets embody at least some aspects of potentially useful concepts of rational approximation, but can be very lacking in application contexts, and rough set theoretical frameworks for cluster validation. While the literature on knowledge from general rough perspectives is rich and diverse, not much work has been done from the perspective of rationality in explicit terms. In this research, the gap is addressed by the present author in variants of high granular partial algebras. Specifically, the nature of optimal concepts of rational approximations is examined, and formalized by her in such frameworks. Graded rough sets are generalized from a granular perspective, and the compatibility of the introduced concepts are studied over it. Further aspects of algebraic semantics of granular graded rough sets are examined. Some incorrect results in graded rough sets in the literature are also corrected.
In the presented study, the problem of interactive feature extraction, i.e., supported by interaction with users, is discussed, and several innovative approaches to automating feature creation and selection are proposed. The current state of knowledge on feature extraction processes in commercial applications is shown. The problems associated with processing big data sets as well as approaches to process high-dimensional time series are discussed. The introduced feature extraction methods were subjected to experimental verification on real-life problems and data. Besides the experimentation, the practical case studies and applications of developed techniques in selected scientific projects are shown. Feature extraction addresses the problem of finding the most compact and informative data representation resulting in improved efficiency of data storage and processing, facilitating the subsequent learning and generalization steps. Feature extraction not only simplifies the data representation but also enables the acquisition of features that can be further easily utilized by both analysts and learning algorithms. In its most common flow, the process starts from an initial set of measured data and builds derived features intended to be informative and non-redundant. Logically, there are two phases of this process: the first is the construction of the new attributes based on original data (sometimes referred to as feature engineering), the second is a selection of the most important among the attributes (sometimes referred to as feature selection). There are many approaches to feature creation and selection that are well-described in the literature. Still, it is hard to find methods facilitating interaction with users, which would take into consideration users' knowledge about the domain, their experience, and preferences. In the study on the interactiveness of the feature extraction, the problems of deriving useful and understandable attributes from raw sensor readings and reducing the amount of those attributes to achieve possibly simplest, yet accurate, models are addressed. The proposed methods go beyond the current standards by enabling a more efficient way to express the domain knowledge associated with the most important subsets of attributes. The proposed algorithms for the construction and selection of features can use various forms of information granulation, problem decomposition, and parallelization. They can also tackle large spaces of derivable features and ensure a satisfactory (according to a given criterion) level of information about the target variable (decision), even after removing a substantial number of features. The proposed approaches have been developed based on the experience gained in the course of several research projects in the fields of data analysis and processing multi-sensor data streams. The methods have been validated in terms of the quality of the extracted features, as well as throughput, scalability, and robustness of their operation. The discussed methodology has been verified in open data mining competitions to confirm its usefulness.
The main focus of this paper is to introduce some aggregation operators namely, Rough-Bipolar Neutrosophic Arithmetic Mean (RBNAM) operator and Rough-Bipolar Neutrosophic Geometric Mean (RBNGM) operator under Rough-Bipolar Neutrosophic Set (RBNS) environment. Besides, we present the concept of score and accuracy functions under the RBNS environment. Further, we propose two multi-attribute decision-making (MADM) strategies based on RBNAM operator and RBNGM operator respectively under the RBNS environment. Finally, we provide a real-life numerical example to validate the proposed MADM strategy.
The authors trace their journey with rough sets since their first interactions with Z. Pawlak. The article is a narration of how work on rough sets was initiated in India, and how it continues to thrive in research groups connected to the authors and others in the country.
Technology improves every day. In order for an established theory to maintain relevance, implementations of such theory must be updated to take advantage of the new improvements. ROSETTA, a framework based on Rough Set theory, was developed in 1994 to exploit Rough Set paradigms in Machine Learning. Since then, much has happened in the field of Computer Technology, and to fully exploit these benefits ROSETTA needed to evolve. We designed and implemented a multi-core execution process in ROSETTA, optimized for speed and modular extension. The program was tested using four datasets of different sizes for computational speed and memory usage, the factors considered the primary limitations of classification and Machine Learning. The results show an increase in computation speed consistent with expected gains. The scaling per thread of memory usage was less than linear after five threads with increases in memory based primarily on the number of objects in the dataset. The number of features in the data increased the base memory needed but did not significantly impact the memory scaling by threads. The multi-core implementation was successful, and ||-ROSETTA (pronounced Parallel-ROSETTA) is capable of fully exploiting modern hardware solutions.
The seminal work of Z. Pawlak [60] on rough set theory has attracted the attention of researchers from various disciplines. Algebraists introduced some new algebraic structures and represented some old existing algebraic structures in terms of algebras formed by rough sets. In Logic, the rough set theory serves the models of several logics. This paper is an amalgamation of algebras and logics of rough set theory. We prove a structural theorem for Kleene algebras, showing that an element of a Kleene algebra can be looked upon as a rough set in some appropriate approximation space. The proposed propositional logic $$\mathcal {L}_{K}$$ of Kleene algebras is sound and complete with respect to a 3-valued and a rough set semantics. This article also investigates some negation operators in classical rough set theory, using Dunn's approach. We investigate the semantics of the Stone negation in perp frames, that of dual Stone negation in exhaustive frames, and that of Stone and dual Stone negations with the regularity property in $$K_{-}$$ frames. The study leads to new semantics for the logics corresponding to the classes of Stone algebras, dual Stone algebras, and regular double Stone algebras. As the perp semantics provides a Kripke type semantics for logics with negations, exploiting this feature, we obtain duality results for several classes of algebras and corresponding frames. In another part of this article, we propose a granule-based generalization of rough set theory. We obtain representations of distributive lattices (with operators) and Heyting algebras (with operators). Moreover, various negations appear from this generalized rough set theory and achieved new positions in Dunn's Kite of negations.
We study decision trees as a means of representation of knowledge. To this end, we design two techniques for the creation of CART (Classification and Regression Tree)-like decision trees that are based on bi-objective optimization algorithms. We investigate three parameters of the decision trees constructed by these techniques: number of vertices, global misclassification rate, and local misclassification rate.
Pawlakian spaces rely on an equivalence relation which represent indiscernibility. As a generalization of these spaces, some approximation spaces have appeared that are not based on an equivalence relation but on a tolerance relation that represents similarity. These spaces preserve the property of the Pawlakian space that the union of the base sets gives out the universe. However, they give up the requirement that the base sets are pairwise disjoint. The base sets are generated in a way where for each object, the objects that are similar to the given object, are taken. This means that the similarity to a given object is considered. In the worst case, it can happen that the number of base sets equals those of objects in the universe. This significantly increases the computational resource need of the set approximation process and limits the efficient use of them in large databases. To overcome this problem, a possible solution is presented in this dissertation. The space is called similarity-based rough sets where the system of base sets is generated by the correlation clustering. Therefore, the real similarity is taken into consideration not the similarity to a distinguished object. The space generated this way, on the one hand, represents the interpreted similarity properly and on the other hand, reduces the number of base sets to a manageable size. This work deals with the properties and applicability of this space, presenting all the advantages that can be gained from the correlation clustering.
The paper presents a new generation of Rseslib library - a collection of rough set and machine learning algorithms and data structures in Java. It provides algorithms for discretization, discernibility matrix, reducts, decision rules and for other concepts of rough set theory and other data mining methods. The third version was implemented from scratch and in contrast to its predecessor it is available as a separate open-source library with API and with modular architecture aimed at high reusability and substitutability of its components. The new version can be used within Weka and with a dedicated graphical interface. Computations in Rseslib 3 can be also distributed over a network of computers.
This article presents similarity based reasoning approach for recognition of compound objects. It contains mathematical foundations for comparators theory as well as comparators network theory. It shows also three different practical applications in field of image recognition, text recognition and risk recognition.
In one perspective, the main theme of this research revolves around the inverse problem in the context of general rough sets that concerns the existence of rough basis for given approximations in a context. Granular operator spaces and variants were recently introduced by the present author as an optimal framework for anti-chain based algebraic semantics of general rough sets and the inverse problem. In the framework, various sub-types of crisp and non-crisp objects are identifiable that may be missed in more restrictive formalism. This is also because in the latter cases concepts of complementation and negation are taken for granted - while in reality they have a complicated dialectical basis. This motivates a general approach to dialectical rough sets building on previous work of the present author and figures of opposition. In this paper dialectical rough logics are invented from a semantic perspective, a concept of dialectical predicates is formalised, connection with dialetheias and glutty negation are established, parthood analyzed and studied from the viewpoint of classical and dialectical figures of opposition by the present author. Her methods become more geometrical and encompass parthood as a primary relation (as opposed to roughly equivalent objects) for algebraic semantics.
We examine double successive approximations on a set, which we denote by L_2L_1, U_2U_1, U_2L_1, L_2U_1 where L_1, U_1 and L_2, U_2 are based on generally non-equivalent equivalence relations E_1 and E_2 respectively, on a finite non-empty set V. We consider the case of these operators being given fully defined on its powerset 𝒫(V). Then, we investigate if we can reconstruct the equivalence relations which they may be based on. Directly related to this, is the question of whether there are unique solutions for a given defined operator and the existence of conditions which may characterise this. We find and prove these characterising conditions that equivalence relation pairs should satisfy in order to generate unique such operators.
This article presents an approach to performing the task of visual search in the context of descriptive topological spaces. The presented algorithm forms the basis of a descriptive visual search system (DVSS) that is based on the guided search model (GSM) that is motivated by human visual search. This model, in turn, consists of the bottom-up and top-down attention models and is implemented within the DVSS in three distinct stages. First, the bottom-up activation process is used to generate saliency maps and to identify salient objects. Second, perceptual objects, defined in the context of descriptive topological spaces, are identified and associated with feature vectors obtained from a VGG deep learning convolutional neural network. Lastly, the top-down activation process makes decisions on whether the object of interest is present in a given image through the use of descriptive patterns within the context of a descriptive topological space. The presented approach is tested with images from the ImageNet ILSVRC2012 and SIMPLIcity datasets. The contribution of this article is a descriptive pattern-based visual search algorithm.