Accuracy based Learning Classifier System (XCS) prefers to generalize the classifiers that always acquire the same reward, because they make accurate reward predictions. However, real-world problems have noise, which means that classifiers may not receive the same reward even if they always take the correct action. For this case, since all classifiers acquire multiple values as the reward, XCS cannot identify accurate classifiers. In this paper, we study a single step environment with action noise, where XCS's action is sometimes changed at random. To overcome this problem, this paper proposes XCS based on Collective weighted Reward (XCS-CR) to identify the accurate classifiers. In XCS each rule predicts its next reward by averaging its past rewards. Instead, XCS-CR predicts its next reward by selecting a reward from the set of past rewards, by comparing the past rewards to the collective weighted average reward of the rules matching the current input for each action. This comparison helps XCS-CR identify rewards that result from action noise. In experiments, XCS-CR acquired the optimal generalized classifier subset in 6-Multiplexer problems with action noise, similar to the environment without noise, and judged those optimal generalized classifiers correctly accurate.
We proposed XCS-VRc 3 that can extract useful rules (classifiers) from data and verify its effectiveness. The difficulty of mining real world data is that not only the type of the input state but also the number of instances varies. Although conventional method XCS-VRc is able to extract classifiers, the generalization of classifiers was insufficient and lack of human readability. The proposed XCS-VRc 3 incorporating "generalization mechanism by comprehensive classifier subsumption" to solves this problem. Specifically, (1) All classifiers of the matching set subsume other classifiers, (2) Abolition of the inappropriate classifier deletion introduced by XCS-VRc (3) Preferentially select classifier with small variance of output in genetic algorithm. To verify the effectiveness of XCS-VRc 3 , we applied on care plan planning problem in a nursing home (in this case, identifying daytime behavior contributing to increase the ratio of deep sleep time). Comparing the association rules obtained by Apriori, and classifiers obtained by XCS-VRc 3 , the followings was found. First, abolishing the inappropriate classifier deletion and comprehensively subsuming promotes various degrees of generalization. Second, parent selection mechanism can obtain classifiers with small output variance. Finally, XCS-VRc 3 is able to extract a small number classifiers equivalent to large number of rules found in Apriori.
This paper introduces a reinforcement learning technique with an internal reward for a multi-agent cooperation task. The proposed methods is an extension of Q-learning which changes the ordinary (external) reward to the internal reward for agent-cooperation. Specifically, we propose here two Q-learning methods, both of which employ the internal reward for the less or no communication. To guarantee the effectiveness of the proposed methods, we theoretically derived the mechanisms that solve the following questions: (1) how the internal rewards should be set to guarantee the cooperation among the agents under the condition of less and no communication; and (2) how the values of the cooperative behaviors types (i.e., the varieties of the cooperative behaviors of the agents) should be updated under the condition of no communication. The intensive simulations on the maze problem for the agent-cooperation task have been revealed that our two proposed methods successfully enable the agents to acquire their cooperative behaviors even in less or no communication, while the conventional method (Q-learning) always fails to acquire such behaviors.
This paper proposes the novel Learning Classifier System (LCS) which can solve high-dimensional problems, and obtain human-readable knowledge by integrating deep neural networks as a compressor. In the proposed system named DCAXCSR, deep neural network called Deep Classification Autoencoder (DCA) compresses (encodes) input to lower dimension information which LCS can deal with, and decompresses (decodes) output of LCS to the original dimension information. DCA is hybrid network of classification network and autoencoder towards increasing compression rate. If the learning is insufficient due to lost information by compression, by using decoded information as an initial value for narrowing down state space, LCS can solve high dimensional problems directly. As LCS of the proposed system, we employs XCSR which is LCS for real value in this paper since DCA compresses input to real values. In order to investigate the effectiveness of the proposed system, this paper conducts experiments on the benchmark classification problem of MNIST database and Multiplexer problems. The result of the experiments shows that the proposed system can solve high-dimensional problems which conventional XCSR cannot solve, and can obtain human-readable knowledge.
This paper proposes high-dimensional data mining technique by integrating two data mining methods: Accuracy-based Learning Classifier Systems (XCS) and Random Forests (RF). Concretely the proposed system integrates RF and XCS: RF generates several numbers of decision trees, and XCS generalizes the rules converted from the decision trees. The convert manner is as follows: (1) the branch node of the decision tree becomes the attribute; (2) if the branch node does not exist, the attribute of that becomes # for XCS; and (3) One decision tree becomes one rule at least. Note that # can become any value in the attribute. From the experiments of Multiplexer problems, we derive that: (i) the good performance of the proposed system; and (ii) RF helps XCS to acquire optimal solutions as knowledge by generating appropriately generalized rules.
This paper focuses on a generalization of classifiers in noisy problems and aims at exploring learning classifier systems (LCSs) that can evolve accurately generalized classifiers as an optimal solution in several environments which include different type of noise. For this purpose, this paper employs XCS-CRE (XCS without Convergence of Reward Estimation) which can correctly identify classifiers as either accurate or inaccurate ones even in a noisy problem, and investigates its effectiveness in several noisy problems. Through intensive experiments of three LCSs (i.e., XCS as the conventional LCS, XCS-SAC (XCS with Self-adaptive Accuracy Criterion) as our previous LCS, and XCS-CRE) on the noisy 11-multiplexer problem where reward value changes according to (a) Gaussian distribution, (b) Cauchy distribution, or (c) Lognormal distribution, the following implications have been revealed: (1) the correct rate of the classifier of XCS-CRE and XCS-SAC converge to 100% in all three types of the reward distribution while that of XCS cannot reach 100%; (2) the population size of XCS-CRE is smallest followed by that of XCS-SAC and XCS; and (3) the percentage of the acquired optimal classifiers of XCS-CRE is highest followed by that of XCS-SAC and XCS.
The correctness rate of classification of neural networks is improved by deep learning, which is machine learning of neural networks, and its accuracy is higher than the human brain in some fields. This paper proposes the hybrid system of the neural network and the Learning Classifier System (LCS). LCS is evolutionary rule-based machine learning using reinforcement learning. To increase the correctness rate of classification, we combine the neural network and the LCS. This paper conducted benchmark experiments to verify the proposed system. The experiment revealed that: 1) the correctness rate of classification of the proposed system is higher than the conventional LCS (XCSR) and normal neural network; and 2) the covering mechanism of XCSR raises the correctness rate of proposed system.
This paper aims at investigating how correct or incorrect opinions are shared among the agents in the weighted network where the relationship among the agent (as nodes of its network) is different each other, and exploring how the agents can be promoted to share only correct opinions by preventing to acquire the incorrect opinions in the weighted network. For this purpose, this paper focuses on Autonomous Adaptive Tuning algorithm (AAT) which can improve an accuracy of correct opinion shared among agents in the various network, and improves it to address the situation which is close in the real world, i.e., the relationship among agents is different each other. This is because the original AAT does not consider such a different relationship among the agents. Through the intensive empirical experiments, the following implications have been revealed: (1) the accuracy of the correct opinion sharing with the improved AAT is higher than that with the original AAT in the weighted network; (2) the agents in the improved AAT can prevent to acquire incorrect opinion sharing in the weighted network, while those in the original AAT are hard to prevent in the same network.
Background Anxiety and mood disorders are the most common mental illnesses, peaking during adolescence and affecting approximately 25% of Canadians aged 14-17 years. If not successfully treated at this age, they often persist into adulthood, exerting a great social and economic toll. Given the long-term impact, finding ways to increase the success and cost-effectiveness of mental health care is a pressing need. Cognitive behavior therapy (CBT) is an evidence-based treatment for mood and anxiety disorders throughout the lifespan. Mental health technologies can be used to make such treatments more successful by delivering them in a format that increases utilization. Young people embrace technologies, and many want to actively manage their mental health. Mobile software apps have the potential to improve youth adherence to CBT and, in turn, improve outcomes of treatment. Objective The purpose of this project is to improve homework adherence in CBT for youth anxiety and/or depression. The objectives are to (1) design and optimize the usability of a mobile app for delivering the homework component of CBT for youth with anxiety and/or depression, (2) assess the app’s impact on homework completion, and (3) implement the app in CBT programs. We hypothesize that homework adherence will be greater in the app group than in the no-app group. Methods Phase 1: exploratory interviews will be conducted with adolescents and therapists familiar with CBT to obtain views and perspectives on the requirements and features of a usable app and the challenges involved in implementation. The information obtained will guide the design of a prototype. The prototype will be optimized via think-aloud procedures involving an iterative process of evaluation, modification, and re-evaluation, culminating in a fully functional version of the prototype that is ready for optimization in a clinical context. Phase 2: a usability study will be conducted to optimize the prototype in the context of treatment at clinics that provide CBT treatment for youth with anxiety and/or depression. This phase will result in a usable app that is ready to be tested for its effectiveness in increasing homework adherence. Phase 3: a pragmatic clinical trial will be conducted at several clinics to evaluate the impact of the app on homework adherence. Participants in the app group are expected to show greater homework completion than those in the no-app group. Results Phase 3 will be completed by September 2019. Conclusions The app will be a unique adjunct to treatment for adolescents in CBT, focusing on both anxiety and depression, developed in partnership with end users at every stage from design to implementation, customizable for different cognitive profiles, and designed with depression symptom tracking measures for youth made interoperable with electronic medical records.
This paper extends the variance-based Learning Classifier System called XCS-SAC, in order to extract two different abstracted level rules (i.e.classifiers). Since XCS-SAC attempts to evolve classifiers whose generality depends on their own parameter, such an attempt results in generating many specific classifiers (i.e.the classifiers having a less number of #). Due to inappropriate generalization, some of classifiers might not be human-understandable. To overcome this problem, our LCS focuses on an extraction of only two different abstracted level rules, both the specific and general rules, to understand a tendency in a given problem. In detail, the specific rules can be only utilized in limited situations but they are very accurate, while the general rules can be widely utilized but they are not accurate. The experimental result shows that our LCS succeeds to extract both specific and general rules appropriately in comparison with XCS-SAC.
This paper focuses on a multi-agent cooperation which is generally difficult to be achieved without sufficient information of other agents, and proposes the reinforcement learning method that introduces an internal reward for a multi-agent cooperation without sufficient information. To guarantee to achieve such a cooperation, this paper theoretically derives the condition of selecting appropriate actions by changing internal rewards given to the agents, and extends the reinforcement learning methods (Q-learning and Profit Sharing) to enable the agents to acquire the appropriate Q-values updated according to the derived condition. Concretely, the internal rewards change when the agents can only find better solution than the current one. The intensive simulations on the maze problems as one of testbeds have revealed the following implications:(1) our proposed method successfully enables the agents to select their own appropriate cooperating actions which contribute to acquiring the minimum steps towards to their goals, while the conventional methods (i.e., Q-learning and Profit Sharing) cannot always acquire the minimum steps; and (2) the proposed method based on Profit Sharing provides the same good performance as the proposed method based on Q-learning.
This paper aims at investigating how correct or incorrect opinions are shared among the agents in the weighted network where the relationship among the agent (as nodes of its network) is different each other, and exploring how the agents can be promoted to share only correct opinions by preventing to acquire the incorrect opinions in the weighted network. For this purpose, this paper focuses on Autonomous Adaptive Tuning algorithm (AAT) which can improve an accuracy of correct opinion shared among agents in the various network, and improves it to address the situation which is close in the real world, i.e., the relationship among agents is different each other. This is because the original AAT does not consider such a different relationship among the agents. Through the intensive empirical experiments, the following implications have been revealed: (1) the accuracy of the correct opinion sharing with the improved AAT is higher than that with the original AAT in the weighted network; (2) the agents in the improved AAT can prevent to acquire incorrect opinion sharing in the weighted network, while those in the original AAT are hard to prevent in the same network.
This paper proposes a novel Learning Classifier System (LCS) which integrates Deep AutoEncoder named DAE to solve high-dimensional problems. In the proposed LCS, DAE starts to compress (encode) an environmental input as a high-dimensional information to an input of LCS as a low-dimensional information and decompresses (decodes) an output of LCS as a low-dimensional information to a system output as a high-dimensional information. Since the compressed inputs are encoded by real value, this paper employs XCSR (i.e., an LCS with real value coding) and combines XCSR with DAE. In order to investigate the effectiveness of the proposed LCS, XCSR with DAE, this paper conducts the preliminary experiment on the benchmark classification problem, i.e., 6-Multiplexer problem. The intensive experiments on the compression from 6 to 5 bits have revealed the following implications: (1) XCSR with DAE performs as well as XCSR even learning from the compressed input data; and (2) XCSR with DAE successfully decodes the compressed rules to extract the rules which are the same as those learned with not compressed input data.
A default hierarchy is set of rules containing one or more exceptions to one or more default rules e.g. all dogs are friendly, except my neighbour's. Default hierarchies were the subject of considerable interest in early Learning Classifier Systems research, but they were abandoned due to the considerable difficulty of solving the credit assignment problems they involve. The most popular Learning Classifier System, XCS, and its derivatives do not support default hierarchies because in XCS each rule must be accurate, whereas in a default hierarchy an overgeneral rule may be overridden by a correct rule. In this work we enable XCS to evolve minimal default hierarchies by allowing two conditions in one rule, but evaluating only the accuracy and fitness of the whole as a whole. This simple step avoids the credit assignment issues faced by earlier systems. We call this XCS-DH. Preliminary evaluation of XCS-DH on a number of Boolean functions indicates a strong tendency to exploit the increased expressiveness of its rules. On some functions we observe slower learning and a larger population size, which we attribute to the increased rule expressiveness, which increases the search space. However, we also observe that in a problem that is particularly suitable for XCS-DH representation, and that is sufficient difficult for XCS, XCS-DH's learning rate is faster than XCS's. We take this as confirmation of the potential of learning default hierarchies with XCS-DH.
This paper focuses on the aircraft landing optimization problem where both the landing routes and the landing order of aircrafts should be optimized to minimize an occupancy time of airport , and proposes its optimization method which is robust to dynamical situations such as weather condition change and other aircrafts’ landing routes change. As a difficulty of this optimization problem, appropriate landing routes of aircrafts change depending on such an environment change. To tackle this problem, this paper proposes the hierarchical evolutionary computation to solve the aircraft landing optimization problem. Specifically, our method firstly generates candidates of main landing route of all aircrafts with their own additional sub-routes, which can be applied into the main routes depending on the current environmental situation. Secondly, our method evolves the good combination of landing routes (including their sub-routes) of all aircrafts to minimize an occupancy time of airport. Through the intensive experiment on a benchmark problem, the following implications have been found: (1) our method successfully generates robust landing routes including some sub-routes,which are flexible depending on environmental situations; and (2) Our method can finds an adequate landing order which contributes to reducing the occupancy time.
This issue is dedicated to Stewart Wilson both in recognition of his achievements and in appreciation for his guidance of the field. Not only has his research transformed the field, but his support, encouragement, and advice have been invaluable to many of us individually and to the community as a whole. Learning Classifier Systems (LCS) are rule-based machine learning algorithms introduced by Holland (see [2, 4–6]) that learn a population of IF-THEN rules that specify ‘‘IF x happens THEN do (or predict) y’’. In 1995 Stewart Wilson, already a leading LCS researcher, published a paper entitled ‘‘Classifier Fitness based on Accuracy’’ [16] that introduced the XCS algorithm. This paper proved a turning point for the field as XCS and its derivatives rapidly became the main focus of LCS research, and continue to be so today. We organised this special issue of Evolutionary Intelligence to mark the 20th anniversary of this landmark paper, and to serve as post-proceedings for IWLCS 2014: the Seventeenth International Workshop on Learning Classifier Systems. The annual IWLCS meetings are the yearly highlight of the LCS calendar and have no doubt contributed to the strong sense of community in the LCS field. For 2015 IWLCS has been rebranded as the more descriptive IWERML: the International Workshop on Evolutionary Rule-based Machine Learning. In 1994, the year before introducing XCS, Wilson introduced ZCS, the ‘‘Zeroth-order Classifier System’’ [15]. Wilson felt the growing complexity of the LCS concept was hindering progress and ZCS was an attempt to strip the LCS down to its essentials while retaining a working system. XCS built on the minimalist ZCS with two radical changes: a switch to accuracy-based fitness and, based on work by Booker [1], the addition of a niche Genetic Algorithm (GA) that strongly favours general rules. This combination was a hit: the niche GA favours general rules but accuracy-based fitness insists strictly that they make good predictions. This combination drives XCS to learn rules that are as general as possible while remaining accurate. Bull’s article in this issue provides much more on the history of LCS prior to and following XCS [2]. The famous ‘‘accuracy-based fitness’’ of XCS needs some explanation. The earlier ZCS featured ‘‘strengthbased fitness’’, in which the fitness of a rule in the GA was derived from its strength, a measure of the amount of reward the rule received. Consequently, fit rules in ZCS are those that receive a lot of reward. In contrast, XCS rules are fit if they make consistently accurate predictions about the reward they receive. This has the counterintuitive consequence that XCS rules that consistently take a bad action can be fit, since they are consistent. However, this poses no problem when it comes time for XCS to choose an action, since that can be done using the magnitude of the reward predicted by each rule. Why does XCS retain rules whose action it does not use? This ‘‘complete map’’ of the state/ & Muhammad Iqbal muhammad.iqbal@ecs.vuw.ac.nz
Sequence labeling is an interesting classification domain where, like normal classification, every input has a class label, but unlike normal classification, prediction of an input's label may depend on the values of other inputs or their classes, and so a learner may need to refer to inputs and classes at different time stamps to classify the current input. This is more difficult because a learner does not know where and how many inputs are needed to classify the current input. Our interest is in learning general rules for sequence labeling. The XCS algorithm is a rule-based knowledge discovery system powered by a genetic algorithm which has often been used for classification. Here we present XCS-SL, an extension of XCS classifier system which can be applicable to sequence labeling. Towards an application of Learning Classifier System (LCS) to sequence labeling, we propose a new classifier condition with memory (called a variable-length condition) and a rule-discovery system for the new classifier condition, which enables XCS to apply it to sequence labeling. In XCS-SL, classification rules (called "classifiers'' here) can include extra conditions on previous inputs, which act as memories. In sequence labeling, the number of conditions/memories needed may be different for each input, hence, using a fixed number of conditions (i.e., fixed-length condition) for all classifiers is not a good solution. Instead, XCS-SL classifiers have a variable-length condition to provide more or less memory. The genetic algorithm can grow and shrink conditions to find a suitable memory size. On two synthetic benchmark problems XCS-SL learns optimal classifiers, and on a real-world sequence labeling task it derives high classification accuracy and discovers interesting knowledge that shows dependencies between inputs at different times. The comprehensively described system is the first application of a LCS to sequence labeling and we consider it a promising direction for future work.
Dirk Thierens合作论文数Department of Information and Computing Sciences, Utrecht University3