Feature engineering is a crucial step in machine learning that provides better data for the learning algorithms to induce robust models, and this effort should be adapted to the capabilities of each algorithm. For example, classifiers that do not perform data transformations (e.g., cluster-based) perform better when the different classes are separated, typically requiring preprocessed data. Other models (e.g., decision trees) can perform several splits in the feature space, easily obtaining perfect results in training data, but have a higher risk of overfitting with unprocessed data. We use the rbd-GP and M3GP genetic programming algorithms to induce new features based on the original features, to be used by shallow and deep decision tree and random forest models. M3GP is wrapped around a learning algorithm, using its performance as fitness. This way, the induced features are adapted to the classifier, allowing us to compare the complexity of the features induced for the different classifiers. We measure the complexity of the induced features using several structural and functional complexity metrics found in the literature, also proposing a new metric that measures the separability of classes in the feature space. Like other authors, we use complexity as an interpretability metric, selecting three models to discuss and validate based on their performance and size. We apply these methods to remote sensing classification problems and solve two tasks that are hard due to the high similarity between the land cover classes: detecting cocoa agroforest and forecasting forest degradation up to one year in the future.
The Semantic Learning algorithm based on Inflate and deflate Mutations (SLIM-GSGP, or simply SLIM) is a variant of Geometric Semantic Genetic Programming (GSGP) designed to generate compact and interpretable models while maintaining the beneficial characteristic of GSGP of inducing an error surface without local optima. To date, no crossover operator has been defined for SLIM and the existing SLIM framework relies solely on two mutation operators: inflate and deflate mutation. This paper introduces two novel crossover operators for SLIM: Swap Crossover (XOSw) and Donor Crossover (XODn). These crossovers capitalize on SLIM’s linked-list representation to facilitate genetic exchange while controlling program size. Experimental results on five symbolic regression problems demonstrate that the new crossover operators often improve fitness and reduce model size when compared to standard SLIM and to GSGP. Our findings establish these operators as solid improvements of traditional GSGP crossover.
Like other machine learning methods, Genetic Programming (GP) frequently faces the issue of overfitting when applied to supervised learning tasks. Traditional regularization techniques, though well-studied, are challenging to apply to GP due to the free-form nature of the evolved models. This work proposes a novel approach that prevents overfitting while inherently improving the interpretability of GP models. It involves a dual optimization process that minimizes loss while penalizing functional complexity using multi-objective selection mechanisms. The improved complexity measure used in this study approximates the mathematical curvature of a function in linear time. While loss minimization is common in GP, penalizing functional complexity is an additional step aimed at evolving robust and smooth functions, less prone to overfitting and potentially more interpretable. Experimental results demonstrate the effectiveness of the two variants of our method, benchmarked against standard GP and two of the seemingly best overfitting-reduction methods found in the literature. By focusing on both loss and complexity, our approach achieves state-of-the-art generalization on difficult problems and a strong feature selection that improves interpretability, making it a unified improvement of GP.
The current trend in machine learning is to use powerful algorithms to induce complex predictive models that often fall under the category of “black-box models”. Thanks to this, there is also a growing interest in studying model explainabil-ity and interpretability so that human experts can understand, validate, and correct those models. With the objective of promoting the creation of inherently interpretable models, we present M6GP. This wrapper-based multi-objective automatic feature engineering algorithm combines key components of the M3GP and NSGA-II algorithms. Wrapping M6GP around another machine learning algorithm evolves a set of features optimized for this algorithm while potentially increasing its robustness. We compare our results with M3GP and M4GP, two ancestors from the same algorithm family, and verify that, by using a multi-objective approach, M6GP obtains equal or better results. In addition, by using complexity metrics on the list of objectives, the M6GP models come down to one-fifth of the size of the M3GP models, making them easier to read by comparison.
Feature engineering is a necessary step in the machine learning pipeline. Together with other preprocessing methods, it allows the conversion of raw data into a dataset containing only the necessary features to solve the task at hand, reducing the computational complexity of inducing models and creating models that are potentially simpler, more robust, and more interpretable. We use M3GP, a wrapper-based feature engineering algorithm, to induce a set of features that are adapted in number and in shape to several classifiers with different levels of predictive power, from decision trees with depth 3 to random forests with 100 estimators and no depth limit. Intuition tells us that classifiers that are restricted in the number of features should compensate for this restriction by using features with a high degree of correlation with the target objective. By opposition, the principle behind the boosting algorithm tells us that we can create a strong classifier using a large set of weak features. This indicates that classifiers with no restrictions should prefer many but weaker features. Our results confirm this hypothesis while also revealing that M3GP induces unnecessarily complex features. We measure complexity using several structural complexity metrics found in the literature and show that, although our pipeline consistently obtains good results, the structural complexity of the induced models varies drastically across runs. Additionally, while the test performance peaks in the early stages of the evolution, the complexity of the feature engineering models continues to grow, with little to no return in test performance. This work promotes using several complexity metrics to measure model interpretability and identifies issues related to model complexity in M3GP, proposing solutions to improve the computational cost of inducing models and the complexity of the final models.
This textbook provides the reader with an essential understanding of computational methods for intelligent systems.
Genetic Programming (GP) has the potential to generate intrinsically explainable models. Despite that, in practice, this potential is not fully achieved because the solutions usually grow too much during the evolution. The excessive growth together with the functional and structural complexity of the solutions increase the computational cost and the risk of overfitting. Thus, many approaches have been developed to prevent the solutions to grow excessively in GP. However, it is still an open question how these approaches can be used for improving the interpretability of the models. This article presents an empirical study of eight structural complexity metrics that have been used as evaluation criteria in multi-objective optimisation. Tree depth, size, visitation length, number of unique features, a proxy for human interpretability, number of operators, number of non-linear operators and number of consecutive nonlinear operators were tested. The results show that potentially the best approach for generating good interpretable GP models is to use the combination of more than one structural complexity metric.
We used a population genomic approach to unravel the population structure, genetic differentiation, and genetic diversity of three widespread wild bee species across the Iberian Peninsula, Andrena agilissima, Andrena flavipes and Lasioglossum malachurum. Our results demonstrated that genetic lineages in the Ebro River valley or near the Pyrenees mountains are different from the rest of Iberia. This relatively congruent pattern across species once more supports the hypothesis of "refugia within refugia" in the Iberian Peninsula. The results for A. flavipes and A. agilissima showed an unexpected pattern of genetic differentiation, with the generalist polylectic A. flavipes having lower levels of genetic diversity (Ho = 0.0807, He = 0.2883) and higher differentiation (F-ST = 0.5611), while the specialist oligolectic A. agilissima had higher genetic diversity (Ho = 0.2104, He = 0.3282) and lower differentiation values (F-ST = 0.0957). For L. malachurum, the smallest and the only social species showed the lowest inbreeding coefficient (F-IS = 0.1009) and the lowest differentiation level (F-ST = 0.0663). Overall, our results, suggest that this pattern of population structure and genetic diversity could be explained by the combined role of past climate changes and the life-history traits of the species (i.e., size, sociality and host-plant specialization), supporting the role of the Iberian refugia as a biodiversity hotspot.
The deployment of 5G technology has drawn attention to different computer-based scenarios. It is useful in the context of Smart Cities, the Internet of Things (IoT), and Edge Computing, among other systems. With the high number of connected vehicles, providing network security solutions for the Internet of Vehicles (IoV) is not a trivial process due to its decentralized management structure and heterogeneous characteristics (e.g., connection time, and high-frequency changes in network topology due to high mobility, among others). Machine learning (ML) algorithms have the potential to extract patterns to cover security requirements better and to detect/classify malicious behavior in a network. Based on this, in this work we propose an Intrusion Detection System (IDS) for detecting Flooding attacks in vehicular scenarios. We also simulate 5G-enabled vehicular scenarios using the Network Simulator 3 (NS-3). We generate four datasets considering different numbers of nodes, attackers, and mobility patterns extracted from Simulation of Urban MObility (SUMO). Furthermore, our conducted tests show that the proposed IDS achieved F1 scores of 1.00 and 0.98 using decision trees and random forests, respectively. This means that it was able to properly classify the Flooding attack in the 5G vehicular environment considered.
Self-supervised learning (SSL) methods have been widely used to train deep learning models for computer vision and natural language processing domains. They leverage large amounts of unlabeled data to help pretrain models by learning patterns implicit in the data. Recently, new SSL techniques for tabular data have been developed, using new pretext tasks that typically aim to reconstruct a corrupted input sample and yielding models which are, ideally, robust feature transforms. In this paper, we pose the research question of whether genetic programming is capable of leveraging data processed using SSL methods to improve its performance. We test this hypothesis by assuming different amounts of labeled data on seven different datasets (five OpenML benchmarking datasets and two real-world datasets). The obtained results show that in almost all problems, standard genetic programming is not able to capitalize on the learned representations, producing results equal to or worse than using the labeled partitions.
Geometric semantic genetic programming (GSGP) and linear scaling (LS) have both, independently, shown the ability to outperform standard genetic programming (GP) for symbolic regression. GSGP uses geometric semantic genetic operators, different from the standard ones, without altering the fitness, while LS modifies the fitness without altering the genetic operators. So far, these two methods have already been joined together in only one practical application. However, to the best of our knowledge, a methodological study on the pros and cons of integrating these two methods has never been performed. In this paper, we present a study of GSGP-LS, a system that integrates GSGP and LS. The results, obtained on five hand-tailored benchmarks and six real-life problems, indicate that GSGP-LS outperforms GSGP in the majority of the cases, confirming the expected benefit of this integration. However, for some particularly hard datasets, GSGP-LS overfits training data, being outperformed by GSGP on unseen data. Additional experiments using standard GP, with and without LS, confirm this trend also when standard crossover and mutation are employed. This contradicts the idea that LS is always beneficial for GP, warning the practitioners about its risk of overfitting in some specific cases.
Vectorial Genetic Programming (Vec-GP) extends regular GP by allowing vectorial input features (e.g. time series data), while retaining the expressiveness and interpretability of regular GP. The availability of raw vectorial data during training, not only enables Vec-GP to select appropriate aggregation functions itself, but also allows Vec-GP to extract segments from vectors prior to aggregation (like windows for time series data). This is a critical factor in many machine learning applications, as vectors can be very long and only small segments may be relevant. However, allowing aggregation over segments within GP models makes the training more complicated. We explore the use of common evolutionary algorithms to help GP identify appropriate segments, which we analyze using a simplified problem that focuses on optimizing aggregation segments on fixed data. Since the studied algorithms are to be used in GP for local optimization (e.g. as mutation operator), we evaluate not only the quality of the solutions, but also take into account the convergence speed and anytime performance. Among the evaluated algorithms, CMA-ES, PSO and ALPS show the most promising results, which would be prime candidates for evaluation within GP.
Divergence in acoustic signals may have a crucial role in the speciation process of animals that rely on sound for intra-specific recognition and mate attraction. The acoustic adaptation hypothesis (AAH) postulates that signals should diverge according to the physical properties of the signalling environment. To be efficient, signals should maximize transmission and decrease degradation. To test which drivers of divergence exert the most influence in a speciose group of insects, we used a phylogenetic approach to the evolution of acoustic signals in the cicada genus Tettigettalna, investigating the relationship between acoustic traits (and their mode of evolution) and body size, climate and micro-/macro-habitat usage. Different traits showed different evolutionary paths. While acoustic divergence was generally independent of phylogenetic history, some temporal variables' divergence was associated with genetic drift. We found support for ecological adaptation at the temporal but not the spectral level. Temporal patterns are correlated with micro- and macro-habitat usage and temperature stochasticity in ways that run against the AAH predictions, degrading signals more easily. These traits are likely to have evolved as an anti-predator strategy in conspicuous environments and low-density populations. Our results support a role of ecological selection, not excluding a likely role of sexual selection in the evolution of Tettigettalna calling songs, which should be further investigated in an integrative approach.
One of the main applications of machine learning (ML) in remote sensing (RS) is the pixel-level classification of satellite images into land cover types. Although classes with different spectral signatures can be easily separated, e.g. aquatic and terrestrial land cover types, others have similar spectral signatures and are hard to separate using only the information within a single pixel. This work focused on the separation of two cover types with similar spectral signatures, cocoa agroforest and forest, over an area in Para, Brazil. For this, we study the training and application of several ML algorithms on datasets obtained from a single composite image, a time-series (TS) composite obtained from the same location and by preprocessing the TS composite using simple TS preprocessing techniques. As expected, when ML algorithms are applied to a dataset obtained from a composite image, the median producer's accuracy (PA) and user's accuracy (UA) in those two classes are significantly lower than the median overall accuracy (OA) for all classes. The second dataset allows the ML models to learn the evolution of the spectral signatures over 5 months. Compared to the first dataset, the results indicate that ML models generalize better using TS data, even if the series are short and without any preprocessing. This generalization is further improved in the last dataset. The ML models are subsequently applied to an area with different geographical bounds. These last results indicate that, out of seven classifiers, the popular random forest (RF) algorithm ranked fourth, while XGBoost (xGB) obtained the best results. The best OA, as well as the best PA/UA balance, were obtained by performing feature construction using the M3GP algorithm and then applying XGB to the new extended dataset.
Supervision is a cross-disciplinary practice among various professional groups. This study focuses on clinical supervision as a practice linked to psychology and psychotherapy. The literature highlights the need to expand and consolidate knowledge in this area. Specifically, in the few existing approaches to research on existential supervision, the need for the systematization of knowledge is clear. The use of qualitative methods is recognized as an approach that is likely to enrich knowledge of supervision. Objective: The aim of this study was to explore the theme of clinical supervision, particularly as it relates to existential psychotherapy, from the supervisor’s perspective to assess insights from the experience of each participant. Method: The three participants are both existential psychotherapists and supervisors that apply the same approach, in group mode, in the context of psychotherapist training. The data were collected using phenomenological interviews. A comprehensive analysis of the transcripts of the interviews was performed using the phenomenological method. Results: Emerging themes presented a general meaning structure that represents eidetic dimensions and how they are related. The eidetic dimensions, relationship and responsiveness, arise in the existential approach as the foundational and promotional aspects of successful supervision.
Search-Based Software Engineering problems frequently have semantic constraints that can be used to deterministically restrict what type of programs can be generated, improving the performance of Genetic Programming. Strongly-Typed and Grammar-Guided Genetic Programming are two examples of using domain-knowledge to improve performance of Genetic Programming by preventing solutions that are known to be invalid from ever being added to the population. However, the restrictions in real world challenges like program synthesis, automated program repair or test generation are more complex than what context-free grammars or simple types can express. We address these limitations with examples, and discuss the process of efficiently generating individuals in the context of Christiansen Grammatical Evolution and Refined-Typed Genetic Programming. We present three new approaches for the population initialization procedure of semantically constrained GP that are more efficient and promote more diversity than traditional Grammatical Evolution.
Understanding patterns of population differentiation and gene flow in insect vectors of plant diseases is crucial for the implementation of management programs of disease. We investigated morphological and genome-wide variation across the distribution range of the spittlebug Philaenus spumarius (Linnaeus, 1758) (Hemiptera, Auchenorrhyncha, Aphrophoridae), presently the most important vector of the plant pathogenic bacterium Xylella fastidiosa Wells et al., 1987 in Europe. We found genome-wide divergence between P. spumarius and a very closely related species, P. tesselatus Melichar, 1899, at RAD sequencing markers. The two species may be identified by the morphology of male genitalia but are not differentiated at mitochondrial COI, making DNA barcoding with this gene ineffective. This highlights the importance of using integrative approaches in taxonomy. We detected admixture between P. tesselatus from Morocco and P. spumarius from the Iberian Peninsula, suggesting gene-flow between them. Within P. spumarius, we found a pattern of isolation-by-distance in European populations, likely acting alongside other factors restricting gene flow. Varying levels of co-occurrence of different lineages, showing heterogeneous levels of admixture, suggest other isolation mechanisms. The transatlantic populations of North America and Azores were genetically closer to the British population analyzed here, suggesting an origin from North-Western Europe, as already detected with mitochondrial DNA. Nevertheless, these may have been produced through different colonization events. We detected SNPs with signatures of positive selection associated with environmental variables, especially related to extremes and range variation in temperature and precipitation. The population genomics approach provided new insights into the patterns of divergence, gene flow and adaptation in these spittlebugs and led to several hypotheses that require further local investigation.