The dynamic multi-mode resource-constrained project scheduling problem (DMRCPSP) is of practical importance, as it requires making real-time decisions under changing project states and resource availability. Genetic Programming (GP) has been shown to effectively evolve heuristic rules for such decision-making tasks; however, the evolutionary process typically relies on a large number of simulation-based fitness evaluations, resulting in high computational cost. Surrogate models offer a promising solution to reduce evaluation cost, but their application to GP requires problem-specific phenotypic characterisation (PC) schemes of heuristic rules. There is currently a lack of suitable PC schemes for GP applied to DMRCPSP. This paper proposes a rank-based PC scheme derived from heuristic-driven ordering of eligible activity-mode pairs and activity groups in decision situations. The resulting PC vectors enable a surrogate model to estimate the fitness of unevaluated GP individuals. Based on this scheme, a surrogate-assisted GP algorithm is developed. Experimental results demonstrate that the proposed surrogate-assisted GP can identify high-quality heuristic rules consistently earlier than the state-of-the-art GP approach for DMRCPSP, while introducing only marginal computational overhead. Further analyses demonstrate that the surrogate model provides useful guidance for offspring selection, leading to improved evolutionary efficiency.
Data-efficient image classification is critical in computer vision with applications across various domains. Although deep convolutional neural networks are successful for image classification, they often require large datasets and high computational resources, which are unsuitable for data-efficient classification tasks. Genetic programming (GP), on the other hand, offers an interpretable, flexible, and efficient alternative to learning features for image classification, particularly when dealing with insufficient training instances. Many multi-objective GP methods control model bloat by limiting tree size or reducing the number of features, but they rarely include objectives specifically aimed at improving generalization. As a result, the evolved models can still overfit the training data. To address these challenges, we propose an improved decomposition-based multi-objective genetic programming (IDMOGP) approach to feature learning in data-efficient image classification. IDMOGP maximizes the classification accuracy with a regularization term and simultaneously minimizes the number of learned features. To reduce overfitting in data-efficient image classification, a novel regularization method is proposed based on Rademacher complexity to improve generalization. In addition, an adaptive global replacement strategy is designed to balance convergence and diversity during evolution. IDMOGP provides a set of trade-off solutions to data-efficient image classification. Experimental results on five different datasets show that IDMOGP achieves superior hypervolume compared to traditional dominance-based multi-objective GP, and higher classification accuracy than non-GP methods. Further analysis demonstrates the effectiveness of the introduced strategies and shows the good interpretability of the IDMOGP approach.
Deep learning is the mainstream method for medical image segmentation, and neural architecture search (NAS) has also been developed for this task. However, existing NAS methods remain limited in their ability to search for high-performance yet lightweight network architectures due to the high computational cost of NAS and the low fidelity of performance evaluation during the search process. In this paper, we propose a novel once-for-all NAS method for medical image segmentation to address these challenges. A carefully designed search space (supernet) incorporating key components of U-shape networks is constructed specifically for medical image segmentation. An effective and efficient hybrid two-stage supernet training scheme is then designed to enhance supernet training while maintaining a balance between performance and computational cost. A multi-objective evolutionary algorithm is leveraged to search for sets of network architectures, which produces high-performing architectures with varying computational complexities, optimized under multiple objectives. We conduct experiments on six widely used medical image segmentation datasets. Compared with existing methods, the proposed method achieves state-of-the-art performance on all six datasets. The searched architectures exhibit an excellent trade-off between performance and computational complexity, which is attributed to the effective multi-objective search.
Aquaculture is a key contributor to Aotearoa New Zealand’s economy, with greenshell mussels representing a major export. As farms move further offshore into more demanding conditions, maintaining float line buoyancy becomes increasingly challenging. Buoyancy failures can lead to mussel loss, highlighting the need for automated, scalable monitoring. Our deep learning-based buoyancy estimation approach offers a cost-effective, automated solution for aquaculture maintenance. Saliency map analyses reveal that neighbouring floats serve as contextual cues, providing motivation for the proposed novel context-aware strategy by considering scale of the context region relative to the target float for image-based assessment. By expanding each float’s background context to include neighbouring floats, we achieve a 7–8
Particle swarm optimization (PSO) has been widely applied to solve complex optimization problems from real-world applications due to its efficient exploration of large solution spaces and the ability to converge towards optimal solutions without requiring gradient information. Common swarm topologies in standard PSO and its variants, e.g., Ring and Star, can be regarded as graphs, where each edge connects only two particles. Such topology structures allow direct interactions only between connected particle pairs, and thus often fail to directly capture the higher-order social relationships that are necessary for navigating complex search landscapes. Therefore, this article proposes a novel PSO variant termed Hypergraph-assisted Particle Swarm Optimization (HPSO). In HPSO, the topology of the particles in a swarm is modeled by a hypergraph, in which hyperedges are used to connect multiple particles. This allows multiple particles within a hyperedge to interact directly. Furthermore, an adaptive hypergraph updating strategy is designed to periodically reconstruct the topology based on cumulative average particle displacement, thereby maintaining swarm diversity throughout the evolutionary process. In the experiments, the effectiveness of HPSO is verified on the IEEE CEC'17 benchmark suite, and the results demonstrate that HPSO achieves promising performance across various types of functions. Furthermore, the ablation experiment demonstrates that HPSO has excellent search capabilities.
New Zealand's aquaculture industry is experiencing significant growth, driven by its carbon efficiency and the lack of spatial constraints on land. Green-lipped mussels, a key farmed species, are grown on long lines supported by buoyant plastic floats. Managing tens of thousands of these floats is an operational challenge for mussel farmers. Accordingly, applying multi-object detection to farm imagery is a promising way to improve crop monitoring and productivity. This paper presents a Genetic Programming (GP)-based method for multi-object detection in mussel farms, designed to address site-specific challenges such as a high density of targets, high levels of partial occlusion, and substantial variation in apparent object size due to camera distance. The proposed method integrates image preprocessing techniques for object localisation and GP techniques for classification. In terms of detection performance, the proposed method surpasses YOLOv12 by achieving an F1 score of 95.2% compared with 88.6%. Although execution speed and bounding box precision could be further improved, the method demonstrates robustness, effectiveness, and practical applicability in mussel farms.
Skin cancer is one of the most prevalent malignant tumours worldwide, and its incidence has continued to climb in recent years. Traditional feature extraction methods often struggle with the high variability and complex patterns in skin cancer images, necessitating more adaptive and automated approaches. This study proposes a genetic programming (GP)-based method with flexible region detection operators (GPFRD) for automatically and flexibly learning discriminative features for various classification tasks of skin cancer images. The proposed GPFRD method integrates preprocessing, region detection, feature extraction, and feature concatenation into a cohesive framework, significantly enhancing flexibility. The newly designed operators precisely localize diagnostically critical regions based on lesion masks while suppressing irrelevant background interference. These operators enable the proposed method to evolve effective feature extraction solutions based on the characteristics of different image datasets. Experimental results on five datasets of varying difficulties demonstrate that the proposed method outperforms the benchmark GP-based method and four traditional feature extraction methods in the majority of cases.
Klaus-Robert Müller is Full Professor for Machine Learning at the Department of Computer Science at Technische Universität Berlin and at the Department of Artificial Intelligence at Korea University, Seoul. Over the years he held the roles as director of the Bernstein Center for Neurotechnology, co-director of the Berlin Center for Big Data and director of the Berlin Machine Learning Center. In 2021, he became director of the Berlin Institute for Foundations of Learning and Data (BIFOLD)—one of 5 permanently funded German national centers for AI. In 2020/2021 and recently in 2024/2025 he was on a short (1-year-long) sabbatical from academia to lead a team at Google Brain rsp. DeepMind as Principal Researcher. In 2012, he was elected to be a member of the German National Academy of Sciences—Leopoldina; in 2017 of the Berlin Brandenburg Academy of Sciences; in 2022 member of the German National Academy of Engineering; and also, in 2017 an external scientific member of the Max-Planck Society (MPII). He received several research awards: among others, in 2014 the Berlin Science Prize awarded by the governing mayor of Berlin; in 2017 the Vodafone Innovation Award; in 2024 the Hector Science award and the Feynman Award; and in 2025 the IEEE CIS Neural Network Pioneer Award. Consecutively from 2019 on he became an ISI Highly Cited Researcher. His research interests are in the field of machine learning, deep learning, explainable AI, and data analysis covering a wide range of theory and numerous scientific (physics, chemistry and medicine) and industrial applications. Google Scholar Citations > 168000, h-index 167.
Segmenting hyperspectral images is challenging because of their high dimensionality, which necessitates the selection of wavelengths to create lower-dimensional multispectral models. Although hyperspectral image analysis studies often involve multiple tasks (e.g., predicting multiple attributes of a sample), existing wavelength selection methods typically optimise for only one. We propose a Genetic Programming (GP)-based approach called Hyperspectral Multitask Segmentation GP (HMSGP). GP solutions are built using hyperspectral-specific and convolutional operators to extract shared feature maps for multiple segmentation tasks within a multi-objective optimisation framework. A hyperspectral image dataset of Pacific oysters (\textit{Magallana gigas}) was created using two hyperspectral sensors that covered the visible and near-infrared ranges. A new framework was developed for the pixel-level fusion of these images by exploiting the overlapping spectral range. The models were trained to perform two segmentation tasks: delimiting the meat area and segmenting anatomical regions. The experimental results demonstrate competitive pixel-wise segmentation performance for both tasks, and this approach can return a diverse set of trade-off solutions. This study lays the groundwork for future research to estimate the nutritional composition of oysters from hyperspectral images and distinct anatomical regions, an important step toward rapid, non-destructive analysis of this high-value seafood product.
Traffic signal control systems represent a large scale, distributed mobile computing environment, where each intersection acts as a spatially distributed computing node that makes decisions based on locally available and dynamically changing information. Recently, Genetic Programming (GP), as a powerful evolutionary machine learning algorithm, is successfully utilized to automatically evolve effective and human understandable traffic signal control policies. The existing GP based methods make decisions at each intersection based only on local observations at the intersection. However, the decisions at different intersections in the traffic network affect each other. Ignoring the interactions and cooperation among them is likely to result in overall performance deficiencies. To address this issue and encourage potential cooperation among the decisions at different intersections, this work 1) proposes a turn movement level communication design that enables neighboring intersections to share structurally aligned communication messages. 2) designs heuristic-based communication messages propagating among intersections; and 3) presents a method to evolve dedicated traffic signal policies that take such communicated messages into account to have better traffic signal control. We evaluate our proposed method on both synthetic and real-world datasets based on the CityFlow traffic simulator. The experiment results demonstrate that each new component in our GP algorithm is effective, and the proposed method significantly outperforms the state-of-the-art traffic signal control methods in large-scale traffic networks.
High-dimensionality is one of the serious real-world data challenges in symbolic regression and it is more challenging if the data are incomplete. Genetic programming has been successfully utilised for high-dimensional tasks due to its natural feature selection ability, but it is not directly applicable to incomplete data. Commonly, it needs to impute the missing values first and then perform genetic programming on the imputed complete data. However, in the case of having many irrelevant features being incomplete, intuitively, it is not necessary to perform costly imputations on such features. For this purpose, this work proposes a genetic programming-based approach to select features directly from incomplete high-dimensional data to improve symbolic regression performance. We extend the concept of identity/neutral elements from mathematics into the function operators of genetic programming, thus they can handle the missing values in incomplete data. Experiments have been conducted on a number of data sets considering different missingness ratios in high-dimensional symbolic regression tasks. The results show that the proposed method leads to better symbolic regression results when compared with state-of-the-art methods that can select features directly from incomplete data. Further results show that our approach not only leads to better symbolic regression accuracy but also selects a smaller number of relevant features, and consequently improves both the effectiveness and the efficiency of the learning process.
Computer vision and machine learning have accelerated the automation of animal re-identification pipelines used in conservation programs worldwide. For species with distinctive markings, such as the spot patterns of giraffes, these automated methods are crucial for research and population monitoring purposes. However, many tools are designed for experts, and their implementation requires substantial technical expertise. Research teams often use specialist software and workflows that are not accessible to the general public. In a zoo setting, visitors lack a simple way to identify an individual animal, and unique features are easily missed by untrained visitors. This study presents a three-part solution: a web interface for zoo visitors to upload photos, a deep learning model for giraffe torso detection, and a fast re-identification method for matching observations to a gallery of known individuals using server-side processing. We compare several re-identification methods (RootSIFT, MiewID, and MegaDescriptor) using a consistent evaluation protocol and report both identification performance and system latency for this closed-set zoo setting. Taken together, this study presents a visitor-facing web system that integrates existing re-identification models into a modular, real-time pipeline for zoo deployment, lowering the barrier to visitor participation and making state-of-the-art re-identification methods more accessible to the general public.
River flow plays a vital role in the hydrologic cycle. In recent years, numerous machine learning approaches, particularly deep neural networks (DNNs), have demonstrated success in forecasting river flow. However, many of these models depend heavily on extensive historical flow data and very specific hydrological features, which are often labor-intensive to gather. Moreover, most existing predictive models, especially DNN-based, are handcrafted, requiring substantial domain expertise and extensive hyperparameter tuning, making the modeling process potentially inefficient. In this study, we propose a novel approach to achieve flexible and efficient river flow prediction that relies solely on openly accessible weather forecast data. A one-dimensional convolutional neural network (1D-CNN) is designed to effectively capture the temporal relationships between weather variables and river flow. Additionally, an efficient evolutionary neural architecture search (NAS) algorithm is developed to automatically discover the best 1D-CNN architectures, thereby improving predictive performance while reducing the need for manual architecture tuning. To evaluate our approach (named AutoNN-Flow), experiments are conducted using the Kaeo River in New Zealand as a case study. Without access to river-specific attributes, AutoNN-Flow achieves highly accurate 7-day and 14-day river flow predictions, greatly outperforming classic machine learning models.
Medical image classification is challenging due to limited labeled instances, high inter-class similarity, and imbalanced class distributions. Although genetic programming (GP) has demonstrated strong potential in general image classification, its application to medical image classification remains underexplored. To address this research gap, this paper proposes a Multi-Objective Genetic Programming with Ensemble Construction (MOGPEC) algorithm for medical image classification. First, a new GP representation is designed to align with typical medical image processing workflows, enabling the evolved solutions to be more consistent with clinical reasoning. A multi-objective optimization framework is then introduced to encourage the generation of diverse solutions across different classes, creating a rich pool of candidate models for subsequent ensemble construction. Finally, an ensemble construction strategy is developed to select and combine complementary GP solutions, thereby enhancing the overall classification performance. Extensive evaluations on multiple medical image datasets demonstrate that MOGPEC outperforms traditional methods, representative GP methods, and deep-learning-based methods in most cases. Ablation experiments in the Supplementary Material further analyze the contributions of the proposed GP representation, multi-objective optimization, and ensemble construction strategy. Visualization of the evolved solutions provides valuable interpretability.
Dynamic Air Traffic Flow Management (DATFM) is vital to modern aviation industries, requiring real-time routing and sequencing aircraft within constrained airspace to ensure safety and efficiency, particularly during unforeseen disruptions. Multitree Genetic Programming (MTGP) provides an automatic design framework to simultaneously learn routing and sequencing policies. However, the coupling interdependence between the routing and sequencing decisions in MTGP tends to produce structurally valid but semantically ineffective offspring, thereby limiting MTGP’s search ability within the heuristic space. This paper presents a novel Multitree Genetic Programming with Behavioral Semantics (MrGPBS) approach that automatically evolves effective reactive scheduling heuristics for DATFM. First, we develop a specialized multitree representation incorporating informative terminals that enable concurrent evolution of routing and sequencing decisions, coupled with a problem-tailored heuristic template that decodes these representations into executable solutions in an online manner. Second, we propose a behavioral semantics-guided evolutionary mechanism that addresses the coupling interdependence issue of existing MTGP methods. Our method preserves well-evolved tree structures by promoting recombination between behaviorally compatible individuals, thereby enabling incremental improvement within the heuristic search space. We evaluate MrGPBS on comprehensive benchmark instances derived from real-world air traffic data, encompassing diverse problem scales and dynamic operating conditions. Experimental results demonstrate that our approach significantly outperforms state-of-the-art methods.
Dynamic flexible job shop scheduling (DFJSS) is a challenging combinatorial optimisation problem that requires effective decision-making under dynamic environments. Although genetic programming (GP) has shown success in automatically learning scheduling heuristics, existing research predominantly addresses dynamic events involving single job arrivals, which does not always reflect real-world situations. In practice, heterogeneous batch arrivals where different jobs arrive simultaneously, are quite common and introduce a new decision-making challenge: handling multiple routing decisions (machine assignments) concurrently. However, this problem has received little attention in the literature. To fill this gap, we first formulate the DFJSS problem with heterogeneous batch arrivals. Furthermore, to effectively coordinate simultaneous routing decisions introduced by batch arrivals, GP is used to evolve scheduling heuristics to prioritise hybrid operation-machine pairs instead of jobs. An update strategy is incorporated to reflect the latest system status after each routing assignment, improving decision quality under dynamic changes. Experimental results across 18 scenarios demonstrate that using routing rules to prioritise pairs achieves the best average rank and outperforms compared methods. The effectiveness of the update strategy is also verified. In general, a single routing rule can be effectively applied to both concurrent and individual routing decisions.
Minimizing the classification error rate and the number of selected features are the two major objectives of feature selection, and they are often in conflict with each other, which is a multiobjective problem. Evolutionary algorithms (EAs) have been widely used for multiobjective feature selection problems. Preselection in EAs is used to improve the sampling quality by selecting only potentially promising candidate solutions for fitness evaluations. However, traditional preselection methods struggle to effectively handle feature selection due to its large-scale combinatorial nature and intricate feature interactions. To alleviate this issue, this article proposes a filter-based performance predictor to preselect feature subsets for subsequent classification fitness evaluations. It uses multiple filter measures to estimate the classification performance of a feature subset, which can explore complex feature interactions and is also insensitive to the dimensionality. Additionally, a correlation coefficient is used to measure the compatibility between the learned performance predictor and the classification performance. Based on the degree of compatibility, a preselection method that considers both the predicted classification performance and the feature subset diversity is proposed, which can preselect promising solutions from multiple candidate solutions and thus improve the feature subset search efficiency. The proposed method is verified experimentally on a total of 18 classification datasets spanning various domains, and the results reveal that it can find feature subsets with better classification performance and converge faster to competitive results compared to state-of-the-art methods.
Multilabel classification (MLC) involves assigning multiple labels to each instance from a predefined set of labels. With the increasing prevalence of multilabel datasets in real-world problems, MLC has become a popular area of research. These datasets frequently contain irrelevant, redundant, or noisy features, highlighting the importance of feature selection in MLC tasks. As a result, numerous multilabel feature selection (MLFS) approaches have been introduced in the literature. Given their effective search capabilities, evolutionary computation (EC) techniques have been adopted to develop MLFS approaches. However, there is a notable absence of a comprehensive survey dedicated to EC-based MLFS approaches to MLC tasks. Although few attempts have been made to address this gap, they do not provide detailed discussions on EC-based approaches. They also tend to overlook the strengths and weaknesses of existing approaches, especially regarding search, optimization, and evaluation processes. This article aims to fill this gap by presenting a comprehensive survey of EC-based MLFS approaches, focusing on the latest advancements, current challenges, and future directions. To be specific, we categorize EC-based MLFS approaches based on different criteria and provide detailed descriptions for each category.
The rising global incidence of skin cancer, particularly Basal Cell Carcinoma (BCC), necessitates the development of automated diagnostic systems that are both accurate and interpretable for clinical adoption. While Deep Learning approaches have shown promise in medical image analysis, their “black box” nature remains a significant barrier to clinical trust and deployment. To address these challenges, this study introduces CoDED (Co-evolutionary Descriptor Engine for Detection), a genetic programming method that enhances existing automated BCC detection through two synergistic advancements. First, the method replaces traditional binary encoding with evolved Local Ternary Patterns (LTP), enabling richer textural feature extraction through adaptive dual-threshold mechanisms. Second, a novel Discriminative Localisation Mapping (DLM) technique provides pixel-level interpretability, revealing the model’s diagnostic rationale through saliency visualisations. Comprehensive experimental evaluation on clinical dermoscopy images demonstrates that CoDED achieves a balanced accuracy of 79.57
Wenlong Fu合作论文数System Applications and Products in Data Processing32