Equation discovery has traditionally focused on regression, where the goal is to recover analytical expressions that model numerical targets. In this paper, we extend this paradigm to binary classification and introduce Equation Discovery for Classification (EDC), a framework that discovers concise symbolic expressions that explicitly define decision boundaries. EDC searches over a configurable grammar of analytical expressions using beam search and optimises equation parameters via dedicated numerical procedures, yielding a single interpretable equation whose sign determines class membership. We design a redundancy-aware grammar that balances expressivity and tractability, enabling the discovery of non-linear decision boundaries while maintaining interpretability. Through experiments on artificial datasets with known generating mechanisms, we show that EDC reliably reconstructs complex target boundaries, including XOR-like and interaction-driven structures, and adapts effectively under increasing levels of noise. Notably, in noisy settings EDC can outperform the original generating equation by approximating the implicit, noise-shifted decision boundary. On UCI benchmark datasets, EDC consistently outperforms existing symbolic classification approaches and other interpretable baselines, while achieving performance competitive with state-of-the-art black-box models. Although computationally more demanding than standard classifiers, we demonstrate that substantial speed-ups can be achieved with reduced search depth and simplified grammars at minimal loss of predictive performance. Overall, EDC provides a principled bridge between symbolic regression and classification, offering a transparent yet expressive alternative to black-box models for applications where interpretability of the decision boundary is essential.
Quantifying the health of civil infrastructure using sensor data remains challenging, as degradation-related signals are typically weak and obscured by dominant environmental and operational effects. In structural health monitoring (SHM), this often results in sensor measurements that are highly periodic or intermittent, while long-term degradation manifests only as subtle drift. This study addresses the problem of extracting meaningful proxies for structural health from such data. We propose monotonicity as a guiding principle, operationalized through absolute Spearman’s rank correlation between sensor values and time. Two complementary methods are introduced. First, subgroup discovery is employed to identify structurally coherent groups of sensors that exhibit significantly elevated monotonicity, enabling the construction of robust health proxies through aggregation. Second, we present Latent Monotonic Feature Discovery (LMFD), a data-driven method inspired by equation discovery, which searches for arithmetic combinations of sensors that yield monotonic behaviour even when individual sensors are predominantly non-monotonic. The methods are evaluated on a two-year monitoring dataset from a Dutch concrete highway bridge comprising strain gauges, geophones, and temperature sensors. Results show that meaningful monotonic proxies can be derived both from naturally monotonic sensor subgroups and from composite features constructed from periodic signals. The proposed approach provides indirect yet interpretable indicators of structural health and offers a principled way to uncover latent degradation trends in long-term SHM data.
The aim of our study is to investigate whether coupling power output (PO) and heart rate (HR) data of semi-professional road cyclists collected in the field is helpful for optimising the training process. Therefore, HR and PO data during all cycling activities were collected from 23 semi-professional road cyclists for 2 years. Weekly cyclist-specific HR response times (recovery, delay and maximal response time) were extracted from models connecting HR and PO. Linear regression was performed between performance, defined as mean PO during a 1- and 10-min indoor time trial (TT) under controlled circumstances, and weekly HR response times. No significant correlations were found between 1-min TT PO and HR response times. In contrast, significant correlations were obtained between 10-min TT PO and recovery time (r = -0.74, p < 0.01), maximal response time (r = -0.70, p < 0.01) and delay time (r = -0.48, p = 0.03). Moreover, linear relationships were found between 10-min TT PO and delay time (r = -0.68, p < 0.01) or maximal response times (r = -0.61, p = 0.02) within 14 days of the performed lab test. This suggests that HR response times are important physiological characteristics related to 10-min TT PO for cyclists.
Many systems in our world age, degrade or otherwise move slowly but steadily in a certain direction. When monitoring such systems by means of sensors, one often assumes that some form of `age' is latently present in the data, but perhaps the available sensors do not readily provide this useful information. The task that we study in this paper is to extract potential proxies for this `age' from the available multi-variate time series without having clear data on what `age' actually is. We argue that when we find a sensor, or more likely some discovered function of the available sensors, that is sufficiently monotonic, that function can act as the proxy we are searching for. Using a carefully defined grammar and optimising the resulting equations in terms of monotonicity, defined as the absolute Spearman's Rank Correlation between time and the candidate formula, the proposed approach generates a set of candidate features which are then fitted and assessed on monotonicity. The proposed system is evaluated against an artificially generated dataset and two real-world datasets. In all experiments, we show that the system is able to combine sensors with low individual monotonicity into latent features with high monotonicity. For the real-world dataset of InfraWatch, a structural health monitoring project, we show that two features with individual absolute Spearman's $ρ$ values of $0.13$ and $0.09$ can be combined into a proxy with an absolute Spearman's $ρ$ of $0.95$. This demonstrates that our proposed method can find interpretable equations which can serve as a proxy for the `age' of the system.
Equation Discovery techniques have shown considerable success in regression tasks, where they are used to discover concise and interpretable models (Symbolic Regression). In this paper, we propose a new ED-based binary classification framework. Our proposed method EDC finds analytical functions of manageable size that specify the location and shape of the decision boundary. In extensive experiments on artificial and real-life data, we demonstrate how EDC is able to discover both the structure of the target equation as well as the value of its parameters, outperforming the current state-of-the-art ED-based classification methods in binary classification and achieving performance comparable to the state of the art in binary classification. We suggest a grammar of modest complexity that appears to work well on the tested datasets but argue that the exact grammar – and thus the complexity of the models – is configurable, and especially domain-specific expressions can be included in the pattern language, where that is required. The presented grammar consists of a series of summands (additive terms) that include linear, quadratic and exponential terms, as well as products of two features (producing hyperbolic curves ideal for capturing XOR-like dependencies). The experiments demonstrate that this grammar allows fairly flexible decision boundaries while not so rich to cause overfitting.
Background The aging population faces numerous health challenges, with sedentary behavior and decreased physical activity being paramount. We explore the physical behaviour of older adults in the GOTO combined lifestyle intervention study and its related immuno-metabolic health effects. Methods The research utilized accelerometers and machine learning to assess physical activity behaviours during a 13-week program of increased physical activity and decreased calorie intake. Subsequently, the association of variation in physical behaviour with immuno-metabolic health parameters is investigated cross-sectionally at baseline and longitudinally using sex-stratified linear regression and linear mixed regression respectively. Results Participants exhibited physical behaviors similar to their age-matched peers from the UK-Biobank. Interestingly, gender-based differences were evident, with men and women showing distinct daily physical behavioural patterns. At baseline, a positive correlation was found between higher physical behavior and a healthier immune-metabolic profile, particularly in men. The longitudinal changes depict an overall boost in activity levels, predominantly among women. While increasing general activity and engaging in intense exercises proved advantageous for physical health, the immune-metabolic health benefits were more pronounced in men. Conclusion The short-term GOTO intervention underscores the significance of regular physical activity in promoting healthy aging even in middle to older age. Gender differences in behavior and health benefits deserve much more attention though. Our results advocate the broader implementation of such programs and emphasize the utility of technology, like accelerometers and machine learning, in both monitoring and promoting active lifestyles among older adults.
BACKGROUND:Training load is typically described in terms of internal and external load. Investigating the coupling of internal and external training load is relevant to many sports. Here, continuous kernel-density estimation (KDE) may be a valuable tool to capture and visualize this coupling.AIM:Using training load data in speed skating, we evaluated how well bivariate KDE plots describe the coupling of internal and external load and differentiate between specific training sessions, compared to training impulse scores or intensity distribution into training zones.METHODS:On-ice training sessions of 18 young (sub)elite speed skaters were monitored for velocity and heart rate during 2 consecutive seasons. Training session types were obtained from the coach's training scheme, including endurance, interval, tempo, and sprint sessions. Differences in training load between session types were assessed using Kruskal-Wallis or Kolmogorov-Smirnov tests for training impulse and KDE scores, respectively.RESULTS:Training impulse scores were not different between training session types, except for extensive endurance sessions. However, all training session types differed when comparing KDEs for heart rate and velocity (both P < .001). In addition, 2D KDE plots of heart rate and velocity provide detailed insights into the (subtle differences in) coupling of internal and external training load that could not be obtained by 2D plots using training zones.CONCLUSION:2D KDE plots provide a valuable tool to visualize and inform coaches on the (subtle differences in) coupling of internal and external training load for training sessions. This will help coaches design better training schemes aiming at desired training adaptations.
In sports, athlete monitoring is important for preventing injuries and optimizing performance. The multitude of relevant factors during the exercise sessions, such as weather conditions, makes proper individual athlete monitoring labour intensive. In this work, we develop an automated approach for athlete monitoring in professional road cycling that takes into account the terrain on which the ride is executed by finding segments with similar elevation profiles. In our approach, the matching is focused on the shapes of the segments. We use 2.5 years of data of a single rider of Team Jumbo-Visma and assess the performance of our approach by determining the quality of the best matches for a selection of 700 distinct segments, consisting of the most representative shapes for the elevation profiles. We demonstrate that the execution time is within seconds and more than ten times faster than exhaustive search. Therefore, our method enables real-time deployment in large scale applications with potentially many requests from multiple users. Moreover, we show that on average our approach has similar accuracy when considering the correlation to a target segment and approximately only has a twice as large mean squared error when compared to exhaustive search. Finally, we discuss a practical example to demonstrate how our approach can be used for athlete performance monitoring.
Through the quantification of physical activity energy expenditure (PAEE), health care monitoring has the potential to stimulate vital and healthy ageing, inducing behavioural changes in older people and linking these to personal health gains. To be able to measure PAEE in a health care perspective, methods from wearable accelerometers have been developed, however, mainly targeted towards younger people. Since elderly subjects differ in energy requirements and range of physical activities, the current models may not be suitable for estimating PAEE among the elderly. Furthermore, currently available methods seem to be either simple but non-generalizable or require elaborate (manual) feature construction steps. Because past activities influence present PAEE, we propose a modeling approach known for its ability to model sequential data, the recurrent neural network (RNN). To train the RNN for an elderly population, we used the growing old together validation (GOTOV) dataset with 34 healthy participants of 60 years and older (mean 65 years old), performing 16 different activities. We used accelerometers placed on wrist and ankle, and measurements of energy counts by means of indirect calorimetry. After optimization, we propose an architecture consisting of an RNN with 3 GRU layers and a feedforward network combining both accelerometer and participant-level data. Our efforts included switching mean to standard deviation for down-sampling the input data and combining temporal and static data (person-specific details such as age, weight, BMI). The resulting architecture produces accurate PAEE estimations while decreasing training input and time by a factor of 10. Subsequently, compared to the state-of-the-art, it is capable to integrate longer activity data which lead to more accurate estimations of low intensity activities EE. It can thus be employed to investigate associations of PAEE with vitality parameters of older people related to metabolic and cognitive health and mental well-being.
EDITORIAL article Front. Sports Act. Living, 25 April 2022Sec. Sports Science, Technology and Engineering https://doi.org/10.3389/fspor.2022.886730
We present a personalized approach for frequent fitness monitoring in road cycling solely relying on sensor data collected during bike rides and without the need for maximal effort tests. We use competition and training data of three world-class cyclists of Team Jumbo–Visma to construct personalised heart rate models that relate the heart rate during exercise to the pedal power signal. Our model captures the non-trivial dependency between exertion and corresponding response of the heart rate, which we show can be effectively estimated by an exponential kernel. To construct the daily heart rate models that are required for day-to-day fitness estimation, we aggregate all sessions in the previous week and apply sampling. On average, the explained variance of our models is 0.86, which we demonstrate is more than twice as large as for models that ignore the temporal integration involved in the heart's response to exercise. We show that the fitness of a cyclist can be monitored by tracking developments of parameters of our heart rate models. In particular, we monitor the decay constant of the kernel involved, and also analytically determine virtual aerobic and anaerobic thresholds. We demonstrate that our findings for the virtual anaerobic threshold on average agree with the results of exercise tests. We believe this work is an important step forward in performance optimization by opening up avenues for switching to adaptive training programs that take into account the current physiological state of an athlete.
In this study, we investigated the relationships between training load, perceived wellness and match performance in professional volleyball by applying the machine learning techniques XGBoost, random forest regression and subgroup discovery. Physical load data were obtained by manually logging all physical activities and using wearable sensors. Daily wellness of players was monitored using questionnaires. Match performance was derived from annotated actions by a video scout during matches. We identified conditions of predictor variables that related to attack and pass performance (p < 0.05). Better attack performance is related to heavy weights of lower-body strength training exercises in the preceding four weeks. However, worse attack performance is linked to large variations in weights of full-body strength training exercises, excessively heavy upper-body strength training, low jump heights and small variations in the number of high jumps in the four weeks prior to competition. Lower passing performance was associated with small variations in the number of high jumps in the preceding week and an excessive amount of high jumps performed, on average, in the two weeks prior to competition. Differences in findings with respect to passing and attack performance suggest that elite volleyball players can improve their performance if training schedules are adapted to the position of a player.
This work introduces RADIUS, a framework for anomaly detection in sewer pipes using stereovision. The framework employs three-dimensional geometry reconstruction from stereo vision, followed by statistical modeling of the geometry with a generic pipe model. The framework is designed to be compatible with existing workflows for sewer pipe defect detection, as well as to provide opportunities for machine learning implementations in the future. We test the framework on 48 image sets of 26 sewer pipes in different conditions collected in the lab. Of these 48 image sets, 5 could not be properly reconstructed in three dimensions due to insufficient stereo matching. The surface fitting and anomaly detection performed well: a human-graded defect severity score had a moderate, positive Pearson correlation of 0.65 with our calculated anomaly scores, making this a promising approach to automated defect detection in urban drainage.
The aim of this thesis is to investigate how non data experts can effectively query a set of frequent itemsets to find interesting patterns in their data set. Frequent itemset mining is a technique to find repeating patterns in data that consists of rows of transactions such as purchases in a supermarket. These patterns consist of items often occurring together in these transactions. This information can be used by domain experts to, for instance, optimize the positioning of items in a store to increase their sales. Currently it is difficult for domain experts to find these frequent itemsets and even if they have them, it is difficult to filter through them. It is difficult to find these itemsets because the tools that can do this are not very accessible for people without a background in data mining. Furthermore, if a domain expert finds these frequent itemsets, the list of them will usually be too large to look through without any help. This thesis tries to make it easier for domain experts to find the frequent itemsets and also to exploit their knowledge by allowing the user to filter through the frequent itemsets using an easy to use query language. To do this, we propose a browser-based tool to obtain the frequent itemsets out of a given data set and allow the user to filter through them with an easy to use query language. This program is run completely in the browser, without offloading anything to a server to ensure safety of their data. We also performed a preliminary user study to evaluate the usability and effectiveness of the query language. From these results we can say that the query language is effective for people with a background in computer science.
ABSTRACT We implemented a machine learning approach to investigate individual indicators of training load and wellness that may predict the emergence or development of overuse injuries in professional volleyball. In this retrospective study, we collected data of 14 elite volleyball players (mean ± SD age: 27 ± 3 years, weight: 90.5 ± 6.3 kg, height: 1.97 ± 0.07 m) during 24 weeks of the 2018 international season. Physical load was tracked by manually logging the performed physical activities and by capturing the jump load using wearable devices. On a daily basis, the athletes answered questions about their wellness, and overuse complaints were monitored via the Oslo Sports Trauma Research Center (OSTRC) questionnaire. Based on training load and wellness indicators, we identified subgroups of days with increased injury risk for each volleyball player using the machine learning technique Subgroup Discovery. For most players and facets of overuse injuries (such as reduced sports participation), we have identified personalized training load and wellness variables that are significantly related to overuse issues. We demonstrate that the emergence and development of overuse injuries can be better understood using daily monitoring, taking into account interactions between training load and wellness indicators, and by applying a personalized approach. Highlights With detailed, athlete-specific monitoring of overuse complaints and training load, practical insights in the development of overuse injuries can be obtained in a player-specific fashion contributing to injury prevention in sports. A multi-dimensional and personalized approach that includes interactions between training load variables significantly increases the understanding of overuse issues on a personal basis. Jump load is an important predictor for overuse injuries in volleyball.
Socioeconomic characteristics arc influencing the temporal and spatial variability of water demand, which arc the biggest source of uncertainties within water distribution system modeling. Improving current knowledge of these influences can be utilized to decrease demand uncertainties. This paper aims to link smart water meter data to socioeconomic user characteristics by applying a novel clustering algorithm that uses a dynamic time warping metric on daily demand patterns. The approach is tested on simulated and measured single-family home data sets. It is shown that the novel algorithm performs better compared with commonly used clustering methods, both in finding the right number of clusters as well as assigning patterns correctly. Additionally, the methodology can be used to identify outliers within clusters of demand patterns. Furthermore, this study investigates which socioeconomic characteristics (e.g., employment status and number of residents) are prevalent within single clusters and, consequently, can be linked to the shape of the cluster's barycenters. In future, the proposed methods in combination with stochastic demand models can be used to fill data gaps in hydraulic models. (C) 2021 American Society of Civil Engineers.
In the past, the training load for athletes was determined by solely relying on the expertise of the coach. Now, with the current technological developments, it is possible to perform objective measurements and investigate what the optimal training load is. In this thesis, we study the relationships between training characteristics and the Rating of Perceived Exertion (RPE) for training sessions in wheelchair tennis. Here, the training attributes are measured by sensors on the wheelchair of the athlete and the Rating of Perceived Exertion is asked from the athlete after each training session. While investigating the data, we found out that there was not enough data for male wheelchair tennis players. This is why we focused on the female athletes for the rest of the thesis. After all the preprocessing was finished, 24 training sessions were suitable for our thesis. We have used several machine learning techniques to model the Rating of Perceived Exertion. First, we applied LASSO regression and regression trees. Moreover, we used subgroup Discovery to investigate cases where big or small values give sub-optimal results. The results showed us the accuracy of the machine learning techniques. The accuracy is defined with the R 2 -score. The R 2 -score score was 0.108 for the LASSO regression and 0.180 for the regression trees. So, the model built with regression trees predicted the training load the most accurate. The parameter α of the LASSO regression model was more stable than the parameter ccp α of the regression trees model. The regression trees model showed three outliers for the ccp α . However, because these parameters have different purposes, they were both stable enough. LASSO regression showed us that the Average Velocity and the percentage of the time the athlete spends in the rotation speedzone of higher than 100 are their most important variables to quantify the relationship between the selected Rating of Perceived Exertion and the training characteristics. The coaches could try to focus more on interval training, because of the fact that average speed was an important factor for predicting the RPE. Also, it would be helpful to include specific upper-body repeated power ability drills in the physical preparation, because heavy rotating of the wheels gave the athletes a heavier training session on a physical level. With Subgroup Discovery, we were able to find out that the subgroup where there have been made less than or equal to 351 turns to the
A population group that is often overlooked in the recent revolution of self-tracking is the group of older people. This growing proportion of the general population is often faced with increasing health issues and discomfort. In order to come up with lifestyle advice towards the elderly, we need the ability to quantify their lifestyle, before and after an intervention. This research focuses on the task of activity recognition (AR) from accelerometer data. With that aim, we collect a substantial labelled dataset of older individuals wearing multiple devices simultaneously and performing a strict protocol of 16 activities (the GOTOV dataset, N = 28). Using this dataset, we trained Random Forest AR models, under varying sensor set-ups and levels of activity description granularity. The model that combines ankle and wrist accelerometers (GENEActiv) produced the best results (accuracy > 80%) for 16-class classification. At the same time, when additional physiological information is used, the accuracy increased (>85%). To further investigate the role of granularity in our predictions, we developed the LARA algorithm, which uses a hierarchical ontology that captures prior biological knowledge to increase or decrease the level of activity granularity (merge classes). As a result, a 12-class model in which the different paces of walking were merged showed a performance above 93%. Testing this 12-class model in labelled free-living pilot data, the mean balanced accuracy appeared to be reasonably high, while using the LARA algorithm, we show that a 7-class model (lying down, sitting, standing, household, walking, cycling, jumping) was optimal for accuracy and granularity. Finally, we demonstrate the use of the latter model in unlabelled free-living data from a larger lifestyle intervention study. In this paper, we make the validation data as well as the derived prediction models available to the community.
In tennis, applying a proper game strategy is an important aspect in performance optimization. In this work, we perform tactical analyses for a specific professional tennis player by using a manually annotated data collection of 4,593 points. Primarily, we will apply Subgroup Discovery to find generic characteristics of successful points in tennis and descriptions that are specific for our player of interest. To demonstrate that easily understandable patterns can be gleaned from our method, that are relatively simple to put into practice, we will focus on finding the most important descriptions of won service points. In general, the most profound characterisation of successful service points in tennis are points that last maximally two strokes. For our specific player, we have found that more service points are won if the player avoids hitting a backhand.
The ECML PKDD 2019 proceedings are dealing with machine learning and knowledge discovery in databases, including innovative applications in this area. They cover topics such as supervised learning; multi-label learning; large-scale learning; deep learning; probabilistic models; etc.
Matthijs Van Leeuwen合作论文数Machine Learning group at the Katholieke Universiteit Leuven4
Arno P.J.M. Siebes合作论文数Department of Information and Computing Sciences, Universiteit Utrecht3
Ad Feelders合作论文数Institute of Information & Computing Sciences Universiteit Utrecht3
Ulf Brefeld合作论文数Institute of Information Systems, Leuphana University of Lüneburg2