Interpretable (or explainable) machine learning models, such as decision trees, play a crucial role in the context of trustworthy AI. However, finding optimal decision trees (i.e., minimum size and maximum accuracy trees) is not a simple task and remains an active area of research. While a single decision tree has limited expressivity, using an ensemble of decision trees can effectively capture the complex structures found in many real-world applications. Many existing tree ensemble methods are greedy and suboptimal, and often suffer from randomness in the tree generation process. In this paper, we introduce DT-sampler, a SAT-based decision tree ensemble which allows explicit control over both the size and accuracy of the sampled trees. We developed a novel SAT-based encoding method that utilizes only branch nodes, resulting in a compact representation of decision tree space. Additionally, standard point predictions made using decision tree ensembles do not offer any statistical guarantee over miscoverage rate. We employ conformal prediction (CP), a distribution-free statistical framework which provides a valid finite-sample coverage guarantee, to demonstrate that DT-sampler is statistically more efficient and produces stable results when compared with random forest classifier. We demonstrate the effectiveness of our method through several benchmark and real-world datasets.
Solving black-box optimization problems with Ising machines is increasingly common in materials science. However, their application to crystal structure prediction (CSP) is still ineffective due to symmetry agnostic encoding of atomic coordinates. We introduce CRYSIM, an algorithm that encodes the space group, the Wyckoff positions combination, and coordinates of independent atomic sites as separate variables. This encoding reduces the search space substantially by exploiting the symmetry in space groups. When CRYSIM is interfaced to Fixstars Amplify, a GPU-based Ising machine, its prediction performance is competitive with CALYPSO and Bayesian optimization for crystals containing more than 150 atoms in a unit cell. Although it is not realistic to interface CRYSIM to current small-scale quantum devices, it has the potential to become the standard CSP algorithm in the coming quantum age. Predicting stable structures of large crystals, unit cells containing tens or even hundreds of atoms, has been a long-standing challenge. In this article, the authors present a combinatorial optimization-based algorithm that leverages the inherent symmetry of crystals, making it potentially achievable, especially with future implementation on quantum hardware.
Message passing neural networks have demonstrated significant efficacy in predicting molecular interactions. Introducing equivariant vectorial representations augments expressivity by capturing geometric data symmetries, thereby improving model accuracy. However, two-body bond vectors in opposition may cancel each other out during message passing, leading to the loss of directional information on their shared node. In this study, we develop Equivariant N-body Interaction Networks (ENINet) that explicitly integrates l = 1 equivariant many-body interactions to enhance directional symmetric information in the message passing scheme. We provided a mathematical analysis demonstrating the necessity of incorporating many-body equivariant interactions and generalized the formulation to N-body interactions. Experiments indicate that integrating many-body equivariant representations enhances prediction accuracy across diverse scalar and tensorial quantum chemical properties.
Multi-Objective Optimization (MOO) is an important problem in real-world applications. However, for a non-trivial problem, no single solution exists that can optimize all the objectives simultaneously. In a typical MOO problem, the goal is to find a set of optimum solutions (Pareto set) that trades off the preferences among objectives. Scalarization in MOO is a well-established method for finding a finite set approximation of the whole Pareto set (PS). However, in real-world experimental design scenarios, it's beneficial to obtain the whole PS for flexible exploration of the design space. Recently Pareto set learning (PSL) has been introduced to approximate the whole PS. PSL involves creating a manifold representing the Pareto front of a multi-objective optimization problem. A naive approach includes finding discrete points on the Pareto front through randomly generated preference vectors and connecting them by regression. However, this approach is computationally expensive and leads to a poor PS approximation. We propose to optimize the preference points to be distributed evenly on the Pareto front. Our formulation leads to a bilevel optimization problem that can be solved by e.g. differentiable cross-entropy methods. We demonstrated the efficacy of our method for complex and difficult black-box MOO problems using both synthetic and real-world benchmark data.
In predictive modelling for high-stake decision-making, predictors must be not only accurate but also reliable. Conformal prediction (CP) is a promising approach for obtaining the coverage of prediction results with fewer theoretical assumptions. To obtain the prediction set by so-called full-CP, we need to refit the predictor for all possible values of prediction results, which is only possible for simple predictors. For complex predictors such as random forests (RFs) or neural networks (NNs), split-CP is often employed where the data is split into two parts: one part for fitting and another for computing the prediction set. Unfortunately, because of the reduced sample size, split-CP is inferior to full-CP both in fitting as well as prediction set computation. In this paper, we develop a full-CP of sparse high-order interaction model (SHIM), which is sufficiently flexible as it can take into account high-order interactions among variables. We resolve the computational challenge for full-CP of SHIM by introducing a novel approach called homotopy mining. Through numerical experiments, we demonstrate that SHIM is as accurate as complex predictors such as RF and NN and enjoys the superior statistical power of full-CP.
Random forest is effective for prediction tasks but the randomness of tree generation hinders interpretability in feature importance analysis. To address this, we proposed DT-Sampler, a SAT-based method for measuring feature importance in tree-based model. Our method has fewer parameters than random forest and provides higher interpretability and stability for the analysis in real-world problems. An implementation of DT-Sampler is available at https://github.com/tsudalab/DT-sampler.
Automated high-stake decision-making, such as medical diagnosis, requires models with high interpretability and reliability. We consider the sparse high-order interaction model as an interpretable and reliable model with a good prediction ability. However, finding statistically significant high-order interactions is challenging because of the intrinsically high dimensionality of the combinatorial effects. Another problem in data-driven modeling is the effect of ``cherry-picking" (i.e., selection bias). Our main contribution is extending the recently developed parametric programming approach for selective inference to high-order interaction models. An exhaustive search over the cherry tree (all possible interactions) can be daunting and impractical, even for small-sized problems. We introduced an efficient pruning strategy and demonstrated the computational efficiency and statistical power of the proposed method using both synthetic and real data.
We present an interpretable machine learning model for medical diagnosis called sparse high-order interaction model with rejection option (SHIMR). A decision tree explains to a patient the diagnosis with a long rule (i.e., conjunction of many intervals), while SHIMR employs a weighted sum of short rules. Using proteomics data of 151 subjects in the Alzheimer's Disease Neuroimaging Initiative (ADNI) dataset, SHIMR is shown to be as accurate as other non-interpretable methods (Sensitivity, SN = 0.84 ± 0.1, Specificity, SP = 0.69 ± 0.15 and Area Under the Curve, AUC = 0.86 ± 0.09). For clinical usage, SHIMR has a function to abstain from making any diagnosis when it is not confident enough, so that a medical doctor can choose more accurate but invasive and/or more costly pathologies. The incorporation of a rejection option complements SHIMR in designing a multistage cost-effective diagnosis framework. Using a baseline concentration of cerebrospinal fluid (CSF) and plasma proteins from a common cohort of 141 subjects, SHIMR is shown to be effective in designing a patient-specific cost-effective Alzheimer's disease (AD) pathology. Thus, interpretability, reliability and having the potential to design a patient-specific multistage cost-effective diagnosis framework can make SHIMR serve as an indispensable tool in the era of precision medicine that can cater to the demand of both doctors and patients, and reduce the overwhelming financial burden of medical diagnosis.
The demand of human identification in a non-intrusive manner has risen increasingly in recent years. Several works have already been done in this context using gait-cycle detection from human skeleton data using Microsoft Kinect as a data capture sensor. In this paper we have proposed a novel method for automatic human identification in real time using the fusion of both supervised and unsupervised learning on gait-based features in an efficient way using Dempster-Shafer (DS) theory. Performance comparison of the proposed fusion based algorithm is done with that of the standard supervised or unsupervised algorithm and it needs to be mentioned that the proposed algorithm is able to achieve 71% recognition accuracy.
Measurement of cognitive load using brain signalsis an important area of research in human behavior and psychology. Recently, there have been attempts to use low cost, commercially available Electroencephalogram (EEG) devices for the analysis of the cognitive load. Due to the reduced number of leads, these low resolution devices pose major challenges in signal processing as well as in feature extraction. In this paper, we investigate the significant leads or channels that are useful for the analysis of the cognitive load. We use a standard matching test and n-back memory test imparting low and high cognitive loads respectively. The investigation is based on the analysis of variance (ANOVA) of Alpha and Theta frequency band signals for various combinations of leads. Comparisons have been done between the previously reported leads and those obtained using a few feature selection algorithms. Results indicate that for a given stimulus, though the significant leads are very much dependent on the subjects, the leads corresponding to the left frontal lobe and right parieto-occipital lobe are in general most significant across majority of subjects for analysis of the cognitive load.
Use of EEG signals in measuring cognitive load is a widely practiced area and falls under Brain-Computer-Interfacing (BCI) technology. However this technology uses medical grade EEG devices that are expensive as well as not user-friendly for regular use. Recent launch of low cost wireless EEG headsets from different companies opens up the possibility for commercialization of BCI and thus drew attention of the research community all over the world. While there are numerous studies on BCI with the use of medical grade devices there are limited numbers of papers reported on those using low cost devices. Moreover, reports on evaluating relative performance of these commercially available EEG devices based on a specific BCI experiment are minuscule. This paper attempts to fill this gap and presents a methodology to compare with various aspects between two widely used low cost wireless EEG devices namely Emotiv and Neurosky for application in cognitive load detection.
This paper proposes an extension of the traditional Fuzzy c-Means algorithm by allowing each component of the datapoints to independently contribute in the decision-making process of determining the cluster membership of the point. The above extension results in an improved accuracy in clustering. The second interesting issue undertaken here is to determine the optimum fuzziness control parameter for stabilization of the cluster centers. Lastly, the proposed extension helps in identifying the important dimensions in characterization of the datapoints. Experimental runs indicate an improvement in accuracy of clustering by the proposed algorithm in comparison to the traditional Fuzzy c-Means, with respect to the measure Fmeasure parameter by 26, 15 and 6 percentage on Colon cancer, Wine and Wisconsin Diagnostic Breast Cancer (WDBC) datasets respectively.
Individuals exhibit different levels of cognitive load for a given mental task. Measurement of cognitive load can enable real-time personalized content generation for distant learning, usability testing of applications on mobile devices and other areas related to human interactions. Electroencephalogram (EEG) signals can be used to analyze the brain-signals and measure the cognitive load. We have used a low cost and commercially available neuro-headset as the EEG device. A universal model, generated by supervised learning algorithms, for different levels of cognitive load cannot work for all individuals due to the issue of normalization. In this paper, we propose an unsupervised approach for measuring the level of cognitive load on an individual for a given stimulus. Results indicate that the unsupervised approach is comparable and sometimes better than supervised (e.g. support vector machine) method. Further, in the unsupervised domain, the Component based Fuzzy c-Means (CFCM) outperforms the traditional Fuzzy c-Means (FCM) in terms of the measurement accuracy of the cognitive load.
Feature selection is an important area of research as it has a tremendous effect on the accuracy and performance of classification algorithms. In this paper we propose an objective function for feature selection, which combines the intra class feature variation and inter class feature distance using a Lagrangian multiplier. The inter class distance is measured using the sum of absolute difference of the ratio of mean and standard deviation for respective classes. The objective function is minimized using Differential Evolutionary (DE) Algorithm where the population vector is encoded using Binary Encoded Decimal to avoid the float number optimization problem. An automatic clustering of the possible values of the Lagrangian multiplier provides a detailed insight of the selected features during the proposed DE based optimization process. The classification accuracy of Support Vector Machine (SVM) is used to measure the performance of the selected features. The proposed algorithm outperforms the existing DE based approaches when tested on IRIS, Wine, Wisconsin Breast Cancer, Sonar and Ionosphere datasets. The same algorithm when applied on gait based people identification, using skeleton datapoints obtained from Microsoft Kinect sensor, exceeds the previously reported accuracies.
This paper presents a novel method of creating synchronized interactive TV applications using Quick Response (QR) code, where QR code is used to tag the broadcasted TV content. The QR code is decoded at the receiver side to communicate with an Internet server for interactivity. The proposed technique can be used for both classical analog TV and digital TV. The method described here can also be applied on presently growing mobile TV and Over-the-Top (OTT) TV. It assumes Internet connectivity on the client. We present interactive TV based distance education as an example of a synchronized application using this system.