Accurate classification of multi-beat electrocardiogram sequences is hindered by frequent confusion between normal and supraventricular beats, severe class imbalance, and computational constraints. To address these challenges, this study proposes a lightweight multi-scale sequential model with channel recalibration. First, a multi-scale convolutional module extracts discriminative features across different temporal resolutions, preserving diagnostically important components such as P and ST segments. Second, a channel-wise attention mechanism adaptively recalibrates feature importance, thereby reducing misclassification between similar beat types. Third, an enhanced loss function combining focal loss and label smoothing alleviates class imbalance and improves generalization. Evaluated on the MIT-BIH Arrhythmia Database under strict inter-patient validation, the model achieves an overall accuracy of 99.64% and yields significant improvement in recognizing supraventricular beats, with an F1-score of 95.6% and precision of 93.5%. Cross-dataset validation on the INCART 12-lead Arrhythmia Database further confirms the generalizability of the proposed architecture, achieving 99.61% accuracy and a macro F1-score of 0.990 on the AAMI three-superclass beat-level classification task. With only 365K parameters, the model maintains low computational cost, making it highly suitable for real-time ECG analysis in resource-constrained environments.
Network intrusion detection systems face two critical challenges that significantly limit their effectiveness. Severe class imbalance leads to poor recognition of minority class attacks, while existing methods fail to adequately model dynamic spatiotemporal dependencies in network traffic. We propose DGT-IDS, a framework with three key innovations. First, we develop a diffusion-guided conditional GAN with multi-conditional control that leverages class labels, sample density, and diffusion timesteps as generation guidance for targeted minority class augmentation. Second, we design a spatiotemporal graph Transformer with semantic-based dynamic graph construction that computes adjacency matrices from feature semantics rather than fixed topologies. Third, we construct a meta-learning adaptive loss that eliminates static hyperparameters through bilevel optimization. Experiments on NSL-KDD and CIC-IDS2017 datasets demonstrate superior performance. On NSL-KDD, we achieve 88.27% F1-score and 90.79% recall. On CIC-IDS2017, we achieve 92.84% F1-score and 93.51% recall. Our method outperforms representative baselines while improving minority class detection.
Change detection (CD) from heterogeneous remote sensing imagery remains a persistent challenge due to the inherent differences in imaging mechanisms across sensors, which complicates the derivation of accurate change maps. Nevertheless, existing CD methods often suffer from limited utilization of the complementary characteristics of heterogeneous data, excessive dependence on labeled samples, and a lack of model interpretability. To address these issues, in this paper, we propose a novel self-supervised transformation (SST) approach for heterogeneous CD by mapping domain consistency in an unsupervised manner. SST integrates self-supervised learning and evidence fusion to automatically mine reliable changed/unchanged pixel pairs and guide the domain transformation across heterogeneous images. Specifically, we introduce an image pixel discretization (IPD) method to enhance inter-object differences, thereby improving the granularity of superpixel segmentation. This is followed by a spectral-textural difference analysis technique that identifies significantly changed, unchanged, and uncertain pixel pairs. To unify heterogeneous data representations, we propose a multi-medium bidirectional pixel transformation (MBPT) strategy based on multivalue transfer estimation, which projects heterogeneous image pairs into a shared feature space. Finally, an evidence fusion with discounting factors-based augmentation enables direct generation of the change map without requiring an intermediate difference image. Comprehensive experiments conducted on multiple real-world heterogeneous remote sensing datasets demonstrate that SST achieves superior performance. Notably, SST achieves a Kappa coefficient of 0.8125 on the Sardinia dataset and 0.8768 on the Gloucester dataset, outperforming several state-of-the-art methods.
Clustering, as a fusion process, involves aggregating similar objects and isolating dissimilar ones, independent of any prior information. Recently, evidential clustering has gained popularity due to its ability to characterize the uncertainty and imprecision of data distribution. However, it remains a major bottleneck of existing evidential clustering methods for clustering imbalanced data, as they cannot effectively detect small clusters (with a few objects). In this paper, we propose a new aggregation-based self-supervised evidential clustering (ASEC) method for dealing with such issues based on the theory of belief functions. Specifically, a cluster density-based aggregation rule is designed first to generate multiple sub-clusters and then fuse them into new singleton clusters, which can effectively detect small clusters of imbalanced data. The new singleton clusters obtained by the aggregation rule serve as prior knowledge. Then, a self-supervised evidential partition rule is developed to fuse the remaining objects into new clusters according to prior knowledge and the K-nearest neighbors (KNNs) technique. In this process, the objects in the overlapping zones of clusters are usually hard to classify, and they are assigned to new meta-clusters to reduce the risk of error. Experiments on several imbalanced datasets demonstrate the effectiveness of ASEC compared to related methods.
In the realm of information fusion, clustering stands out as a common subject and is extensively applied across various fields. Evidential clustering, an increasingly popular method in the soft clustering family, derives its strength from the theory of belief functions, which enables it to effectively characterize the uncertainty and imprecision of data distributions. This survey provides a comprehensive overview of evidential clustering, detailing its theoretical foundations, methodologies, and applications. Specifically, we start by briefly recalling the theory of belief functions with its transformations into other uncertainty reasoning theories. Then, we introduce the concepts of soft data, partitions, and methods with an emphasis on data and partitioning within the theory of belief functions. Subsequently, we summarize the advancements and quantitative evaluations of existing evidential clustering methods and provide a roadmap to help in selecting an appropriate method based on specific application needs. Finally, we identify the major challenges faced in the development and application of evidential clustering, pointing out promising avenues for future research, including theoretical limitations, applicable datasets, and application domains. The survey offers a structured understanding of existing evidential clustering methods, highlighting their theoretical underpinnings, practical implementations, and future research directions. It serves as a valuable resource for researchers seeking to deepen their understanding of evidential clustering.
Over-sampling methods concentrate on creating balanced samples and have proven successful in classifying imbalanced data. However, current over-sampling methods fail to consider the uncertainty of produced samples, potentially altering the data distribution and impacting the classification process. To address this issue, we propose a distribution assessment-based multiple over-sampling (DAMO) method for classifying imbalanced data. We first introduce a multiple over-sampling method based on distribution assessment to create different forms of synthetic samples. The core is quantifying the inconsistency of data distribution before and after sampling as a constraint to guide multiple over-sampling, thereby minimizing the data shift and characterizing the uncertainty of produced samples. Then, we quantify the local reliability of the classification results and select several imprecise samples with low local reliability that are indistinguishable between classes. Neighbors serve as additional complementary information to calibrate the results of imprecise samples, thereby reducing the likelihood of misclassification. The calibrated results are combined by the discounting Dempster-Shafer fusion rule to make a final decision. DAMO's efficiency has been demonstrated through comparisons with related methods on various real imbalanced datasets.
The issue of clustering missing data is a hot topic that remains challenging. This is because imputation-based clustering methods that provide inaccurate estimations may cause uncertainty and imprecision, leading to a negative impact on clustering accuracy. Moreover, missing values may result in incomplete objects that are indistinguishable between clusters, especially when critical attribute values are lost. To address these issues, a partial distance evidential clustering (PEC) method is introduced in this study. Specifically, the (complete or incomplete) object is preliminarily partitioned as certain or uncertain according to evidential clustering based on the partial distance without any imputation strategy. In this manner, we can fully model the data structure and avoid the negative effects of inaccurate estimations on the clustering process. In addition, an uncertain object with missing values is input multiple times to obtain different complete versions, modeling the uncertainty of the estimations. The clustering results of the versions are combined to make a decision under the Dempster–Shafer theory framework. In this process, the objects in the overlapping zones of clusters are usually indistinguishable between clusters, and they are assigned to meta-clusters to characterize imprecision and reduce the risk of errors. Experiments on a synthetic toy and several real-world datasets demonstrate the effectiveness of the proposed PEC method compared with related methods.
Recently, many infrared small target detection (IRSTD) methods have relied on U shaped neural models with encoder-decoder architectures based on Convolutional neural networks or Transformers. However, the former have a limited receptive field, restricting them to local perception and making them less effective at handling global information. On the other hand, the computational complexity of Transformers grows quadratically with the sequence length, requiring significant computational resources, which makes both training and inference challenging. To address this, we introduce the Spatial-BiDirectional Mamba Network (SBMambaNet), which leverages Spatial-BiDirectional Mamba blocks (SBMBs) working in parallel with Skip Connections. Specifically, it consists of two key components: (1) Spatial-BiDirectional Cross Mamba (SCM) captures local spatial information in sequences, enhances the exchange between local spatial features and global channel data, models bidirectional sequences for global context, and fuses features via gated mechanisms to strengthen longrange dependency and information flow, and (2) Mixed-Scale Spatial-Gate Feed-forward Network (MSGFN), which enhances feature discriminability through a multi-scale strategy and cross-space-channel information interaction, supporting the efficient transfer of beneficial information. Furthermore, we further develop the CNN-Mamba Channel-wise Fusion module (CMCFusion) to guide the effective fusion of multi-scale channel information and pixel-level spatial relationships at the decoder stage, thereby resolving ambiguities. Our model demonstrates robust segmentation and generalization performance in the field of infrared image segmentation, and experimental results on public datasets also highlight its exceptional efficiency.
Fuzzy clustering is still a hot topic because it can calculate the support degrees of an object belonging to different clusters to characterize uncertainty. However, it remains a challenge to detect clusters of arbitrary shapes, sizes, and dimensionality. What is worse, some objects are indistinguishable (imprecise) when they are in the overlapping regions of different clusters. To address such issues, this article investigates a belief-based fuzzy and imprecise clustering (BFI) method, which can detect arbitrary clusters and provide the behavior (support) of objects to these clusters. Moreover, BFI can assign each imprecise object to a meta-cluster, defined as the union of specific clusters, to characterize (partial) imprecision. The proposed BFI can significantly reduce the risk of misclassification, and the effectiveness is validated in image processing (e.g., image segmentation and classification) and several benchmark datasets by comparing it with some typical methods.
Classification of incomplete data remains a challenging task since the distribution of training and test sets may be inconsistently caused by missing values. To address such a problem, this paper investigates an incomplete data transfer calibration classification (IDTC) method based on Dempster-Shafer theory to improve the accuracy by modeling uncertainty and imprecision due to missing values. The proposed IDTC method consists of three aspects. First, incomplete samples are imputed by complete neighbors, and the trained basic classifier then classifies the test sample to obtain the preliminary result. Second, a data transfer framework is presented to map the test sample to the training set. Afterward, the mapped training samples optimize multiple calibration matrices to calibrate the preliminary classification result. The calibration matrix is learned by minimizing the deviation between the classification results of mapped training samples and the truth. Third, each calibrated classification result is considered as a piece of evidence under Dempster-Shafer theory. As a result, these pieces of evidence with different discounting factors are fused to make the final decision. Finally, the effectiveness of the proposed IDTC is widely validated on real datasets by critically comparing to other typical methods.
Classification of missing data based on estimation is still challenging since existing methods relying on one imputation strategy fail to consider the diversity of different attribute distributions. In this case, there are inevitably some “bad” estimations at the attribute level, reducing the performance of classification. This article proposes a mixed-type imputation method (MTI) to classify missing data under the theory of belief functions (TBF) via two quality matrices to address this problem. The proposed MTI method has the advantages of making estimations as close to the truth as possible at the attribute level while reducing the negative impact of possible bad estimations on the classification. Specifically, the first matrix used to impute missing values can characterize the different supports of multiple imputation methods for estimating various attributes. The other matrix used to perform the classification task can extract the reliabilities of estimations on the different classes. The validity has been demonstrated in the final decision support based on the TBF, famous for characterizing uncertainty and imprecision, for example, caused by missing values.
Classifying incomplete data remains a challenging task, as missing values can provide uncertain and imprecise information that reduces classification performance. To address this issue, we proposed a hybrid imputation-based optimal evidential classification (HOEC) method for missing data under the Dempster-Shafer theory framework. The proposed HOEC method can capture uncertainty and imprecision during imputation and classification procedures. Specifically, a hybrid imputation strategy was developed to estimate the missing values in the training and test sets by combining single and multiple imputations. Thus, we obtained accurate estimations and captured their uncertainties. An optimal evidential partition rule was then designed to adaptively submit an incomplete sample to a singleton class or meta-class under the Dempster-Shafer theory framework. Therefore, we can capture the imprecision caused by missing values and reduce classification errors. Experiments on several incomplete datasets demonstrated the effectiveness of the HOEC method compared with related methods.
Oversampling methods concentrate on creating a balanced dataset by generating samples, widely utilized in classifying imbalanced data. However, current oversampling methods overlook the uncertainty in the samples produced, potentially shifting the data's distribution and adversely affecting the classification outcomes. To address this problem, we introduce a multi-oversampling with evidence fusion (MOEF) method for imbalanced data classification based on Dempster-Shafer theory. We first design a multi-oversampling strategy to produce various balanced datasets, characterizing the uncertainty of generated samples. Then, we develop a discounting fusion rule based on the inconsistency of data distribution post-oversampling, thereby mitigating the adverse effects of data distribution alterations on classification. Extensive testing on various imbalanced datasets indicates that the proposed MOEF method exhibits more satisfactory performance than other related methods.
Aiming at the problems such as low accuracy of pedestrian detection and poor effect of small target pedestrian detection by traditional detection model, an improved YOLOv7 pedestrian detection algorithm is proposed. First of all, in order to improve the positioning accuracy and reduce the target missing detection rate, the NWD module for small target pedestrian detection is used to improve the IOU. Mean-while, the improved CBAM attention mechanism is introduced in this paper, and the optimized C3-Faster module is introduced into the Backbone of YOLOv7 to improve the model accuracy. Combined with the improved CBAM attention mechanism, a new lightweight and high-precision CB $A$ MC3Fast module is proposed. To test model performance, a new data set is built based on CUHK Occlusion Datase data set. Several comparison experiments are designed to verify the effectiveness of the algorithm. The experimental results on the self-established data set showed that compared with the previous model, the improved model achieved 2.20% improvement in Precision, 1.97% improvement in Recall, and 1.34 % reduction in box_loss, thus reducing the false detection rate. The weight size of the pro-posed algorithm is only 40.8MB, which has a smaller model volume and higher detection accuracy, so that the algorithm can be adapted to more scenarios, and can play an important role in the fields of traffic management, epidemic prevention and temperature measurement and video surveillance.
Over-sampling approaches focus on generating samples to balance the dataset and have been widely applied in classifying imbalanced data. However, existing approaches do not take into account the uncertainty of generated samples, which may alter the data distribution and introduce uncertain information into the classification process. To tackle this issue, we propose a multiple adaptive over-sampling approach (MAO) for classifying imbalanced data based on evidence reasoning. First, we construct balanced training sets through multiple adaptive over-sampling for the minority class, which characterizes the uncertainty of over-sampling. Then, we define the intra- and inter-class inconsistency of data distribution after over-sampling to quantify the weights of different classifiers trained by various balanced subsets, weakening the negative impact of changes in data distribution on classification. Finally, we employ neighbor information to revise the results of samples that are hard to classify correctly, to avoid the risk of misclassification caused by uncertain synthetic samples to some extent. The effectiveness of MAO has been verified on several real imbalanced datasets by comparing it with other related approaches.
The classification analysis of incomplete data is an important and challenging topic in machine learning. Many approaches have been devised to cope with incomplete data. They usually consider that the training and test sets are independent and identically distributed, but the distribution of training and test sets may be inconsistent caused by missing values in applications. In this paper, we propose a novel evidential classification approach to address such a problem based on the Dempster-Shafer theory. First, attributes with missing values are combined with other high correlation attributes to generate different subsets, and they are imputed by K nearest neighbors (KNNs) in subsets. Basic classifiers trained by edited subsets are employed to classify the test sample. Second, the mixed discounting factor composed of the importance and reliability factors is designed to calibrate classification results of the test sample. The importance is evaluated by the difference of distribution between training and test subsets, and the reliability is quantified by minimizing the deviation between classification results of training samples and the truth. Different classification results are combined with mixed discounting factors by the Dempster-Shafer (DS) fusion rule thereby making the final decision. We conduct extensive experiments with several real incomplete data, and the results show that the proposed approach yields more promising and stable performance with respect to other typical approaches.
Remote sensing image change detection remains a challenging task. Most existing approaches are based on fully supervised learning, but labeled data are so scarce for change detection. It is difficult to exhibit high detection performance with a limited amount of labeled data. In this paper, we propose a semi-supervised Label Propagation (SSLP) approach for multi-source remote sensing image change detection. First, a clustering label propagation (CLP) method is designed to cluster pre and post images, respectively, and assign pseudo labels to unlabeled pixel pairs that have similar mapping relationships to labeled pixel pairs. Second, a pixel density metric is investigated to filter out the data with low density and retain the data with high density, which can ensure the reliability of the propagated data. Third, a secondary expansion method based on pixel neighborhood is used to generate enough training data for training a classifier. Finally, the effectiveness of SSLP is validated on three real datasets by comparing to other related methods.
针对传统分类模型在处理不平衡数据时会侧重于大类而忽略小类的问题,提出了一种复合可靠性分析下的不平衡数据证据分类方法,通过评估分类模型的全局可靠性和局部可靠性来提升模型对每个不平衡测试样本的分类能力.首先,对大类多次降采样,采样后的数据与小类组成多个训练子集,用这些子集训练得到多个分类模型,通过最大均值差异度量采样前后数据分布的差异性得到不同分类模型的全局可靠性.其次,利用待测样本在训练集中的近邻来评估其分类结果的局部可靠性,待测样本与其近邻具有相似的数据分布和空间结构,分类模型对近邻的分类结果与真实类别偏差越小,其局部可靠性就越大.最后,在证据推理框架下,全局可靠性与局部可靠性组合为复合可靠性因子对不同分类模型得到的分类结果进行折扣,将部分概率值分配给完全未知类来表征数据类别的不确定性,用 Dempster-Shafer(DS)规则融合多个折扣后的分类结果做决策分析.实验结果表明:所提方法对KEEL和UCI数据库的 12 个不平衡数据分类结果的平均FM 为 80.18%,GM 为 87.24%,相较于其他不平衡数据分类方法中最优结果分别高出 8.1%和 4.99%.所提方法的有效性在不平衡数据分类中得到了证实.
The classification analysis of missing data is still a challenging task since the training patterns may be insufficient and incomplete in many fields. To train a high-performance classifier and pursue high accuracy, we learn a credal classifier based on an optimized and adaptive multiestimation (OAME) method for missing data imputation on training and test sets. In OAME, some incomplete training patterns are estimated as multiple versions by a global optimization method thereby expanding the training set. On the other hand, the test pattern is adaptively estimated as one or multiple versions depending on the neighbors. For the test pattern with multiple versions, the corresponding outputs with different discounting factors (weights), represented by the basic belief assignments (BBAs), are fused for final credal classification based on evidence theory. The discounting factor contains two aspects: the importance and reliability factors that are used, respectively, to quantify the importance of the edited version itself and to represent the reliability of the classification result of the version. The effectiveness of OAME is widely validated on several real datasets and critically compared to other related methods.