
Machine learning models are widely used in computer vision and classification tasks. However, imbalanced classification biases predictive models toward larger classes, reducing predictive performance for minority classes. To address this challenge, we propose meta-adaptive resampling selection plus plus (MARS + +), a stability aware meta-adaptive framework that chooses the best method based on the data. It can select from different types, such as resampling, hybrid, and algorithm-based methods. The proposed method uses inner cross-validation to test each option. It selects methods based on both performance and stability by combing predictive performance and variability across folds (mean − λ·std). It also avoids using any method if none gives clear improvement. MARS + + is tested on multiple benchmark datasets with varying imbalance levels, including both tabular and image data. We use two classifiers, the random forest and logistic regression. The results show that no single method is always the best. MARS + + still gives results close to the best choice in most cases. It also avoids the large drops in performance seen in some other methods. Statistical tests support the effectiveness of adaptive selection. This shows the importance of selecting methods based on the data. In addition, we provide a detailed analysis of method selection, and also examine cases where no method is selected. Finally, the study analyzes how close the results are to the best possible performance. It also examines how dataset features, such as imbalance and sample size, affect the method selection. These results suggest that selecting methods based on stability worked well on the evaluated datasets. They also support using an adaptive imbalance handling strategy.
Major Depressive Disorder (MDD) affects people worldwide, although electroencephalography (EEG) provides an objective neurophysiological basis for depression assessment, existing graph-based EEG frameworks suffer from several limitations. These include single-threshold brain network construction, decoupled spatiotemporal modeling, noise-sensitive population graph formation, and inadequate handling of inter-subject domain heterogeneity. To address these challenges, a Multi-Granularity Domain-Partitioned Hypergraph Spatiotemporal Network (MG-DPHSN), a novel unified end-to-end framework for EEG-based depression detection, is proposed. MG-DPHSN incorporates three complementary modules. First, a Multi-Granularity Masked Relational Learning (MG-MRL) module constructs a hierarchy of Binary Brain Networks across progressively refined threshold levels, preserving weak yet functionally relevant cortical connections. Second, an Attention-Gated 3D Graph Convolution (AG-G3D) module integrates a learnable channel-frequency attention matrix into a unified graph convolution operation, enabling simultaneous aggregation of spatial, temporal, and spectral information. Third, a Hypergraph-Embedded Secondary Subject Partitioning (H-SSP) module employs Maximum Mean Discrepancy (MMD)-based clustering to partition subjects into biologically coherent sub-domains, thereby restricting hypergraph message passing to distributionally similar subjects and mitigating inter-subject domain shift. Experiments conducted on the OpenNeuro (ds003478) demonstrate the effectiveness of the proposed framework. MG-DPHSN achieves 87.50
In recent years, constrained multi-objective optimization problems(CMOPs) remain challenging due to the complex structure of feasible regions, the difficulty of balancing convergence and diversity, and the lack of adaptive operator scheduling mechanisms. To address these issues, this paper proposes a hierarchical reinforcement learning–based subtask-coordinated scheduling method for constrained multi-objective evolutionary algorithm (HRL-SCMOE). The proposed framework employs a two-level architecture, where a high-level agent dynamically schedules subtasks–such as forward-oriented exploration, feasibility-driven exploitation, and diversity guidance–according to the environmental state, while a low-level agent adaptively selects variation operators tailored to each subtask. Both agents are trained using Double Deep Q-Networks (Double DQN) and Prioritized Experience Replay (PER) to enhance stability, sample efficiency, and value estimation reliability. Moreover, the algorithm constructs a set of collaborative information pools targeting different search objectives to maintain balanced exploration between feasible and infeasible regions. An adaptive reward mechanism and soft target updates are also incorporated to improve robustness in hierarchical policy learning. Experimental results on three benchmark test suites and four real-world application domains demonstrate that the proposed method consistently outperforms nine state-of-the-art constrained multi-objective evolutionary algorithms (CMOEAs) in terms of convergence, feasibility, and diversity, thereby confirming its effectiveness and strong general applicability.
This paper addresses the coordinated simultaneous-arrival path planning problem for multiple amphibious unmanned aerial vehicles (UAVs) operating under heterogeneous speed constraints in complex amphibious environments. Unlike conventional UAVs, amphibious UAVs must traverse both aerial and aquatic domains, which imposes distinct speed constraints and dynamic adaptability requirements. The objective is to generate collision-free, smooth trajectories that enable all UAVs to reach a common target simultaneously while respecting individual speed limits, avoiding terrain obstacles, and preventing inter-vehicle collisions. To solve this problem, we propose a novel algorithm, termed NEL_MSCPSO (neighborhood elite learning-based multi-strategy cooperative particle swarm optimization). The algorithm integrates hierarchical neighborhood reconstruction with cross-subswarm elite learning, weighted centroid-guided follower updates, fully adaptive parameter adjustment, a hybrid Gaussian-Cauchy mutation scheme with elite protection, adaptive dimensional mutation on the global best, and an enhanced differential evolution-based terminal replacement strategy. Comprehensive experiments on the CEC2022 benchmark suite demonstrate that NEL_MSCPSO achieves the lowest total rank sum among twelve state-of-the-art metaheuristic algorithms. Ablation studies confirm the necessity of each component. More importantly, the algorithm is successfully applied to four multi-amphibious UAV path planning scenarios of increasing spatial complexity, consistently producing feasible trajectories that strictly satisfy simultaneous-arrival constraints under heterogeneous speed profiles. These results demonstrate the superior engineering feasibility and robustness of NEL_MSCPSO for amphibious UAV coordination tasks.
Multimodal large language models offer strong potential for adaptive and explainable support in creative learning, but many existing systems mainly focus on content generation rather than structured guidance, progress tracking, and interpretable feedback. This paper presents TD-GAE, a Multimodal Large Language Model framework for adaptive and explainable guided art learning. The proposed method combines staged guidance using a directed acyclic graph, learner-state estimation, retrieval-augmented pedagogical context, structured feedback actions, confidence-gated critique, and safety controls to preserve learner agency and originality. Experiments were conducted on six public datasets covering sketch recognition, sketch-photo retrieval, aesthetic assessment, and fine-art style analysis, with comparisons against prompt-only, retrieval-based, planner-only, and agentic baselines. Results show that TD-GAE improves sketch interpretation by 10.2 percentage points on Quick, Draw! and 8.7 points on TU-Berlin over the prompt-only baseline. It also achieves a 22.2
Document-level relation extraction (DocRE) finds relations across a whole document. It often needs evidence from several sentences. It also needs to link repeated entity mentions and follow multi-hop clues. Many large language model (LLM) based DocRE methods use fixed relation descriptions. They make little use of confident errors or samples with missing labels. We propose SPO-PA, which combines self-correcting prompt optimization (SPO) and preference alignment (PA). SPO compares LLM predictions with reference annotations. It then revises relation descriptions, clarifies relation boundaries, and adds role constraints. PA uses false negative (FN) conflicts as the main supervision signal. It also tests false positive (FP) conflicts under similar constraints. PA applies source-specific constraints to these signals. The constraints help smaller DocRE models separate correct relations from wrong candidates. On Re-DocRED, SPO improves F1 to 26.23 PA_FN gives the most stable gains on ATLOP, KD-DocRE, DREEAM, and DAATF. The largest gain is 2.89
In post-disaster environments, the failure of terrestrial communication infrastructure necessitates the rapid deployment of unmanned aerial vehicles (UAVs) as aerial base stations to restore wireless connectivity. This paper addresses the joint UAV activation-and-placement problem in continuous space, with the objective of minimizing the number of deployed UAVs while satisfying coverage and minimum-separation constraints. To solve this problem, we propose a Hybrid K-means Quantum-Inspired Evolutionary Algorithm (HKQEA) that combines K-means-guided initialization, a calibrated penalty-based feasibility objective, non-elitist evolutionary search, and a quantum-inspired learning update. Experimental results over 50 independent runs show that HKQEA attains a best fully feasible solution with 8 UAVs, while achieving average values of 98.94
Multi-view multi-label classification has attracted increasing attention because it can characterize complex real-world data from multiple perspectives. However, the simultaneous presence of missing views and labels often causes incomplete feature information and uncertain data distributions, hindering the extraction of discriminative semantics. To address this problem, we propose a unified framework termed Reliable Cross-View Neighborhood Relation Transfer and High-Order Semantic Structure Mining (ReCHSM). Specifically, we design a reliable cross-view completion mechanism that transfers neighborhood relations from observed views to reconstruct missing ones. To reduce erroneous transfer caused by cross-view heterogeneity, the mechanism evaluates candidate-neighbor reliability and employs a lightweight nonlinear network to refine the reconstructed features. To further extract structural information from the completed representations, we introduce a random-walk-based strategy that propagates the repaired local neighborhoods to capture high-order semantic dependencies. Furthermore, considering that individual views provide different discriminative information and that the quality of completed views may vary, we develop a label-semantic-consistency-driven module to assess view quality and adaptively assign fusion weights. Extensive experiments on five public datasets demonstrate the effectiveness and competitive performance of ReCHSM.
UAV-assisted mobile edge computing (MEC) provides a flexible way to process computation-intensive tasks for ground users, but the open air-to-ground transmission links expose offloaded data to eavesdropping risks. Existing secure offloading schemes often apply fixed protection policies to all tasks, which may introduce unnecessary overhead for low-sensitivity data while providing insufficient adaptation for highly sensitive data. To address this issue, this paper proposes a rating-aware graded security offloading framework for UAV-assisted MEC networks. Each task is assigned a normalized sensitivity rating, and the rating is mapped to a security-strength coefficient that affects encryption/decryption delay, security-related energy consumption, and offloading cost. Based on this model, the offloading decision among local computing, UAV-MEC computing, and cloud computing is jointly optimized with transmission power and computing resource allocation. The resulting mixed-integer non-convex problem is solved by a tailored alternating optimization algorithm. Simulation results under the same baseline parameter settings show that the proposed framework maintains a higher offloading ratio and lower system cost than the fixed-security and no-rating baselines in the considered hover-based UAV-MEC scenario.
The rapid build-out of electric-vehicle (EV) charging infrastructure has created a protocol-rich cyber-physical attack surface that spans many independently operated charging networks. Anomaly detection across such networks is most effective when operators pool experience, yet session-level charging telemetry is commercially sensitive, privacy-regulated, and non-IID, so it cannot be centralized. Federated learning (FL) removes the need to share raw data, but plain FL leaks information through shared updates, is fragile under poisoning by malicious participants, and provides no verifiable accountability among mutually distrustful operators. We present PFAD-BC, a privacy-preserving, Byzantine-resilient, and blockchain-anchored federated anomaly-detection framework tailored to multi-operator Charge Point Operator (CPO) consortia. PFAD-BC couples an Open Charge Point Protocol (OCPP)-aware feature pipeline and an LSTM-autoencoder local detector with differentially private training under Rényi-DP accounting, a reputation-weighted cosine-similarity aggregation rule that tolerates up to f < K/3 Byzantine operators, and three permissioned-ledger contracts that record update commitments, enforce declared per-operator privacy-cap admission, register aggregation attestations, and maintain reputation state. Across two datasets—ACN-Data and the labeled-attack CICEVSE2024 benchmark—and nine baselines, PFAD-BC attains F1 = 0.736 at ϵ = 1.0 on ACN-Data (within 0.006 of the non-private centralized ceiling) and F1 = 0.896 on CICEVSE2024, leading every privacy-protected baseline. It sustains F1 ≥ 0.68 under sign-flip, Fang, and LIE attacks at a 40
Computational cost during model deployment can be reduced through knowledge distillation (KD) which transfers knowledge from a teacher model to a lightweight student model while maintaining performance. However, KD methods based on a single teacher often provide limited knowledge diversity and may transfer less informative features. Furthermore, the application of KD for plant disease classification in low-light noisy environments has received little attention. To address these limitations, we propose a KD framework using heterogeneous teachers with adaptive feature alignment (KDHT-AFA) that integrates an adaptive feature distillation switch (AFDS) to selectively align intermediate representations. For effective knowledge transfer, we propose a lightweight custom student model with 5.16 million parameters. The model incorporates a regional statistical hybrid attention (RSHA) to enhance important features semantically while increasing sensitivity to local structure and contrast. Additionally, we introduce a new self-collected real low-light plant disease dataset (ReLL-PDD v1) acquired under realistic agricultural conditions. For performance evaluation, we considered two open-source datasets, PlantVillage and the Potato Leaf Disease dataset, along with the ReLL-PDD v1 dataset. Our model achieved a mean accuracy of 86.88
Scheduling AI inference across heterogeneous edge-cloud chips requires balancing energy, latency, cost, and thermal feasibility. This study presents a chip-aware scheduling framework, rather than a new multi-objective reinforcement learning theory. A 53-dimensional state describes directed-acyclic-graph tasks, CPU/GPU/NPU/FPGA status, network conditions, and queue slack. A fuzzy controller adjusts energy, latency, and cost priorities, while a hybrid-action Proximal Policy Optimization policy jointly selects the offloading target, physical chip, and dynamic-voltage-and-frequency-scaling coefficient. Dependency, deadline, thermal, and bandwidth constraints are enforced through action masks and residual penalties. In traffic-video and industrial-inspection simulations, the framework achieved 5.8 TOPS/W, 132 ms average latency, a normalized cost coefficient of 0.17 per task, and 85.2
Face recognition has been widely applied in identity authentication systems, but its vulnerability to various presentation attacks poses significant security risks. As a result, face anti-spoofing (FAS) has become one of the key technologies for ensuring the reliability of such systems. Most existing domain- generalized FAS (DGFAS) approaches rely on domain distribution alignment to learn cross-domain invariant representations. However, many of these methods independently model each sample while overlooking semantic consistency among cross-domain samples and different regions within each sample, making them vulnerable to semantic drift. Moreover, some existing methods introduce additional pixel-level supervision signals, such as pseudo-depth maps and binary masks, which increase annotation costs and limit their generalizability. To address these challenges, we propose DSCM-FAS, a novel framework that combines inter-sample and intra-sample semantic consistency to enhance robustness and cross-domain generalization, without requiring auxiliary supervision. Specifically, we design a Dual Semantic Consistency Module (DSCM). Across samples, contrastive learning is leveraged to learn discriminative representations, followed by the adaptive construction of a high-confidence similarity adjacency graph via statistical thresholding. The resulting graph is then fed into a graph convolutional network (GCN) to strengthen cross-domain semantic consistency, thereby improving the model's robustness to domain shifts. Within each sample, the intermediate embedding features are first partitioned into multiple patches, and Laplacian Regularization is introduced to constrain the semantic relationships among different patches. This effectively suppresses local noise interference and promotes the learning of more robust and semantically consistent feature representations. Extensive experiments on four public FAS datasets demonstrate that DSCM-FAS consistently outperforms state-of-the-art methods under various protocols. These results validate the effectiveness of our DSCM-FAS in improving the cross-domain generalization of FAS models.
Multimodal misinformation often reuses authentic images while shifting key contextual details such as time, location, identity, or event description, making verification harder than ordinary image-text matching. We study this problem under a bounded-evidence setting, where verification relies on limited external evidence and, when available, early social traces collected within a fixed post-publication window. We propose SA-GGCoT, a three-stage framework that uses checkpoint-aware state tracking to localize conflicting evidence, LLMs to convert selected evidence into structured rationales, and a heterogeneous evidence graph to verify evidence consistency. Experiments on Weibo, PHEME, and MR ^2 show consistent gains over protocol-matched baselines, with ablation and grounding analyses supporting the contribution of each component.
This paper tackles the challenges of affordance segmentation in complex multi-object, multi-function scenarios, particularly addressing issues such as occlusion, boundary ambiguity, and semantic confusion. To overcome these challenges, we propose a comprehensive solution that integrates model architecture design, attention mechanisms, and real-time inference optimization. Specifically, we develop two flexible frameworks based on Multi-Shift Hybrid Attention (MSHA) modules configured in both sequential and parallel forms to enhance segmentation performance. For high-precision requirements, we introduce SMSHA-Segmenter, a two-stage system that employs a sequential MSHA-enhanced decoder alongside RTMDet as the object detector. For real-time applications, we present RTMAff, an end-to-end extension of RTMDet that embeds both sequential and parallel MSHA modules to improve semantic consistency and multi-scale adaptability. Experimental results on the IIT-AFF dataset demonstrate the effectiveness of our approach: the two-stage SMSHA-Segmenter achieves an accuracy of 84.62
Deep neural networks achieve state-of-the-art performance across computer vision, natural language processing, and scientific discovery, yet their internal decision-making mechanisms remain difficult to interpret. Post-hoc explanation methods such as SHAP, saliency maps, and attention visualization provide empirical insights but offer limited mathematical guarantees about learning dynamics and representation formation. This survey takes a complementary, theory-grounded perspective, examining interpretability through four classical machine learning (ML) frameworks, each grounded in a mature mathematical theory. These are kernel methods (functional analysis), sparse representations (convex optimization), matrix factorization (linear algebra), and manifold learning (differential geometry). We synthesize and critically evaluate 203 papers (2016–2026), organizing them by the theoretical principle each framework contributes rather than by application domain. Neural tangent kernel (NTK) methods provide rigorous convergence and generalization guarantees but depend on idealized infinite-width assumptions; sparse representation methods yield interpretable-by-design architectures but face reconstruction and approximation trade-offs; matrix factorization reveals an implicit bias toward low-rank structure that helps explain generalization; and manifold learning offers geometric tools for analyzing representation structure, class separability, and generative behavior. Beyond reviewing each framework, we establish explicit mathematical bridges between them, arguing that they offer complementary views of a common phenomenon, namely that gradient-based optimization often drives deep networks toward low-dimensional, interpretable structure. We further situate these frameworks among adjacent interpretability paradigms, including mechanistic, concept-based, and attribution-based methods, and delineate the validity conditions and failure modes that determine when classical theory yields reliable rather than misleading explanations. The review concludes with open problems and directions for scalable, theory-grounded interpretability.
Conventional evidential deep learning (EDL) represents evidence strength and vacuity but typically converts the learned evidence into a single-label prediction. This may conceal class ambiguity when multiple classes remain plausible and lead to unreliable precise decisions. To address this problem, this paper proposes an evidential set-valued classification framework, termed EDL-SC. First, the decision space is extended from singleton labels to both singleton and class-pair actions, allowing class ambiguity to be explicitly represented. Second, a set-aware supervision term is incorporated into evidential learning to align the learned evidence with the expanded decision space. Third, a maximum expected utility rule is used to select between precise and set-valued predictions, with a risk preference parameter controlling the trade-off between true-class coverage and prediction imprecision. Extensive experiments demonstrate that EDL-SC consistently improves true-class coverage over precise EDL while maintaining selective set-valued outputs. Compared with randomized Adaptive Prediction Sets, it produces more compact prediction sets and achieves a more favorable utility-based balance between coverage and decision specificity. Further analyses show that MEU inference enables the transition from forced precise predictions to set-valued decisions, while set-aware learning helps refine when such decisions should be issued. Overall, EDL-SC provides a framework that aligns evidential learning with MEU-based set-valued decision making, thereby supporting the explicit representation of class ambiguity and controllable decision making under uncertainty.
Manual visual assessment of tea leaf diseases is labour-intensive, subjective, and difficult to scale across large plantation environments. Accurate and deployable image-based recognition is important for timely tea crop management. To address the accuracy–efficiency trade-off in mobile tea leaf disease classification, this study proposes a lightweight deployment-aware framework based on ECA-MBConv-Net and a hybrid Dung Beetle Optimizer–RIME algorithm, termed HDBO-R. The proposed network integrates Efficient Channel Attention into MBConv-style depthwise separable blocks to enhance disease-relevant channel representation with limited computational overhead. HDBO-R searches architectural and training configurations using a deployment-aware objective that jointly considers macro-F1, parameter count, FLOPs, and CPU inference latency. Experiments were conducted on three datasets: the authors’ Sri Lankan tea leaf disease dataset, an external Indian tea leaf disease dataset, and a Taiwan tomato leaf dataset for cross-crop evaluation. On the SLTea dataset, ECA-MBConv-Net achieved 99.33
Hyperspectral image (HSI) band selection often suffers from band redundancy and information loss due to the neglect of spectral characteristics. To address this issue, this paper proposes a semi-supervised multi-objective particle swarm optimization with boundary selection and curve fitting (MOPSO_BSCF) method for band selection in hyperspectral image classification. The method establishes a dual-objective evaluation framework: the criterion based on Fisher discriminant analysis enhances class separability, while the criterion based on information entropy improves information richness and reduces band redundancy. MOPSO_BSCF performs adaptive band boundary grouping through local density estimation combined with K-means clustering to suppress high-redundancy interference. Subsequently, multi-curve fitting and cross-validation are applied to the Pareto front to select candidate solutions on the favorable side of the fitted front as the global best solution, thereby accelerating convergence. Experiments on three hyperspectral datasets using SVM and KNN included established and recent band-selection methods, controlled comparisons with three MOPSO variants, component ablations, robustness analyses, and statistical significance testing. MOPSO_BSCF achieved the highest mean OA, AA, and Kappa for most evaluated dataset, classifier, and band-subset combinations. The Wilcoxon tests identified significant improvements in several settings and comparable performance in others. The ablation results further supported adaptive curve-model selection, early curve-fitting activation, and multi-start K-means boundary construction. These findings demonstrate competitive and generally stable band-selection performance under the evaluated conditions.
Post-detection Internet of Things (IoT) botnet mitigation requires identifying non-flagged endpoints that may require proactive quarantine while limiting unnecessary service disruption. This study presents a prospective exposure-risk scoring framework for quarantine decision support in software-defined networking (SDN)-based IoT environments. Per-second network telemetry is transformed into communication graphs and lightweight endpoint features. A first-stage anomaly signal provides the current suspected-compromise context, while prospective exposure is defined by whether a currently non-flagged endpoint communicates with a Stage-A-flagged endpoint at a subsequent prediction horizon. This design separates the prediction target from same-second suspicious-neighbour adjacency. Standard graph neural network models are evaluated against direct-neighbour and topology-agnostic baselines, while quarantine thresholds are selected on validation data to maximise recall subject to a false-positive-rate budget of 0.05. Experiments using CICIoT2023 packet-capture data and CICIoT-DIAD2024 show that the value of graph-based scoring depends on the traffic representation and evaluation scenario. On CICIoT2023-PCAP, graph models achieved substantially higher mean AUPRC than the topology-agnostic baseline, although performance varied across random seeds. On CICIoT-DIAD2024, the ordinary split showed no consistent graph-learning advantage, whereas under the evaluated scenario-disjoint Mirai condition, all three graph models achieved higher AUPRC than both the topology-agnostic MLP and the direct-neighbour rule. The findings demonstrate the potential of prospective exposure-risk scoring for controlled post-detection quarantine while highlighting the effects of telemetry representation and distribution shift.