
Student distraction remains a critical barrier to effective learning in classroom environments, yet existing detection methods relying on manual observation or post-session feedback are inherently subjective, time-intensive, and ill-suited for real-time adaptive teaching. This study presents the Artificially Intelligent Technology for Student Engagement Assessment in Classroom Environments (AI-TEACH), a novel co-teacher paradigm that automatically identifies and quantifies student distraction through multimodal fusion of asynchronized audio and video streams captured via classroom surveillance cameras. AI-TEACH integrates YOLO-NAS and ByteTrack for real-time student detection and tracking, MediaPipe for behavioral cue extraction, and Silero-VAD with Wav2Vec2-SVM for audio-based emotional distraction classification; a BiLSTM network fuses these multimodal features to generate per-incident severity scores, and a session-wide engagement report accessible through an interactive instructor dashboard. The framework was validated on a curated dataset of classroom recordings and further evaluated through a controlled experiment involving 200 students divided into experimental and control groups, achieving 90
Autonomous vehicle routing in complex urban environments suffers from high computational overheads, energy inefficiency, and increased collision risks using traditional routing algorithms. These limitations necessitate innovative approaches to address scalable and real-time decision-making requirements for multi-agent systems. This research proposes a reinforcement learning-based Multi-agent Federated Deep Q-Network (MAQDRL) model, which integrates federated learning with QuadTree spatial partitioning and indexing framework for collision detection. The model employs a decentralized training approach to optimize routing efficiency and reduce energy consumption. Each agent of MAQDRL trains the Deep Q-Network (DQN) independently using local state information. Model parameters are periodically synchronized through federated learning to align policies while preserving data privacy. The QuadTree partitions the environment dynamically based on agent density, focusing computational resources on high-interaction areas to ensure efficient collision detection and proximity querying. Advanced preprocessing techniques like state normalization, action encoding, and reward scaling stabilize training and enhance policy convergence. The framework further employs an adaptive exploration-exploitation strategy and decentralized decision-making, allowing agents to collaboratively achieve optimized routing while mitigating collisions and reducing energy consumption in complex, high-density scenarios. Experimental results demonstrate that the proposed MAQDRL is a promising solution for advancing autonomous vehicle routing systems. MAQDRL achieves a 97.8
Personalized and precise classification of neurodegenerative diseases (Alzheimer’s and brain tumors) is crucial for early diagnosis and patient care. Brain tumors and Alzheimer’s disease are challenging to classify due to their varied shapes and features. In this work, we propose a novel four-block residual attention model for the classification of brain lesions and Alzheimer’s disease in magnetic resonance imaging (MRI). The 4-block attention model integrates a channel attention module (CAMB) and a spatial attention module (SAMB) into each of its four blocks. We also implemented the Residual Attention Network (RAN) to improve overall disease information further and achieve better classification and generalization. An ablation study was conducted to determine the optimal number of blocks, testing RAN models with 2, 3, 4, and 5 blocks for training and validation. The proposed 4-block RAN model achieved an accuracy of 96.8
In Unmanned Aerial Vehicle-assisted Mobile Edge Computing (UAV-MEC), dynamic workloads and limited onboard energy pose significant challenges for efficient task scheduling and long-term mission sustainability. Cognitively-inspired computing paradigms provide an intelligent solution by enabling UAVs to perceive environments, learn from experience, and make adaptive decisions. This paper proposes a TS-Diff (Two-Stage Diffusion Policy) framework for joint task offloading, trajectory planning, and energy harvesting. A brief Soft Actor-Critic pre-training stage first constructs an exploratory experience memory buffer to address the cold-start issue of diffusion models. A Diffusion Policy Actor is then employed to iteratively generate robust continuous control actions, forming a perception–decision–action loop for adaptive UAV control. Experimental results show that TS-Diff achieves a final average return of approximately -145, improving performance by about 20
This review aimed to explore the integration of Quantum Computing (QC) with Artificial Intelligence (AI) subsets such as Machine Learning (ML) and Deep Learning (DL), addressing the computational demands posed by the exponential growth of visual data. It identifies key challenges such as interdisciplinary complexity, lack of standard benchmarks, scalability, integration barriers, and the theoretical-practical gap in quantum applications. The review systematically examines existing literature on the application of quantum algorithms in areas including image processing, Natural Language Processing (NLP), Transfer Learning (TL), Federated Learning (FL), networking, cybersecurity and the finance sector. It highlights the usage of quantum principles like superposition and entanglement to accelerate computations, optimize models, and enhance data security in ML/DL frameworks. Findings indicate that integrating QC with ML/DL offers faster convergence, improved optimization, secure decentralized learning, and efficient handling of large-scale and complex data. Specific improvements are observed in TL and FL approaches, NLP accuracy, cryptographic robustness, and performance in medical diagnostics and autonomous systems. QC holds transformative potential in enhancing ML/DL capabilities across domains. Despite existing challenges such as error mitigation and integration complexity, its combination with classical learning methods opens new frontiers for research in AI-driven sectors. Future studies should focus on bridging theoretical and application-level gaps while creating standardized evaluation frameworks.
Retrieval-based dialogue systems aim to select a proper response according to multi-turn conversational history. Persona-based conversation utilizes prior knowledge to maintain persona consistency, enhancing retrieval accuracy. However, reference-based evaluation relies on high-quality data annotation, which is costly and time-consuming. To address this, we discover persona-centric metamorphic relations to infer test samples from annotated data, without additional annotation cost. Benefiting from this, this work efficiently evaluates the robustness of personalized dialogue models regarding persona consistency. Specifically, we discover three types of metamorphic relations from three aspects: self-persona, partner-persona, and response, to automatically derive new test samples . Then the inherent inference relations between originals and derivatives allow for robustness evaluation. Using this evaluation methodology, our work assesses three widely used training paradigms: non-pretraining, fine-tuning after pre-training, and prompt learning, in personalized dialogue retrieval to observe whether these paradigms are more robust or exhibit the same flaws as the other two paradigms. Our experimental results, based on the three discovered metamorphic relations with consistent outputs reveal that prompt learning is more robust than training from scratch and fine-tuning. While traditional reference-based validation and natural language processing methods achieve competitively high retrieval accuracy (Hits@1 up to 87.4
Multimodal emotion recognition is critical for affective computing applications, but it faces growing privacy risks when handling sensitive user data, while generic models often fail to adapt to individual emotional characteristics. Existing methods rarely address both privacy protection and personalization simultaneously. This paper proposes HLPCMF, a novel multimodal emotion recognition model that integrates hypergraph learning with pairwise cross-modal fusion. Specifically, we construct multimodal hyperedges and temporal hyperedges to model high-order relationships among utterances, enabling the capture of complex interaction patterns beyond traditional pairwise graphs. A dual-stream gated attention network (DSGAN) is introduced to reduce information redundancy among nodes. Furthermore, we design a Transformer-based pairwise cross-modal fusion (PCMF) mechanism, which treats each unimodal feature as an anchor in turn and performs pairwise fusion with other modalities to extract deep emotional interaction information. Experiments on standard multimodal datasets show that the framework maintains high emotion recognition accuracy while effectively protecting user privacy. It outperforms generic models in personalized scenarios, achieving better alignment with individual emotional expression habits. The proposed HLPCMF model effectively captures multimodal emotional cues and enhances fine-grained emotion recognition in conversational contexts. The integration of hypergraph learning and pairwise cross-modal fusion provides a robust framework for modeling complex dependencies in multimodal dialogue emotion recognition tasks.
The early detection of dementia is a major challenge because prodromal cognitive decline is subtle and variable among individuals. The majority of currently available approaches are based on cross-sectional population cut-offs, expensive neuroimaging studies or cumbersome multimodal deep learning pipelines, limiting the interpretability and external validity for pragmatic deployment in clinical practice. To alleviate these deficiencies, we develop a new AID onset detection which is referred to as AIDCMEDD-PBSD that stands for Artificial Intelligence-Driven Cognitive Monitoring based on Personalized Baseline Shift Detection (PBSD) to identify early D stage. In contrast to traditional classifiers, the presented approach estimates an individual’s cognitive and behavioral baseline, and subsequently identifies prolonged deviations over time thereby facilitating fine-grained longitudinal surveillance. AIDECAMEDD-PBSD consists of a set of low-cost multimodal signals such as short speech elicitation tasks, cognitive microtests, and passive smartphone behavioral markers. Resilient statistical feature normalization with median and MAD is paired with cumulative sum-based shift detection to uncover sustained cognitive differences. The feature-level alerts are then merged through a simple and interpretable logistic regression model, yielding a clinical actionable dementia risk score as well as individual explanations of the features. Experimental results on benchmark speech datasets and longitudinal cognitive assessment data show that the proposed model achieves 94.6
The rapid evolution of AI-generated synthetic media, called deepfakes, has raised substantial concerns regarding digital misinformation, security breaches, and public trust. Existing centralized detection systems are often limited with privacy risks, weak generalizability, and reduced robustness when exposed to multimodal manipulations across audio, text, and video data. There is always a human cognitive capacity to fuse and contrast cues across sensory modalities while judging authenticity. This research proposes a privacy-preserving, federated learning-based deep-fake detection framework that facilitates secure, decentralized training across heterogeneous devices. The proposed framework leverages Fed-DFakeNet for localized model training, enabling robust feature extraction from multimodal datasets without transmitting raw data. It incorporates MogDetNet, a knowledge distillation-based fusion module that aligns multimodal features for improved generalization and accuracy. Furthermore, the mCreamFL aggregation strategy introduces a contrastive representation ensemble and encrypted communication mechanism, ensuring data integrity and optimal performance while preserving user privacy. Comprehensive experiments are conducted using three benchmark datasets FoR (audio), SDFVD (video), and TweepFake (text) demonstrating superior performance. The framework achieves classification accuracies of 98.32
Medical image segmentation, which delineates anatomical structures and pathological regions at the pixel level, plays an important role in computer-aided diagnosis. Existing segmentation methods are typically developed and evaluated under a single-dataset setting, and their performance often degrades when applied to images from different medical centers, scanners, or acquisition protocols due to substantial domain gaps. While cross-dataset collaborative learning methods can train a unified model from multiple datasets, they generally do not explicitly distinguish domain-specific appearance variations from domain-invariant semantic information, placing a heavy burden on shared layers when the domain gap is large. To address this issue, we propose MDCL-UNet, a supervised multi-domain collaborative learning framework based on domain feature disentanglement. MDCL-UNet adopts a two-branched encoder in which a domain-specific branch equipped with the proposed Domain Style Instance Normalization (DSIN) module and a domain-invariant branch jointly disentangle domain features. A domain adversarial classifier and an MMD-based domain decoupling loss are used to ensure that the two branches learn complementary and orthogonal representations. Instead of discarding domain-specific features as in conventional domain generalization methods, a Domain Fusion Attention Module (DFAM) is further introduced to adaptively fuse domain-specific style cues with domain-invariant semantic features through attention-based integration, enabling effective reuse of cross-domain information. Extensive experiments on retinal vessel, optic disc/cup, and abdominal multi-organ segmentation datasets demonstrate that MDCL-UNet achieves consistently higher Dice scores and lower HD95 distances than single-dataset baselines and existing cross-dataset collaborative learning methods. Moreover, MDCL-UNet maintains stable training and favorable scalability with increasing numbers of domains. The source code will be available at https://github.com/lqr41710085/MDCL-UNET .
Effective management and analysis of watershed systems are complicated by uncertainty arising from complex, interdependent interactions among watershed elements. This work proposes a novel and efficient decision-making model based on q-Fractional Fuzzy Sets (q-FrFS), an advanced extension of fuzzy set theory. The primary objective is to improve the representation of ambiguity and imprecision in decision-making scenarios and criteria. The Maclaurin Symmetric Mean (MSM) is used to derive several new aggregation operators, including q-Fractional Fuzzy MSM (q-FrFMSM), q-Fractional Fuzzy Weighted MSM (q-FrFWMSM), q-Fractional Fuzzy Ordered Weighted MSM (q-FrFOWMSM), and q-Fractional Fuzzy hybrid Weighted MSM (q-FrFHWMSM). The operators are applied within a multi-attribute group decision-making (MAGDM) framework to successfully incorporate expert opinions. When used to assess watershed management options, the suggested MAGDM methodology shows greater consistency, robustness, and discriminative capacity than current methods in ambiguous situations. To ensure the stability and reliability of the proposed framework, a sensitivity analysis is used. Finally, the proposed framework provides a solid and practical solution for handling challenging decision-making situations in a context of uncertainty. It could be useful in environmental management and other areas and could suggest avenues for future research to refine the model and address current limitations.
While clinical dialogue summarization systems can alleviate documentation burdens, current approaches focus solely on informational content like symptoms and diagnoses. These systems often overlook the patient’s emotional signals, such as fear, anxiety, or anger. Although Large Language Models (LLMs) capture these nuances, high computational costs and data privacy risks restrict their real-time clinical integration. To overcome these obstacles, this study proposes an innovative and hybrid framework based on Emotion-Aware Knowledge Distillation. In this study, the empathy and reasoning capabilities of the Gemini teacher model are distilled via synthetic data generation strategies and transferred to a resource-efficient, compact proposed student model (fine-tuned BioBART-v2). The developed model improved the signal-to-noise ratio in clinical dialogues through a patient-centric filtering strategy and reduced politeness bias from 44.49
Motor imagery (MI)-based brain-computer interfaces (BCIs) rely on accurate decoding of electroencephalography (EEG) signals. However, the non-stationary nature of EEG signals and the complex relationships among discriminative patterns make MI classification a challenging task. To address these issues, this paper proposes a Multi-Scale Convolutional Attention-Guided Capsule Network (MSC-AG-CapsNet) for MI-EEG classification. The proposed framework consists of a multi-scale convolutional module and an attention-guided capsule network module. Specifically, the multi-scale convolutional module employs parallel temporal convolutions with different receptive fields to capture EEG representations at multiple temporal scales. Subsequently, the extracted feature sequence is transformed into primary capsules and aggregated through an attention-guided capsule aggregation mechanism, which replaces conventional dynamic routing with a feed-forward aggregation strategy. In this way, informative relationships among capsule representations can be effectively exploited while avoiding iterative routing operations. Extensive experiments conducted on the BCI Competition IV-2a and IV-2b datasets demonstrate the effectiveness of the proposed framework. Ablation studies further verify the contributions of both the multi-scale convolutional module and the attention-guided capsule aggregation mechanism. The results indicate that MSC-AG-CapsNet provides a competitive and efficient solution for MI-EEG decoding. The code is available at https://github.com/wiaobang/MSC-AG-Capsnet.
Cross-subject emotion recognition based on multimodal physiological signals has broad application prospects in fields such as human-computer interaction and health monitoring. However, the dual challenges of multimodal heterogeneity and cross-subject differences severely restrict the generalization ability of the model. The existing methods mainly focus on static decoupling and the study of emotional correlation, but lack attention to dynamic interaction and emotional causal relationships. To address these issues, this paper proposes a Meta-causal Bidirectional Decoupled Learning (MBDL) framework. This framework consists of a bidirectional interaction disentanglement (BID) module and a meta-causal learning (MCL) module. Specifically, the BID module decomposes the representation into three subspaces of emotion, identity and modality, and allows structured mutual modulation between them, which simulates the complex interaction of multiple factors in the process of emotion generation; Secondly, the MCL module not only enables the model to adapt quickly when the new agent only provides a small number of samples by meta learning, but also the causal inference mechanism imposes stability constraints through counterfactual generation, forcing the model to learn stable causal relationships rather than spurious correlations. Extensive experiments on the DEAP and MAHNOB-HCI public datasets have shown that the MBDL framework significantly outperforms the existing state-of-the-art methods under the cross-validation protocol, and has strong generalization performance in cross-subject experiments.
UAV-Ground visual tracking is to achieve robust tracking by leveraging the discriminative information from both UAV and ground views. The existing approach uses a multi-view collaborative model to associate and fuse target features from different views by calculating the appearance similarity. However, it fails in challenging scenarios due to cross-view spatial misalignment caused by ignored geometric relations. To handle this problem, we propose a robust UAV-Ground tracker based on the novel Geometric Relation Prediction Transformer (GRPT), which leverages the coordinate offset of the target between two views to achieve accurate collaborative modeling. Moreover, we design a SRA strategy to adaptively correct the location of search regions for cross-view spatial alignment. We evaluate our method on public dataset UGVT, achieving 82.5
Fungal infections pose a growing global health threat exacerbated by the limited efficacy and rising antimicrobial resistance of conventional antifungal agents. Antifungal peptides (AFPs) emerge as promising alternatives due to their multimodal mechanisms of action and favorable toxicity profiles. To address the resource-intensive nature of traditional experimental screening, we present a multimodal deep learning framework that synergistically integrates autoencoder (AE) and convolutional autoencoder (CAE) architectures by leveraging one-hot encoding, multiple sequence information. Our innovative approach combines reconstruction losses from both AE and CAE with classification loss to optimize feature representation and enhance generalization capabilities. Rigorous evaluation on our independent test dataset, Antifp_Main, Antifp_DS1, Antifp_DS2 demonstrated superior performance with average accuracy of 91.71
Background: Hope Speech (HpS) has been proposed as a strategy to shift online moderation from purely punitive approaches toward positive reinforcement, amplifying expressions of aspiration, resilience, and support. Despite substantial research in psychology, philosophy, and linguistics highlighting the multifaceted nature of hope, most computational models treat HpS as a coarse, binary phenomenon, neglecting the linguistic and cognitive nuances that underpin real-world hopeful discourse. Objective: We introduce C-Hope, a linguistically grounded framework for fine-grained Hope Speech Detection (HpSD) that differentiates between distinct notional categories of hope (Counterfactual, Desire, Belief, Plan) and considers temporal orientation, modality, and commitment level in speaker stance. Our goal is to improve the interpretability and practical utility of HpSD models, addressing the gap between broad affective categories and the complex reality of hope in discourse. Methods: We re-annotate existing HpS datasets with the C-Hope scheme, assembling a fine-grained, linguistically-grounded benchmark capturing degrees of speaker commitment in hope speech, comprising 4,370 English texts. C-Hope distinguishes five main groups of linguistic markers, including modal verbs, propositional attitude verbs, mood, tense, and grammatical constructions. We evaluate both fine-tuned transformer models and prompted large language models (LLMs), designing structured prompts that incorporate theoretical insights into the nature of hope. Results: Our findings show that models trained on the C-Hope scheme outperform binary HpS models, demonstrating improved detection of nuanced categories such as counterfactual and plan-based hope. Moreover, prompt-based LLMs leveraging structured linguistic knowledge achieve competitive results, suggesting that grounding AI models in linguistic theory is a promising avenue beyond purely data-driven approaches. Conclusion: Our study demonstrates the viability of a linguistically informed, multi-category framework for HpSD, paving the way for future research integrating fine-grained semantic distinctions and speaker stance modeling. The new benchmark corpus and methods provide tools for developing more transparent, ethical, and context-aware NLP systems.
Emotional recognition plays a pivotal role in understanding human interactions, shaping decision-making processes, and enhancing communication across diverse domains. With technological advancements, there is an increasing demand for systems that can accurately interpret emotional cues, driving innovations in healthcare, marketing, and human-computer interaction. This study addresses the gap between human intuition and algorithmic analysis by evaluating the performance of human annotators and machine-learning algorithms in emotion recognition tasks. Utilizing the CMU, CAM3D, and BEAST datasets, we investigated the capabilities of human annotators to identify emotions from facial expressions and hand gestures. Concurrently, we applied the YOLOv8 algorithm to the same datasets to assess its efficacy under three distinct scenarios. Performance metrics including accuracy, precision, recall, F1-score, and mAP were employed to rigorously evaluate and compare the results. Our findings indicate that YOLOv8 achieves high accuracy on less complex datasets, such as CAM3D, but encounters difficulties with the nuanced and compound emotions present in the CMU dataset. This underscores the inherent challenges in emotion recognition tasks and highlights the potential of integrating multimodal inputs and human expertise to enhance machine-learning systems. Recommendations include diversifying datasets and fostering interdisciplinary approaches to address the algorithmic limitations. This research contributes substantially to machine learning by addressing a critical application problem: the development of accurate, context-aware emotion recognition systems. By offering empirical insights and methodological advancements, this study establishes a foundation for improving the classification, prediction, and real-time application of learning methods in fields such as healthcare, robotics, and virtual reality. The results not only advance computational approaches to learning but also demonstrate their impact on solving real-world challenges, aligning with the journal’s commitment to robust empirical studies and practical applications.
Multilingual language models achieve strong results in emotion classification; however, their robustness under culturally complex and linguistically diverse conditions remains limited. Standard evaluations rely on structured datasets that do not capture phenomena such as code-switching, idiomatic expressions, and indirect affective language, which are common in real-world communication. This study introduces a culturally oriented evaluation framework to analyze the behavior of mBERT, XLM-R, and a fine-tuned variant (XLM-R-FT) under non-standard linguistic conditions. The evaluation combines a multilingual dataset with a subset of manually annotated samples incorporating regional variation in Spanish, Portuguese, and Italian. Results show consistent performance degradation on culturally marked inputs, with reductions of up to 10 F1 points. Fine-tuning improves stability but does not eliminate sensitivity to implicit and idiomatic expressions. Interpretability analysis using LIME and SHAP indicates that predictions are often driven by lexically salient tokens, leading to misinterpretations. The proposed framework enables evaluation beyond standard benchmarks and highlights the need to incorporate cultural and linguistic variability in multilingual emotion analysis.
With the popularization of Internet of Things devices, the volume of data generated at the network edge has grown explosively. As a distributed machine learning solution, Federated Edge Learning (FEL) in mobile edge computing (MEC) provides an effective way to protect privacy by training locally and sharing only model parameters rather than the original data. However, in practical applications, FEL faces two core challenges: First, the data heterogeneity of each edge device can lead to deviations in the global model; The second is how to design an effective incentive mechanism to encourage nodes with limited resources to continuously participate in training. Our aims to simultaneously address these two major challenges by optimizing the local training and global aggregation processes to comprehensively enhance the efficiency and performance of FEL. For this reason, we propose a hybrid model. Firstly, the Stackelberg Stackelberg game model is adopted to describe the relationship between aggregators and edge devices. Meanwhile, the existence of Nash equilibrium is theoretically proved to ensure the stability of the model. Secondly, we propose a novel algorithm named Group Relative Policy Optimization Based on Hierarchical Mean-Field Theory (OGRPO-HMF), which can jointly optimize the local training of nodes and the global model aggregation of servers. We validate the effectiveness and generality of our approach through extensive experimentation on various FEL tasks, showcasing significant performance gains. Extensive experiments on benchmark FEL datasets demonstrate the superior performance of our proposed algorithm, improving the global test accuracy by up to 2.26% and reducing the global test loss by up to 66.27% compared to the state-of-the-art counterparts.