
Breast cancer remains one of the leading causes of mortality among women worldwide, making accurate tumor segmentation across medical imaging modalities essential for early diagnosis and treatment planning. Although Swin Transformer–based U-Net architectures have shown promise, existing methods often lack effective mechanisms to exploit encoder–decoder interactions and multi-scale tumor characteristics. In this study, we propose a novel Swin Transformer–based U-Net architecture whose primary contribution lies in a Swin-Enhanced Cross Attention (SECA)-driven decoding strategy, rather than the use of a Swin encoder alone. The proposed framework employs a Swin Transformer encoder to capture long-range dependencies and global contextual information, while introducing SECA modules in the decoder to explicitly model cross-attention between encoder and decoder features, enabling precise alignment of global semantics with fine-grained spatial details. To further address tumor size and shape variability, a transformer-guided multi-scale decoder based on Inception modules is integrated. In addition, a hybrid loss function combining Dice, Focal, and Boundary losses is adopted to mitigate class imbalance, enhance learning from hard samples, and improve boundary delineation. The proposed model is extensively evaluated on four benchmark datasets, CBIS-DDSM, BreastDM, BCSS, and the Breast Ultrasound Images dataset, spanning mammography, DCE-MRI, histopathology, and ultrasound modalities. Experimental results demonstrate consistent performance gains over U-Net, U-Net++, and existing Swin U-Net variants, achieving improvements of up to +2.1
The integration of deep learning in healthcare and human-centered systems has enabled significant advancements in predictive analytics. However, concerns regarding interpretability, reliability, and safety remain critical barriers to adoption. This paper proposes SCLAF-FV (A SHAP-Guided CNN–LSTM–Attention Framework with Formal Verification for Reliable Injury Detection), a human-centric framework for injury risk prediction that combines deep learning, explainable artificial intelligence (XAI), and formal verification. SCLAF-FV integrates convolutional and recurrent neural networks with an attention mechanism to capture both spatial and temporal patterns in physiological signals. To enhance transparency, SHAP (SHapley Additive exPlanations) is employed to interpret model predictions and quantify feature importance. Furthermore, the Marabou formal verification framework is utilized to ensure that SCLAF-FV satisfies predefined safety and robustness constraints. Experiments conducted on the Ninapro dataset demonstrate that the proposed approach achieves superior predictive performance compared to baseline models, while also providing interpretable and verifiable outputs. This work contributes toward bridging the gap between high-performance machine learning models and their safe deployment in real-world scenarios.
H.265/HEVC(High Efficiency Video Coding) involves a lot of computational complexity, the main factor being recursive block partitioning and PU mode decision. In this paper, a novel online threshold model based square PU partition mode early detection scheme has been proposed, all 2 N square PU modes, including both the merging and AMVP modes, as well as the skip mode. Based on the statistical analysis of the RD cost of the best PU partition mode of the previously encoded CUs, we established a threshold model for early detection of CUs where the 2 N square PU partition mode is optimal. If less than a threshold, the evaluation of the non-square PU partition modes is entirely skipped and the square PU partition mode is determined optimally. Through experiments, the proposed method proved to be superior to previous methods in terms of negligible coding efficiency and quality degradation compared to HM, especially for video sequences with different motion characteristics.
Transcripts of lecture videos may have topic transitions that are subtle and gradual, making segmentation difficult for computers. Current segmentation methods face limitations: lexical cohesion methods rely on overlap between words on the surface level without capturing deep meaning, transformer methods break a single long-form transcript into chunks to impair coherence, and methods based on structure cannot handle the casual nature of a transcript from a lecture video. This work introduces the Thematic Segmentation of Lecture Video Transcripts, a four-component hybrid technique involving deep semantic encoding in a long-context semantics model, neural segmentation on multiple scales using contrastive learning, graph refinement, and reinforcement learning using quality-driven rewards. Our experiments on the Synthetic Lecture Transcript Benchmark (210 transcripts) as well as on several real-world lecture datasets demonstrate that this proposed method achieves F1 = 0.78 ± 0.01, Pk = 0.27 ± 0.01, WD = 0.25 ± 0.01, outperforming the evaluated baseline methods by an average of + 10.4
This study develops a multimedia training system for cardiopulmonary resuscitation (CPR) that couples HRNet-based pose estimation with ST-GCN + + skeleton action recognition to deliver real-time audio-visual feedback. The system segments streaming video into short clips, classifies CPR actions on the fly, and triggers process-aware prompts when the execution order or posture deviates from the standard procedure. This study is framed as a controlled feasibility study. The model was evaluated using randomized segment-level splits. The system achieves an action classification accuracy of 0.93, recall of 0.93, a balanced accuracy of 0.91, and a macro-averaged F1-score of 0.91. A quasi-experimental study was conducted with 60 participants divided into two groups of 30. The entire experiment was completed over approximately eight hours, including participant briefing, pre-test, instructor-led instruction, group-based practice, and post-test assessment. Simulated operation performance was evaluated by a blinded CPR instructor based on action sequence correctness and hesitation duration. The intelligent system group significantly outperforms the traditional training group in skill execution (mean score: 89 vs. 72, p < 0.001, rank-biserial r = 0.606), reports lower cognitive load, achieves a System Usability Scale score of 78.6, and expresses greater satisfaction with their training experience. The results suggest that skeleton-based action recognition with real-time multisensory feedback can support practical skill acquisition and learner confidence in a controlled feasibility setting. The proposed lightweight architecture is implemented on commodity hardware, suggesting its potential applicability to remote or resource-constrained settings and possible extension to other skill-based domains after further validation and runtime profiling.
Human Activity Recognition (HAR) using sensor-based multimodal time-series data from wearable devices has gained significant attention in recent years, primarily due to privacy concerns associated with vision-based approaches. Modern wearables such as smartphones, smartwatches, and smart glasses are equipped with inertial measurement unit (IMU) sensors that generate rich time-series signals widely used for HAR applications. Developing accurate and reliable HAR systems for recognizing Activities of Daily Living (ADLs) requires access to large-scale, diverse, and high-quality datasets. However, existing datasets often suffer from key limitations, including limited demographic coverage, low sampling rates, and small participant pools. This study makes two major contributions. First, we introduce HumCareADL, a novel and diverse multimodal HAR dataset specifically designed to address these limitations by capturing data from participants across varied age groups and activity contexts. Second, we propose a hybrid multi-branch CNN-LSTM architecture, termed HyMCL-Net, which leverages the complementary strengths of Convolutional Neural Networks (CNNs) for spatial feature extraction and Long Short-Term Memory (LSTM) networks for temporal dependency modeling. Extensive experiments demonstrate that HyMCL-Net achieves an average accuracy of 85.27
Content-based image retrieval (CBIR) relies on feature extraction techniques such as color, shape, and texture to retrieve visually similar images. Traditional CBIR methods often struggle to extract these different features, leading to unsatisfactory retrieval results. In this regard, we introduce a multimodal CBIR framework that combines handcrafted color histograms with deep texture and shape features obtained using a customized variant of the VGG16 network. To generate compact and comparable representations, the network is adapted by removing the fully connected layers and applying Global Average Pooling to the convolutional outputs. These descriptors are then grouped via unsupervised k-means clustering to facilitate more discriminative image grouping within the feature space. Our system is optimised using Nash game theory transforming our problem into a multi-criteria optimization process. Each feature extraction method is considered an independent player in a non-cooperative game, with each one proposing the cluster that contains images that are visually similar to the query image. The game follows an iterative process at each step, the intersection of all the similar images becomes the new dataset, and the same process is reapplied until a Nash equilibrium is reached, where no player can further improve the results.
The explosive growth of large-scale video archives from surveillance networks, online platforms, and personal devices has made efficient and semantically rich video retrieval a critical challenge. Existing approaches based on deep multimodal embeddings have significantly improved retrieval accuracy. However, they often lack scalability, modularity, and system-level integration with indexing and metadata management. In this work, we present a modular and scalable pipeline for semantic video indexing and retrieval, tailored to person-centric search. The proposed architecture decouples a web-based front-end from a back-end organized into two pipelines. The indexing pipeline performs video chunking, person detection and tracking, crop selection, metadata enrichment, and semantic vectorization. The retrieval pipeline supports visual, textual, and hybrid queries, including face-based matching. The system leverages YOLO11 and BoT-SORT for real-time person detection and tracking, SigLIP2 for multilingual vision-language embeddings, and InsightFace for face recognition, storing all representations in a vector database with rich, traceable metadata. We further fine-tune the SigLIP2-SO400M-Patch14-384 checkpoint on a curated mixture of person-centric image-text datasets and evaluate the resulting model on the RSTPReid benchmark. Experimental results show that our approach achieves state-of-the-art performance on the RSTPReid benchmark under the considered setting. In particular, it achieves competitive Recall@k performance with respect to recent text-based person search methods and significantly improves mean Average Precision, reaching 0.68 against a best competing value of 0.54.
Offline signature verification is a traditional method used to authenticate a person since ancient times. Though it has been studied for over 40 years, it remains a challenge for the research community due to two major issues: intra-person variation and inter-person similarity. To address these issues, it is crucial to determine stability-driven, consistent features that are stable, repeatable, and influential discriminatory features across genuine signature samples of the same signer. In this paper, a Multi-Phase Fusion Architecture is proposed, where fusion at multiple stages is employed to develop a reliable signature verification system. The primary aim of this work is to identify a compact set of consistent and stable discriminatory features from offline signature images that repeatedly appear across genuine signature samples of the same individual. The first phase of the proposed method combines global texture features with fine-grained texture features to effectively capture the discriminative characteristics of handwritten signatures. Proceeding with the second phase to determine the consistent features, multiple feature selection techniques encompassing both filter-based and wrapper-based methods were explored. More emphasis is placed on stability analysis, which evaluates the consistency and robustness of the selected features through stability metrics across varying datasets. Evaluating the consistency of features ensures that the model generalises well on unseen and challenging signature samples in real-world scenarios. In the third phase, an ensemble learning technique is used to achieve enhanced reliability in final decision-making. The proposed method is evaluated on multilingual signature datasets: CEDAR, BHSig260 (Hindi, Bengali), MCYT-75, and UTSig, achieving accuracies of 100
The classification of tomato leaf diseases is critical for advancing smart agriculture; however, deep learning model performance is often limited by low-data regimes. This study introduces and evaluates a Combined Data Augmentation (CDA) strategy that integrates traditional augmentation techniques (TDA) with data generated via a Stable Diffusion model (GDA) to improve classification performance under severely limited data conditions, specifically with only 20 training images per class. Experiments were conducted on subsets of the PlantVillage ( D^PV ) and PDR2018 ( D^PDR ) datasets, utilizing EfficientNet-B0 as the backbone across three training configurations: transfer learning with partial freezing, training from scratch, and fine-tuning all layers. The statistical significance of the results was assessed using the Wilcoxon signed-rank test with Holm correction and Cohen’s d_z effect sizes across 15 cross-validation folds. Findings indicate that Combined Data Augmentation (CDA) provides its most decisive benefit under the from-scratch configuration, rescuing the model from training collapse on both datasets ( D^PV : +21.95 percentage points, p_holm = 0.009; D^PDR : +7.91 percentage points). Under pretrained configurations, the improvement is modest on D^PV and absent or negative on D^PDR , where augmentation strategies including GDA were found to significantly reduce accuracy. Further analysis demonstrates that the effectiveness of generative augmentation is highly dependent on the Strength parameter; Strength = 0.35 is the only regime preserving performance parity with the Baseline, while higher values introduce semantic drift and statistically significant performance degradation. A learning rate sensitivity analysis at η _0 = 10^-3 confirmed that these conclusions are stable across hyperparameter choices. This study provides empirical and statistical evidence characterizing the conditions under which combined data enhancement strategies are beneficial, neutral, or detrimental in limited data contexts, and informs the development of robust plant disease diagnostic systems.
Skin cancer is a common and possibly deadly illness if it goes undetected in its early stages. Recently, progress in deep learning techniques has significantly improved the effectiveness of models used for classifying skin cancer. This work presents SkinEnsemNet, a transfer-learning-based heterogeneous ensemble framework that integrates VGG16, DenseNet201, and EfficientNetB0 for skin lesion classification using the HAM10000 and ISIC-2019 datasets. By combining feature representations from architecturally diverse CNN backbones through feature-level fusion and dataset-specific classification layers, the framework aims to exploit complementary spatial, hierarchical, and computationally efficient representations of dermoscopic images. The ensemble leverages the complementary strengths of these models: VGG16 captures localized spatial hierarchical features, DenseNet201 provides feature re-usage and hierarchical feature extraction, and EfficientNetB0 balances accuracy and computing cost. The preprocessing stage of the pipeline comprises normalization and class imbalance handling by random oversampling. Following feature extraction from the base models, global average pooling, feature concatenation, and classification using a stack of dense layers with dropout regularization are performed. On the held-out image-level test partitions, SkinEnsemNet achieved accuracies of 98.85
With the rapid expansion of digital media, ensuring authenticity and security has become fundamental. Audio Watermarking (AW) provides an effective solution to embed the information while maintaining imperceptibility and robustness. However, the prevailing studies overlooked the overlapping segments during watermark embedding, thus resulting in high computational overhead. Therefore, this article proposes a Robust Audio Watermarking System with Spatial-and-cross scaled gaussian error linear unit-based deep convolutional neural network and Blancmange curve cryptography for Security (RAWS-SBS) for attaining high security and robustness. Primarily, the FlatTop window (FT) technique segments the audio signal while ensuring non-overlapping segments. Afterward, by utilizing Polar Cartesian-based Discrete Wavelet Transform (PC-DWT), the segmented frames are decomposed; also, matrix formation is done. Then, the features are extracted. The Gamma Functional Energy Valley Optimization (GFEVO) determines the optimal location to embed the watermarks. Next, the matrix is transformed into a 2-D format. Thereafter, the watermark image is encrypted using the Blancmange Curve Cryptography (BCC) algorithm, and the encrypted bit is converted into a vector bit. Lastly, the audio watermarked signal is obtained. The Spatial-and-Cross with Scaled Gaussian Error Linear Unit-based Deep Convolutional Neural Network (SC-SGELU-DCNN) is employed to extract the watermarks. The experimental outcomes stated that the proposed method obtained a minimum Bit Error Rate (BER) of 0.03232. The GFEVO had 4.256 Perceptual Evaluation of Speech Quality (PESQ). Further, the proposed approach had a maximum accuracy of 97
Given the widespread impact of earthquakes, there is a pressing need for more effective methods capable of leveraging multimodal disaster data. This research proposes a multimodal prediction framework for classifying earthquake-related tweets. The proposed multimodal architecture first captures contextual dependencies in tweet text and fuses them with high-level semantic features learned by deep convolutional networks, resulting in a more discriminative joint representation that enhances the classification performance. The effectiveness of the proposed framework is tested using the CrisisMMD multimodal dataset. The dataset contains tweets collected during real-world natural disaster events. In thisstudy, we analyze earthquake related data and CrisisMMD which includes 2,743 text and image samples from two earthquakes data: the Mexico earthquake and the Iraq-Iran earthquake. The evaluation is based on performance metric, including accuracy, precision, recall, F1-score, and statistical significance scores. Comparison with deep learning models used in previous studies shows that the proposed framework is one of the high-performing investigations. The experimental results demonstrate that the proposed framework achieves an accuracy of approximately 88
Generative artificial intelligence (AI) provides a potential way to reduce the time and cost required for creating three-dimensional (3D) assets for cultural heritage virtual reality (VR) applications. However, its practical value depends not only on generation speed but also on whether the generated assets can be optimized, integrated, and deployed on standalone VR hardware while maintaining acceptable user experience. This study proposes a low-cost AI-to-VR content generation pipeline for a Taiwanese three-courtyard (Sanheyuan) cultural heritage tour. The pipeline combines open-source Trellis3D-based asset generation, Blender-based geometric and texture optimization, Unity OpenXR integration, and deployment on Meta Quest 3S. A total of 43 supplementary cultural objects were generated and optimized for static VR presentation. The measured object-generation time ranged from 25 to 135 s (M = 59.2 s, SD = 26.7 s), with kitchen and dining utensils showing the shortest average generation time (M = 46.6 s). The deployed system maintained a stable target frame rate of at least 72 FPS with interaction latency below 20 ms on Meta Quest 3S. A user evaluation with 50 valid questionnaire responses was conducted using a five-point Likert scale. The revised constructs showed acceptable to excellent internal consistency, with Cronbach’s alpha values ranging from 0.847 to 0.970. The construct means ranged from 3.72 to 3.87, indicating moderately positive perceptions of immersion, usability, comfort, usefulness, model realism, and system performance. The results suggest that open-source generative AI can support low-cost asset creation for cultural heritage VR, although post-processing, model-detail consistency, and broader participant validation remain necessary.
Vision Transformer (ViT) has achieved significant success across various vision tasks. However, due to their lack of inductive bias and high computational cost, many lightweight architectures have incorporated both Convolutional Neural Networks (CNNs) and Transformers. The problem is that transformer and convolution operations result in different feature distributions. This discrepancy causes an additional optimization challenge during joint training. Furthermore, the trade-off between performance and latency remains a challenge in mobile environments. In this work, we propose DoTCoM, which can optimize the distribution differences between convolution and transformer. The proposed DoTCoM adopts the Quarter-Inverted Bottleneck (QIB) block, a combination of the Quarter Groups (QG) convolutional block and the Inverted Bottleneck. The QIB block is designed to compensate for the insufficient receptive fields of the Inverted Bottleneck and achieves an optimal trade-off between parameters and performance. DoTCoM achieved outstanding performance by employing the following methods: the proposed Co-optimization Bias (Co-Bias) aligns convolutional and transformer feature distributions by reducing variance discrepancies through learnable per-channel bias parameters, stabilizing training and improving accuracy. DoTCoM could be lightweight and achieved significant improvements compared with existing lightweight ViT models. On ImageNet-1K, the proposed DoTCoM-L, DoTCoM-B, DoTCoM-S, and DoTCoM-T achieved 81.6
Human Gait, or the pattern of walking, has emerged as a pivotal area of study across rehabilitation, healthcare, and diagnostic applications. This systematic review presents a comprehensive and in-depth survey of state-of-the-art computational methods for human gait analysis. The novelty of this work lies in its integrative perspective covering methodologies based on kinematics, kinetics, and computational intelligence, which are typically studied in isolation. Kinematics-based methods emphasize motion characteristics such as joint angles and segment trajectories, kinetics-based methods focus on forces such as ground reaction forces and torque, and computational intelligence approaches leverage AI and machine learning to model, recognize, and predict gait abnormalities. The review spans literature published between 2010 and 2023, encompassing over 240 high-quality papers published in reputed journals and conferences indexed in the Scopus database. Papers were selected based on the presence of the keyword "Human Gait," methodological rigor, and their relevance to various application domains. This paper details gait parameters, classification of gait abnormalities, machine learning frameworks, and the use of public datasets and tools for gait analysis. Furthermore, it discusses applications in clinical diagnostics, sports monitoring, biometric identification, and humanoid robotics. Emphasis is placed on both low- and high-level gait parameters and their role in detecting pathological conditions. The review also explores the impact of gait technology across multiple domains and highlights open challenges. This review will serve as a valuable resource for researchers and practitioners involved in gait analysis, rehabilitation technologies, and artificial intelligence.
Context-aware multimedia systems require structured representations of visual content to support intelligent scene understanding and recommendation. However, most existing movie recommender systems primarily rely on high-level metadata and do not exploit fine-grained cinematic attributes that contribute to visual storytelling. To address this limitation, this paper presents CarMovie, a context-aware, content-based recommendation framework that models cinematic scenes as structured contextual instances defined by location, time, and mood, and associates them with visual attributes including color schemes and shot types. A context-annotated cinematic dataset is constructed from real movies using a structured annotation protocol to support contextual scene analysis and recommendation. To enable efficient and interpretable recommendation, the framework incorporates a Minimum-Attribute-Difference (MAD) ranking mechanism that performs reference-based, attribute-level comparison using binary feature encoding. Unlike conventional similarity measures based on global vector comparison, MAD prioritizes contextual consistency through structured attribute matching while maintaining low computational complexity. A mobile prototype system is implemented to demonstrate the practical applicability of the proposed framework. The approach is evaluated using a two-stage experimental setting involving general and personalized recommendation scenarios with real cinematic data. Results demonstrate consistently high precision under the controlled evaluation setting, while recall, F-measure, and accuracy improve from 24.67
This paper proposes a novel neutrosophic correlation coefficient that explicitly exploits the intrinsic indeterminacy of Neutrosophic Sets (NSs) to model complex uncertainty in decision-making problems. The fundamental properties of the proposed measure are systematically established, including a weighted correlation coefficient, two complementary closeness indices, and a tunable composite index that enhances flexibility in preference aggregation. Building on this foundation, a refined TOPSIS framework is developed within a neutrosophic environment, yielding an effective NS-TOPSIS methodology. The proposed approach is applied to the prioritization of electric vehicles (EVs) under linguistic and indeterminate information. Comprehensive comparative analysis with existing multi-criteria decision-making (MCDM) methods confirms the consistency, robustness, and superior discriminative power of the proposed method, demonstrating its suitability for complex decision scenarios characterized by uncertainty and imprecision.
Computer-aided detection (CAD) systems are increasingly used as screening and decision-support tools to assist clinicians in the early detection and classification of diseases using medical imaging modalities such as computed tomography (CT), X-ray, and magnetic resonance imaging (MRI). However, excessive computational complexity, overfitting, and redundant feature representations generally hinder the actual usability of current DL-based Capsule Network (CapsNet) models. To solve these issues, in this research, an improved capsule network (ICN) is proposed for automated medical image categorization. The proposed ICN utilizes an optimized convolutional neural network (OCNN) for efficient feature extraction and a modified capsule architecture for classification. OCNN employs lightweight convolutional layers with Leaky Rectified Linear Unit (ReLU) activation and L1 regularization to improve feature learning, reduce redundant features, and alleviate overfitting. Further, an intermediate Secondary Capsule (S-Caps) layer is proposed between the Primary Capsule (P-Caps) layer and the Digit Capsule (D-Caps) layer to improve the hierarchical feature representation while keeping the computational efficiency. The proposed architecture was evaluated on multiple medical imaging datasets, including MRI, X-ray, and CT modalities. The experimental results proved that the ICN model achieved 6.66