
The rapid adoption of extended reality (XR) is driving the development of immersive educational services and increasing the need for systematic design approaches. From a Design for Extended Reality (DfXR) perspective, this study investigates authentic XR use-cases to develop an empirically grounded framework that synthesises design principles supporting educational service innovation, learner engagement, user adoption, and organisational implementation. The study is based on qualitative findings from open-ended provider and user-reflections, structured queries where use-case descriptions have been gathered. It explores a series of authentic use-cases, delving into both design and user experience aspects. The data was collected through firsthand experimental testing, use-case workshop reflections, and insights from the lead designer. In addition, two technology providers added use-case insights together with two classes of engineering design students. The findings uncover use-case design insights and their value concerns. Past research has shown that XR service solutions deliver added value to various stakeholders, employees, and customers on multiple fronts. The study identifies critical aspects for establishing user adoption proficiency and highlights that delivery value deviates between stakeholders, benefits, and practical use. This study contributes to the emerging DfXR discourse by synthesising findings from authentic XR use-cases into an empirically grounded framework for educational service design. The proposed framework demonstrates how pedagogical, technological, and organisational design considerations can be systematically integrated to support immersive learning, user adoption, and organisational implementation. By shifting the focus from technological capability to intentional design, the study provides practical guidance for developing scalable and effective XR educational services across higher education and professional learning contexts.
Detecting motor faults in real time is essential for timely maintenance, reducing downtime, and preventing potential losses. This article proposes a low-cost, real-time bearing fault diagnosis technique for multiple machines using deep learning by developing an efficient and lightweight model that accurately classifies bearing faults while minimizing computational requirements. This technique involves gathering vibration data with an MPU9250 sensor and an ESP32 microcontroller attached both machines housing (Induction Motor and Brushless DC Motor), sending the both machines vibration data wirelessly through Message Queuing Telemetry Transport (MQTT) Protocol to a Raspberry Pi microcontroller. The real-time vibration data is then transformed into images through the Short-Time Fourier Transform (STFT) technique. These images are subsequently utilized as input for a Light Weight Customized Convolution Neural Network (LWCCNN) model integrated into the Raspberry Pi microcontroller for predicting bearing faults for multiple machines. Finally, the proposed model is compared with MobileNet-V2 and Inception-V4 in terms of accuracy, model memory size, training time, No. of parameters and model robustness. The proposed model achieved a higher classification accuracy of 99.58
This study proposes a clinically informed, energy-aware Internet of Things (IoT) architecture for wearable vital-sign monitoring, addressing literature-reported limitations and healthcare-professional requirements. A sequential methodology was adopted. A bibliometric analysis of 310 Scopus-indexed articles published between 2019 and 2024 identified frequently monitored physiological parameters, sensing approaches, and recurrent system limitations. A survey of 76 healthcare professionals captured expectations regarding real-time monitoring, alerts, usability, data access, security, and system integration. These findings were translated into a three-tier wearable IoT architecture incorporating National Early Warning Score 2 (NEWS2)-based risk assessment and a machine-learning-assisted Intelligent Energy Management System (ML-IEMS). Retrospective validation used an initial cohort of 1500 VitalDB cases, with final policy-level simulation on 276 independent test cases totaling 1058 h and comparison against always-on monitoring and a rule-based IEMS baseline (Rule-IEMS). Heart rate, respiratory rate, and blood pressure were the most investigated parameters, while key gaps concerned autonomy, motion robustness, validation size, and multi-parameter integration. The survey confirmed demand for multi-parameter sensing, energy autonomy, and real-time data access. In simulation, estimated battery autonomy increased from 71 h with always-on monitoring to 155 h with Rule-IEMS and 159 h with ML-IEMS. ML-IEMS preserved episode sensitivity of 1.000 for NEWS2 ≥ 5 events, with a false-alarm burden of 1.06 alerts/h, indicating that future work should further optimize alert specificity. The proposed architecture links literature-derived limitations and healthcare-professional needs to a simulation-supported ML-IEMS policy, improving autonomy and signal fidelity while maintaining clinical sensitivity.
The visual symptoms of a particular disease on the leaves or the crop exhibit similarities regardless of the specific crop being examined. Despite the extensive research on disease classification in crops, there are limited studies on classifying diseases without explicitly incorporating crop identity as an input feature. This study explores a stacked ensemble approach to improve the predictive performance of deep learning models on diseased crop leaf datasets. Eleven deep learning models were trained to classify crop leaves into six labels and evaluate them using accuracy, training duration, and confusion matrices. The top five models were selected based on performance and subjected to McNemar’s statistical test to identify diverse model pairs. Subsequently, three ensemble models were formed using specific architectures as base learners and Logistic Regression as the meta-learner. The three ensemble models were assessed on three test datasets—one random test set comprising 30
2D video based Indian classical dance (ICD) identification using deep networks (DN) has been associated with anomalies like unequal spatial and temporal distribution of dancing subjects with respect to the manifesting pose sequences. Overcoming these challenges was materialized using complex spatial feature extraction and temporal motion representation algorithms. One excellent solution offered and accepted was increased dimensionality with 3D skeletal representation of ICD pose dynamics that can be directly captured using Microsoft Kinect sensor. In this work, ICD 3D skeletal poses (s3DICD) were captured for 5 different lyrical song descriptions from Kuchipudi using 5 different dancers at various stages of their learning cycle. In each frame a pairwise 3D joint-distance features are computed and a radial basis function (RBF) kernel maps them with inter-frame differences characterized by the proposed Joint Motion Kernel Maps(JMKM). The JMKM’s are color coded maps(CCMs) that represents 3D skeletal motions as bounded spatio-temporal images. Temporal lengths of dance poses are subjective to human dancer and the underlying pose characteristics. Therefore, the resulting the JMKMs also vary in width in accordance with the sequence length. We introduce Variable Overlapping Patch Modulation (VoPM), an adaptive patch-tokenization strategy that adjusts stride while maintaining a fixed patch size. This ensures a consistent number of tokens are supplied to a standard Vision Transformer (ViT) classifier for identification 3D dance poses. Experiments were conducted to evaluate the proposed JMKM-VoPM-ViT pipeline on s3DICD, NTU RGB+D, and Let’s Dance datasets.
To meet the stringent dimensional tolerances required for lithium-enriched ceramic granules, the Karlsruhe Institute of Technology (KIT) has devised the KALOS (KArlsruhe Lithium OrthoSilicate) technique, in which a steady molten jet is deliberately destabilised by an adjustable excitation frequency. Lithium-enriched ceramic granules are a cornerstone for the tritium-breeding blankets that will be employed in future fusion power plants, and their quality hinges on precise control of the jet-breakup process. In this work we present a novel high-speed camera-based measurement and control platform that automatically monitors and regulates pebble production. Image-processing algorithms extract droplet size, position, and inter-droplet distance distributions in real time. For this purpose, both classical image-processing methods and deep neural networks are investigated. Experimental results demonstrate that the system delivers accurate measurements and reliably adjusts the driving frequency to maintain target pebble dimensions. Thus, the proposed computer-vision system enables closed-loop control of the KALOS process, paving the way for robust, large-scale ceramic pebble fabrication.
Biosensors for disease diagnosis are composed of a bio-receptor to capture target biomolecules, a transducer to convert chemical phenomena into optical or electrical signals, and a detector to receive the signals. The performance of a biosensor depends not only on the physical sensitivity of the bio-receptors and transducers, but also on the postprocessing performance of the sensor signals. However, there have been very few studies on postprocessing of sensor signals compared with the studies on biochip itself. This study proposes a drastic reduction in assay duration of biosensors by post-processing sensor signals using a time-frequency signal analysis system. The system includes: (1) time-series signal preprocessors; (2) frequency analyzers; and (3) correlation analyzers with target biomolecules. The experiment focused on a liposome-immobilized cantilever biosensor developed to detect aggregated and fibrillated α-synuclein(αSyn), a major causative protein of Parkinson’s disease. The sensor signals from a built-in strain gauge R(t_i) were acquired for multiple samples with different αSyn concentrations. The system was utilized to split the sensor signal into small segments, preprocess them with low-pass filters and autocorrelators, convert them into frequency spectra, and then examine their correlation with αSyn concentration. The frequency spectra revealed clear difference among different αSyn concentration even in the shorter segments. Several segments revealed significant correlation ( r≥ 0.95 ) in the DC and the 10–20 mHz bands. Such good correlations were found in the segments of 560 s and 1,120 s, which are much shorter than the conventional assay duration, demonstrating the potential for drastic reduction in assay duration.
In high-dimensional datasets, many variables may be irrelevant or redundant, creating difficulties in data analysis, and interpretation. Unsupervised feature selection techniques address this problem by identifying relevant features without relying on labeled data. This paper introduces a novel unsupervised feature selection method that captures nonlinear relationships among variables using a neural network with a single hidden layer. The approach incorporates a sparsity regularization term to enable feature selection, along with a Laplacian regularization term that preserves the local geometric structure of the original data space. Experimental evaluations on different real-world datasets demonstrate the effectiveness of the proposed method compared to other methods in the literature, further highlighting the benefits of integrating Laplacian regularization.
In the rapidly evolving landscape of Cyber-Physical Systems (CPS), secure and resilient user authentication is essential for safeguarding sensitive data and ensuring operational integrity across domains such as smart transportation, healthcare, and industrial automation. This study proposes a robust multimodal biometric authentication framework that integrates facial and speech modalities at the feature level to overcome the limitations of unimodal systems in terms of accuracy and spoof resistance. Facial features are extracted using Gabor Wavelet transforms to capture texture-rich spatial details, while speech characteristics are represented through Constant Q Cepstral Coefficients (CQCC) to model dynamic vocal patterns. To enhance data integrity and prevent unauthorized tampering, an echo-hiding watermarking technique is employed to securely embed biometric features within audio signals. The fused feature vector is then classified using a Convolutional Neural Network (CNN), optimized for pattern recognition in high-dimensional feature spaces. The proposed system is evaluated on publicly available AVSpeech dataset, achieving an accuracy of 97.11
This study examines the effectiveness of enhanced multilingual embedding models in improving retrieval performance for Turkish medical text data. We consider two specific medical applications in Turkish language: the TUS examination, a standardized medical assessment featuring exam questions, and Clinical QA, which involves authentic patient-physician interactions. By implementing multi-stage fine-tuning protocols on domain-specialized models, we provide detailed performance assessment and explore cross-domain transfer capabilities of the trained models. Our findings indicate that domain-specialized models improve in-domain retrieval relative to generic models, and that systematic optimization through our multi-stage pipeline yields measurable gains in retrieval precision. For instance, domain-specific fine-tuning improves TUS retrieval performance from 0.69 to 0.79 in P@1 and from 0.77 to 0.85 in MRR, while Clinical QA fine-tuning with hard-negative sampling improves P@1 from 0.33 to 0.39 and MRR from 0.41 to 0.48 relative to the vanilla encoder. In addition, reranking improves P@1 from 0.788 to 0.823 in our evaluated setting, corresponding to a 4.5
Effective diet management is crucial for preventing non-communicable diseases, including heart disease, diabetes, and obesity. This paper introduces NutriVision, an enhanced system featuring an interactive chatbot that delivers real-time dietary guidance and personalized recommendations. NutriVision enables users to inquire about nutritional content, receive customized recipe suggestions, and track nutritional goals through natural language interaction. By integrating user data such as health conditions, dietary preferences, and nutritional requirements, it generates tailored responses that enhance user engagement and decision-making. With the help of computer vision and machine learning, NutriVision accurately identifies food items and estimates quantities from smartphone-captured images, providing instant nutritional analysis that includes both macronutrient and micronutrient details with a 94
Temporal deepfake detection models often suffer from high cross-fold variance and optimization bias, reducing their reliability in real-world forensic applications. Highly correlated sequential video data exacerbates this instability, causing traditional adaptive optimizers to overfit to specific data splits. To address this, we propose Momentum-Adaptive Switching (MAS) and its decoupled weight-decay variant, MASW. These proposed hybrid optimization strategies blend the adaptive gradient scaling of Adamax with the regularizing heavy-ball momentum of SGD to stabilize training trajectories. The framework utilizes a spatial backbone (evaluated on both MobileNetV2 and ResNet50V2), a BiLSTM with self-attention for temporal modeling, and a Kolmogorov–Arnold Network classifier. To prevent state collapse during inference under imbalanced folds, a robust Hidden Markov Model with Random Over-Sampling and class-weighted transition penalties is employed. Experimental results demonstrate that the proposed MAS framework significantly reduces cross-fold variance and outperforms modern optimizers like AdamW. Notably, controlled ablation studies reveal that MAS provides superior temporal stability, reducing prediction variance and label flip rates across varied sequence lengths, while mitigating catastrophic state collapse in downstream probabilistic modeling.
Applications that handle vast, real-time data streams originating from data repositories, social media platforms, sensor networks, mobile devices, and web-based sources require highly scalable computational frameworks. Parallel implementations that utilize the computational and storage resources of high-performance computing systems and cloud platforms support scalable big data analytics at present and will facilitate extreme-scale data analysis in the near future. As the amount of data outsourced continues to expand at an exponential rate, researchers are becoming increasingly interested in the problem of data deduplication in an effort to find workable and efficient solutions for cloud data centers. High redundancy in the cloud’s storage memory is a significant challenge. Data duplication in the cloud has been addressed using traditional algorithms, but doing so effectively is challenging since it calls for a solution to manage the duplication on an application-by-application basis in order to prevent the leakage of private information. Many industries, especially commercial ones, have come to rely on cloud storage. When it comes to securing the cloud and preventing hacking, scalability is seen as a crucial issue. One of the primary strengths of the cloud paradigm is its inherent scalability, which notably distinguishes it from more advanced forms of outsourcing. Auto-scaling represents a significant capability of cloud computing, permitting resources to be dynamically adjusted in accordance with changing demand levels. This research proposes an Automatic Scalable Cloud Environment based on Resource Demand and Time Limited Utilization with Data Deduplication (ASCE-RDTLU-DD) model for providing better Quality of Service (QoS) to the cloud users. The proposed model demonstrates superior performance in resource scalability and deduplication when compared with traditional approaches.
Image steganography has gained significant attention as a technique for secure data hiding, with Pixel Value Differencing (PVD) being one of the widely studied approaches. However, conventional PVD based methods often suffer from limitations such as the Falling-Off Boundary Problem and the Incorrect Extraction Problem, which may affect embedding reliability and image quality. To address these issues, this paper proposes a hybrid steganographic scheme that integrates Diagonal Pixel Value Differencing (DPVD), Local Binary Pattern (LBP), and Least Significant Bit (LSB) embedding within a unified 3 × 3 block framework. In the proposed scheme, secret data are first embedded in diagonal pixel pairs using a remainder based DPVD strategy, followed by additional embedding in the remaining pixels through LBP guided LSB substitution. This coordinated two stage embedding mechanism exploits both pixel difference information and local texture characteristics, enabling higher embedding capacity while preserving image quality. Experimental results show that the proposed method achieves an average Peak Signal-to-Noise Ratio of 52.47 dB at an embedding rate of approximately 2.2 (bit per pixel) bpp. Statistical and histogram based analyses indicate that the scheme introduces only minor distortion, suggesting reduced detectability under the evaluated statistical measures. Overall, the proposed DPVD–LBP–LSB framework provides an effective and practical approach for high capacity image steganography, with comprehensive evaluation against advanced steganalysis methods remaining for future work.
Deep neural network (DNN) approaches to discourse coherence achieve strong performance using distributional and contextual representations but often fail to explicitly model coherence levels, thereby limiting the explainability of their assessments. This limitation is particularly evident in Transformer-based DNNs, whose internal decision-making processes remain largely opaque. We address this gap by exploring multi-level discourse coherence through two approaches: traditional machine learning (ML) with explicit feature engineering and DNNs with architectural specialization. We employ explicit lexico-syntactic and discourse-level features to build interpretable ML models, whereas DNNs capture coherence patterns through architecture-driven representation learning. Extensive experiments on the Classification for Assessing Discourse Coherence task compare our models with state-of-the-art approaches, evaluating the impact of linguistic features and cross-domain generalization. Feature-based ML models—particularly ensemble learning methods combined with resampling and dimensionality reduction—deliver strong, interpretable performance but show limited generalization across domains. By contrast, Transformer-based DNNs, especially hybrid Transformer architectures that effectively integrate semantic and syntactic information, achieve high performance and demonstrate superior generalization across diverse writing styles. Feature-importance analysis shows that lexico-syntactic and discourse-level features significantly influence coherence prediction, with domain-dependent variations, underscoring the need for domain-aware modeling. Moreover, coherence across diverse writing styles is best captured through referential consistency. This study underscores the complementary strengths of feature-based ML models for interpretability and Transformer-based DNNs for robust cross-domain discourse coherence prediction.
In recent times, we tend to witness the quickest transformation of digital-health-care from a conventional hospital-centric system to a distributed patient-centric system. In this context, the Internet-of-Medical-Things (IoMT) is a game changer. This review aims to examine the role of IoMT systems in improving the predictive precisions and on-time response management of non-communicable or chronic-diseases. Additionally, it offers insights into the utilization of IoMT technologies and their impacts on the management of different chronic-diseases. The analysis highlights diverse approaches within IoMT frameworks, ranging from sensor integration to advanced Point-of-Care Testing (POCT) and wearable devices, elucidating their roles in improving health-care delivery for chronic-disease management. The findings extracted from this investigation lay the groundwork for forth-coming progressions in chronic-disease management within the IoMT sphere, underscoring its pivotal contribution to reshaping health-care delivery in managing chronic condition and enhancing patient outcomes in this regard. A thorough on-line search across multiple repositories including PUBMED, Web-of-Science, and the IEEE-Digital-Library was conducted to identify publications relevant to chronic-disease management within the IoMT domain. This search yielded a total of 1936 articles. After careful screening, 144 articles met the eligibility for inclusion in the review. Additionally, relevant articles cited within the selected publications were also taken into consideration.
In the rapidly evolving field of the Internet of Things (IoT), traditional systematic literature review (SLR) guidelines, such as PRISMA, often fall short in addressing the unique technical requirements and interdisciplinary challenges associated with IoT research. This paper introduces TEFA(S) (Things, Edge, Fusion, Application, Security), a novel framework specifically designed to enhance the effectiveness of SLRs in IoT by updating and extending PRISMA. TEFA(S) stands for Things, Edge, Fusion, Application, and Security—components critical to the comprehensive assessment of IoT literature. In addition to revising the guidelines for literature search, identification, and eligibility, TEFA(S) incorporates a structured quality assessment mechanism, including a tiered relevance model for evaluating the ‘Security’ component. This framework supports both qualitative and quantitative assessment, allowing flexibility while maintaining rigor. A dedicated evaluation of TEFA(S) against existing SLR methodologies further clarifies its relevance in technical domains. Empirical validation demonstrates the utility of TEFA(S), while its component-based architecture ensures adaptability across diverse IoT domains, such as smart healthcare, Industrial IoT (IIoT), and cybersecurity. Researchers conducting SLRs in the IoT field are strongly encouraged to adopt TEFA(S) alongside updated PRISMA practices to improve the relevance, reliability, and generalizability of their SLRs.
Generative Artificial intelligence is one of the emerging technologies in the 21st century, and this article delves into the fascinating relationship between Generative AI and the Internet of Things (IoT), exploring how they come together in various domains. By thoroughly reviewing existing literature, this study sheds light on the practical applications of Generative AI in IoT, with a specific focus on areas such as smart homes, healthcare, industrial IoT, automotive, and smart cities. Each of these applications is examined meticulously, highlighting how the integration of Generative AI and IoT can bring about improved control and automation, heightened operational efficiency, optimized urban infrastructure, and more. Furthermore, the article exhibits a prototype implementation of a generative AI-based voice assistant for the Internet of Things as a real-world example. The fundamental algorithms are covered in detail in this paper, including voice recognition, GPT-4 response creation, and the speech-to-text conversion of created text. Concluding the study, the article provides insights into the fascinating potential of generative AI in IoT. It highlights the opportunity for tailoring a future world that is linked and intelligent.
Data leakage, where semantically or visually similar samples exist across training and test splits, continues to threaten the reliability of object detection benchmarks. The D-LeDe method was recently proposed as a statistical technique for detecting data leakage, showing promise in initial applications. This paper extends the D-LeDe method to evaluate its robustness and generalizability through a focused experimental study. We revisit the KITTI dataset which was previously identified as leakage-prone and introduce a new application of D-LeDe on SODA10M. The core investigation centers on whether visually similar images, measured using perceptual hashing, are the primary cause of data leakage indications captured by D-LeDe. To this end, we progressively remove visually similar image pairs from the test sets of both datasets and observe changes in the Relative Increase Rate, the key decision metric of D-LeDe. Results show that even after removing about 50
This study investigates user correction, that is, challenging others who post misinformation on social media; communication styles (direct or indirect) when doing so; and the influence of relational, psychological, and competency factors on that. It specifically examines the influence of altruism, empathy, self-esteem, and social media use competency in two distinct cultural contexts: the United Kingdom (UK) and Arab Peninsula countries. We collected data from 686 participants (367 British and 319 Arabs) using an online survey supplemented with vignettes. The findings revealed an intention–behaviour gap in user correction in both cultural contexts. Participants in both groups preferred indirect communication. Multivariate Multiple Regression (MMR) analysis revealed that altruism and social media use competency predicted both the willingness and actual participation in user correction in both cultural groups, whereas self-esteem predicted them only in the UK context. Empathy showed no significant prediction in both cultural contexts. Further MMR analysis of communication style preferences considering the variations in gender similarity (different, same), social status (higher, lower, equal), and social distance (distant, close) between the participants and the misinformation posters showed that altruism predicted preference for direct and indirect styles across relational contexts in both samples, while self-esteem played a more prominent role in the UK sample. In the Arab sample, social media use competency predicted style preferences, particularly in different gender and social status variation contexts. These findings highlight the importance of considering the interplay between cultural, relational, and psychosocial factors when designing socio-technical interventions to promote user correction and combat the spread of misinformation.