
In healthcare, accurately mapping materials from purchase orders into the SAP system is a critical but challenging operational task, often fraught with inefficiencies and errors. To address the limitations of manual handling and conventional automation, we introduce a multi-modal Retrieval-Augmented Generation framework. The proposed system is designed to capture the contextual complexity of healthcare supply chains by integrating information from heterogeneous sources, including textual descriptions, product specifications, and images. We evaluate the framework on a dataset of 230 real-world purchase orders, demonstrating significant improvements in both efficiency and accuracy compared to existing approaches. Furthermore, the system enhances trustworthiness by incorporating explainable AI mechanisms, ensuring transparency in decision-making. Our findings highlight the potential of the proposed framework to provide an intelligent, scalable, and reliable automation solution for material management workflows in healthcare facilities.
The integration of artificial intelligence (AI) and data-driven decision-making in education has the potential to enhance learning experiences, improve student outcomes, and support educators in personalized instruction. However, many implementations of learning management systems lack predictive analytics, adaptive learning mechanisms, and proactive intervention strategies. To address these challenges in Malaysia, we developed SINAR, an AI-powered educational support system designed to provide personalized learning interventions, predictive performance analytics, and intelligent recommendations. SINAR incorporates deep learning models, large language models (LLMs), and explainable AI techniques to enhance student monitoring, intervention planning, and content recommendations. This paper provides a pilot expert validation-based case study, focusing on educators’ perceptions of its usefulness, ease of use, and effectiveness. The results indicate strong acceptance, with experts recognizing SINAR’s ability to improve job efficiency, productivity, and decision-making. The system’s learning intervention functions, predictive analytics, and recommendation capabilities were rated as highly significant. Additionally, most experts expressed willingness to adopt SINAR and recommend it to others, demonstrating confidence in its potential impact. This preliminary study provides early evidence of SINAR’s feasibility and acceptance, demonstrating its potential to bridge existing gaps in personalized and data-driven learning. Moving forward, continued refinements will focus on enhancing system adaptability, explainability, and scalability across educational environments.
With limited hardware capabilities the deployment of multi-object tracking in near real-time remains a challenging task. To address this, a combination of YOLOv10 and DeepSORT for object tracking and detection is proposed which can run on CPU-only devices. This paper carefully selects the key parameters such as the resolution and removes irrelevant objects from an image. This work achieved 86
Assessing the ethical risk of large language models (LLMs) is a critical challenge for Human-Computer Interaction (HCI). A standardized, theory-based method for ethical risk assessment remains scarce. To address this gap, we present the Ethical Risk Scoring System (ERSS), a framework grounded in a multi-theoretic ethical foundation and operationalized through the FACTOR9 determinants. ERSS consolidates the most widely discussed concepts of ethical risk drawn from both contemporary literature and all major ethical theories—including deontological, utilitarian, to environmental ethics. It then models the intrinsic functional relationships among these ethical dimensions to define the Ethical Integrity Quotient (EIQ)—a standardized index designed to quantify and compare the overall ethical quality of an artificial system. We introduce an innovative HCI-based methodology which utilizes the new conversation capabilities of the Chatbots built on the LLMs. We demonstrate this technique by evaluating five widely LLMs, each exhibiting distinctive ethical, interactive and design characteristics. To our knowledge, no existing quantitative framework integrates such a multi-theoretic foundation of ethical risk analysis with a fully operationalized, HCI-driven conversational interrogation methodology.
Tea leaf diseases can have detrimental impacts on tea production. Convolutional neural networks (CNNs) are increasingly being utilized in agriculture to identify and classify plant diseases automatically. Working in similar lines, this paper introduces an attention-based deep CNN for robust and efficient detection of multiple tea leaf diseases under complex field conditions. The designed CNN utilizes a stack of Mobile Inverted Bottleneck Convolution (MBConv) blocks, integrating Squeeze-and-Excitation (SE) modules as a feature extractor, with a Shuffle Attention (SA) block at the end to effectively capture contextual information, further enhancing the low-level features. The proposed pipeline employs the Segment Anything Model (SAM)-based zero-shot pre-processing step to enhance image contrast, remove background, and facilitate more accurate disease feature extraction, even under variable illumination and background clutter. To validate the efficacy of the proposed framework, we conducted extensive experiments on a comprehensive tea leaf dataset (teaLeafBD) containing six disease classes and a healthy class. The proposed model, with 76,560 trainable parameters, 299.08 KB memory footprint, and 0.406 GFLOPs, achieved a mean 5-fold cross-validation accuracy of 92.90
Efficient inventory control underpins modern supply chain networks by maintaining demand-supply balance, reducing costs, and ensuring timely fulfillment. However, traditional manual or basic automated systems face challenges such as human error, delayed visibility, overstocking, and difficulties in managing cluttered warehouses with stacked items, varying orientations, and scales. These issues lead to inaccurate tracking, operational delays, and higher expenses. This paper proposes a hybrid vision-based framework that integrates an ensemble of YOLOv8 detection–segmentation models with self-supervised DINOv2 Vision Transformer embeddings for precise object localization, counting, and classification in inventory scenes. The YOLOv8 ensemble ensures robust detection under diverse conditions, while DINOv2 offers semantically invariant representations for reliable recognition. In practical settings, the framework enables automated monitoring in environments such as manufacturing warehouses, retail stockrooms, and healthcare facilities by integrating seamlessly with enterprise systems to support proactive replenishment and cost optimization. Overall, this scalable approach enhances supply chain intelligence by addressing traditional limitations and promoting automation-driven inventory operations.
This study explores the application of multivariate pattern analysis (MVPA) on fMRI data to classify osteoarthritis patients and healthy controls using functional connectivity from the middle frontal gyrus. The dataset, sourced from the paper “Brain connectivity predicts placebo response across chronic pain clinical trials”, focused on placebo-only and treatment-only studies. Preprocessing steps included normalization, smoothing, and extraction of time-series data, followed by the use of tangent space embedding to compute connectivity matrices. Feature selection and classification were performed using models such as SVM and logistic regression. The middle frontal gyrus was selected based on prior evidence linking it to placebo response and chronic pain processing. Our models showed promising classification accuracy, particularly in the placebo group. The findings align with previous research demonstrating the capacity of MVPA to decode pain-related brain patterns and predict clinical outcomes. This study highlights the potential of using brain connectivity patterns as neurobiomarkers for chronic pain.
To address the growing changes in environmental factors and lifestyle threats to respiratory health, this study presenting a deep learning–based hybrid framework for precise detection and classification of lung cancer from chest CT images. The proposed model integrates VGG19 and ResNet152, that is utilizing most important features extraction capability to enhance diagnostic reliability. To further improving optimization and stability of model, we make a hybrid of this model with six nature-inspired algorithms—Greylag Goose Optimization, Crested Porcupine Optimization, Lotus Effect Algorithm, Polar Light Optimization, Ant Lion Optimizer, and Walrus Optimization. Using a Kaggle Chest CT dataset comprising 1,000 images divided into 70
Human gaze contains a rich amount of information about the human attention on the scene. Thus, it is often used as a robust and rapid method of interaction with the computer. Predicting future gaze location in advance can provide an advantage in multiple use cases, including, but not limited to, human-computer interaction, designing efficient user interface, analyzing human behavior, image or video compression. However, prediction of gaze poses a unique challenge because of the dynamic nature of the human vision system and spatio-temporal nature of the data. We explore the applicability of multimodal CNN-LSTM network to predict future gaze location in videos. We compare the results of our model on the Coutrot and EGTEA gaze dataset against existing statistical models. The results show that modeling eye gaze with spatiotemporal model performs better than models that use per-scene analysis to predict eye gaze locations.
Complexity of ecological data and the limited availability of intuitive analytical tools to support decision-making have posed a challenge to biodiversity monitoring effectiveness in tropical forests. Digital dashboards provide a way to visualize and integrate biodiversity information, but many remain underutilized due to insufficient attention to usability and interpretability. This paper presents a human-centered usability assessment of the Biodiversity Dynamics Assessment Dashboard developed for the Pasoh Forest Reserve, Negeri Sembilan, Malaysia. The dashboard functions as an intelligent augmentation system by transforming raw ecology data into cognitively supportive visualization for decision-makers and researchers by aggregating multivariate ecological data such as forest composition, species dynamics, and carbon indicators. Participants from environmental and governmental agencies evaluated the dashboard using the importance of ratings for key functions and metrics and System Usability Scale (SUS) score. The findings contribute to the design of human-centered biodiversity dashboards and offer design implications that enhance interpretability, transparency, and decision confidence in biodiversity monitoring systems.
With availability of computing, memory and AI there is a surge in research related to cognitive science and cognitive architecture. This has led to increase in new design, development and deployment of cognitively enabled systems in the environment. The cognitively enabled system are deployed in the land, sea, air, outer space and cyber space in the form of autonomous systems or unmanned systems. Often human operators have to co-work with these autonomous systems such as unmanned ground vehicle, unmanned under water vehicle and unmanned aerial vehicles which in turn is increasing the cognitive load of the human operator in the loop. The increased use of cognitively enabled unmanned systems has fueled the research and development of a failsafe and robust cognitive architecture which meets both the safety critical and mission critical criteria. In this paper we propose a generic self-aware cognitive architecture for an autonomous system which answers two critical questions viz. Where am I? amp; what is my next state? The proposed architecture performs SLAM (Simultaneous Localization and Mapping) in a SWaP (Size weight, area and Power) constrained environment. Continuously meeting the mission requirements of a cognitive system in operation.
This Research investigates the effectiveness of an Audio Spectrogram Transformer (AST) using a moderate negative mining technique to detect depression from speech automatically. Addressing the significant challenge of class imbalance in the DAIC-WoZ speech depression dataset, we implemented a moderate negative mining technique to selectively sample negative instances for balancing the dataset, thereby enhancing model learning and overall performance. The approach leverages transfer learning by fine-tuning a pre-trained AST model on the DAIC-WoZ dataset. Experimental evaluation demonstrates that the proposed method achieves robust classification results, attaining a macro F1 score of 0.7353 and a test accuracy of 76.43
Affective computing enhances human–computer interaction by enabling systems to recognize and respond to emotional states. Ear-centered sensing offers a practical approach due to comfort, stability, and unobtrusiveness. This paper reviews advances in ear-based electroencephalography (EEG), ear-based photoplethysmography (PPG), and their integration with facial expression analysis for robust affective state recognition. We highlight contributions in on-device deep learning, time-frequency feature extraction, and multimodal pipelines for real-time stress, fatigue, and emotion monitoring.
Brain Computer Interface (BCI) technology provides a communication link between the human brain and external devices, bypassing regular neuromuscular pathways and offering a wide range of smart home applications. This paper aims to develop a non-invasive and efficient BCI system that detects the user’s eye-open and eye-close states using electroencephalogram (EEG) signals. In this direction, we have used sub-band characteristic Response Vector (sub-band CRV) based feature extraction method which represents the direction in which brain’s static energy is concentrated. It captures the intricate inter-channel dependencies by computing a correlation matrix across EEG channels and applying eigenvalue decomposition to generate a low-dimensional description of brain activities. These sub-band CRV features are used to train a suitable classifier that distinguishes between the open and closed state of the eyes which are then mapped to control various smart home appliances. The results show that the proposed approach efficiently distinguishes between eye open and eye closed brain states, paving the way for a reliable and responsive BCI-driven home automation framework.
For recommendation systems, the problem of class imbalance arises from the fact that interactions from the minority class – i.e., premium purchases, niche preferences, rare ratings – are highly under-represented and as a result predictions are skewed and fail to capture patterns that are commercially interesting. Common oversampling methods are unsuitable for recommendation scenarios, as they create fake interactions from which the inherent impossible-to-solve structure of user-item matrices are learned by the recommender system, overturning the existing extreme sparsity (>90
Cognitive load monitoring plays a crucial role in designing accessible and adaptive interfaces, particularly for visually impaired persons (VIPs) whose sensory processing strategies may impose different mental demands compared to sighted individuals. This work implements a real-time EEG-based system that quantifies workload through a Cognitive Score (CS) aligned with industrial standards, where CS<50 denotes no-load. EEG signals are acquired during an auditory gamified learning intervention system, preprocessed to remove artifacts, and segmented for feature extraction using power spectral density, entropy measures, and engagement indices. These features are mapped to continuous CS values, enabling differentiation of low, medium, and high workload levels. A dedicated web application that provides an interactive platform for uploading EEG data, performing automated signal analysis, and visualizing results through dynamic plots of CS and spectral features. Experimental evaluation demonstrates high responsiveness and reliable discrimination of workload states, offering a scalable solution for assistive neurotechnology and inclusive cognitive state research.
Spatial queries are often a concatenation of many simple queries expressed in an unstructured way. Formulating and posing spatiotemporal queries in a structured manner and posing them on a database is cognitively challenging unlike structured query language posed on a normalized databases. In this paper we propose an AI Agent which tries to resolve queries which are unstructured and more close to natural language query posed by human. Navigation is a cognitively intensive task. Humans have developed applications that make this task less cumbersome. We try to study this by posing queries in natural language and observe results to see how a formal systems approach to navigation via map applications differs to a human-based approach.
Cryptojacking, defined as the covert exploitation of victim resources for the purpose of unauthorised cryptocurrency mining, represents a growing threat to organisational infrastructure. The present paper presents a computational framework for modelling and analysis of cryptojacking attacks through multi-order adaptive behavioural analysis and the role of support of it by Human-AI interaction. The approach involves decision-actions made by internal mental models and learning or forgetting of these mental models individually and sharing them between an AI Coach and employee. A structured What-If analysis further shows risk reduction when the AI-Coach is enabled.
Electronic health records (EHRs) contain highly sensitive personal and medical information, making privacy preservation a critical concern when leveraging cloud computing for machine learning (ML) analytics. The widespread adoption of cloud-based healthcare systems has introduced significant challenges in maintaining data confidentiality during both training and inference phases of ML models. Traditional cloud architectures fail to provide adequate service level agreements regarding availability, integrity, and confidentiality of sensitive health- care data. This paper presents a novel secure cloud architecture that enables privacy-preserving machine learning on EHRs while seamlessly integrating web-based analytics tools for healthcare professionals. Our proposed system employs a multi-layered security approach incorporating homomorphic encryption, secure multi-party computation, differential privacy, and federated learning techniques. The architecture supports real-time inference through privacy-preserving artificial neural networks while maintaining HIPAA and GDPR compliance. Experimental results demonstrate that our system achieves 94.2
Accurate brain MRI segmentation remains a challenging task due to the presence of noise, intensity non-uniformity (INU), and tissue boundary ambiguity. Traditional Fuzzy C-means (FCM) variants either incorporate spatial information or perform bias field correction but seldom integrate both while modeling high-order uncertainty. To address these limitations, we propose a Bias-Corrected Picture Fuzzy C-means algorithm with Spatial regularization (S_BCPFCM). The proposed method employs the picture fuzzy set framework to represent the acceptance, rejection, and hesitation degrees of pixels, effectively handling uncertainty near tissue boundaries. A bias field correction term is introduced to compensate for INU, while an adaptive spatial regularization term preserves edges and suppress noise. Extensive experiments were conducted on the BrainWeb simulated MRI dataset with varying noise levels (5 α =10,30,50 ). The performance was evaluated using fuzzy performance index (FPI), modified partition entropy (MPE), Xie-Beni (XB) index , average segmentation accuracy (ASA) , and dice score (DS). Experimental result demonstrate that S_BCPFCM consistently outperforms existing FCM variants, achieving higher segmentation accuracy and robustness in challenging noisy and biased conditions.