
Enzyme function classification plays a critical role in understanding biological processes, drug discovery, and protein annotation. This paper presents a computational pipeline that leverages ESM2, a transformer-based protein language model to generate contextual embeddings from raw amino acid sequences. We explore strategies to address class imbalance and evaluate the embeddings on two supervised learning architectures: a MLP and a deeper custom neural network. Our observations demonstrate that the MLP model with oversampling achieves the best performance, achieving a test accuracy of 93.5% and macro F1-score of 91% outperforming deeper architectures and class-weighted loss. Our findings suggest that even embeddings generated from a lightweight transformer combined with effective imbalance handling techniques can provide an efficient solution for enzyme function classification.
This paper studies the stability of perturbed continuous-time Markov chains (CT-MCs) by establishing stability criteria and robustness indicators. The robust stability criteria are derived from the spatial relationships between reachable sets and invariant subsets. Moreover, a robustness indicator is formulated to quantify the ability of each transition to resist perturbations affecting the stability, where lower values signify more critical transitions. Finally, a biological case study demonstrates the applicability of the theoretical result. This framework not only facilitates stability assessments for perturbed CT-MCs but also identifies critical transitions, the monitoring of which can enhance system resilience.
The rapid evolution of large language models (LLMs) has made machine-generated text (MGT) increasingly indistinguishable from human writing, posing challenges in authenticity detection. In this study, we present a novel hybrid framework that integrates the structural sensitivity of string kernels with the semantic depth of transformer embeddings to detect AI-generated content. We propose four complementary methods including Attention-Augmented Kernels and a Custom Kernel Function to capture both linguistic structure and contextual nuance. Our evaluation across eight diverse datasets, featuring texts from GPT-3.5, GPT-4, DeepSeek, and Kimi, demonstrates that the Transformer-Guided N-gram Selection and Custom Kernel Function consistently outperform traditional baselines in accuracy and computational efficiency. This framework offers a scalable, interpretable, and robust solution for real-world MGT detection across varied generation styles.
Segmenting targets with ambiguous features in medical images poses a persistent challenge for quantitative analysis. This is particularly true for stool segmentation in abdominal X-rays—vital for the Stool Volume Score (SVS) in constipation management—where conventional single-channel deep learning methods often struggle. This study investigates if incorporating expert-annotated intestinal gas masks as an explicit, additional input channel to a U-Net model enhances stool segmentation accuracy and SVS reliability. An ablation study compared a dual-channel U-Net (X-ray + gas mask) against an identical single-channel baseline (X-ray only), using the same architectures and training protocols. Performance was assessed using the Dice coefficient, Precision, Recall, and SVS analysis (Pearson correlation, Bland-Altman agreement). The dual-channel model showed an improved Dice coefficient (0.726 vs. 0.710) and Recall (0.706 vs. 0.675) compared to the single-channel baseline, while Precision was slightly higher for the baseline (0.748 vs. 0.747). Crucially, the enhanced segmentation by the dual-channel model yielded a stronger correlation between predicted and ground-truth SVS (Pearson r=0.852 vs. 0.818) and excellent SVS agreement (Bland-Altman bias: 0.01, limits of agreement: -0.06 to 0.07). Explicit gas mask guidance thus improves overall stool segmentation, primarily by achieving a more comprehensive capture of stool regions, which in turn enables more reliable SVS quantification. This dual-channel strategy offers a promising approach for objective fecal load assessment in clinical practice.
The development of artificial intelligence (AI) technologies for reasoning based on big data is rapidly advancing day by day. Moving beyond large language models (LLMs), recent technological trends show the emergence of large vision models (LVMs), indicating that the application scope of AI is expanding at an accelerated pace. Currently, most AI services are implemented through software technologies. However, from the perspective of energy saving and environmental pollution, it is a crucial turning point where a shift to hardware-oriented AI technology must take place. Hardware-oriented AI aims to move away from the conventional series high-speed operation, towards low-power computing technologies that maximize computational concurrency. In order to achieve this goal, changes in computing architecture are necessary, with semiconductor memory technology playing a central role. Simultaneously, recent research indicates that high-speed, large-scale computation systems naturally lead to increased system temperatures, which can produce gases harmful to human health. Although these goals differ in terms of the original starting points, all these technological objectives share a common aim of low-power and small-number computing. This paper examines next-generation AI computing technologies based on large-capacity memory technologies, specifically dynamic random-access memory (DRAM) and flash memories built on relatively mature Si fabrication processing suitable for chip production, evaluating pattern recognition accuracies.
Accurate traffic sign recognition (TSR) is critical for autonomous vehicles, especially in diverse driving environments. While datasets such as the German Traffic Sign Recognition Benchmark (GTSRB) have been extensively studied, limited attention has been paid to specific regional challenges, especially in Malaysia. This study explores traffic sign recognition in Malaysia using a composite-scale convolutional neural network (EfficientNet). The Malaysian Traffic Sign (MTS) dataset contains signs with cultural uniqueness, language, and design differences, which are underexplored compared to global datasets. EfficientNet-B5 was selected for its balance between accuracy and efficiency. Enhancements including image resizing, architecture depth expansion, and data augmentation were applied. Results show that EfficientNet-B5 achieves significant improvements, especially on the MTS dataset, demonstrating the potential for scalable, real-time TSR for autonomous driving systems in Malaysia.
Adversarial attacks against object detectors have often been limited by the use of visually conspicuous or synthetic patterns, reducing their practicality in real-world scenarios. In this work, a novel framework is proposed to generate naturalistic adversarial textures by transforming authentic Southeast Asian clothing designs through a structured pipeline. The method is built upon 2D-to-3D UV mapping, dual K-means clustering for pattern simplification, and differentiable Voronoi-based optimization to enable gradient-based attacks while preserving visual realism. The effectiveness of the approach is evaluated using dynamic human models rendered with 360° viewpoint rotations and varying body geometries based on the SMPL framework. Detection confidence from YOLOv3 is consistently suppressed, with an average value of 0.3966 observed across five body models. Strong generalization is demonstrated under diverse poses and viewpoints, confirming the robustness of the adversarial textures. These findings suggest that visually plausible garments can be exploited to achieve stealthy adversarial effects, and future work may explore extensions to handle challenging lighting conditions and real-time applications.
Analyzing vulnerabilities in software supply chains is critical for mitigating security risks. The Software Bill of Materials (SBOM) provides transparency by listing all components and dependencies. Additional sources, such as the National Vulnerability Database (NVD), the Open Web Application Security Project (OWASP), and the Common Vulnerabilities and Exposures (CVE) system, enhance vulnerability tracking. However, relying on a single source leads to incomplete analyses, as vulnerabilities in dependencies may still pose threats even when the primary software appears secure. Therefore, we require a comprehensive process to provide information on dependency vulnerability and an understanding of general software vulnerabilities.To address the above-mentioned problems, we introduce SOMVE (SbOM and cVE), a process based on Machine Learning (ML) model to integrate SBOM and CVE. This integration provides comprehensive vulnerability information, which is not available by SBOM or CVE individually otherwise. As such SBOM provides 5 metrics and CVE contains 21 metrics. Therefore, we identify and extract the matching CVIDs that are common between both datasets. The Random Forest (RF) model used in SOMVE traces the vulnerabilities (using CVIDs) missed by focusing exclusively on the SBOM data. Specifically, SOMVE includes key details such as vulnerability scores, impacts, descriptions, and assessments. A rigorous experiment on SOMVE shows 97% accuracy in mapping the CVE IDs from both sources. Finally, SOMVE generates a consolidated JSON report with 11 metrics that highlight the importance of comprehensive vulnerability analysis to secure software supply chains.
The increasing use of unmanned aerial vehicles (UAVs) in areas like surveillance, environmental monitoring, and disaster response underscores the urgent need for high-quality imaging. Unfortunately, the limitations of onboard sensors often lead to poor-quality, low-resolution aerial images, which can compromise how accurately we interpret scenes. This paper introduces an improved image super-resolution method that utilizes the Real-ESRGAN architecture, specifically tailored to enhance vertical UAV imagery for better visual clarity. By employing Residual-in-Residual Dense Blocks (RRDB) along with a relativistic discriminator, the model successfully reconstructs high-frequency textures and minimizes noise artifacts in aerial images. We created a custom dataset featuring synthetic degradations to mimic real-world UAV conditions. The model was tested using Peak Signal-to-Noise Ratio (PSNR), Root Mean Square Error (RMSE), and Perceptual Index (PI) across 15 test images, demonstrating notable improvements compared to baseline ESRGAN versions. This research not only advances the field of aerial image enhancement but also showcases a practical approach for integrating deep learning-based super-resolution into real-time UAV applications.
Kerangas forests (tropical heath forests) are nutrient-poor, fire-prone ecosystems threatened by deforestation and land-use change. Monitoring these post-disturbance landscapes is crucial for ecological restoration but remains challenging due to field inaccessibility and the high resource demands of conventional deep learning. This study evaluates lightweight CNNs (MobileNetV1/V2) for classifying Kerangas imagery into three ecological succession stages. Using transfer learning and domain-specific data, twelve RGB and grayscale models were assessed by accuracy, latency, RAM, and flash usage for edge deployment. MobileNetV2 with 160×160 RGB and α =0.75 reached 95.8% accuracy, matched by a 96×96 grayscale model with α =0.35—which reduced memory by over 97% and achieved sub-second inference time.
Liquid chromatography-tandem mass spectrometry (LC-MS/MS) serves as a key tool for the test of lipophilic substances in laboratory medicine and is widely employed in the analysis of coenzyme Q10 (CoQ10) and 25-hydroxyvitamin D (25OHD). In this paper, fuzzy concept was applied to improve the LC-MS/MS methods used for CoQ10 and 25OHD detection. The focus was placed on selecting the optimal mobile phase for CoQ10 analysis and examining the differences between LC-MS/MS and chemiluminescence immunoassay (CLIA) methods for 25(OH)D measurement. Through screening various organic phase combinations and employing fuzzy inference, the optimal mobile phase ratio for CoQ10 test is determined to be methanol and isopropanol at a ratio of 8:2. Additionally, fuzzy logic was employed to analyze the variations in 25OHD concentrations across different sexes and age groups. The results showed that women aged 30–40 exhibited greater differences in 25(OH)D levels compared to other groups. This study shows that the use of fuzzy concepts can enhance the adaptability and accuracy of LC-MS/MS detection, offering a novel approach to the analysis of lipophilic substances.
Few-shot remote sensing scene classification (FSRSSC) tackles the challenge of recognizing novel scene categories with only a limited number of labeled examples, which heavily relies on pre-trained transferable deep representations. Recent advances have widely explored contrastive learning to improve global feature representations for this task. However, we argue that global feature representations often fail to capture the fine-grained local features crucial for distinguishing remote sensing scenes, which typically include numerous small-scale, densely distributed objects. To address this limitation, we propose Dense Supervised Contrastive Learning (DSCL), which applies supervised contrastive learning at the patch level to improve local feature discriminability. By optimizing a dense pairwise similarity loss across local patches, DSCL significantly boosts generalization in few-shot scenarios. Experiments on three benchmark datasets (UCM, WHU-RS19, and NWPU-RESISC45) demonstrate that DSCL achieves competitive performance compared to recent state-of-the-art methods in both 5-way 1-shot and 5-shot settings, highlighting the potential of local-level feature learning for FSRSSC.
Sequence classification is a core task in computational genomics with wide-ranging applications in gene regulation, functional annotation, and disease prediction. In this study, we propose a novel hybrid deep learning architecture that combines Vision Transformers (ViTs) with Convolutional Neural Networks (CNNs) for classifying human non-TATA promoter sequences. To harness the representational power of vision-based models, we introduce a multi-channel Hilbert Curve encoding technique that transforms linear DNA sequences into 2D image-like grids, spatially preserving k-mer relationships and regulatory motifs. This spatial restructuring enables CNNs to extract local features while allowing ViTs to capture global dependencies via self-attention. Evaluated on a balanced dataset of 36,131 promoter and non-promoter sequences (251 bp each), the proposed model achieves a validation accuracy of 90.28%, along with strong precision and recall across both classes. Our approach demonstrates that image-based sequence representation, when paired with hybrid architectures, offers a powerful alternative to traditional sequence modeling.
The integration of artificial intelligence (AI) and computer vision in agricultural scenarios provides a significant advancement in crop monitoring, autonomous navigation, and disease detection. The following research proposed a system to improve agricultural automation incorporated with deep learning-based object detection models such as RCNN, ResNet50, and DenseNet121. Using computer vision models, this research aims to optimize the identification of potential obstacles and crops in farming environments. Moreover, this research justified the role of autonomous ground robots for navigation, path planning, and real-time decision making. The following study evaluates existing methodologies and presents an improved framework that combines deep learning architectures with robotic perception systems to enhance agricultural automation. Furthermore, a Soil Science Rover Test Module (SSRTM) has been proposed in this study, which is responsible for in-situ soil analysis focusing on moisture percentage, pH levels, and nutrient composition. Experimental results justified the effectiveness of the proposed systems in real-life scenarios. This research contributes to the growing field of AI-driven precision agriculture, which will eventually come up with intelligent and fully autonomous farming systems.
Despite their many potential applications, graph neural networks (GNNs) are challenging to deploy on low-resource devices like Internet of Things nodes and mobile platforms due to their high energy and processing demands. We propose MLC-GNN, a lightweight architecture that dynamically optimizes depth, precision, and channel width during inference. MLC-GNN uses a middle-layer controller based on real-time mutual information estimate and a PID-based energy guard to offer efficient adaptation under strict energy and latency constraints. Extensive experiments on four benchmark datasets and five baselines demonstrate that MLC-GNN achieves competitive or superior accuracy and precision, while reducing energy consumption and inference delay by up to 50%. These results show the utility of MLC-GNN in enabling energy-efficient GNN inference on edge devices.
This study proposes MAP-CARE, a radiomics-based framework for predicting future surgical intervention in patients with asymptomatic carotid artery plaques using multi-modal MRI. Unlike prior research that focuses on stroke risk prediction, MAP-CARE directly targets the clinical decision of whether carotid revascularization, such as endarterectomy or stenting, will be indicated within one year, even before the onset of neurological symptoms. Radiomic features are extracted from 3D Turbo Spin Echo (3D-TSE) and Time-of-Flight (TOF) MR sequences to capture both structural and hemodynamic characteristics of carotid plaques. Feature selection is performed using principal component analysis and mutual information. Four machine learning models (logistic regression, support vector machine, LightGBM, and random forest) are trained and evaluated using stratified five-fold cross-validation. A total of 36 carotid arteries from 18 patients were analyzed. The combination of mutual information-based selection and a random forest classifier yields the best performance, achieving an AUC of 0.933. To enhance model interpretability, SHapley Additive exPlanation (SHAP) is applied to identify important geometric and texture-based features. By bridging non-invasive imaging with real-world clinical outcomes, MAP-CARE provides an interpretable and actionable tool for proactive treatment planning in asymptomatic carotid stenosis.
Understanding pediatric brain maturation is essential for detecting neurodevelopmental disorders. However, quantitative spatiotemporal analysis remains limited in clinical practice, particularly using computed tomography (CT). This study proposes a perturbation-based deep learning framework that estimates brain age from pediatric CT scans while revealing region-wise contributions across developmental stages. A 3D ResNet architecture is trained on pediatric CT images to perform age regression and to enable region-wise interpretability through perturbation analysis. Anatomically guided perturbation analysis is conducted, consisting of lobe-wise masking to evaluate major regions. The proposed method was validated using head CT data from 201 pediatric patients aged 0 to 47 months with five-fold cross-validation stratified by age. The model achieved high accuracy with an RMSE of 5.205 months, an MAE of 4.007 months, and a Pearson correlation coefficient (r) of 0.925. The analysis reveals a dynamic posterior to anterior shift, with occipital lobe dominance in early infancy, emerging parietal lobe contributions in mid-infancy, and increasing frontal and temporal lobe importance in later stages. To our knowledge, this is the first CT-based study to investigate perturbation-driven region-wise interpretability in pediatric brain age estimation. The proposed framework offers a clinically meaningful approach for early developmental assessment.
Aspiration pneumonia has high incidence and mortality rates among the elderly. Continuous patient monitoring is essential for prevention, but it remains challenging. This study proposes a system to predict aspiration during sleep based on deep learning. The system captures facial images of sleeping subjects using infrared cameras. Time series feature values based on the facial action coding system are computed using the artificial intelligence library MediaPipe. The deep learning model composed by a convolutional neural network and a long short-term memory model detects aspiration during sleep from the feature values. We conducted an experiment on nine healthy subjects to obtain facial image data. We used these images to train the model and evaluate its prediction accuracy. The results presented that the model could predict two types, normal and pained expressions, with high accuracy. This work is expected to improve the safety of elderly people and reduce the workload of health and care professionals.
Accurate evaluation of lumbar loads stands as an essential requirement for both understanding lower back risks and injury prevention in occupational and clinical scenarios. This paper examines existing methods used to analyze lumbar burden while outlining their integration with modern human pose estimation algorithms. Precise measurements emerge from traditional methods like electromyography, finite element models, and wearable sensor systems yet their practical use faces challenges due to invasiveness, computational requirements, and placement sensitivity. Real-time human pose estimation has achieved advancements through methods like RTMPose, OpenPose and HRNet which offer non-invasive scalable solutions for real-time motion analysis. This paper reveals the advantages and shortcomings in both domains while identifying the value of integrating pose estimation frameworks with lumbar load analysis for better accuracy in real-world usage. The assessment reaches its conclusion by evaluating present roadblocks while suggesting future research agendas to combine motion tracking systems with biomechanical modeling frameworks.
Emotion recognition plays a key role in human-computer interaction(HCI) and intelligent systems. This study proposes a multimodal approach that combines facial expressions and speech information to improve classification accuracy. VGG16-based CNN and 1D-CNN are used for facial and speech recognition, respectively. A weighted late fusion strategy integrates both outputs to make the final prediction. Experiments using KDEF and SAVEE datasets demonstrate that the proposed method outperforms unimodal models, particularly for neutral, fear, and disgust emotions, while reducing gender bias. The results highlight the effectiveness of multimodal fusion in enhancing emotion recognition.