
Multimodal learning that integrates Color Fundus Photography (CFP) and Optical Coherence Tomography (OCT) can enhance retinal disease recognition by combining surface appearance with depth-resolved structural cues. However, the limited availability of synchronized 2D-3D CFP-OCT datasets hinders the deployment of multimodal AI in real-world screening settings where OCT devices are often unavailable. This study investigates whether synthetic 3D OCT volumes can serve as a practical auxiliary modality to support multimodal classification in fundus-only scenarios. We synthesize 3D OCT from fundus images using three generative paradigms-a 3D GAN (DCGAN/WGAN-GP style), a 3D VAEGAN, and a fundus-conditioned latent diffusion model (3D-DDPM with classifier-free guidance)-and integrate the generated volumes into an EyeMoST- based uncertainty-aware fusion framework. We evaluate the pipeline on RFMiD and ODIR-5K across multiple backbone configurations. Empirically, diffusion-generated OCT yields the most stable fusion performance. In contrast, synthetic OCT does not consistently outperform strong fundus-only models due to domain gap and imperfect biomarker preservation.
Edge-vision systems on resource-constrained platforms require low-latency and energy-efficient front-end processing. Conventional gradient-based edge operators continue to rely on multiply-accumulate (MAC) operations, which can increase logic utilization, switching activity, and power consumption in eld-programmable gate array (FPGA) implementations. This study introduces a kernel-aware arithmetic-selection framework, supported by synthesized evidence, for energy-efficient FPGA-based edge detection. The framework combines XNORpopcount-based arithmetic mapping with selective approximate computation. Its central design principle is to replace MAC operations with exclusive-NOR-population count (XNORpopcount) when the kernel structure is suitable for binary-friendly arithmetic, and to apply approximation only to the remaining adder-dominated stages when a complete XNOR-popcount mapping is not practical. Under this rule, Prewitt-like operators are mapped to an XNORpopcount datapath over their active nonzero taps. In contrast, Sobel-like operators are realized using a hybrid datapath that combines binary matching, shift-add processing, and approximate accumulation. The resulting framework shows that multiplier removal is the dominant source of hardware savings, while approximate arithmetic provides a controlled secondary optimization. Overall, the proposed approach establishes a structured design methodology for low-power FPGA edge-detection architectures on embedded platforms.
An intelligent honeypot system designed to mimic legitimate websites using Sequence-to-Sequence (Seq2Seq) learning and Deep Q-Learning. The system generates realistic, contextually appropriate responses to attacker queries, prolonging interactions and providing insights into malicious behaviors while safeguarding actual systems. The Seq2Seq model, trained on HTTP request-response pairs, enables the honeypot to produce responses that closely resemble those of real servers, enhancing its ability to deceive attackers. Deep Q-Learning optimizes engagement by selecting the most effective responses through a custom reward function, balancing realism and interactivity to maximize session length. Performance was evaluated using metrics such as Response Realism Rate (RRR), Semantic Consistency Accuracy (SCA), and Average Session Length (ASL). The honeypot achieved an RRR of 92.3%, an SCA of 89.7%, and a 94.5% Optimal Response Selection Rate (ORSR). These advancements increased ASL by 143.5%, from 3.2 to 7.8 exchanges, reflecting prolonged attacker engagement. By integrating Seq2Seq and Deep Q-Learning, this honeypot demonstrates significant improvements in generating convincing responses and sustaining interactions. These results contribute to modern cybersecurity by providing a practical and theoretical framework for developing next-generation honeypots capable of deceiving attackers and gathering actionable intelligence.
The subjective and unstructured nature of coffee tasting notes creates a significant data annotation bottleneck, limiting the application of computational methods in sensory science. This study introduces a novel end-to-end framework for automating multi-label coffee flavor classification, integrating LLM-driven annotation with transformer-based classification and local interpretability analysis. Google's Gemini-1.5-Flash was employed to perform evidence-based annotation on 8,327 expert evaluations authored by certified Q Graders, generating a high-quality training dataset across 17 flavor categories. Critically, the Q Grader provenance of the source texts enables a reverse-mapping validation framework: statistically significant and directionally coherent correlations between LLM-derived labels and expert quantitative scores (Floral r = 0.32, Roasted r = −0.25, all p <0.001) provide implicit ground truth without requiring separate annotation effort. Reliability analysis revealed substantial consistency for concrete physical descriptors (mean κ = 0.68) but notably lower agreement for abstract sensory concepts such as Mouthfeel (κ = 0.10), identifying two distinct reliability regimes that define the framework's operational boundaries. A ne-tuned BERT model trained on these annotations outperformed a TF- IDF baseline, achieving a Micro F1-score of 0.9164 versus 0.8763 and a Hamming Loss of 0.0640 versus 0.0969. To address severe class imbalance (up to 125×), Focal Loss (γ = 2.0) combined with per-class threshold optimization successfully recovered detection of rare defect categories, improving Macro-F1 from 0.656 to 0.719. Local interpretability analysis via LIME further confirmed that model predictions align with domain-expert sensory reasoning. These results demonstrate that an LLM-driven annotation pipeline offers a scalable, transparent, and effective solution to the data bottleneck in sensory science, establishing a robust methodological foundation for interpretable classification across other sensory-driven domains.
Online Social Networks (OSNs) have emerged as a major source of digital evidence in cybercrime investigations, abuse detection, and online incident analysis. Publicly available data, such as posts, comments, reactions, and user interactions, provide critical insights into suspicious activities and behavioral patterns. However, extracting and analyzing such data in a forensic context remains challenging. Existing social media data acquisition approaches primarily rely on web scraping or API-based techniques that produce unstructured outputs (e.g., raw text and images) without preserving the relationships among entities such as users, posts, and interactions. This results in the loss of contextual information essential for forensic analysis. Furthermore, data collection and analysis are often performed separately, resulting in delayed investigations and limited real-time insight. To address these limitations, this paper focuses on two key forensic requirements: (i) efficient and structured acquisition of social media evidence and (ii) correlation-aware analysis of interactions. We propose FEDIS, a unified forensic data acquisition and analysis system that integrates a hybrid DOM-TAO data model, graph-based representation, and parallel keyword-based search. Experimental results demonstrate that FEDIS achieves complete relationship preservation, improves data collection completeness, and significantly reduces search latency compared to traditional approaches, making it practical for real-world social media forensic investigations.
This paper investigates transfer learning for developing a multi-speaker Thai text-to-speech (TTS) system under low-resource conditions, with a focus on cross-lingual knowledge transfer from source languages with different phonological and prosodic characteristics. The model is pre-trained on speech data from Thai, Mandarin Chinese, and English, and subsequently fine-tuned using Thai speech from three target speakers, each with only 1530 minutes of data, covering both male and female speakers. Both objective and subjective evaluations consistently demonstrate that Thai pre-training achieves the best overall performance. Among the cross-lingual models, transfer learning from Mandarin Chinese outperforms transfer from English, yielding lower F0 RMSE (20.79 vs. 21.65) and higher MOS scores (3.55 vs. 3.44). In addition, speaker-dependent analysis indicates that speaker gender has a noticeable influence on synthesis quality under limited data conditions, suggesting that acoustic similarity between pre-training data and target speakers can affect the effectiveness of knowledge transfer. However, this factor is secondary to linguistic and prosodic similarity. Natural speech achieves a MOS of 4.93, while the Thai, Mandarin Chinese, and English-pre-trained models obtain MOS scores of 3.65, 3.55, and 3.44, respectively. These results highlight the importance of linguistic proximity in cross-lingual multi-speaker TTS, particularly tonal and prosodic similarity between source and target languages. Overall, the study confirms that transfer learning is an effective approach for low-resource Thai TTS and that tonal source languages provide more beneficial knowledge transfer than non-tonal, accent-based languages.
Precise quantification of foot morphology is critical for clinical diagnostics, biomechanics, and personalized orthotic design. Conventional anthropometric methods often remain labor-intensive or dependent on specialized hardware, necessitating more efficient predictive frameworks. This study develops and validates a series of machine learning (ML) models to predict nine essential foot anthropometric parameters. Leveraging a dataset of 544 independent foot samples from 272 participants, encompassing high, normal, and flat arch types. The models were evaluated using a 10-fold cross-validation strategy to ensure robust generalizability. Our pipeline integrates correlation-based feature selection with hyperparameter-optimized regression algorithms, including XGBoost, Random Forest, Support Vector Regressors, Neural Networks, and Linear Regression. The results demonstrate high predictive fidelity, with Mean Absolute Errors (MAE) consistently remaining below 0.5 cm. This level of precision meets the 0.5 cm clinical tolerance threshold established through expert consultation for in-sole production, while simultaneously aligning with international footwear sizing increments, thereby confirming the framework's practical utility in real-world manufacturing. Although parameters such as the length from heel to midfoot area and the length from heel to distal metatarsal head achieved exceptional precision (MAE of 0.012 cm and 0.026 cm, respectively), predicting arch height remains a notable challenge. This research underscores the necessity of optimal feature engineering and algorithm selection in automating foot morphometric assessment.
The rapid growth of the restaurant industry in Thailand has intensified the importance of online reviews, which significantly shape customer perceptions and influence business performance. Sentiment analysis has emerged as an effective computational approach for extracting customer opinions from such reviews; however, multi-class sentiment classification in Thai remains challenging due to the language's non-segmented structure and the issue of class imbalance. This study investigates three hybrid deep learning modelsWangchanBERTa-MLP, WangchanBERTa- CNN, and WangchanBERTa-BiLSTMby integrating WangchanBERTa, a Thai-specific pre-trained language model, with different neural architectures. Using a balanced dataset of restaurant reviews obtained through SMOTE, the models were evaluated based on accuracy, precision, recall, and F1-score. The experimental results show that WangchanBERTa- BiLSTM performed the best overall, achieving an accuracy of 85.22% and significantly improving the classification of neutral and positive sentiments compared to the BERT-based models and other hybrid methods.
High Efficiency Video Coding (HEVC) and its successors, such as Versatile Video Coding (VVC), offer substantial bitrate reductions, yet challenges remain in preserving visual fidelity under bandwidth and computational constraints. This paper proposes a deep learning-based super-resolution (SR) framework that operates natively in the YUV color space, eliminating costly RGB-YUV conversions and integrating seamlessly with modern video compression pipelines. We develop two convolutional network architectures trained on YUV-formatted video data: a full 3-channel model and a lightweight two-stream variant that separately processes luminance (Y) and chrominance (UV) channels using compact subnetworks. The proposed method enhances both full-frame and region-of-interest (ROI) quality, outperforming conventional HEVC baselines in terms of rate-distortion efficiency. Evaluations on diverse video sequences demonstrate significant bitrate savings and effective ROI preservation, with the lightweight model offering a practical solution for AI-driven applications in resource-constrained environments.
Efficient log reduction is critical for Security Operations Centers (SOCs) and Managed Security Service Providers (MSSPs), which must store, analyze, and retain massive volumes of event data while satisfying compliance requirements and controlling operational costs. Traditional pipelines often retain redundant or low-value records, leading to excessive storage overhead and slower analytics. This study evaluates five Vector-based log-reduction methods: lter-based selection, eld pruning, event sampling, template hashing, and a combined pruning + sampling profile. The evaluation uses more than 3 million log records from two well-known public intrusion datasets, CIC-IDS2017 and UNSW-NB15, to measure efficiency, throughput, and attack coverage under the same experimental setup. Compared with a baseline Filebeat pipeline, the proposed Vector-based approach improved throughput by 45%, reduced outbound traffic by 80%, and maintained 98% attack coverage. The results show that a substantial proportion of raw logs is redundant and can be trimmed without compromising essential evidence or analytic clarity. Template hashing preserved fidelity with moderate CPU cost; although it required slightly more processing than filtering or pruning, it still consumed fewer resources than the baseline. We repeated each test three times to ensure consistent results and validated the findings through ClickHouse queries at the sink layer. We also release the scripts and benchmark data to support reproduction and extension. Overall, the benchmark demonstrates how log-reduction design can improve operational efficiency while preserving analytic fidelity.
Neural Image Assessment (NIMA) has become a widely adopted approach for blind image quality assessment (BIQA), yet it remains sensitive to simple spatial transformations such as horizontal ips. Such variation can lead to inconsistent predictions, even when the perceived visual content remains largely unchanged. To address this issue, we introduce Flip-Robust Neural Image Assessment (FR-NIMA), an enhanced training strategy that enhances the spatial robustness of BIQA models. Instead of modifying network architectures, FR-NIMA incorporates a flip-consistency regularization term that penalizes discrepancies between the predicted quality distributions of an image and its horizontally flipped counterpart. Two variants-one-branch and two-branch formulations-are explored, both introducing no additional model parameters. FR-NIMA is evaluated across four CNN backbones (MobileNetV2, VGG19, Xception, InceptionV3) and one Vision Transformer (ViT-Small) using the LIVE dataset and two additional test sets representing distinct scene types. Performance is assessed using complementary metrics, including the Test Loss (EMD2), Absolute Flip Gap (|FlipGap|), Flip-Consistency Win Rate (FCWR), Average Flip- Gap Delta (AFGD), and Average Flip-Gap Ratio (AFGR). Experimental results demonstrate that FR-NIMA effectively reduces ip-gap magnitude and variability while maintaining comparable test accuracy across all back- bones. These findings establish FR-NIMA as a simple yet effective framework for enhancing the stability, spatial consistency, and trustworthiness of deep IQA models.
Multimodal Question Answering (MQA) in large language models (LLMs) requires adaptive modeling of modality relevance across conversational turns. However, existing approaches rely on static fusion strategies that treat modalities uniformly and fail to capture dynamic modality importance. To address this limitation, we propose DynaRoute, a dynamic memory routing framework for LLM-based MQA. DynaRoute integrates a Bi-LSTM-based conversational memory to model evolving dialogue context and a query-conditioned routing mechanism that dynamically assigns modality relevance at each interaction step. The resulting representations are processed by an LLM-based decoder to generate context-aware responses. Experiments on four benchmarks-VQA-v2, GQA, VisDial, and A-OKVQA-demonstrate consistent improvements over unimodal, static fusion, and mixture-of-experts baselines. DynaRoute achieves an improvement of up to 8.7% under noisy conditions and 6.5% under clean settings, while also obtaining the highest multi-turn consistency (68.9) and robustness (81.3) scores. These results highlight the effectiveness of memory-aware dynamic routing and establish DynaRoute as a principled framework for conversational multimodal question answering.
Long-Term Evolution (LTE) provides low-latency, high-data-rate services, which are essential for delay-sensitive applications such as video streaming and online gaming. Despite this, user mobility among cells can degrade network performance, so efficient handover management is crucial to maintain Quality of Service (QoS). Traditional handover mechanisms use static control parameters, such as hysteresis margin and time-to-trigger, that are not flexible for working with users' dynamic mobility or a range of user trajectories. In this paper, we present a learning-based optimised data-driven approach for LTE handover decision support. An XGBoost model trained with Hyperopt to learn the relationship between user movement angle and handover performance parameters. Interpretable if-then rules are developed to modify the handover control parameters adaptively. Experimental results further show that the performance of the fixed-parameter solutions depends on the maximum handover delay and the mean time to handover, including the minimum handover rate, indicating that a single configuration is unlikely to provide the best performance across all mobility scenarios. The solution offers an efficient, scalable, and interpretable decision-support system to improve LTE handover efficiency in dynamic wireless networks.
Thai medicinal plants are essential to traditional healthcare and local livelihoods. However, many Thai medicinal plants have similar morphological characteristics such as shape, colour, and texture. This problem leads to misidentification and misclassification. Image classifiers utilizing convolutional neural networks (CNNs), which are a class of deep learning models, provide a scalable substitute for manual classification. This study aims to evaluate and compare the performance of three CNN architectures (DenseNet-121, EfficientNet-B3, and MobileNetV2) for classifying 10 species of Thai medicinal plants. The dataset comprises 5,000 leaf images representing 10 species (500 images per species). This study partitioned the dataset into 80% training set and a 20% test set. To enhance model generalization, we applied data augmentation techniques-specifically rotation, flipping, and colour manipulation. Furthermore, we utilized TensorFlow and Keras on Google Colab with GPU acceleration to train the models. Evaluation metrics include accuracy, precision, recall, F1 score, model size, inference time, and CPU utilization. The results highlight a trade-off between accuracy and efficiency: DenseNet-121 achieved the highest accuracy at 96.0% and a Matthews Correlation Coefficient (MCC) of 0.9558. Statistical analysis confirmed that DenseNet-121 significantly outperformed the other architectures (p < 0.05), albeit with a higher inference time (579.22 s). Notably, EfficientNet-B3 and MobileNetV2 both achieved an accuracy of 93.4%, with MobileNetV2 performing the best in terms of model size (11.07 MB) and inference time (3.86 s). In conclusion, DenseNet-121 is the most accurate model, while MobileNetV2 is best suited for real-time applications due to its lightweight and rapid inference time. EfficientNet-B3 offers an optimal balance between accuracy and computational efficiency.
With the advancement of intelligent transportation systems and large-scale urban video surveillance technologies, vehicle image retrieval based on textual descriptions has become increasingly important. Although MCANet exhibits effectiveness in multi-scale feature alignment, significant limitations persist in accurate fine-grained semantic matching. To address this, we present VhAR-Neta modular cross-modal retrieval architecture that enables independent design and flexible combination of feature interaction mechanisms, enhancement strategies, and supervisory signals. This framework employs ResNet-50 as the visual representation extractor and BERT as the textual semantic encoder, and innovatively introduces a cross-modal attention unit that establishes explicit associations between image representations and linguistic descriptions. Concurrently, the system establishes small-object aware auxiliary supervision through binary classification tasks targeting discriminative fine-grained vocabulary, thereby directing the network toward distinctive microscopic semantic units. Empirical evaluations on the public benchmark dataset T2I-VeRi indicate that the optimal configuration achieves 75% Top-1 retrieval accuracy, with Top-10 recall covering 85% of relevant samples. After introducing the feature enhancement module and small-object supervision mechanism, the cumulative matching rate for Top-5 and Top-10 both reached 85%, demonstrating improved retrieval robustness.
Carbon monoxide (CO) is a harmful gas from incomplete fuel combustion, often found in motor vehicle emissions. Prolonged exposure can cause serious health issues or death. While existing Internet-of-Things (IoT) systems monitor CO levels, most lack predictive capability. One prior study used an Artificial Neural Network with limited accuracy (79%). To address this, a new IoT-based CO prediction model is proposed using a Long Short-Term Memory (LSTM) algorithm. The model predicts future CO concentrations based on seasonal patterns, empowering users to anticipate and proactively respond to potential exposure. By leveraging Edge-to-Cloud architecture, this approach enables low-power edge devices to send data to the cloud for accurate forecasting without local model deployment. Based on the evaluation, the model achieved 98.42% accuracy, outperforming previous approaches by 19.42%. It also showed superior performance against other algorithms, with the lowest MAE (0.026305), MSE (0.016004), RMSE (0.126506), and the highest R² (0.997647). Evaluation with AIC and BIC confirmed its reliability, scoring zero after MinMax scaling. The model demonstrates a substantial advancement in predictive CO monitoring, giving users actionable insights to protect health and safety.
An ontology is a widely used knowledge base for representing domain knowledge. Developing a knowledge-representing ontology is difficult, as it requires both domain and engineering expertise. Yet, such ontologies are essential for enabling intelligent systems to comprehend real-world knowledge through structured concept networks. In the Thai context, ontology research remains limited due to the scarcity of structured resources, standardized schemas, and annotated corpora for automatic knowledge extraction. This study addresses this gap by proposing a pattern-based methodology for ontology generation and instance extraction from Thai semi-structured medicine data, providing an alternative to resource-intensive deep-learning methods. The proposed approach identifies patterns of collocated Thai text and builds a collocation tree of word sequences, in which shared sequences represent ontological properties and variable sequences represent instance values. The method was applied to two complementary Thai medicine datasets, namely I-Med (a hospital dispensing-record database) and Pobpad (a public health-information website), to generate and integrate ontology components. These templates were transformed into ontological properties and converted into RDF/OWL format to produce a standard ontology usable for querying and reasoning. The generated ontology achieved high performance (Precision = 0.97, Recall = 0.90, F1 = 0.91) and received favorable assessments from domain experts. The results indicate that the proposed approach can effectively extract structured knowledge from Thai semi-structured text and produce a reliable ontology suitable for medical knowledge representation, providing a data-driven foundation for future Thai intelligent systems.
Pedestrian Attribute Recognition (PAR) is an important component of intelligent surveillance systems. Ground-to-aerial cross-domain PAR and the effect of UAV flight conditions remain largely unexplored. This work investigates whether PAR models trained on ground-level CCTV datasets can be applied to UAV imagery and quantitatively analyzes the impact of UAV elevation angle and horizontal distance on attribute recognition performance. Five CNN models are trained on two CCTV datasets with different characteristics: UPAR, a large and diverse dataset, and TPAD, a homogeneous dataset, using a multi-label classification framework with positive class weighting to handle class imbalance. Cross-dataset evaluation on CCTV data leads to the selection of RegNet and ConvNeXt. For ground- to-aerial evaluation, selected models are evaluated on three UAV datasets: UAV-Human, AG-VPReID, and a self-collected UAV-PT1 dataset, where UPAR-trained models achieve a mean attribute accuracy of 62.23-65.48%, while TPAD-trained models perform worse. RegNet achieves comparable performance to ConvNeXt with significantly lower computational complexity, making it more suitable for UAV deployment. Attribute-level analysis shows that UpperBodyLength, LowerBodyLength, LowerBodyColor, and Backpack are more reliably recognized. Further analysis using UAV-PT1 shows increasing the horizontal distance from 25 m to 50 m reduces accuracy by 11.3712.11%, and a high elevation angle of 50◦causes a significant performance drop, providing an evaluation of ground-to-aerial PAR and the impact of UAV flight parameters on attribute recognition.
This study proposes a framework for a geospatial internet speed test platform, specifically designed to assess the broadband network experience in Thailand. The novelty of this research lies in its comprehensive approach to broadband internet quality assessment, uniquely measuring internet performance from ISPs sources to end-users using custom-designed reTerminal hardware devices based on Raspberry Pi 4. Tests were conducted involving four major broadband service providers. Data were collected from 30 speed test devices, which were installed at the source of each service provider network, and 200 internet user devices. The total number of data records collected was 25,000. The results indicated that the Download Percentage Average was 64.30%, while the Upload Percentage Average was 71.84%. The average latency and jitter were 11.82 ms and 16.53 ms, respectively. All speed test parameters at the source of the ISPs network were found to reach almost 100% compared to the speed test network devices of the ISPs. This means that ISPs provide internet quality according to the standards set by the NBTC. These results can assist the NBTC for setting relevant policies or strategies to improve the quality of broadband internet services at user's sites.
Malnutrition is a serious condition caused by nutrient deficiency that poses a high risk to toddler growth and development, potentially leading to long-term health problems or even death if left untreated. Early detection of malnutrition symptoms is crucial to enable prompt and appropriate medical interventions. This study aims to develop an expert system capable of diagnosing malnutrition diseases quickly, accurately, and efficiently, particularly as a knowledge-based decision support tool in toddler healthcare. The method used is Case Based Reasoning (CBR), which applies experiences from previous cases to solve new ones. The system processes data consisting of 22 symptoms and 8 types of malnutrition diseases, supported by a database of 22 real cases. Each symptom is associated with the likelihood of a disease based on its similarity to previous cases. Performance evaluation results show an accuracy of 80% and a sensitivity of 85.7%, indicating that the system is fairly reliable in recognizing positive cases (REUSE) and providing appropriate diagnoses. In conclusion, the CBR- based expert system can serve as an effective diagnostic aid for medical personnel in quickly identifying malnutrition in toddlers, thereby supporting more efficient and targeted decision-making.