
Emerging technologies evolve dynamically in response to societal needs and global disruptions, reflecting the interplay between technological advancement and public perception. This study demonstrates how the identification and prioritization of emerging technologies are influenced by shifts in public attention as captured through Wikipedia activity during the COVID-19 pandemic. By combining Wikipedia statistics and social network analysis, the proposed hybrid framework enables the detection and evaluation of technologies from both social and technical perspectives. The empirical analysis of point-of-care testing and e-healthcare reveals contrasting driving factors-market-oriented vs. technology-oriented dynamics-illustrating how emerging technologies develop under a major global event. These findings provide evidence that open, data-driven approaches can support governments and affiliated organizations in identifying emerging technologies in a timely manner and managing them. Furthermore, the study makes methodological contributions to data-driven decision making, suggesting that future research should incorporate heterogeneous data sources from diverse domains to strengthen evidence-based predictive foresight.
Blockchain networks (BCNs) deployed in a cloud infrastructure support secure service transaction but may be compromised due to the rapid advancement of quantum computing. The proposed approach presents a zero-trust enabled, post-quantum blockchain framework to enhance security and trust in federated multi-domain cloud network services. A zero-trust approach performs continuous authentication of services and ensures the integrity of operational metrics. It leverages security by integrating a lattice-based post-quantum cryptography (PQC) algorithm that uses the N-th degree truncated polynomial ring method optimized with the Toom-Cook approach, highlighting security against quantum threats with computational efficiency. This study incorporated a PQC-verified inter-ledger smart contract to enable federated trust in a multi-ledger design, which ensures secure cross-ledger interoperability. The experimentation showed reduced average time-to-write of 16.4 seconds by Quorum with 821 polynomial degree and average instantiation time of 4.39 seconds for log retrieval in the network. The validation time recorded by the proposed framework is 57.89% less than that of the classical blockchain at high transaction loads. Quorum achieved the lowest per-call overhead with latency of 38 milliseconds and throughput of 787 transactions per second (TPS), outperforming Hyperledger Fabric and Ethereum. The proposed framework garnered an average empirical federated trust score of 91.3%.
This study explores the challenge of aspect-based sentiment analysis (ABSA), aiming to enhance the identification and interpretation of fine-grained sentiment information in textual data. ABSA holds practical value in e-commerce, social media, and market analysis. In this paper, we propose a composite deep learning model (CDLM) that integrates bidirectional long short-term memory networks (BiLSTM) and deep pyramid convolutional neural networks (DPCNN) to capture contextual relationships among words, improving feature representation and classification efficiency. To improve information retrieval precision, we introduce an attention mask matrix that emphasizes salient terms while suppressing irrelevant content. A self-attention mechanism is employed to capture global dependencies between different parts of the sentence, enhancing the model's ability to interpret complex syntactic structures. Finally, we incorporate a graph attention network with a dynamically updated adjacency matrix to adaptively refine the graph structure, thereby enhancing the extraction of semantic information from sentences and improving sentiment classification performance, particularly in cases involving multiple sentiment elements. Experiments demonstrate that our model achieves competitive results across three benchmark datasets. By integrating several advanced deep learning techniques, this study demonstrates potential for practical engineering applications, such as user sentiment feedback analysis. The source code is publicly available on GitHub (https://github.com/fortytwo1111/cdlm).
The expansion of the attack surface due to the proliferation of cloud environments has driven a rapid increase in advanced cyber threats. Recently, visualization-based deep learning malware detection systems that convert input into grayscale images have been researched to counter these threats. However, deep learning models are vulnerable to adversarial attacks that mislead malware classifiers into classifying the attacks as benign due to their sensitivity to minor perturbations. Such vulnerability poses a critical security threat in cloud environments where accurate malware detection is vital for protecting sensitive data. Existing purification methods suffer from over-purification, which damages input data features, causing information loss. To address this problem, this study proposes a generative adversarial network-based multi-teacher distilled purification (GAN-MDPuri) scheme that distills knowledge from teacher models. The GAN-MDPuri scheme employs the competitive learning structure of GANs to achieve balanced knowledge distillation from a purification teacher network that removes perturbations and a reconstruction teacher network that restores normal sample input. Unlike existing knowledge distillation approaches that require complex weight adjustment mechanisms, the proposed scheme dynamically balances distinct knowledge during GAN training. The student model trained using this approach removes perturbations while preserving critical input features. The validation of the GAN-MDPuri scheme demonstrates that it achieved a high average adversarial malware detection accuracy of 0.9588 on large-scale datasets. Furthermore, the proposed model detected adversarial malware, reducing the adversarial attack success rate by an average of 0.3297 compared with existing purification schemes.
Lip reading, or visual speech recognition, interprets spoken words by analyzing lip movements and facial cues. While visual-only and audio-only systems show promise, their performance degrades significantly in noisy or unconstrained real-world environments. This study proposes an enhanced multimodal lip reading framework combining visual and audio information to improve recognition accuracy and robustness. The visual stream employs a custom convolutional neural network (CNN) for spatial feature extraction from mouth regions, with the audio stream using Mel-frequency cepstral coefficients modeled with stacked long short-term memory (LSTM) layers. The extracted features are fused and processed through bidirectional LSTM layers to capture temporal dependencies. An evaluation on the lip reading in the wild (LRW) dataset demonstrates that our proposed audio-visual fusion model substantially outperforms unimodal baselines, achieving 87.3% accuracy with notable robustness under noisy conditions. Five-fold cross-validation confirms model reliability and generalization capability, with ablation studies validating the contribution of each component. Our proposed CNN-BiLSTM framework demonstrates the effectiveness of multimodal learning for enhanced lip reading performance, providing a computationally efficient solution that is practically deployable in assistive technologies and challenging acoustic environments on resource-constrained devices. While the current implementation focuses on isolated word recognition, this work lays the foundation for future continuous lip reading systems.
To address the nonlinear and time-delay issues in nitrogen oxides (NOx) emission prediction of coal-fired boiler selective catalytic reduction (SCR) systems, this study proposes a time-delay gated recurrent unit (TD-GRU) model integrating cross-correlation lag analysis and stratified KMeans sampling. Pearson correlation analysis is first used to identify key parameters that are highly correlated with NOx concentration. This is followed by cross-correlation analysis to determine the optimal lag of each feature, enabling precise time-delay feature reconstruction that captures the dynamic behavior of NOx emissions. A stratified KMeans clustering method based on NOx concentration intervals is then employed to select representative training samples, effectively reducing redundancy while preserving operational diversity. Finally, a GRU-based prediction model is developed and compared with the LSTM, DNN, and ELM models. Experimental results show that the proposed TD-GRU model achieves over 95% prediction accuracy and superior performance across all metrics, accurately capturing the dynamic evolution of SCR outlet NOx concentrations and providing a reliable approach for the intelligent control and optimization of SCR systems.
Spiking neural networks (SNNs), as biologically plausible models, have been applied in domains with temporal characteristics. To overcome the limitations of SNNs in capacity and representational power, deep SNNs have gradually become a research focus. However, as the scale of data and models rapidly expands, the computational resources required by deep learning algorithms have approached the limits of classical computers. Fortunately, with the ongoing development of quantum information science, quantum technology holds promise for overcoming the bottleneck of computational resources. To explore new solutions, we propose a hybrid quantum spiking residual network model (HQSResNet). Unlike conventional deep SNNs, HQSResNet integrates a spiking residual architecture with parameterized quantum circuits, enabling efficient classification by leveraging limited quantum resources. We evaluate the classification performance of HQSResNet on the Fashion-MNIST dataset under different quantum circuit parameter settings, and optimize the model to achieve high classification accuracy. In noise robustness tests conducted on MNIST, KMNIST, and Fashion-MNIST datasets, the model demonstrates superior classification performance under noise interference compared to the classical spiking ResNet, achieving an accuracy of 92.2% on the MNIST dataset with uniform noise of intensity 0.9. Experimental results indicate that with the incorporation of quantum techniques, HQSResNet exhibits enhanced robustness to noise.
This study addresses inefficiencies in multi-source heterogeneous edge data fusion, including memory-resident data fragmentation, redundancy accumulation, and single-thread processing bottlenecks. We propose a memory-mapping-based concurrent fusion architecture tailored for unmanned aerial vehicle (UAV) ground station systems. This framework enables direct memory-indexed parsing through mapped locations, supports atomic frame operations, and utilizes a thread pool-based executor that slides unidirectionally along temporal baselines to delegate batch tasks. By integrating architectural innovation, domain-specific optimization, and practical engineering considerations, the proposed method delivers a significant advancement in UAV link data processing. Experimental results and comparisons demonstrate that it achieves a better balance in heap memory utilization, CPU utilization, and throughput, leading to more reasonable memory usage, shorter task times, improved schedulability, and scalability. In practical engineering applications, the approach enhances the efficiency of information integration and analysis by up to 20 times, significantly reduces data redundancy, and eliminates redundant time alignment, thus achieving efficient multi-threaded parallel data fusion.
The Internet of Medical Things (IoMT) has continued to revolutionize the healthcare sector, resulting in reduced costs and improved quality of treatment. To facilitate real-time remote patient monitoring, biosensors frequently interact with medical systems over the public Internet. This connectivity, however, exposes the IoMT environment to a myriad of privacy and security threats. Although past research has proposed numerous security techniques to address these challenges, most of these solutions fail to provide an adequate balance between security and performance. In this paper, we propose a robust IoMT security scheme that leverages elliptic curve cryptography and one-way hashing operations to achieve low communication, energy, and computation overhead. Formal security analysis using a random oracle model demonstrates the robustness of the negotiated session keys. In addition, semantic security analysis confirms that our scheme mitigates various typical IoMT cyber threats, such as desynchronization, forgery, impersonation, and ephemeral secret leakage attacks. A comparative performance evaluation verifies that the proposed protocol incurs the lowest execution, energy, and communication costs. Specifically, it reduces computation and energy costs by 19.04%, increases supported functionalities by 69.23%, and lowers communication overheads by 8.7%. These efficiencies make our scheme ideal for deployment in IoMT devices with limited energy, processing and communication capabilities.
With the development of Internet of Things (IoT) technology, traditional human-computer interaction methods are gradually shifting toward multi-modal interaction. Users can now interact with devices more naturally and efficiently through multi-modal information such as speech, images, text, and video. Against this backdrop, the efficient retrieval of multi-modal data has become a critical issue that requires in-depth research and exploration. This paper introduces IoT-MIR-MMRS, a novel multi-modal information retrieval method specifically tailored for IoT environments. The proposed approach leverages a data aggregation algorithm that ranks and aggregates retrieval results from edge devices, ultimately delivering more accurate and comprehensive outcomes to users. Experiments on the FG-Xmedia dataset validate the effectiveness of IoT-MIR-MMRS, demonstrating that it significantly outperforms existing methods in both one-to-one cross-modal retrieval and one-to-many cross-modal retrieval tasks. The IoT-MIR-MMRS method is promising in terms of advancing the application of multi-modal data retrieval in IoT-based human-computer interaction, further improving interaction efficiency and user experience.
As smart cities continue to evolve, the exponential growth of sensor data and real-time processing demands poses significant computational challenges. In response, we combine quantum computing and artificial intelligence to propose a quantum convolutional neural network (QCNN) for image classification, with specific validation in intelligent transportation systems. Unlike traditional neural networks, our approach encodes image features into quantum states using parameterized rotation gates and leverages quantum entanglement through controlled operations to extract hierarchical features with dramatically reduced parameters. We comprehensively evaluate the proposed QCNN across three diverse benchmarks: FashionMNIST, GTSRB, and CIFAR-10. Experimental results demonstrate that QCNN achieves competitive or superior performance in structured, domain-specific tasks, notably achieving 90.87% accuracy with real-world GTSRB traffic signs while using 95.6% fewer parameters than classical baselines. This parameter efficiency directly addresses the resource constraints of edge devices in smart city IoT networks. However, comprehensive cross-dataset analysis reveals that quantum convolutional operations excel in tasks with clear geometric patterns but face challenges with highly complex, unstructured natural images, providing actionable guidance for practical deployment. These findings demonstrate that QCNN is most promising for specialized smart city applications-traffic sign recognition, medical imaging, and industrial inspection-where structured visual patterns dominate and computational efficiency is critical. This work bridges the quantum computing theory and practical urban infrastructure, offering a pathway toward compact, energy-efficient intelligent systems for next-generation smart cities.
The increasing deployment of Internet-of-Things devices amplifies the need to reconcile data-driven benefits with stronger privacy guarantees, especially where data are distributed across heterogeneous owners. Federated learning (FL) mitigates raw data exposure, but existing approaches struggle to satisfy per-user privacy preferences and maintain high global model utility under non-IID conditions simultaneously. To address this issue, this paper proposes FLDP-GT, a game-theoretic framework that formalizes the privacy-utility optimization for FL systems with non-IID data distributions. Unlike conventional uniform privacy budgeting, FLDP-GT achieves user-centric adaptation by embedding a three-layer optimization mechanism: clients dynamically select personalized differential privacy budgets through utility-aware best-response strategies; the server hierarchically aggregates models using gradient contribution evaluation to minimize global utility loss; and the Nash equilibrium ensures provable convergence to stable tradeoffs between individualized privacy preservation and collective model performance. An experimental evaluation on three datasets demonstrates that FLDP-GT preserves privacy more flexibly and attains higher aggregate utility than methods that apply a single, uniform privacy budget.
Early screening for diabetic retinopathy (DR) is essential for preventing visual impairment. However, conventional methods often struggle to extract fine-grained local details from complex and overlapping lesion regions. Accordingly, this study proposes a diagnostic approach that employs a Siamese network with multi-scale patch cross enhancement to effectively integrate regional lesion specifics with holistic contextual information. Based on the Siamese architecture, a two-stage, multi-scale feature extraction framework is developed, enabling semantic comprehension using only image-level annotations. Secondly, an iterative clustering algorithm is designed to screen high-discriminant local lesions and generate dynamic feature weights, so that the model can adapt to different pathological manifestations. Furthermore, the multi-scale cross enhancement strategy facilitates a synergistic integration of regional fine-grained information with holistic representations, thereby boosting scoring precision and consistency. Empirical assessments on the IDRiD and DDR datasets reveal that our approach surpasses current leading methods, achieving kappa and accuracy rates of 72.7% and 78.2%, respectively. These outcomes further corroborate the efficacy of the proposed framework for DR grading, particularly in managing and interpreting intricate lesions, and offer fresh perspectives for early DR screening and diagnosis.
The use of generative artificial intelligence (AI) in Korean language education is growing, yet beginner learners remain understudied. This study developed and evaluated a writing correction template designed for beginner level learners of Korean as a second language. The research was conducted in three stages. First, two prompt strategies-zero-shot chain-of-thought (CoT) and Few-shot-were compared in terms of accuracy, consistency, and learner-friendliness; the Few-shot method produced more stable and comprehensible feedback. Second, a prompt-based proficiency classification module was integrated using graded vocabulary and grammar lists from the National Institute of Korean Language. Third, 500 beginner texts were analyzed using the template and compared with human-annotated errors. The template achieved precision of 0.948 +/- 0.002, recall of 0.649 +/- 0.003, and F1-score of 0.771 +/- 0.003 across five repeated runs, demonstrating stable performance despite large language model (LLM) nondeterminism. False negatives frequently involved particle omission and written-style register errors, whereas false positives primarily resulted from fluency-oriented over corrections. Cohen's kappa (0.56) indicated moderate agreement with human annotation. These findings suggest that a carefully engineered template can partially approximate human judgment, support level-appropriate feedback, and offer a scalable foundation for AI-assisted writing correction for beginner learners of the Korean language.
Detecting Insider threats poses a major challenge due to the limitations of traditional security systems in identifying behavioral anomalies. This study investigated how stress-inducing factors-such as time pressure, workload, and task complexity-affected insider behavior and evaluated the effectiveness of non-invasive detection techniques. A meta-analysis of 30 empirical studies focused on behavioral changes under stress and on the performance of various detection methods including keystroke dynamics, heart rate variability, eye tracking, and electroencephalogram. The findings indicate that keystroke dynamics and heart rate variability are among the most accurate techniques, with highly stressed individuals being up to three times more likely to commit security violations. These results support the integration of stress-aware, non-invasive monitoring into organizational security systems. The study enhances human-centered cybersecurity by validating behavioral monitoring under stress and advocating multimodal approaches to real-time insider threat detection.
This study presents the multidimensional style anchoring-dynamic feedback calibration (MSA-DFC) algorithm, designed to mitigate style drift and enhance consistency in generative artificial intelligence (AIGC) outputs across diverse generation modalities and long-sequence tasks. The algorithm constructs a text-visual multidimensional style primitive model that integrates semantic mood vectors, syntactic structure matrices, color distributions, and composition rules. These features are mapped into a unified cross-modal representation space through a modality adaptation network. A domain-adaptive style anchoring mechanism improves style adaptation accuracy, while a dynamic feedback calibration module suppresses style drift during long-sequence generation. To evaluate the algorithm, a hybrid dataset combining expanded public resources and self-constructed multi-scenario data was developed, along with a dual-dimensional evaluation system incorporating objective performance and subjective experience metrics. Experimental results show that MSA-DFC improves style similarity by 29.4% and reduces the style drift rate to as low as 3.2% compared to baseline methods. The algorithm also achieves a user satisfaction score of 88.7 and reduces task completion time by 27.4%, with all improvements being statistically significant (p<0.001). The proposed method outperforms mainstream models such as StyleGAN3 and ChatGLM-6B (fine-tuned) across multiple scenarios. This work addresses the core challenge of style control in AIGC and establishes a quantitative correlation between style consistency and user experience, providing both technical and theoretical support for the practical deployment of AIGC systems.
Educational assessment is crucial to master the students' ranking accurately. Yet, conventional assessment methodologies often falter in reliability and efficiency. This paper introduces a pioneering blockchainenhanced scheme that revolutionizes educational assessment. Firstly, we proposed a hybrid blockchain-based network model, when the cross-institutional data access happens, the proposed model can fortify the veracity of assessment data. By instituting stringent joining criteria, our model curtails participation from nodes with low contributions, thereby enhancing overall network reliability. Furthermore, we introduced an innovative data storing and retrieving mechanism, significantly amplifying operational efficiency by categorizing blockchains into three types. Ultimately, a blockchain-based educational assessment system was developed based on the designs mentioned above, filling the blank of a practical and usable blockchain application in a real scene. The experiment results demonstrated the exceptional practicality of our system. This work not only bridges the gap in current educational assessment tools but also sets a new benchmark for future academic integrity and efficiency.
Rapid and reliable delivery of essential supplies is critical in the aftermath of natural disasters. This study addresses the humanitarian vehicle routing problem with drone, a dynamic optimization challenge in humanitarian logistics. We propose a multi-objective mixed-integer linear programming model that simultaneously minimizes the total delivery completion time, cumulative routing risk, and number of unserved high-priority locations. To solve this problem efficiently, we develop the improved non-dominated sorting genetic algorithm III (INSGA-III), which contains a multi-objective balancing module that dynamically adjusts optimization priorities in response to changes to the disaster scenario, such as the emergence of new demand locations (e.g., temporary shelters and makeshift hospitals) and unexpected disruptions (e.g., road blockages). The study conducts extensive experiments on diverse benchmark instances that indicate that the INSGA-III outperforms the NSGA-II and the standard NSGA-III in terms of convergence and solution quality. Specifically, it achieves higher average hypervolume (of up to 95.17%) and reduces the average total travel time to 1,387.03 units, compared with 1,439.11 units for the NSGA-III. Lower standard deviations across performance metrics further validate the robustness of the proposed approach. These findings underscore the INSGA-III's potential to enhance decision-making and operational efficiency in humanitarian logistics.
Brain-computer interfaces (BCI) leverage neurophysiological features like electroencephalography (EEG) to enhance user-computer interaction. EEG's high temporal resolution and unobtrusiveness make it ideal for BCI systems, facilitating real-time interaction with minimal latency. Prior studies employed EEG across domains, including health and security. Researchers expanded EEG data within natural language processing and information retrieval (IR), facilitated by machine learning methods. While the primary intention of using EEG has been to enhance the search experience, the nature of EEG data collection may capture sensitive information about the subjects, such as their identities, beyond the intended use. This raises ethical and privacy concerns, which have yet to be addressed. This work explores the detection of participants' identities from their EEG. A case study was formulated, incorporating EEG data from 40 participants engaging in the single-session simultaneous capacity (SIMKAP) experiment. Each participant's EEG recordings were utilized to perform subject classification. Results show that deep learning models identify 40 individual subjects that can be revealed with an accuracy of up to 58.8% (SD 5%). Based on these findings, the benefits of EEG within IR, ethical dilemmas presented by its use, and potential solutions must be addressed before wider development and implementation of further EEG-IR systems.
This study leverages high-quality data from soccer matches to derive a better understanding of the elements contributing to a team's success. Initially, we classically analyzed extensive soccer logs from major European leagues and international tournaments. We have defined a team's technical performance through a vector of features, including goalkeeping, intercepts, tackles, dribbles, and more. The objective of the paper is to classify the state of a match with the labels of win, defeat, and draw. We have not made predictions about any future outcomes, but did focus on understanding the characteristics of the data itself to identify patterns and trends of the match. In doing so, we guessed the match result only after having collected the above feature data for the entire match duration. Thus, our scenario poses a classification problem. We compare different models (SVM, Logit, XGB, and MLP), the last one outperforming the others. Moreover, as a brand-new approach, we analyzed the logs by considering matches of different increasing duration. In particular, the lengths of the matches were the terms of an arithmetic series with a common difference of 5 minutes. In doing so, we have provided a dynamic approach that labels the match outcome every 5 minutes, using an MLP to track the accuracy of the state over the time. The findings have revealed an improved detection of draws, and highlighted that the model accuracy is higher in the early stages but decreases as the match progresses. In both approaches, explainable AI techniques have identified the key predictive features, offering insights into how technical features influence success dynamically throughout a match.