
With the increasing demand for automation in the patent domain, generating Chinese patent abstracts faces challenges such as weak alignment between patent figures and claims, as well as insufficient consistency in domain-specific terminology. This study proposes a cross-modal dual-channel fusion method for patent abstract generation, aimed at enhancing the accuracy of text-image understanding and abstract generation, thereby advancing patent automation technologies. This paper introduces a patent abstract automatic generation model, PCA-PAG, which features a dual-channel cross-modal fusion method. By extracting visual representations with a frozen BLIP-2 encoder, mapping the global visual vector into learnable visual prefixes, and incorporating text-vision cross-attention, the method achieves both global semantic guidance and fine-grained component alignment. It uses conditional cross-entropy and penalty mechanisms to optimise generation quality and is validated through various evaluation metrics. Compared with the strongest baseline, PCA-PAG improves BERT-F1 by 3.22 percentage points while maintaining a favourable performance-efficiency trade-off. It confirms that the dual-channel strategy offers a feasible and practically promising technological pathway for patent cross-modal abstract generation without significantly increasing inference costs.
Vision language models (VLMs) are multimodal, generative AI models that can understand and process modalities beyond text, including image, video and audio. Over the past several years, VLMs have remarkably advanced and are continuing to advance at a rapid pace. A large number of VLMs have been released, with a wide range of capabilities and architectures. The literature is vast with academic papers, technical blog posts, YouTube videos and news articles. The objective of this article is to provide a comprehensive guide to understanding essential aspects of VLMs, including the capabilities, use cases, architectures, training methods, benchmarks, challenges and R&D directions.
This paper presents a smart cloud-based framework for virtual museum exhibitions that integrates artificial intelligence, 3D printing, and accessibility technologies to enhance the personalisation and inclusivity of cultural heritage experiences. The system operates on a scalable cloud infrastructure, employing AI modules to adapt virtual tours to individual user preferences and accessibility needs. Interactive components such as speech recognition, tactile 3D models, and colourblind and child-friendly modes make the framework suitable for diverse audiences, including people with disabilities. Quantitative testing demonstrated stable system performance, with an average scene-generation time of less than 5 seconds and a streaming latency of less than 200 milliseconds. A pilot user study with 30 participants, including children and visually impaired users, indicated high satisfaction (mean score 4.6/5) and positive feedback on accessibility. The proposed framework contributes to the digital transformation of cultural heritage institutions by combining adaptive AI-driven interaction with tangible engagement through 3D printing. Future developments will focus on optimising performance, enhancing print quality, and expanding multimodal accessibility features.
In vehicular ad hoc networks and internet of vehicles, stress and frustration while driving can negatively impact safe driving. Thus, managing drivers' stress levels is crucial for improving safety. In this work, we introduce an intelligent system based on fuzzy logic (FL) to evaluate safe driving level in a vehicle edge computing (VEC) environment. We implement the proposed system in cascade considering three modules: driver anxiety level (DAL) module, driver mental status (DMS) module and safe driving evaluation level (SDEL) module. We carried out many simulations to evaluate the performance of each module for different parameters. Simulations results show that DAL is good for drivers aged between 30 and 50 years, but it tends to decrease for young drivers less than 30 or older drivers more than 50. While DMS is good when the driving time is shorter.
Cloud computing enables elastic resource provisioning, yet multi-cloud systems face challenges in policy heterogeneity and SLA inefficiency. Existing solutions inadequately resolve policy conflicts and SLA management in multi-cloud environments. This paper proposes the policy-driven SLA management in multi-cloud (PSM-MC) framework, integrating policy coordination, dynamic SLA negotiation with game-theoretic bargaining, and adaptive admission control using reinforcement learning. Key innovations include real-time policy conflict resolution via Pareto-optimal compromises, hybrid SLA monitoring combining proactive prediction with reactive adjustments, and Nash bargaining-based resource allocation to balance provider-user tradeoffs. Experimental results demonstrate PSM-MC achieves 15-30% higher virtual resource utilisation and 20% improved SLA compliance compared to existing approaches in large-scale multi-cloud testbeds.
In artificial intelligence (AI) services, a vast amount of data is amassed from various services and devices into data centres (DCs). Numerous users share the data by issuing transactions. Consequently, the electricity consumption of DCs increases by the proliferation of AI services. Hence, a control method to maintain data integrity and improve the throughput of transaction processing while reducing the electricity consumption of servers has to be realised for AI services. In this paper, a multi-version energy efficient role ordering (MVEERO) scheduler is newly proposed to maintain data integrity and improve the throughput of transaction processing while reducing the electricity consumption of servers. In evaluation, the execution time of transactions and the electricity consumption of a server cluster in the MVEERO scheduler are shown to be maximally reduced 31% and 13%, respectively, to the energy-efficient role ordering in virtual machine environment (EERO-VM) scheduler which is previously proposed in our studies.
The internet of things and its applications have been exponentially coming to the forefront over the past years. Waste management is one of the areas where IoT has a potential of making a significant difference to our everyday lives. While a variety of smart waste management solutions has already been studied and applied, majority of them are focuses on larger-capacity bins with lower frequency of collection. Contrariwise, this work concentrates on a lower-volume bins with daily frequency of collection - bins administered by Municipality of Bratislava, mainly located in the city centre. We study the current waste management situation, as well as potential costs and benefits of a particular smart waste management solution. A solution provided by a Bratislava-based provider of smart waste management solutions, is used as an example of possible IoT applications to the discussed bins. We conclude that given our assumptions, implementation of such a solution would be beneficial. Based on the results of our analysis, we also define limitations of the solution that may be overcome through adoption of recommended measures.
High-performance computing platforms accelerate rendering application execution by efficiently distributing workloads across clusters of computing hosts. Priority-based scheduling offers a simple and effective mechanism for computing resource allocation, often aligned with pay-per-use models. Traditional priority calculation methods often overlook inter-user parameters, such as competing user priorities and system scales. This paper presents an enhanced adaptive resource allocation strategy that introduces two normalisation techniques: priority scaling and weight sharing. By balancing fairness and responsiveness, the proposed method allows short jobs to complete earlier and avoids queue congestion, resulting in a more efficient and user-friendly environment for rendering workloads with diverse job types and priority levels. Experimental results show that this adaptive approach significantly reduces waiting times with marginal impact on the completion time of high-priority tasks.
The explosive growth of web services complicates developer selection. While service representation is key for intelligent management, existing methods rely on textual semantics or network structure alone, often neglecting deep multi-feature fusion and text imbalance or absence across nodes. This paper proposes a transformer model empowered by heterogeneous networks to unify context-aware text and heterogeneous structure encoding. Heterogeneous structure information is incorporated into each transformer layer to capture node/edge information, handling nodes with or without text. A fully-connected attention mechanism integrates representations from text-rich neighbours, textless neighbours, and the node's own content at each layer. To fully fuse features, a specialised transformation matrix projects different node types into a shared latent space. Experiments show our method outperforms the strongest baselines by nearly 1% in LogLoss and 2% in AUC.
The measurement of online service reputation based on ordinal preferences has been proposed to address the issue of unreliable reputation measurement results due to inconsistent user evaluation criteria. When users' complete ordinal preferences are unavailable, these methods ignore unknown preferences or use collaborative filtering to predict preferences without verifying the accuracy of preference prediction, leading to an untrustworthy service reputation. This study proposes an approach that models users' complete preferences using the conditional preference networks (CP-Nets) and then measures service reputation by aggregating CP-Nets. The approach designs an adaptive Tabu search algorithm to learn users' CP-Nets efficiently and aggregating all the CP-Nets using the ranked pairs method. The service reputation ranking is then deduced from the aggregated CP-Net. Experimental results on real datasets show that the proposed method is more efficient compared to existing methods, with more accurate preference prediction, and the reputation ranking is more consistent with user preferences.
Customer loyalty is closely related to the development of e-commerce platforms, but due to the non-contractual characteristics of e-commerce users and the poor performance of traditional customer data analysis, the phenomenon of customer churn is more prominent and obvious. Therefore, based on this, a combination prediction model is proposed to analyse customer data, which optimises indicator data on the basis of recency, frequency, and monetary models. By adding emotional feature indicators, customer rating indicators, and introducing an improved K-value clustering algorithm, the problem of customer churn prediction is analysed. Subsequently, principal component analysis is used to reduce the dimensionality of the data and combined with adaptive boosting algorithms to better ensure classification accuracy. The results show that the overall accuracy of the combined algorithm on the dataset is above 98%, significantly superior to other algorithms, and its recall results for non-churn customers are also above 97%, with a mean absolute percentage error of less than 2%. The specific stability and fitting are good, with the overall accuracy value and consistency coefficient basically below 0.02. This combination prediction model can effectively provide reference value for e-commerce operators to improve customer relationship management and reduce customer churn issues.
Effectively managing the vast data generated by sensor networks has become crucial with the rapid spread of IoT devices under strict resource constraints. This research introduces a framework that integrates explainable artificial intelligence (XAI), digital twins (DT), federated learning (FL), and multi-agent reinforcement learning (MARL) to optimise energy use and task distribution in distributed IoT environments. The proposed method, called explainable digital twin with federated multi-agent RL (XDT-FMARL), balances computational load through federated training and intelligent offloading between constrained IoT devices and edge servers. DT predicts short-term operational metrics such as battery state, processor load, and network delay using linear regression and moving averages. Guided by XAI, MARL agents select adaptive offloading or local processing strategies, enhancing interpretability and trust. Experiments show that XDT-FMARL maintains device batteries above 80% and applies responsive offloading under high load, while single-agent models default to uniform local processing with limited adaptability.
In recent years, air pollution emission statistics in Taiwan have indicated that mobile sources contribute the largest share of pollution in metropolitan areas. As a result, developing an intelligent, fully automated system for identifying high-emission diesel vehicles has become a critical necessity. The development of a high-pollution vehicle detection system utilises a one-stage architecture. A YOLOv4-based neural network module has been implemented to extend its application. The real-time image data of each vehicle is processed and optimally integrated through data augmentation, algorithm optimisation, and image enhancement. This enables the system to effectively identify high-pollution vehicles from complex, real-time traffic flow imagery and capture relevant vehicle information. Using these inputs, the system demonstrates enhanced adaptability to diverse environments. Furthermore, the system uses multiple image datasets for augmentation, incrementally improving accuracy. The simulation results indicate that the system achieves a detection accuracy that exceeds 91%, particularly for high-polluting diesel vehicles.
Continuous-time reinforcement learning (CTRL) and discrete-time reinforcement learning (DTRL) have been successfully applied to various tasks. However, CTRL has garnered less attention, with research primarily targeting algorithmic enhancements and giving limited consideration to model parameters. This neglect of model parameters hampers reproducibility and performance tuning. In this paper, we conduct a large-scale experimental analysis of hyperparameter tuning in CTRL based on the Hamilton-Jacobi DQN (HJDQN) algorithm. We aim to improve CTRL by minimising performance variations from irreproducibility and misunderstandings, while reducing waste of computational resources. Experimental results on four Mujoco tasks reveal that larger gamma values yield better performance, with optimal average rewards achieved at sampling intervals h of 0.2, 2.0 and 4.0, respectively. Additionally, experiments in the PandaGym environment demonstrate that the sampling interval's impact on performance aligns with its effect on the Q-value function, and smaller Lipschitz constraint constants facilitate agent learning.
With the rapid development of computing and networking, individual users are increasingly relying on cloud storage and computing services. To ensure the integrity of data in the cloud, many audit schemes have been proposed. Yang's scheme combines fuzzy identity-based signatures to achieve dynamic revocation and public audit, while Wang's scheme utilises chaotic systems and smart contracts to achieve non-interactive audit and fair payment. However, these schemes enhance functionality but ignore security. This paper conducts a systematic cryptographic analysis of the cloud audit schemes proposed by Wang and Yang. Firstly, Yang's scheme contains errors in algebraic usage, making it vulnerable to forgery attacks. We verify this through a specific attack scheme. Secondly, Wang's scheme is susceptible to data deletion attacks. We first conduct a detailed analysis of Wang's scheme, identify its security issues, and verify our conjecture by proposing a method of data deletion attack.
Recently, there are many ways to enable computer network communication such as cell phone networks (4G, 5G), WiMAX, WiFi, Bluetooth, WMNs, M2M, P2P and so on. Also, the relay stations can enhance communication coverage, but they are fixed, expensive and have a limited capacity to handle certain number of users or devices. In this paper, we propose a cloud-fog-edge platform for control of moving omnidirectional access point (MOAP) Robot and improvement of communication environment. The MOAP robot moves omnidirectionally and functions as an access point (AP). We estimated the optimal position of MOAP robots in public spaces such as train stations and airports. Furthermore, we conducted experiments to evaluate the position of MOAP robots.
This study uses virtual reality technology to create digital teaching materials for primary school students on marine ecology and environmental education. These materials enhance learning motivation while conveying information about endangered marine organisms. Through an immersive VR diving experience, students can observe the habits and threats faced by these creatures. The study includes four games that focus on ecological issues and pollution along Taiwan's east and west coasts, developed in collaboration with an elementary school in Danshui using the ADDIE instructional design model. A pre-test was conducted before introducing the VR system to evaluate progress in two classes. Results showed that the class using the VR system improved more in the post-test than the one receiving the traditional way. Although the improvement was not statistically significant, feedback from both students and teachers indicated a positive response, suggesting that this approach benefits students with weaker foundational knowledge.
Bipartite graphs have been widely applied in data mining to represent data relationships, such as in e-commerce recommendation systems. Graph neural networks (GNNs), with their powerful ability to process structured data and explore higher-order information, have become the state-of-the-art method for recommendation problems. Recommendation systems increasingly rely on graph structures to represent relationships between users and items, like user click behaviours and purchase records. Through graph convolutional networks (GCNs), these structures capture connections between users and items, integrating structural information (e.g., user-item links) with node features (e.g., user preferences and item attributes) for accurate recommendations. This study combines improved simplified swarm optimisation (iSSO) with bipartite graph convolutional networks and eye-tracking technology to explore user preference behaviour, called iSSO-BGCN. We construct a node-feature bipartite graph, using iSSO's optimisation capabilities and natural gradient descent to train the model. Trials validate its ability to deliver precise recommendations.
Research is progressing in semantic search, in which the meaning of a sentence is represented as a numerical vector called an embedding and sentences are searched for based on their similarity. The ability to process queries combined using logical operators is essential for semantic search. We developed an appropriate technique for processing queries containing logical operators with reference to fuzzy set operations. In this technique, AND and OR between queries take the minimum and maximum similarity of the search results, respectively. A NOT operation on a query subtracts the similarity to the query from 1 and scales the result to obtain the desired search results. We devised an example-based semantic search method that obtains search results that match the user's intention as closely as possible based on the positive and negative example sentences that should and should not be included in the search results, respectively, as specified by the user.