
In order for autonomous mobile robots to play an active role in our daily lives, the robots must be able to recognize an unknown 3D environment in real time. To address this challenge, we introduce an innovative method for object tracking using 3D point clouds as input. Our approach leverages the Growing Neural Gas (GNG) algorithm for real-time environment recognition in unfamiliar settings, coupled with object detection using Axis-Aligned Bounding Box (AABB), and Intersection over Union (IoU) for precise tracking. Our method achieves real-time object tracking in an unknown environment, processing in approximately 33 milliseconds. The proposed method showcases performance with an average IoU value of around 0.77, which was sufficient for object tracking. Finally, we validate the efficacy of this method by conducting two experiments in both a simulated environment using a depth camera and a real environment using a 3D-Lidar.
Meta-heuristic algorithms require a large amount of computational memory. The use of probabilistic models has been suggested as an alternative to the compact method of particles. In this study, a novel lightweight multi-strategy particle swarm algorithm is proposed. The use of a new probabilistic model is proposed to replace the commonly used normal distribution. The proposed lightweight algorithm is evaluated on the CEC2014 test set with excellent results.
To detect the suspect poisoned data in the training phase, most backdoor defenses rely on a prevalent assumption, i.e., the feature separability between poisoned and benign samples. However, this assumption can be bypassed by novel adaptive attacks, which merge the features of poisoned and benign samples. In this paper, we contrast these adaptive attacks and propose a so-called Local-Feature-Powered Defense (LFPD), which leverages a local feature algorithm to measure samples' similarity in the image space and uses it to guide the training process to increase the feature sepa-rability between poisoned and benign samples. Then, our LFPD detects the outliers in the training dataset as poisoned samples and removes the backdoor by unlearning them. Finally, we compare our LFPD with five existing defenses, and our experimental results demonstrate that LFPD outperforms them in defending against adaptive attacks.
In-silico toxicity prediction plays a key role in the health industry and research. Machine learning has been widely rewarded for its high efficiency in this field, but the methods are mostly expertise driven and with limited growth. After mitigating many practical problems smartly, deep learning has also been entrusted with high expectations in toxicity prediction. In this work, we attempted to predict compound toxicity by leveraging a graph representation of the compounds and a neat graph-learning framework. Key atomic and bond features were captured by the graph representations, and different graph-learning techniques in the framework were investigated. As evaluated on the Tox21 data, the graph-learning framework possesses a good potential in attaining state-of-the-art performances and handling imbalanced data in toxicity prediction tasks.
The security of academic credentials is increasingly at risk due to cyberattacks and credential fraud. Traditional verification systems rely on centralised databases, creating single points of failure and privacy concerns. This paper explores Zero-Knowledge Proofs (ZKPs) with blockchain to enhance educational data security. We propose an optimised ZKP protocol for education, improving lightweight infrastructure, scalability, credential revocation, and ease of use. By refining existing techniques, our approach enhances secure, privacy-preserving credential verification, creating a resilient educational data system.
Registration Data Access Protocol (RDAP) is designed to replace the classic WHOIS protocol by providing Representational State Transfer (RESTful) web services and tiered access control, enabling more secure and controlled data access based on the client's authorization level. Despite these improvements, the de-centralized management of registration metadata by various internet registries and registrars presents a challenge for users needing to obtain login credentials from multiple RDAP servers. This study examines the federated authentication mechanism proposed by RFC 9560 [8], identifying vulnerabilities within the current WHOIS/RDAP architecture and proposing countermeasures to enhance security and efficiency in the domain registration ecosystem.
This paper aims to forecast short-term solar power generation one hour ahead to enable precise power dispatching. Typically, time-series generation data is input into deep learning networks such as Long Short-Term Memory (LSTM), Convolutional Neural Networks (CNN), and Deep Neural Networks (DNN) for training. In this study, weather features are preprocessed and input into deep learning models alongside the generation data. First, the historical data's weather features and their correlation with future power generation are analyzed to select seven key features: power generation, temperature, relative humidity, sunshine duration, global solar radiation, UV index, and cloud amount. Subsequently, temperature, relative humidity, and UV index values are clustered into seven categories using k-means clustering. The average solar power generation is calculated based on forecasted weather features, and several neural networks are established for different ranges of power generation for training and testing. The research results show that using the LSTM-DNN model with grouped data reduces the RMSE value by 10% compared to the non-grouped LSTM-DNN model. Additionally, the LSTM-CNN-DNN model with grouped data reduces the RMSE value by 12% compared to the non-grouped LSTM-DNN model. These results demonstrate that the proposed method can reduce forecasting errors and assist in power dispatching.
In recent years, deep learning has achieved remarkable success in various fields, including biology and medicine. However, the interpretability and robustness still face challenges, as erroneous predictions can lead to serious consequences in critical applications. This paper introduces a new framework for estimating model uncertainty that can enhance the decision-making process. Compared to existing uncertainty estimation methods like Bayesian neural networks and Monte Carlo dropout, this paper provides a more direct way to model uncertainty by directly estimating the output variance through a dual-module neural network architecture. Without the need for multiple forward passes or complex Bayesian inference, it can reduce computational intensity. In specific, this framework consists of a prediction module for input processing and an uncertainty module for estimating prediction uncertainty. We demonstrate the efficiency of our method through applications in medical image segmentation and protein-ligand binding affinity prediction. By incorporating model uncertainty estimation, the proposed framework not only improves the accuracy of deep learning models, but also significantly reduces the computational burden which is especially beneficial for processing large datasets.
This paper presents a comprehensive method for dataset construction, utilizing 3D object detection to automatically label objects detected by LiDAR sensors and synchronizing multi-sensor labeling through coordinate calibration, thereby automatically generating image and radar datasets that support various learning algorithms. Initially, the cocalibration from the camera, radar, and LiDAR sensors is conducted to standardize the coordinate system based on the LiDAR. The camera output includes image information, encompassing object depth and related data. The radar sensor, particularly in automotive applications, returns data on the position of objects in front of the vehicle. Further, the Hungarian Algorithm is employed to analyze the association between radar and camera-detected objects. The proposed collaboration process with software workflow for automatic dataset generation with multi-sensors is detailed in this study. Finally, the preliminary results from sensor fusion over single-sensor modalities to object detection applications are presented to facilitate the efficient and rapid development of our approach to multi-sensor dataset generation, which is still extremely limited to the optical counterparts in autonomous vehicle environments.
The integration of computer vision technology into factory production offers significant improvements in efficiency, quality, and safety. This study explores the applications and benefits of multi-camera systems and AI models in monitoring and analyzing production processes. Key applications include quality control, process optimization, worker safety, and preventive maintenance. By utilizing real-time data, these technologies enhance product quality, optimize operations, and prevent accidents. Despite challenges such as occlusion, varying illumination, and the need for robust AI models, advancements in 3D point clouds and AI frameworks like YOLOS and Google Mediapipe show preliminary and promising results. This research demonstrates the transformative potential of computer vision in modernizing traditional manufacturing, leading to more efficient, safer, and higher-quality production environments. Future work will focus on refining these technologies and developing prototype systems for widespread industrial application.
Road marking signs are an important part of the road network. It gives information to the driver to help them understand the road conditions and improve driving safety. Fast and accurate road marking sign detection is a popular research topic. It is a basis for an efficient advanced driver assistance system (ADAS) and autonomous driving. In this research, we merged two datasets, The Taiwan Road Marking Sign Dataset (TRMSD) and the Taiwan Road Marking Sign Dataset at Night (TRMSDN), and relabeled the images to create the new dataset. The latest dataset provides more variation in lighting conditions, providing robustness to the trained model. YOLOv9, the latest YOLO, is trained and compared with YOLOv8 in the detection performance. The experiment shows that YOLOv9 performs better by achieving 0.921 precision, 0.925 recall, and 0.961 mAP50.
Rice diseases are one of the major factors affecting rice production. Traditionally, the identification and assessment of rice diseases have been done manually by experts and farmers, which is time-consuming and cannot provide real-time disease prediction, leading to significant losses. Therefore, there is a need for faster and more efficient methods for rice disease recognition. In this paper, we propose a solution for rice disease detection using artificial intelligence in image recognition. We investigate the performance of the visual geometry group (VGG) deep learning algorithms in recognizing rice diseases. Simulation results show that the VGG16 and VGG19 pretrained models achieve validation accuracies of 67% and 72.5%, respectively.
Kidney cancer currently has the 14 th highest incidence rate globally. Abnormal growths such as tumors and cysts in the kidney can be malignant and lead to cancer. The current method employed by nephrologists is manually locating the tumors and cysts from CT scans. However, due to the high incidence rate of kidney cancers, this time-consuming approach puts pressure on nephrologists and may lead to misdiagnoses. To improve the diagnostic process, this study proposes an automatic segmentation-guided classification method that can identify and differentiate between normal kidney areas and abnormal kidney areas that contain tumors, cysts, or both. Thereby, the model can detect abnormal kidney regions. In the proposed method, firstly, TotalSegmentator segmentation model's kidney masks are used to crop kidney regions from the abdominal CTs while generating 2D axial images. Then, ResNet 101 is used to classify the images into ‘normal’ and ‘abnormal’ classes. The proposed method has a high performance in all metrics with a score of 0.9626 and 0.8670 in recall and precision, respectively.
There's growing recognition of how machine learning can revolutionize the precision and swiftness of clinical diagnoses by improving the classification of white blood cells. However, a machine learning model specialized in white blood cell classification demands access to a large trove of well-annotated data. This poses a challenge, as in the medical field, annotating data comprehensively comes with a high cost in terms of both money and time. To circumvent this, domain adaptation (DA) techniques are employed to leverage similar cell images from related domains, thereby boosting the model's performance in white blood cell classification in specific contexts. Among the myriad DA approaches available for this application, including TCA-based and deep learning methods, all typically suffer from lengthy training durations. Addressing this challenge, we introduce a novel DA method inspired by the random vector functional link (RVFL) network, a type of neural network characterized by its random weights. Our proposed method, named Domain Adaptation Random Vector Functional Link (DA-RVFL), capitalizes on the efficiency of RVFL networks to enhance white blood cell classification. This innovative approach shows some insights for future research in the efficient and effective classification of white blood cells.
The stability of the power supply chain is crucial for maintaining the long-term development and cost-effectiveness of the power industry. Although existing equipment can maintain power supply in the short term, interruptions in the supply chain of components, such as geopolitical tensions, policy changes, limited supplier numbers, or poor management, may have an impact on the long-term stable operation and cost control of the power system. This article proposes a risk identification technique based on the Neo4j graph database, aiming to discover potential risk points in the supply chain in a timely manner through the analysis ability of the graph database. By constructing a supply chain knowledge graph(PSSCKG), this study aims to enhance the risk warning capability of the supply chain, optimize spare parts management and cost control, and thus support the sustainable development of the power system.
Single-cell multi-omics sequencing technology has made significant advances, which provides rich information for cell type identification. As a label-free data, clustering methods have been employed on multi-omics data to realize cell type identification. Although effectively integrating data from different omics facilitates the clustering tasks, most of existing methods designed for specific types of single-cell multi-omics data lack generalization capabilities. Therefore, this paper proposes a Graph Contrastive Learning framework for clustering single-cell multi-omics data (scGCL), which can be applied to various multi-omics datasets. Specifically, Our scGCL includes two modules, i.e., the Graph Generation (GG) module and the Graph Contrastive (GC) module. The GG module leverages k-nearest neighbors and random walk to build graphs based on omics data, capturing the global information and interactions among cell samples. The GC Module utilizes contrastive learning to extract discriminative feature-level and cluster-level representations. Moreover, considering the heterophily within single-cell multi-omics data, the backbone of GC module has been adjusted to enhance the representation learning. Experimental results on three single-cell multi-omics datasets demonstrate the superiority of our proposed scGCL.
This study utilizes a hyperspectral imaging system for in vitro imaging of blood vessels. Image preprocessing techniques, including Savitzky-Golay filtering, median pooling, gamma correction, CLAHE, and bottom hat transformation, are utilized to smoothen the images and improve the contrast between blood vessels and skin. The study introduces an iterative process that integrates an Automatic Target Generation Procedure (ATGP) for blood vessel extraction, an Orthogonal Subspace Projection (OSP), and an Edge-Preserving Filter for target detection and boundary extraction. The process will visually demonstrate the pathways of the vascular structures from the images. The iterative process was applied to compare the original and enhanced images. The results showed that the original images' average IOU and Dice coefficients were 0.11 and 0.19, respectively, while the enhanced images showed improvements to 0.29 and 0.44, respectively. Furthermore, when evaluated using the skeleton similarity metrics, the curve similarity accuracy of the original and enhanced images was 0.71 and 0.89, respectively. These results confirm the significant effectiveness of the image enhancement and iterative process fusion method (IEPF) in improving vascular segmentation performance.
In recent years, the number of farmers in Japan has been decreasing, and the automation of agriculture using robots, AI, and IoT, known as “smart agriculture,” is being promoted. In this context, Toyama city is promoting a project to build a robot for automatic weeding between plants in a perilla fields. In this research, we are in charge of the system part that detects weeding regions and informs the weeding mechanism. In the field, a wide variety of weeds form vegetation, and it is necessary to prevent food crops from being cut. Therefore, we detect crop regions and the rest of the area should be the weeding regions. In this study, we construct a system to detect crop regions by transfer learning the object detection model YOLO-v8, which is a convolutional neural network (CNN), and evaluate the system in terms of detection accuracy and processing speed.
Standing in line poses a significant challenge for the visually impaired. Existing queue guidance systems, which recognize line locations and estimate personal orientation, require highspecification computers, making them impractical for everyday use. This study proposes a smartphone-based queue guidance system that leverages the device's camera and sensors to estimate personal orientation efficiently. By optimizing computationally efficient algorithms for smartphone hardware, we aim to create an accessible, cost-effective solution. This system will enhance the independence and mobility of visually impaired individuals, providing real-time guidance to navigate queues seamlessly. Initial testing indicates promising usability and accuracy, paving the way for broader implementation.
Drug-drug interactions (DDIs) refer to the synergistic or antagonistic effects between different drugs. Synergistic effects can enhance therapeutic efficacy, while antagonistic effects may reduce efficacy or even trigger adverse reactions, worsening the patient's condition. Most existing DDI prediction studies overlook the features of chemical bonds between atoms within drug molecules and the interaction types between drug pairs, which to some extent limits the accuracy and reliability of the predictions. In view of this, we have designed a novel framework called KGE-DDI, which leverages knowledge graph embedding (KGE) technology to enhance the prediction of DDIs. Specifically, for individual drug molecules, KGE-DDI first integrates the features of atoms and chemical bond. Then, by introducing attention mechanisms and graph neural network technology, it learns representations of the drug molecules. For pairs of drug molecules, KGE-DDI models their interactions utilizing atom-level Pearson correlation matrices. Finally, it predicts the interactions between drug pairs by employing our designed scoring function. We conducted comprehensive experiments under both transductive and inductive settings, and the results demonstrate the effectiveness and superiority of KGE-DDI in three scenarios: (existing drug, existing drug), (new drug, new drug), and (new drug, existing drug).