
This paper presents a novel system for simultaneous underwater acoustic (UWA) positioning and communication aimed at enhancing underwater drone operations. The proposed method utilizes one of the communication schemes, orthogonal signal division multiplexing for UWA positioning. An experimental study was conducted in a coastal area to validate the system’s performance in actual sea conditions. The results demonstrated that the proposed system achieved a mean positioning error of 1.38 meters and error-free communication in more than 95% of cases. This system is anticipated to serve as a foundational technology for enhancing the capabilities of underwater drones.
This research investigates the influence of UI components on decision-making in E-commerce services, with a focus on decision fatigue. Through depth interviews and experiments within the apparel E-commerce genre, we analyzed the effects of different decision-making experience models featuring varying levels of button component differentiation. Emotional influences were assessed using the Profile of Mood States (POMS). The results revealed significant changes in "vigor" and "confusion," suggesting that differentiating UI components can positively influence decision-making experiences. The research provides guidelines for UI design to mitigate decision fatigue and improve user experiences.
In recent years, the development of autonomous driving technology has increased the need to acquire information about the surrounding area. Accordingly, LiDAR sensors are used to acquire a 3D point cloud of the surrounding environment and detect vehicle areas and other object as 3D bounding boxes. However, high-precision LiDAR sensors are very expensive, and it is difficult to install them in all vehicles. A method to generate 3D point clouds from images has been proposed to improve this problem, but the generated point clouds are subject to noise and missing parts, resulting in low detection accuracy of 3D bounding boxes. Therefore, this method improves the accuracy of 3D bounding box detection by applying appropriate noise reduction and completion processing to point clouds obtained from stereo images.
A direction finding algorithm based on point cloud data analysis is described for a planar antenna array. Cumulative phases are calculated by sequentially adding inter-element phase differences to free from wrapping between -π and π. The cumulative phases are transformed into propagation delays as the point cloud data. The norm vector, which is perpendicular to an approximated plane of the point cloud data, is obtained as the eigenvector with the minimum eigenvalue through eigenvalue decomposition of a 3 × 3 covariance matrix. Simulation results show that the angle of arrival is successfully estimated in environments with a single wave of arrival.
This paper explores the integration of Artificial Intelligence (AI), the Internet of Things (IoT), and Human-Computer Interaction (HCI) to enhance data analysis and decision-making in healthcare, particularly during the COVID-19 pandemic. Through a mixed-methods approach, including a literature review and case study analysis, the study highlights how AI-driven analytics and IoT data streams, exemplified by the COVID-19 Data Lake and WHO Health Alert platform, provide real-time insights and improve decision-making efficiency. The success of these systems relies on effective HCI integration, robust data validation, and comprehensive user training. Challenges such as data accuracy and system adaptability are identified, with future research directions focusing on advanced AI techniques and the development of sophisticated HCI frameworks. By addressing these challenges, the integration of AI, IoT, and HCI can further transform healthcare and enhance public health outcomes.
Low computation cost is crucial for a seamless experience on mobile consumer devices with limited resources. This study presents an efficient attention module for deep learning-based image super-resolution under low computing cost. We propose partial enhanced spatial attention (PESA) to achieve efficient and high-performing attention modules, which draws inspiration from partial convolution for feature extraction. Utilizing an efficient super-resolution network, our approach is assessed on two super-resolution datasets and contrasted with other attention strategies. PESA obtains the lowest computing cost and model parameters, as well as the top quantitative results on both datasets.
This study aims to integrate blockchain technology into agricultural systems to establish an open crop environment monitoring system. The advancement of smart agriculture systems contributes to improving productivity and quality; however, centralized data management and distrust can lead to issues with data integrity. To address these concerns, we propose a monitoring system that utilizes blockchain smart contracts to ensure the transparency and integrity of crop environment data, thereby resolving issues of distrust in data management. This system will enhance trust in crop data and contribute to the digital transformation of agriculture, increasing the efficiency of the production process.
Haptic stimuli, including vibratory stimuli, to upper body parts effectively elicit emotional responses. The external auditory meatus can be a target region due to the concentration of vagus nerves. We investigated whether vibratory stimuli to the external auditory meatus lead to a relaxed state. Ten participants experienced two conditions: with and without vibratory stimuli. In the without-vibration condition, a still contactor was placed at their ear. The condition with vibratory stimuli was reported to be more relaxing than the without-vibration condition. However, several indices computed from electrocardiograms, including heart rate and its variation, and the ratio of low- to high-frequency components of RR intervals (LF/HF), did not exhibit differences between before and after the stimulation. This research could aid in the development of emotionally relevant consumer electronics products.
According to the World Health Organization, mental health refers to the good psychological well-being of a person. Due to genetical, environmental, biological, and psychological reasons an individual’s capability to cope with day-to-day work may reduce resulting in various mental issues like anxiety, depression, and bipolar disorder. With the rapid growth of social-media penetration, people tend to share their real-life updates. Social media data is a good source which reflects one’s mood, tension, and feelings. These psychological behaviours could vary over time, indicating the importance of the time factor. Monitoring a person’s post behaviour allows more dynamic and contextual understanding of a person’s mental state beyond static assessments. This paper provides a narrative review of social media anxiety detection using time series analysis over the last decade. Here we answer the research question on, what are the existing methods, challenges, and future directions on social media time series anxiety detection.
In recent years, the accuracy of automatic speech recognition (ASR) for major languages has been greatly improved by pre-training methods using large spoken language resources. However, practical ASR technology has not yet been realized to cover the large and rich variety of regional dialects of the Japanese language. This study focuses on the adaptability of two state-of-the-art large pretrained models for building a unified ASR model for Japanese dialects. We present results from adapting these models using a total of several dozen hours of Japanese dialect speech. We compare models optimized for each dialect region, including dialect region identification, with models adapted without distinguishing between dialect regions. By comparing these two different learning processes, we investigate how various adaptation methods impact ASR performance for Japanese dialects.
Companies and organizations collect and analyze various data types from service provision to enhance customer satisfaction. The challenge lies in efficiently utilizing only the necessary data while safeguarding individual privacy. This paper proposes the method to protect data privacy through aggregation using mathematical operations, preserving the data’s utility for decision-making. The major contribution is to map original values to specific numerical values, which helps to hide the real values during aggregation. The effectiveness of this method is shown by evaluating the root mean square error (RMSE), which decreases as more aggregation data is used. Our proposed method reduces the error to less than 0.1.
This study presents a support system for rehabilitating patients with mental illnesses using data from wearable devices. Utilizing the Fitbit Sense 2 smartwatch, the system collects daily life data and displays it graphically for both patients and physicians, enhancing traditional diagnostic methods. The system comprises a sensor unit for data collection, an information processing unit, and an information presentation unit. A pilot study with four patients and one physician over three months showed effective data acquisition and significant utility in depression treatment. This integration of sensing data with traditional diagnostics offers valuable insights for mental health care.
With the rising need for electric vehicles, more and more manufacturers participate in legislating the standard of charging stations, such as ISO 15118 and IEC 61851. They concentrate on the safety and security of electricity and communication. However, charging station manufacturers and electric vehicles might only partially comply with the standards when implementing the products. It can lead to potential vulnerabilities for attackers to exploit the charging stations or electric vehicles. Therefore, we construct a security testbed to identify the potential risks in the ISO 15118 standard. We successfully emulate a scenario to perform a sniffing attack that can capture plaintext traffic and a man-in-the-middle attack that modifies the messages between the charge station and the electronic vehicle.
Cultural heritage interpretation holds a critical role in enriching the comprehension and admiration of historical sites in small towns by visitors. Nevertheless, the lack of adequate human resources in remote areas may pose a challenge to the effective interpretation of heritage. In this research, a new methodology is suggested, which utilizes Retrieval-Augmented Generation (RAG) technology to deliver personalized cultural heritage narratives that are customized to suit the preferences of tourists. Through the consolidation of extensive language models with retrieval capabilities, the system can connect visitors' existing knowledge with the local attractions, objects, and activities, thereby nurturing a stronger bond with the abundant history of the location. The proposed system helps overcome the challenges of reaching diverse audiences and highlighting the cultural importance, while also reducing the need for specialized and trained staff members. A trial run employing the Llama3 model and resources linked to Nanliao village in Penghu Island showcases the system's proficiency in producing pertinent and culturally sensitive outputs. The research shows the potential benefits of combining advanced language analysis technology with explanations of cultural heritage. This could provide a scalable and efficient way to enhance the experiences of visitors to small towns.
Preventing diseases in crops and fruit trees is an urgent issue in the fight against global food shortages. If a disease occurs, it is highly possible that it may have spread throughout the farm before being detected by the farmer. Therefore, in this study, to facilitate early disease forecasting, we propose using a sensing agricultural robot that can meticulously observe environmental alterations that could serve as potential disease indicators. In addition, because the forecast results for a different farm showed a decrease in recall compared to the farm where the images were collected, we employed explainable artificial intelligence (XAI) for feature analysis. We confirmed an improvement in the recall through simple feature selection.
In this paper, we propose a novel method for detecting Deepfakes in video content such as news broadcasts. Deepfake is the techniques to replace parts of images and audio to a high degree using deep learning, and its quality is improving as computing power increases. It can be used to create malicious videos, such as replacing the faces of politicians or company CEOs to instruct fraudulent transactions, raising concerns about the risk of propaganda and fraud, and measures are needed. Previous research has mainly implemented Deepfake content detection using machine learning, but challenges remain, such as low accuracy and difficulty in quickly responding to new Deepfake techniques. In this study, we propose a method that embeds a watermark resistant to linear processing, image processing, and collusion attacks into both the facial and background regions of a single frame in the video. The results of applying our proposed method to 1,000 facial images showed that we could detect Deepfake videos with 100 % accuracy.
In recent years, with the spread of devices with various screen sizes, the demand for resizing methods is increasing. Seam carving (SC) removes visually unimportant areas (hereinafter referred to as "seam") and resizes the image without losing its impression. However, SC for video has the problem of misalignment of the subject because the position of the seam is different in each frame. Also, previous methods that solve this problem have the problem of distortion of the subject. Therefore, the proposed method enables video SC with reduced misalignment and distortion by calculation of seams considering subject’s movement and reduce of seam position changes.
With the resurgence in travel demand post-COVID19, enhancing access to tourism sites is critical. Web-based information can be complex, making destination selection difficult. This work uses image search to find short videos and related spots to generate tourist attraction introduction videos. It conveys attractions more realistically than text and images. Image search finds tourist spots similar to an input photo, reducing decision time. We collect and label short tourist videos, determine tags for user-input photos, and use cosine similarity to extract relevant spots.
To enhance driving safety, in this work, an effective driver monitoring system (DMS) is developed by YOLOv7-tiny for object detection and by Dlib keypoints for head pose estimation. The proposed design is implemented on Nvidia Jetson Xavier device, and the system demonstrates the high precision and recall, and then it achieves a mean average precision of 98.2%. The experimental results highlight the performance of the proposed non-contact DMS design for improving driver’s attentiveness and reducing accidents.
This paper proposes an Artificial Intelligence (AI) self-harm behavior detection system called ThermalEye. By using AI to recognize the skeleton of human body in low-resolution thermal images, and self-harm behavior recognition algorithms, the ThermalEye system can instantly recognize the self-harm behaviors of patients in the psychiatric seclusion rooms of the hospital. Experimental results demonstrate that the ThermalEye system can significantly improve the accuracy of skeleton recognition in low-resolution thermal human images and accurately and immediately detect when patients engage in self-harm behavior, thereby improving the hospitalization safety of people with a mental health condition.