
This study focuses on enhancing impurity detection in laver images through advanced data augmentation and tailored dataset design. The irregular texture of laver and the microscopic size of impurities present significant challenges for traditional detection methods. To address these issues, a comprehensive dataset was constructed using StyleGAN to generate 10,000 synthetic laver images, coupled with traditional augmentation techniques for impurity data. The YOLOv8 model was employed for impurity detection, achieving a precision score of 0.96 under IoU 75, a substantial improvement from the 30
Recently, 3D data generation techniques using neural networks are widely used in a variety of fields. In our study, we propose an efficient interactive annotation system by combining two 3D models: a high-quality 3D scene model generated by 3D Gaussian splatting and a model with surface information such as a mesh model. The system allows users to annotate the 3D space with intuitive operations. This system is expected to promote the effective use of 3D data and information sharing, and to be deployed in a wide range of scenes such as living environments.
Personal Protective Equipment (PPE) detection is essential for strengthening intelligent surveillance systems and ensuring worker compliance with safety protocols in modern manufacturing facilities. A YOLO-based PPE detection model provides an efficient solution for resource-limited environments, supporting real-time operation at high speeds. This work proposes a Multi-Perspective Attention (MPA) module to enhance YOLO11n’s performance in PPE detection, enabling its feature extractors to focus on critical elements. Hence, the improved YOLO11n achieves superior results compared to other methods on the PPE Detection (PPED) and Color Helmet and Vest (CHV) datasets, operating at 19.94 and 28.09 frames per second (FPS) on an Intel Core i7-9750H CPU and NVIDIA Jetson Orin Nano GPU devices, respectively, without significantly increasing parameters or computational cost.
This paper describes a method for creating human action sentences from video information, which aims to infer the purpose of human actions from video information. To develop a flexible human-robot communication method that enables robots to infer their own roles from the appearance of people, it is necessary to observe and convert human behavior into data. Therefore, we have created sentences describing human behavior by combining an action recognition model and an object detection algorithm, and we have organized the spatial relationships between objects and people. In Verification 1, a simple one-sentence explanation of the behavior was provided based on the distance between the two-dimensional (2D) coordinates of the person and the object taking the action from the video image information. Through the verification, we have identified issues in the generation of behavioral sentences. In Verification 2, we spatially organized their relationship based on the change in the three-dimensional (3D) coordinates of the object and the person. Through this validation, we were able to obtain spatial information on human behavior and object access records. In the future, by integrating these methods, we aim to generate an in-spatial action log that links human behavior and spatial information.
Human fall detection has been studied and applied in health care and human monitoring systems. Falls often occur in the elderly, and people with neurological and musculoskeletal diseases. It can also appear in heavy work environments, slippery surfaces, or human carelessness. Falls cause many serious health injuries, especially to the spine and brain. Developing tools to warn of falls early is essential and reduces many unnecessary risks. This paper proposes vision-based human fall detection and warning systems using person detection pre-trained models. Thereby, evaluating and comparing the performance between three nano versions of the YOLO (You Only Look Once) architecture, including YOLOv5n, YOLOv8n, and YOLOv11n. The YOLOv8n model achieved the best speed at 109.06 fps (FPS) on an NVIDIA GeForce GTX 1080Ti 11GB GPU as a real-time test result. The demo video can be accessed at this link: https://bit.ly/4flU7Xv .
Due to the continuous digitization of cultural heritage, many users can now experience their own museums. However, most digitization work is currently done manually, making it inefficient in terms of time and cost. This manual process underscores the necessity of automation to keep pace with the speed of digitization. Cultural heritage items are categorized into various groups, but the metadata information for each item varies, making it difficult to infer specific classification criteria. This study aims to automate classification by utilizing visual data. We addressed the issue of imbalanced datasets by generating a training dataset for cultural heritage using generative models. A dictionary was created to use appropriate prompts randomly, ensuring consistency in image generation. To validate the performance of the training model using the generated dataset, a multi-label classification model is used. The results demonstrate the potential for automating classification using visual features derived from text in the cultural heritage domain.
In deep learning, a phenomenon has recently been reported, in which the generalization performance is rapidly improved by learning a model until losses become almost zero during learning, and by continuing learning for a further large number of epochs after over-fitting. This is known as “Grokking”. In addition, “Flooding” is a method to improve recognition accuracies by learning a model to control that its loss becomes non-zero during a learning process. In this paper, we report on the effect of applying Flooding in a situation where Grokking occurs during learning a model, using a controlled environment with specific datasets, deep learning models, and hyper-parameters. In the evaluation experiments, we have confirmed that the application of Flooding sometimes improves recognition accuracies even in situations where Grokking occurs.
Expanding deep learning models to incorporate additional organ classes in medical image segmentation often requires costly retraining and significantly increases the risk of catastrophic forgetting of previously learned knowledge. Continual learning enables models to acquire new knowledge while retaining previously learned information progressively and has emerged as a promising solution to this challenge. In addition, to continual learning, another promising approach is multidata set training, which allows combining multiple organ classes from different datasets and training a deep learning model. This approach exposes the model to diverse datasets during training, significantly improving its ability to generalize to unseen data. The proposed study provides a detailed comparison of the three aforementioned approaches. We performed extensive experiments on two multi-organ segmentation datasets, FLARE22 and BTCV, to analyze the performance of each method. The multi-dataset training approach achieved an average Dice Similarity Score of 90.58
Robust inference against fluctuations in illumination conditions remains a significant challenge in image sensing. Existing works aim to achieve illumination-invariant features by estimating optical properties such as reflectance. However, accurate estimation using RGB images, where the spectral distribution is encoded into three dimensions, is inherently difficult. To address this limitation, approaches leveraging hyperspectral (HS) images, which provide richer wavelength information than RGB images, have been developed for reflectance estimation. Nonetheless, these methods often require the presence of a gray card with known diffuse reflectance within the image. Furthermore, the absence of publicly avail- able datasets combining reflectance values with real RGB or HS images constrains advancements in this field. To overcome these challenges, we propose an intensity-based illumination and reflectance estimation grounded in the Retinex theory. This approach eliminates dependency on gray cards by being guided by objects with known reflectance within the image, enabling illumination- invariant inference across diverse datasets. Additionally, we construct an HS-reflectance dataset and validate its utility. Through experiments, we evaluate our method and analyze its impact on classification using estimated reflectances. The results demonstrate the efficacy of the HS-reflectance dataset, our method, and its potential to enhance illumination robustness in visual systems.
Diffusion models have emerged as state-of-the-art generative frameworks, capable of producing high-quality realistic images. However, achieving precise and localized editing within specific regions of an image while preserving global consistency remains a significant challenge. In this paper, we propose a mask-based region-specific editing framework that leverages the latent space of diffusion models. Our approach enables fine-grained semantic modifications within a designated region of interest (ROI) while ensuring the unmasked regions remain unaffected. To achieve this, we compute the Jacobian matrix 𝐉_ℳ,t restricted to the masked region and employ Singular Value Decomposition (SVD) to identify the most influential directions in the latent space. These directions are further refined via nullspace projection to eliminate undesired changes outside the ROI. By efficiently manipulating the latent space along the identified directions, our method achieves localized edits without requiring external labels, retraining, or additional supervision. Extensive experiments demonstrate that our framework produces visually consistent and semantically meaningful edits across diverse attributes, masks, and scale levels. The proposed method provides a mathematically principled, and highly controllable solution for localized image editing, advancing the capabilities of diffusion models in practical applications.
In Japan, the number of elderly people who go missing has been increasing year by year. A delay in finding a missing person can lead to death. However, there are limits to the early detection that can be achieved through search efforts conducted by volunteers using visual observation. Therefore, there is hope that a method that utilizes information from security cameras and person matching technology can be used to track the whereabouts of missing persons and help with the early detection. In recent years, there has also been interest in person matching technology that utilizes 2D pose estimation. However, there is a problem that the person matching accuracy decreases due to the decrease in the pose estimation accuracy caused by the blocking of body parts due to the shooting angle of the security camera or the person’s clothes. This paper proposes a method to improve the person matching accuracy by focusing on the parts that are susceptible to a decrease in pose estimation accuracy and by masking the pose estimation information. As a result of the evaluation, the average accuracy rate improved by 3.7
Projection mapping is a technology that enables the display of visual content onto 3D objects using projectors. To dynamically change the projected content in response to varying real-world conditions, a projector-camera system is typically required, which combines a projector for displaying visual contents with a camera for capturing the target environment. However, estimating the geometric relationship between the projector and camera is a complex and time-consuming process, yet it is essential for accurately overlaying projection images onto real-world objects. To address this issue, a method involving the alignment of the optical axes of the projector and camera using a half-mirror has been proposed, which simplifies the alignment of their coordinate systems. However, this approach significantly reduces the intensity of the projected light, as half of the light is blocked. This research proposes a technique to pseudo-align the position and orientation of the projector and camera. Our method involves capturing two images using two cameras and synthesizing a novel viewpoint image at the optical center of the projector. This approach enables accurate projection onto 3D objects without the need for projector-camera calibration and avoids light loss during the process.
The occurrence of red tides causes mass mortality of farmed fish and leads to significant damage to the fishing industry, making early detection crucial. Currently, the detection and counting of phytoplankton, which are responsible for red tides, are mainly conducted through visual inspection using microscopes. However, this process requires time, effort, and knowledge of species identification. Therefore, in this study, we attempted to detect and count red tide plankton from microscope images of phytoplankton using the object detection method DINO.
This research addresses the automation of wood grain sensory evaluation and judgment rationale generation using a small dataset. Wood grain sensory evaluation involves assessing sensory elements such as color and pattern, which leads to significant variations in evaluator judgments. This challenge complicates the establishment of consistent evaluations, standardization of evaluations, and evaluator training. To address this issue, we construct a dataset using data collected from an expert wood grain evaluator and fine-tune a Vision-Language Model (VLM) to automate wood grain sensory evaluation and generate judgment rationales. However, existing VLMs are found to have insufficient capabilities in extracting local image features, which is crucial for this type of evaluation. To improve local feature extraction capabilities, we modify the architecture of the existing VLM. Specifically, we integrate features from the intermediate layers of a CLIP ViT-based model with those from a Convolutional Neural Network (CNN), creating a model that can capture both local and global image characteristics. By fine-tuning this model, we aim to enhance the accuracy of wood grain sensory evaluation and rationale generation.
In image-based authentication systems designed to protect physical assets, the lighting conditions under which images are captured typically remain consistent. Therefore, if an image captured under different lighting conditions is input into the system, it can be considered an indication of unauthorized access. This paper presents a system to discriminate whether the image used for authentication and the image registered in the authentication database is captured under differing lighting conditions using hyperspectral imaging. Specifically, we construct a dataset for this task and train a feature extractor for the identification based on metric learning, enabling robust and accurate discrimination of lighting conditions.
The decline in the number of farmers and the consequent loss of agricultural expertise poses a major challenge in passing on important skills to the next generation. The art of cherry mirror-packing, which requires many years of training and dexterity, is a typical example of this challenge. In this study, we research and develop a skill acquisition system to support the transmission and acquisition of the mirror-packing technique. Two expert and one novice packers from Yamagata Prefecture, one of Japan's major cherry-producing regions, were filmed and the 3D hand landmark coordinates of their hands were tracked using Mediapipe. Angular changes between the finger joints were calculated as a time series while the cherries were being packed. To handle time series of various lengths, kNN classification models were trained using DTW similarity distance, and separate models were created for the left and right hands. Based on the test results, classification achieved an accuracy of 0.86 for the left hand and 0.78 for the right hand. These results highlight the feasibility of automatic expertise evaluation for skill transfer and acquisition systems.
The digitization of cultural heritage artworks and crafts has become a critical strategy for their preservation, actively promoted by governments worldwide. One approach to this digitization is the generation of free-viewpoint images. These images, created through scene reconstruction based on multi-view image acquisition, not only enable detailed reproduction of intricately shaped crafts but also allow interpolation of unknown viewpoints, offering broad potential applications in fields such as education, tourism, and research. However, securing specialized personnel to manage the processes leading up to multi-view image acquisition remains a significant challenge. This research focuses on developing optimized environments for multi-view image acquisition by leveraging neural network-based methods. By defining environmental settings, such as camera arrangements, as known configurations, it becomes possible to easily acquire the necessary multi-view images without requiring specialized knowledge. By evaluating the reconstructed images, we propose shooting environments suitable for multi-view image acquisition, aiming to support the creation of digital content and enhance its utility.
Multi-media processing has achieved great success based on semantic segmentation. Semantic segmentation can be viewed as pixel clustering based on semantic prototypes. However, existing methods focus more on consistent semantics while ignoring the consistency in vision, making this task challenging. Motivated by the success of discrete visual representation learning, we propose Multi-group Visual Semantic Centroid (MVSC) to better cluster the pixels while maintaining consistent semantics of the dense features for any image encoder. Specifically, we randomly initialize multiple groups of prototypes as multi-groups in visual space. The visual features are also randomly split into the same groups and forced to be aligned with the corresponding prototypes. Then these visual prototypes are projected into the semantic space and supervised by the same classifier as the dense features. Compared with existing methods, MVSC further considers the visual space and thus facilitates the task. Experimental results on COCO-Stuff show great improvements compared with previous methods.
This paper presents a novel method for grasping objects with varying stiffness using an underactuated hand and a stereo camera. In factories, robots are required to handle a wide variety of objects. Tasks such as grasping soft objects without causing damage are particularly important in industries like food processing. While many existing approaches equip robotic hands with sensors, such as force or pressure sensors, these methods are unsuitable for food items due to hygiene concerns. To address the challenges of grasping various objects without causing damage or dropping them, underactuated hands that can conform to object shapes have gained attention. In this study, we propose a method for controlling an underactuated hand using only a stereo camera as an external sensor. First, the target object is detected using a background subtraction method. Next, the contact between the hand and the object is detected. Then, the object is grasped with appropriate force, calculated based on four elements: the centroid shifts of the hand and the object, the deformation rate of the object, and the occlusion rate of the hand. Finally, drop detection is performed to ensure the object is not dropped during pick-and-place tasks. Experiments were conducted using six different objects to validate the proposed method.
Person identification can be performed using temporal features extracted from video sequences of body sway captured by an overhead camera. We propose a method for extracting such temporal features in a manner that is robust to headwear variation. When people are wearing headwear, such as a cap or helmet, their head shapes observed by the camera change significantly according to the type of headwear. Existing methods for person identification cannot achieve high accuracy in the presence of headwear variation because the features used by existing methods are strongly affected by the changes in head shapes. To extract temporal features that are not influenced by headwear variation, we measure the time-series signals representing body sway by estimating the center positions from head shapes. Moreover, we propose a learning-based low-pass filter to remove the components that are uninformative from the frequency components of the time-series signals, while retaining the informative components. Experimental results show that our temporal features significantly enhance the accuracy of person identification in the presence of headwear variation, compared with the use of existing features.