We are developing a Cyber-Physical System (CPS) for earthwork sites, called ROS2-TMS for Construction, utilizing OPERA, an autonomous construction platform under development by the Public Works Research Institute. In this paper, as a case study of the system’s field implementation, we automated the loading of cohesive soil into a hopper during soil improvement work using a hydraulic excavator. Experimental results demonstrated that the system was able to dynamically determine excavation positions and achieve continuous operation for more than one hour. Furthermore, we conducted an additional slope-collapsing experiment using a hydraulic excavator, which demonstrated the applicability of the proposed system to diverse earthwork tasks.
The construction industry faces severe labor shortages, driving the need for robotic automation solutions. However, effective deployment of construction robots requires robust environmental perception capabilities, particularly accurate identification of diverse objects in complex, dynamic construction environments. Closed-set object detection methods are limited to predefined categories, proving inadequate for the highly varied object types encountered on construction sites. This paper introduces EVLOD (Ensemble Vision-Language Open-vocabulary Detection), an ensemble framework that integrates multiple state-of-the-art vision-language models to enable open-vocabulary object detection in construction scenarios. EVLOD employs a voting-based fusion strategy that combines predictions from GroundingDINO and GroundingDINO-CLIP detectors, utilizing their complementary strengths while mitigating individual model weaknesses. The ensemble approach incorporates confidence voting, object name voting, and bounding box voting to produce reliable detections with reduced false positives. Evaluated on a comprehensive dataset of 825 Unmanned Aerial Vehicle (UAV)-captured construction images with 5,020 annotated objects, EVLOD achieves an Average Precision (AP) of 0.49 when Intersection over Union (IoU) equals 0.5, representing a 36.1% improvement over the best-performing baseline. The method effectively reduces detection noise from 5,495 to 3,232 detections. Qualitative analysis reveals primary limitations in detecting small-scale objects and low-contrast elements.
Labor shortages driven by aging workforces have increased demand for robotic automation on construction sites, yet existing LLM-based systems represent task plans in formats that field operators cannot readily inspect or verify, and provide no mechanism for targeted modification once execution has begun. This paper presents a human-verifiable execution-time DAG refinement framework in which the active plan is rendered as a color-coded node-edge graph that operators can inspect and confirm at each step. Natural-language corrections are translated into targeted edits of the pending subgraph, with completed subtasks locked and in-progress subtasks paused only when necessary for a validated update. A refinement taxonomy and validation pipeline enforce execution-state consistency before each update is dispatched. Evaluation in simulation, a 12-case error-correction study spanning three complexity levels, and deployment on an edge device confirm practical feasibility. A controlled user study further shows that the DAG-based operator interface achieves 94.2
In this study, we present a novel autonomous excavation method that achieves high efficiency under varying soil conditions. This method consists of two main steps, including first estimating the density of the soil and then generating an optimal excavation path based on the estimated density. The proposed method estimates soil density by taking advantage of the bulking phenomenon, which refers to an increase in the volume of excavated soil. This estimation relies solely on 3D point-cloud data obtained before and after excavation. Using the estimated soil density, an optimal excavation path is generated by applying a genetic algorithm in a physics simulator that replicates both the hydraulic excavator and the target ground. The algorithm explores a range of paths over multiple generations to find one that maximizes efficiency. The effectiveness of the proposed method was verified through simulations and field experiments. In particular, field experiments conducted in soft soil showed that the proposed method improved excavation efficiency by 27.7% compared with a baseline method using fixed parameters.
Earthwork operations face increasing demand, while workforce aging creates a growing need for automation. ROS2-TMS for Construction, a Cyber-Physical System framework for construction machinery automation, has been proposed; however, its reliance on manually designed Behavior Trees (BTs) limits scalability in cooperative operations. Recent advances in Large Language Models (LLMs) offer new opportunities for automated task planning, yet most existing studies remain limited to simple robotic systems. This paper proposes an LLM-based workflow for automatic generation of BTs toward coordinated operation of construction machines. The method introduces synchronization flags managed through a Global Blackboard, enabling multiple BTs to share execution states and represent inter-machine dependencies. The workflow consists of Action Sequence generation and BTs generation using LLMs. Simulation experiments on 30 construction instruction scenarios achieved up to 93% success rate in coordinated multi-machine tasks. Real-world experiments using an excavator and a dump truck further demonstrate successful cooperative execution, indicating the potential to reduce manual BTs design effort in construction automation. These results highlight the feasibility of applying LLM-driven task planning to practical earthwork automation.
In this study, we propose a system to detect changes in three-dimensional (3D) space for autonomous plant visual inspection by a mobile robot. The videos captured by a mobile robot during past inspections are compared with the videos obtained during the current inspection using both pose information and the acquired images. To ensure robustness against changes in shooting conditions, change detection is executed employing deep learning techniques. Subsequently, the detected information is projected onto a 3D space to localize the changes. To verify the effectiveness of the proposed method, experiments were conducted both in a real plant environment and a simulated indoor plant environment. The results of the outdoor experiments showed that the proposed system achieved image pair determination, change detection, and integration into a 3D space. The results of the indoor experiments and evaluations confirmed that the proposed method for image pair determination was suitable based on considerations of detection accuracy and computation time.
In this paper, we examine the adaptability of a multi-robot coordination system in which distributed autonomous robots respond to unexpected situations based on functional expressions using large language models (LLMs). In recent years, there has been growing interest in systems where multiple autonomous distributed construction machines (robots) collaborate and adaptively perform tasks in open and unknown environments, including disaster sites. In open environments, unforeseen events that cannot be predicted in advance may occur, and it is challenging to address these events solely with existing model-based approaches. In this paper, we leverage the high environment comprehension capabilities of foundation models to understand unforeseen situations and develop a system that enables adaptive coordinated actions by flexibly integrating the functions of multiple robots using LLMs. Additionally, we account for the embodiment (interactions between the robots and their environment). Our experiments demonstrate that the designed system is capable of adaptively responding to three types of unforeseen situations, including path obstructions caused by either an obstacle or a robot. In cases of path obstructions caused by obstacles of varying weights, the system can exhibit appropriate obstacle removal behaviors by reflecting the torque capability of the robot as one aspect of its embodiment.
This study proposes a method for change detection by extracting small image pairs from recorded video pairs in patrol inspections of oil plants using a mobile robot equipped with a camera. While change detection methods often perform well in controlled laboratory environments, their performance tends to degrade in plant environments. This is because, in plant environments, the regions of change caused by anomalies within the images are often small, and the environments are structurally complex. In this study, to address this issue, we adopt an approach that extracts small image pairs from the target and reference images for change detection. Additionally, we develop a method for extracting small image pairs that considers the consistency of three-dimensional coordinates based on multi-view stereo. As a result, the proposed method achieved high-precision change detection, demonstrating its potential for application in patrol inspections of oil plants using a mobile robot.
When soil spillage occurs during excavation due to the use of excavators, it must be cleaned up, leading to a decline in labor productivity and a negative impact on the working environment. In this study, we aimed to reduce the amount of soil spillage generated when using automated excavators operating without human intervention. The excavation process consists of penetrating, dragging, and scooping operations, and it was observed that soil spillage occurred most frequently during the scooping operation. Therefore, a method to optimize the excavation operation with the goal of reducing soil spillage was proposed. In the scooping operation, we proposed a back motion that moves the bucket towards the front of the excavator. The proposed excavation operation reduced the amount of soil spillage, as confirmed through physical simulation experiments and experiments using an actual excavator.
In this paper, we propose a novel approach to detection and localization of abnormal sound sources for robotic inspection in oil refineries. Such environments are difficult environments with high noise from multiple machines and where swift detection of anomalies is critical. The rarity of anomalies hinders the gathering of a balanced training dataset for the common supervised learning approach. Our previous work, based on autoencoders, bypassed this issue but lacked the ability to locate the abnormal sound source. Our proposed method first learns a spatial map of the normal sounds, allowing to predict what sound should be present at each robot position. This enables a detection based on a comparison between the predicted and observed sound. Localization can then be conducted based on this comparison using optimization. Experiments conducted in laboratory conditions showed the effectiveness of the proposed method. Additionally, experiments in field conditions in an actual oil refinery further showed the potential of the proposed method.
Tracked robots are useful in disaster recovery missions owing to their high traversability. However, even tracked robots find traversing terrains with obstacles such as rocks difficult. Particularly, unfixed obstacles pose problems. Therefore, this study aims to generate motion for tracked robots to go over unfixed obstacles on slopes by utilizing a reinforcement learning (RL) approach. In RL, rewards are important to make robots learn optimal policies for various tasks. Hence, a new definition of going over an unfixed obstacle is proposed. Based on the proposed definition, a sparse reward is designed to motivate the robot to go over the obstacle while avoiding the large sliding-down phenomenon responsible for the failure of going over and the large directional deviation of the robot. In our experiment, a robot is trained in a dynamic simulator and it attempts to go over a spherical unfixed obstacle under various environmental conditions: radius of the obstacle, slope degree, and initial position of the obstacle. We verify that the success rate converges to a high value and the robot successfully generates motions to go over the spherical unfixed obstacle.
This study addresses the urgent need to monitor the water depth of landslide dams, which themselves have the potential to cause downstream flooding. A drone-deployable water depth sensor was developed to enable this monitoring. The sensor employs a simple mechanism that calculates water depth based on the number of reel rotations, enabling long-term, fixed-point observations despite its compact size and lightweight design. The experiments demonstrated the sensor's accuracy, operational duration, and the feasibility of its deployment scenarios.
In this study, we propose a method for detecting changes on the outer surface of pipes using inspection videos captured by an inspection robot. It is critical to detect anomalies on the outer surface of pipes during patrol inspections. Anomalies are defined as deviations from the normal state and should be detected as areas that have changed from the normal state. Therefore, for appropriate maintenance of the plants, it is crucial to perform change detection by comparing videos that capture the past normal state with those capturing the current state. The problem with detecting changes from videos is deciding which frame to compare. We therefore propose sequential filtering to determine image pairs based on the position of the images and their similarity. We then apply a deep learning method to perform change detection. An indoor simulated plant environment has been constructed to test the efficacy of the proposed method. Experiments and evaluation results showed that the proposed method outperformed an autoencoder. The proposed method also achieved an F1 score of 0.880 for change detection in the inspection videos by introducing sequential filtering, which prevented mismatching of image pairs and reduced computational costs.
Most research on automating gravel pile transport using wheel loaders has been performed primarily through simulations. Thus, studies should evaluate the usefulness of automatic gravel pile transportation by demonstrating it with an actual wheel loader. This study demonstrates automatic driving control using a retrofitted 3-ton wheel loader for gravel pile transportation. The driving model of a retrofitted wheel loader, in which multiple control systems are interlocked, is considered a simple control model with one input and one output for the pedal and vehicle velocity as well as for the steering wheel and steering angular velocity. In this study, we propose a simple and practical method for constructing a driving model via simple response analysis using an actual machine by constructing a feedforward control model based on control input/output using step responses. In this study, feedforward control was applied to the translation of the vehicle, which has a large dead time. By generating the path following the target point from the vehicle predicted position after the dead time from the driving model, the appropriate control input value calculation considering the dead time was performed. By applying the proposed method to a retrofitted wheel loader in a real environment and evaluating the control performance through control experiments, the effectiveness of the proposed method in practice was demonstrated.