Deep learning models often suffer from performance degradation when applied to construction sites that differ from the source domain due to their sensitivity to data distribution shifts. Although methods such as transfer learning, domain adaptation, and synthetic data generation have been explored to improve generalization, collecting and annotating data from new target domains remains a labor-intensive bottleneck. This study presents a self-training-based framework to generate training data for construction object detection in unlabeled target domains. The method identifies moving objects using optical flow estimation, propagates class labels through iterative self-training, and synthesizes realistic training images via image inpainting and copy-paste augmentation. Experimental results from four visually distinct construction scenes demonstrate that the proposed method significantly improves detection performance without relying on manually labeled target data. These findings contribute to advancing automated and scalable domain adaptation techniques for vision-based construction monitoring.
Although computer vision technology has shown great potential, its reliability can significantly degrade in the target domain where the model is applied. Collecting and labeling training data from the target domain can address this issue; however, it is a tedious and time-consuming task. To address this issue, this paper presents a novel method generating training data for construction site monitoring. The proposed method consists of extracting moving objects, classifying each region by comparing their features to the target classes, and then assigning class labels. The newly labeled data is copied and pasted to a clear background of the target domain where its foregrounds are removed by image inpainting. Experiments were conducted on construction site videos captured in far-field monitoring environments. The proposed method can significantly reduce the amount of effort required for data collection and labeling, thereby increasing the efficiency of developing robust computer vision models for construction site monitoring.
Multiple Object Tracking (MOT) has potential applications in construction site safety man- agement, particularly for the individualized assessment of scaffold workers' safety status. However, its accuracy can be degraded due to inconsistent detection results and insuffi- cient object association capabilities. This paper addresses these challenges by proposing an object tracking method, named as the Shallow Cascaded Buffered Intersection over Union (Shallow C-BIoU). This non-machine learning-based tracker employs a color hash- ing technique to enhance tracking performance. Employing rigorous evaluation metrics, the experimental results underscore the efficacy of the proposed method. Compared to state-of- the-art algorithms, the Shallow C-BIoU method improved association performance by 7.56%, reducing 52.44% of falsely assigned tracking IDs. Consequently, this paper contributes to the development of reliable object trackers, thereby advancing monitoring technologies for construction management purposes.
Context-Adaptive CCTV Pan-Tilt-Zoom method for Personal Protective Equipment Detection Seokhwan kim, Minwoo Jeong, Minkyu Koo, Taegeon Kim, Hongjo Kim Pages 768-775 (2024 Proceedings of the 41st ISARC, Lille, France, ISBN 978-0-6458322-1-1, ISSN 2413-5844) Abstract: PPE items, including hardhats, hooks, harnesses, and straps, are critical for fall prevention. Ongoing research in construction safety has focused on using deep learning models to detect Personal Protective Equipment (PPE) worn by high-altitude workers. Despite efforts using computer vision-based models for safety monitoring, small object detection, such as hooks and straps, remains challenging due to image resolution issues. This study introduces a novel technique using mobile CCTV cameras controlled by an automated Pan-Tilt-Zoom (PTZ) algorithm to enhance the detection of small-sized PPE. The method leverages the size gap between worker and PPE. In a zoomed-out state with a short focal length, the system identifies the worker's bounding box (b-box), then zooms in with a longer focal length for precise PPE detection. When encountering multiple workers, the system applies predetermined zoom-in rules. Experimental results demonstrated a significant increase in detection accuracy for the small PPE: hook detection improved from 39.8% to 88.3%, and strap detection from 49.4% to 71.8%, as measured by an mAP of 50. This encouraging performance improvement suggests that automated PTZ control technology could enhance the effectiveness of safety monitoring. Keywords: Construction safety, PTZ CCTV control, monitoring, PPE detection, Small object detection DOI: https://doi.org/10.22260/ISARC2024/0100 Download fulltext Download BibTex Download Endnote (RIS) TeX Import to Mendeley
Zero-shot Learning-based Polygon Mask Generation for Construction Objects Taegeon Kim, Minkyu Koo, Jeongho Hyeon, Hongjo Kim Pages 81-88 (2024 Proceedings of the 41st ISARC, Lille, France, ISBN 978-0-6458322-1-1, ISSN 2413-5844) Abstract: For construction sites monitoring, the use of segmentation-based computer vision technology has been proposed. In such environments, the main technical challenge is the generation of data for training the segmentation model. The training data for a segmentation model involves polygon annotation of objects within an image, which is a time-consuming task. To address this issue, this study proposes a new approach that uses the YOLOv8 object detection model to predict bounding box labels and inputs these into a Segment Anything Model (SAM) to automatically generate polygon label data. The performance of the YOLOv8 model exceeded 80%, and the automatic generation of polygon labels through SAM resulted in an IoU range of 55-86%, producing high-quality mask label data. This approach significantly reduces the time, labor, and cost associated with the labeling process. Keywords: Polygon label generation, Instance segmentation, Zero-shot learning DOI: https://doi.org/10.22260/ISARC2024/0012 Download fulltext Download BibTex Download Endnote (RIS) TeX Import to Mendeley
Image segmentation-based applications have been actively investigated. However, it is non-trivial to prepare polygon annotations. Previous studies suggested pseudo label generation methods based on weakly supervised learning to lessen the burden of annotation. Nevertheless, the quality of pseudo labels could not be ideal due to target object characteristics and insufficient data size in the construction domain, as identified in this study. This study proposes a fusion architecture, SESC-CAM, to address the challenge, building upon weakly and self -supervised learning methods. The proposed architecture was validated on the AIM dataset, and the generated pseudo labels recorded a mIoU score of 64.99% and 67.65% after the refinement by using a conditional random field, and outperformed its predecessors by 11.29% and 9.14%. The refined pseudo labels were used to train a segmentation model and recorded a 74% mIoU score in semantic segmentation results. The findings of this study provide insights for automated training data preparation.
The performance of deep learning models can be significantly degraded on unseen data that has different visual characteristics compared to a domain where training data was collected. A simple and obvious way to maintain the performance of deep learning models is to prepare training data again in a new domain where target objects and backgrounds have different appearances compared to the original. However, it is not a trivial task considering time and efforts required in data preparation. To address this issue, this study proposes a pseudo label generation method from images that can automatically collect video clips for objects of interest and assign labels. The proposed method consists of a moving object detector to extract target objects in images and a classifier to assign labels on the extracted regions. The findings of this study provide important knowledge for construction site monitoring in securing the performance of computer vision models in various environments.
If water trash exceeds the allowable load of a trash barrier, water trash barriers could be destroyed and the spilled waste negatively impacts the environment. Therefore, it is essential to measure the trash load in the water infrastructure to process collected water trash in a timely manner. However, there has been little investigation about how to monitor water trash in an automated way. To fill the knowledge gap, this study presents detailed investigation of water trash monitoring methods based on object detection models. To verify effective detection models and their performances, a new dataset is established, called the Foresys marine debris dataset. The dataset consists of a total of 6 water trash categories (Plastic, Vinyl, Styrofoam, Paper, Bottle, and Wood). State-of-the-art detection models were employed to test their performance, such as YOLOv3, YOLOv5, and YOLOv7 pretrained on the COCO dataset. The experiments showed that the detection models could achieve decent performance with proper amount of training image data; a number of training data required to secure decent performance varies by target class. The findings of this study will give a fresh insight for developing an automated water trash management system.
To facilitate road image data collection, participatory sensing has been proposed in the literature utilizing a dashboard camera of a normal vehicle. It is not trivial to identify road cracks in such crowdsourced images due to the dynamic natures of photographing conditions which results in inconsistent the image quality. Although previous studies presented promising ways to identify road damages using deep convolutional neural networks (CNN), the performance is insufficient to be implemented in practical monitoring purposes. This study investigates core problems in improving the road crack segmentation performance by applying state-of-the-art segmentation models based on CNN and transformer architectures. Using a benchmark dataset, it was found that coarse annotation on crowdsourced images is detrimental to the performance evaluation and further development of participatory sensing-based monitoring technology. Interestingly, segmentation models could be trained by training data with coarse annotation. This study will give a fresh insight of advancing the knowledge in participatory sensing-based infrastructure monitoring.
Deep learning models, due to their high sensitivity to training data distributions, may suffer from performance reduction when applied to construction sites different from the source domain where the training data originated. To overcome the laborious processes of data re-collection and annotation, this paper introduces a self-training strategy to generate training data for construction object detection attuned to the target domain. The proposed method produces training data by: (1) employing optical flow estimation to detect moving objects, (2) leveraging self-training to propagate existing labels to unlabeled data, and (3) utilizing copy-paste augmentation and image inpainting to generate target domain-specific training data. Experimental results from four different scenes substantiate the efficacy of the proposed method in boosting the performance of object detectors within new target domains. The findings provide a fresh perspective on the generalizability of deep learning models for construction site monitoring, thereby enhancing the automation potential of construction management.