Construction remains one of the most hazardous global industries, with worker fatalities occurring approximately every 99 min. However, safety is intrinsically linked to broader project performance; a truly safe construction site is hazard free, organized, and efficient, thereby minimizing waste and resource depletion. In response, a growing body of research has explored the use of advanced technologies, such as machine learning, robotics, wearable, and immersive visualization, to synergize safety improvements with sustainable building practices. In this article, we systematically review 165 peer-reviewed articles in the era of artificial intelligence (AI) that capture how technologies have evolved to address the dual imperatives of safety and sustainability in response to technology-driven transformation. To analyze this inflection, the article develops a four-domain taxonomy spanning AI, visualization, surveillance, navigation and collaboration, and wearable technology each mapped against both earlier and recent research studies. Crucially, we highlight emerging dual-utility applications, such as the use of "green digital twins" for real-time lighting optimization and UAVs for building envelope energy auditing. By mapping this transition, this article provides a grounded perspective on current capabilities, identifies research practice gaps, and supports informed decision-making for the implementation of future-integrated safety and sustainability technologies. The insights derived are aimed at supporting researchers, technology developers, and industry professionals to guide future integration of advanced, sustainable construction safety practices. Our review identifies recurring themes in the recent studies, including the integration of real-time sensing, multimodal data fusion, and user-centered system design. It also highlights unresolved challenges related to cross-project generalizability and long-term field adoption. By offering this structured analysis, the review aims to contribute a domain-wise benchmark of progress, capturing the recent technological inflection in construction safety and sustainability, and outlining future opportunities to support researchers, technology developers, and industry professionals.
LAser Detection And Ranging (LADAR) is increasingly used in fields like urban planning, autonomous driving, and object recognition, creating a need for better data processing tools. The ATLAS software now incorporates SuperPoint Transformer (SPT) to boost processing speed. A deep learning-aided lasso tool, SPLAT, was developed to improve object selection. It uses 3D region growing techniques, requiring a user to click on an object to generate a seed point. The SPT features are used to define similarity metrics, enabling iterative region expansion for object segmentation. This method enhances LADAR segmentation in the ATLAS labeling suite.
Accurate segmentation of the sky region is crucial for various applications, including object detection, tracking, and recognition, as well as augmented reality (AR) and virtual reality (VR) applications. However, sky region segmentation poses significant challenges due to complex backgrounds, varying lighting conditions, and the absence of clear edges and textures. In this paper, we present a new hybrid fast segmentation technique for the sky region that learns from object components to achieve rapid and effective segmentation while preserving precise details of the sky region. We employ Convolutional Neural Networks (CNNs) to guide the active contour and extract regions of interest. Our algorithm is implemented by leveraging three types of CNNs, namely DeepLabV3+, Fully Convolutional Network (FCN), and SegNet. Additionally, we utilize a local image fitting level-set function to characterize the region-based active contour model. Finally, the Lattice Boltzmann approach is employed to achieve rapid convergence of the level-set function. This forms a deep Lattice Boltzmann Level-Set (deep LBLS) segmentation approach that exploits deep CNN, the level-set method (LS), and the lattice Boltzmann method (LBM) for sky region separation. The performance of the proposed method is evaluated on the CamVid dataset, which contains images with a wide range of object variations due to factors such as illumination changes, shadow presence, occlusion, scale differences, and cluttered backgrounds. Experiments conducted on this dataset yield promising results in terms of computation time and the robustness of segmentation when compared to state-of-the-art methods. Our deep LBLS approach demonstrates better performance, with an improvement in mean recall value reaching up to 14.45%.
As climate change accelerates the transformation of polar regions, accurately monitoring changes in these critical environments has never been more urgent. In the case of the Arctic, delineating three vital classes - snow, open water and meltponds - is essential for tracking the influence of climate change on the ecosystem, and at large the global sea levels. In this paper, we explore the innovative use of Spiking UNet, an architecture that combines the technology of Spiking Neural Networks (SNNs) with the widely recognized UNet framework, to perform pixelwise classification on high-resolution imagery. This method follows the UNet architecture, where each block utilizes a double convolutional layer with ReLU and employs the Leaky Integrate-and-Fire (LIF) and Integrateand-Fire (IF) function as the activation for the final layer. The addition of LIF/IF in the final layer changes the computation paradigm from rate-based (ANN) to event-based (SNN), even if the rest of the model is CNN. The LIF and IF neuron brings in time dynamics, spiking behavior, and membrane potential accumulation, which are core neuromorphic properties. Both models are trained on a simulated environment, Norse. The performance of the model is evaluated on Healy-Oden Trans Arctic Expedition (HOTRAX) and NASA's Operation Icebridge datasets. Insight into both Spiking UNet models, is presented by performance comparison.
Deep learning models often suffer from performance degradation when applied to construction sites that differ from the source domain due to their sensitivity to data distribution shifts. Although methods such as transfer learning, domain adaptation, and synthetic data generation have been explored to improve generalization, collecting and annotating data from new target domains remains a labor-intensive bottleneck. This study presents a self-training-based framework to generate training data for construction object detection in unlabeled target domains. The method identifies moving objects using optical flow estimation, propagates class labels through iterative self-training, and synthesizes realistic training images via image inpainting and copy-paste augmentation. Experimental results from four visually distinct construction scenes demonstrate that the proposed method significantly improves detection performance without relying on manually labeled target data. These findings contribute to advancing automated and scalable domain adaptation techniques for vision-based construction monitoring.
Identification of precancerous polyps during routine colonoscopy screenings is vital for their excision, lowering the risk of developing colorectal cancer. Advanced deep learning algorithms enable precise adenoma classification and stratification, improving risk assessment accuracy and enabling personalized surveillance protocols that optimize patient outcomes. Ultralight Med-Vision Mamba, a state-space based model (SSM), has excelled in modeling long- and short-range dependencies and image generalization, critical factors for analyzing whole slide images. Furthermore, Ultralight Med-Vision Mamba's efficient architecture offers advantages in both computational speed and scalability, making it a promising tool for real-time clinical deployment.
3D perception has revolutionized the field of Computer Vision and ushered in an era where its use, specifically point clouds, is ubiquitous. However, despite its common place, the challenge of efficiently labeling point clouds in a quick manner remains a significant hurdle. We introduce leveraging the use of superpoints as part of an innovative labeling suite to augment current labeling techniques with state-of-the-art methods that improve efficiency and accuracy. Using superpoints (clusters of points with local features), the labeling suite shifts focus from a previously labor-intensive workflow that operates on individual points to a feature-guided workflow. This novel approach empowers annotators, especially with expansive point clouds, to engage with intuitive, meaningful entities, significantly reducing cognitive load and improving annotation quality by enhancing both the speed and accuracy of the annotators. Additionally, superpoints provide access to local geometric properties at multiple levels, allowing for an extension of classes via subclassification within a region of interest. Experiments performed confirm and demonstrate that the use of superpoints accelerates the labeling process, in large part by having the ability to utilize superpoints after unsupervised labeling for manual corrections, and quick subclassification leading to superior labeling and workflow outcomes. Furthermore, not only do superpoints set a new standard for data labeling in 3D environments, promising a future of both rapid and accurate data labeling, but they also pave a pathway forward for scalable point cloud annotation.
Rapid changes in polar regions due to global warming demand precise monitoring of surface melt dynamics, especially the formation of meltponds on the sea ice surface. As meltponds play a critical role in climate change by changing the under-ice light environment and associated productivity, identifying them in the Arctic is crucial. This work presents an Ultralight VisionMamba UNet segmentation framework that leverages state-space models (SSM) representing Mamba for meltpond localization in the Arctic area. At the heart of the architecture is the Parallel Vision Mamba (PVM) Layer, which enhances feature processing by dividing the channels and simultaneously feeding them into several Mamba modules. Mamba which is built upon SSM models is used as an efficient alternative to traditional convolutions or Transformers because these models operate with linear complexity relative to input size. The approach of integrating SSMs into the UNet architecture with the PVM layer can capture the complex spatial dynamics of meltpond formation in the imagery while reducing computational overhead. We trained the Ultralight VisionMamba model using high-resolution aerial images of the Arctic region captured during NASA's Operation IceBridge and the Healy-Oden Trans Arctic Expedition (HOTRAX). The Ultralight VisionMamba framework showed superior performance in segmenting meltponds compared to other state-of-the-art approaches.
As global warming drives climate change, interest has increased in the Arctic on the development of meltponds in seasonal sea ice. The lack of extensive annotated data on Arctic sea ice poses a significant challenge for training deep learning models aimed at predicting meltpond dynamics. This study employs a separable convolution diffusion model, a type of generative model, to create synthetic Arctic sea ice data. Diffusion models can create new and realistic data based on the image features of the original dataset by learning the data distribution and transitioning from a simple to a more complex distribution. By transforming a simple noise distribution with a series of denoising operations, it can generate high-fidelity synthetic images. These high-fidelity images contain similar image feature distributions found in the training data. Separable convolutions simplify the convolution operation by breaking it down into two smaller convolutions, one applied across the width and the other across the height of the image. Incorporating separable convolutions into the diffusion model maintains its performance while improving efficiency compared to traditional convolution methods, making it faster and more adept at producing synthetic images. We utilized high-resolution aerial images of the Arctic region collected during NASA's Operation IceBridge project (Digital Mapping System L1B Geolocated and Orthorectified data) obtained in 2016, to train the generative model. The synthetic images generated by the model are validated and compared with the original images using the UNet architecture, a deep learning model designed for precise image segmentation by utilizing skip connections between the encoder and decoder.
In this paper, we present a novel hybrid method for accelerated sky region segmentation that integrates Convolutional Neural Networks (CNNs) with the Lattice Boltzmann Level-Set (LBLS) model. By leveraging DeepLabV3+, Fully Convolutional Network (FCN), and SegNet, our approach guides the active contour and extracts regions of interest. We employ a local image fitting level-set function to characterize the region-based active contour model and utilize the Lattice Boltzmann Method (LBM) to achieve rapid convergence. This deep LBLS segmentation method not only preserves precise details of the sky region but also demonstrates improved robustness and computational efficiency. Experimental results on the CamVid dataset show that our proposed method outperforms state-of-the-art techniques, achieving a mean recall value improvement of up to 14.45%. Our method contributes to advancements in sky region segmentation, essential for applications such as object detection, augmented reality, and virtual reality.
We present a method to learn a diverse group of object categories from an unordered point set. We propose our Pyramid Point network, which uses a dense pyramid structure instead of the traditional 'U' shape, typically seen in semantic segmentation networks. This pyramid structure gives a second look, allowing the network to revisit different layers simultaneously, increasing the contextual information by creating additional layers with less noise. We introduce a Focused Kernel Point convolution (FKP Conv), which expands on the traditional point convolutions by adding an attention mechanism to the kernel outputs. This FKP Conv increases our feature quality and allows us to weigh the kernel outputs dynamically. These FKP Convs are the central part of our Recurrent FKP Bottleneck block, which makes up the backbone of our encoder. With this distinct network, we demonstrate competitive performance on three benchmark data sets. We also perform an ablation study to show the positive effects of each element in our FKP Conv.
We present a novel Fourier camera, an in-hardware optical compression of high-speed frames employing pixel-level sign-coded exposure where pixel intensities temporally modulated as positive and negative exposure are combined to yield Hadamard coefficients. The orthogonality of Walsh functions ensures that the noise is not amplified during high-speed frame reconstruction, making it a much more attractive option for coded exposure systems aimed at very high frame rate operation. Frame reconstruction is carried out by a single-pass demosaicking of the spatially multiplexed Walsh functions in a lattice arrangement, significantly reducing the computational complexity. The simulation prototype confirms the improved robustness to noise compared to the binary-coded exposure patterns, such as one-hot encoding and pseudo-random encoding. Our hardware prototype demonstrated the reconstruction of 4kHz frames of a moving scene lit by ambient light only.
SPIE is working with SAE International to develop lidar measurement standards for active safety systems. This multi-year effort aims to develop standard tests to measure the performance of low-cost lidar sensors developed for autonomous vehicles or advanced driver assistance systems, commonly referred to as automotive lidars. SPIE is sponsoring three years of testing to support this goal. We discuss the second-year test results. In year two, we tested nine models of automotive grade lidars, using child-size targets at short ranges and larger targets at longer ranges. We also tested the effect of high reflectivity signs near the targets, laser safety, and atmospheric effects. We observed large point densities and noise dependencies for different types of automotive lidars based on their scanning patterns and fields of view. In addition to measuring point density at a given range, we have begun to evaluate the point density in the presence of measurement impediments, such as atmospheric absorption or scattering and highly reflective corner cubes. We saw dynamic range effects in which bright objects, such as road signs with corner cubes embedded in the paint, make it difficult to detect low-reflectivity targets that are close to the high-reflectivity target. Furthermore, preliminary testing showed that atmospheric extinction in a water-glycol fog chamber is comparable to natural fog conditions at ranges that are meaningful for automotive lidar, but additional characterization is required before determining general applicability. This testing also showed that laser propagation through water-glycol fog results in appreciable backscatter, which is often ignored in automotive lidar modeling. In year two, we have begun to measure the effect of impediments to measuring the 3D point cloud density; these measurements will be expanded in year three to include interference with other lidars.
Although computer vision technology has shown great potential, its reliability can significantly degrade in the target domain where the model is applied. Collecting and labeling training data from the target domain can address this issue; however, it is a tedious and time-consuming task. To address this issue, this paper presents a novel method generating training data for construction site monitoring. The proposed method consists of extracting moving objects, classifying each region by comparing their features to the target classes, and then assigning class labels. The newly labeled data is copied and pasted to a clear background of the target domain where its foregrounds are removed by image inpainting. Experiments were conducted on construction site videos captured in far-field monitoring environments. The proposed method can significantly reduce the amount of effort required for data collection and labeling, thereby increasing the efficiency of developing robust computer vision models for construction site monitoring.
Construction safety monitoring increasingly relies on machine learning-based models due to their strong learning capability. However, these models' performance often degrades when they encounter shifts in data distribution. To address this issue, a tracking algorithm can be used for the detection model to enhance detection performance and ensure high-quality safety monitoring. Most importantly, utilizing such method can lead to robust personalized workplace risk assessment and customized safety training to workers. This paper proposes a novel object tracking approach, called hashing-supported cascaded buffered intersection over union (HC-BIoU) tracking, which addresses low tracking performance due to visual domain shifting. The experimental results demonstrate that the proposed method achieves significant improvements in mean average precision (9.3%) and association accuracy (42.6%) compared to a YoloV8 object detector and a Cascaded Buffered IoU (C-BIoU) tracker, tested on a video sequence of a scaffold dismantlement scene with 1,983 frames and 9,606 ground truth bounding boxes.
With the need for explainable AI, several visualization methods have been developed to explore neural networks. Our proposed work seeks to isolate and generate image features of input data using any neural network. We employed a modified gradient ascent algorithm with a smoothed gradient loss function for image-to-image feature generation.
Ming-Jung Seow合作论文数27