Digitalization of historical buildings is essential for their preservation and dissemination. However, existing digital models struggle to simultaneously achieve realistic visualization, semantic enrichment, and user-friendly interaction. Additionally, heritage scenarios face challenges of specialized semantics and data unavailability. Therefore, this paper proposes Heritage-3DGS, an MLLM-driven language-embedded 3DGS framework for generating realistic and semantically enriched digital models that integrate domain-specific terminology and enhance user engagement while minimizing manual effort. It comprises four steps: (1) collection of on-site images and component textual descriptions; (2) SAM-MLLM-based component segmentation, generating semantic masks with only one manually annotation per component; (3) optimized language-embedded 3DGS to efficiently reconstruct 3D semantic field enriched with domain-specific knowledge; and (4) chatbot integration for open-vocabulary and fuzzy searches. Validation experiments on two cathedrals in Guangzhou and Hong Kong, China, achieved average 84.15 % mIoU for domain-specific semantic field reconstruction and 0.92 SSIM for realistic scene representation, demonstrating its effectiveness and applicability.
Deep learning (DL)-based semantic segmentation of point clouds for construction site holds great promise for improving site management and progress monitoring. However, there is currently a scarcity of publicly available benchmarks that specifically capture the unstructured and dynamic characteristics of active construction phases (e.g., temporary facilities and partially built structures). To bridge this gap, this study constructs an expert-annotated ConSite dataset, which contains 13 categories of structural and temporary components commonly found on construction sites. Building on this dataset, this paper proposes SiteNet, a DL model designed specifically for semantic segmentation of point clouds in construction scenes. SiteNet integrates several key modules, including an adaptive transfer learning strategy, a Directional Feature Encoding Module, a video-based point cloud augmentation method, and a smoothed focal loss function. Experimental results show that SiteNet achieves an overall accuracy of 95.73% and a mean Intersection over Union of 86.32%, outperforming representative DL baselines such as Point Transformer. Ablation experiments and cross-validation further confirm the effectiveness and generalization ability of the proposed model. This study contributes to automation in construction by enabling intelligent construction site monitoring, progress tracking, and digital twin development through automated point cloud understanding.
Existing digitization of buildings with reflective glass facades suffers from geometric reconstruction distortion, unrealistic view-dependent texture rendering, and difficulties in object-based semantic enhancement. Therefore, we propose RefGlass-GS, a fusion framework that enables end-to-end UAV-based photorealistic, semantic, and interactive digitization of reflective glass facades. The contributions include: (1) proposing an individual glass panel segmentation method based on maximum a posteriori estimation with structural regularities, robust to severe reflection and background interference; (2) formulating a UAV viewpoint planning optimization function that maximizes the coverage of view-dependent appearance for sufficient data capture; (3) developing an optimized Gaussian Splatting framework with a Reflection MLP, a novel deferred shading function, and two enhanced regularization terms for effective modeling of high-frequency near-field reflections; (4) introducing a standardized data organization paradigm for structuring GS-based representations into object-based models, facilitating interactive facility management on digital twin platforms. Experiments on real-world reflective glass facade scenes validate the effectiveness and superiority of the proposed method. Specifically, the glass panel segmentation achieves an improvement of 0.1927 in mIoU over SOTA methods, and only our method enables instance-level panel extraction. The UAV view planning improves novel view synthesis for reflective facades by 13.15 dB in PSNR compared to commercially used nap-of-the-object planning methods. The RefGlass-GS modeling outperforms SOTA Gaussian Splatting approaches for reflective scenes with an average improvement of 5.08 dB in PSNR.
Mechanical, electrical, and plumbing (MEP) systems are critical for delivering essential services and ensuring comfortable environments. To improve the management efficiency of these complex systems, digital twins (DTs) that reflect the as-is conditions of facilities are increasingly being adopted. To generate DT models, laser scanners are widely used to capture as-built environments in the form of high-resolution images and dense 3D measurements. However, existing scan-to-BIM methods primarily produce basic geometric models, lacking detailed descriptive attributes of the components. To address this limitation, this paper proposes an informative DT model generation method for MEP systems based on fine-grained object recognition and object-aware scan-vs-BIM. The proposed method adopts a few-shot learning strategy to detect target objects in complex 3D environments and identify their family types based on vision foundation models. Following this, the association between as-designed components and as-built installations is formulated as a bipartite graph matching problem, which is solved using the Hungarian algorithm. This enables the automated updating of as-designed models into as-built DT models. Notably, the proposed association method is robust and applicable to components with significant installation deviations, a common challenge in MEP systems. The feasibility of the proposed approach was validated through experiments conducted on two construction sites in Hong Kong. Results demonstrated that the proposed approach significantly enhanced the accuracy of the scan-vs-BIM of MEP systems, thereby enabling informative DT model generation.
High-quality 3D reconstruction of existing buildings is essential for their maintenance, restoration, and management. Effective view planning for image collection significantly impacts the quality of photogrammetry-based 3D reconstruction. Intricate building structures, such as the overhangs, protrusions, and concave regions, can lead to under-sampled regions with traditional view planning methods, while excessively increasing the number of views require substantial computational resources and data collection efforts. To address these issues, this paper proposes a novel exploration-then-exploitation view planning strategy to achieve high-quality building reconstruction with minimal views. Firstly, the UAV no-fly regions and building attention regions are identified through semantic and geometric analysis of the images and coarse model during the exploration stage. Then, a novel optimization fitness function is mathematically formulated, considering building attention regions and reconstruction influential factors, including distance, incidence angle, parallax angle, and overlap. Furthermore, a modified sparrow search algorithm is proposed with the improved optimization mechanism and the integration of view planning physical model, enabling effective generation of optimal viewpoint set. Finally, the collision-free shortest trajectory is designed, allowing the UAV to collect images and reconstruct a high-quality model during exploitation stage. Experiments in virtual and real-world scenarios validate the effectiveness of our proposed modified SSA mechanism and the view planning strategy. Results demonstrate that the modified SSA achieves higher convergence accuracy and speed compared to the original SSA, PSO and GA. Our strategy can generate more accurate and complete 3D reconstruction models with the same or fewer captured images compared to commonly used and state-of-the-art strategies.
High spatio-temporal resolution street-level air pollution (SLAP) estimation is essential for urban air quality management, yet traditional methods face significant challenges in capturing the detailed spatial and temporal variability of pollution. Methods relying on fixed monitoring networks provide limited spatial coverage, while those utilizing mobile monitoring campaigns, despite their flexibility, often suffer from data sparsity and temporal incompleteness. To address these limitations, we propose a Two-Step Machine Learning Gap-Filling Framework employing a Multi-task Graph-based XGBoost (MTGXGB) model to enhance SLAP resolution. This framework expands high-resolution pollution estimation from a purely spatial perspective to a spatio-temporal view and effectively addresses data gaps. Our approach achieves spatial resolutions of 30-200 m and hourly temporal resolutions, capturing both short- and long-term variations in PM2.5 concentrations. Applying this framework to London's urban environment, we identify critical pollution hotspots and uncover correlations between SLAP, traffic speed, and urban environmental features. Additionally, the derived uncertainty maps provide actionable insights for optimizing mobile monitoring strategies. This study advances machine learning methodologies for spatio-temporal SLAP estimation and highlights the potential of high-resolution spatio-temporal SLAP data to inform policy-making, such as Low Emission Zones (LEZs), thereby demonstrating its practicality and scalability for urban air quality management.
Reflective glass facades are a popular choice in modern urban architecture, yet their 3D reconstruction and semantic digitalization pose significant challenges due to two key issues: (1) the view-dependent textures of reflective glass adversely impact geometric and texture reconstruction; and (2) texture interference, coupled with the lack of comparative materials, complicates the accurate segmentation of glass panels. Conventional 3D representations, such as meshes or point clouds, are limited to static textures and struggle to effectively demonstrate dynamic textures. Furthermore, existing glass segmentation methods, including zero-shot approaches like SAM and deep learning- based technologies, are insufficient for segmenting individual glass panels across an entire glass curtain wall. To address these challenges, this study proposes a novel framework for generating realistic and semantic 3D digital models of reflective glass facades using drone imagery. The scientific contributions include (1) a 2D glass panel segmentation method leveraging structural regularities and back-projection, adaptable to multi- view images with varying panel sizes; (2) the RefGlass-3DGS algorithm for generating 3D models with view-dependent textures and semantic annotations; and (3) an optimized UAV view-planning strategy based on full ray coverage, ensuring comprehensive capture of dynamic textures. Validation on a Hong Kong office building demonstrates the framework's effectiveness in reconstructing and semantically digitalizing reflective glass facades.
This study proposes a novel framework integrating Building Information Modeling (BIM) and Geographic Information Systems (GIS) with real-time crowd analytics from Closed-Circuit Television (CCTV) for quantitative walkability assessment. The framework extends open data standards (IFC and CityGML) to model infrastructural and pedestrian flow attributes comprehensively. A walkability scoring mechanism quantifies route quality based on accessibility, efficiency, and physical comfort, differentiating among pedestrian groups, such as individuals sensitive to weather conditions or carrying belongings. Implemented at the Hong Kong University of Science and Technology (HKUST), results indicate that the framework effectively captures variations in walkability scores due to directional differences (uphill vs. downhill), crowd conditions, and operational constraints like facility closures. Statistical tests confirm significant differences in walking costs across these scenarios with variations of up to 30%, demonstrating the framework’s robustness and practical utility for real-time, human-centric urban infrastructure planning.
Assessing building photovoltaic (PV) potential is crucial for urban energy planning and achieving Net-Zero Energy Building (NZEB) goals. However, urban-scale assessment still faces limitations. Most studies have neglected the importance of facade PV potential. Moreover, they often treat buildings as isolated entities, overlooking the impact of interactions between buildings, such as shading effects. To address these shortcomings, this study proposes a novel Shadow-Attention Graph Neural Network (SAGNN) to accurately predict solar irradiation for large-scale urban buildings. Analyzing 1.08 million buildings in New York City at a 1-meter spatial resolution, the study explores the potential for achieving NZEB and Net-Zero Electricity Building (NZEL) status. Results show that building rooftops and facades have annual PV power generation potential of approximately 29,851.9 GWh and 32,062.2 GWh, respectively. However, the PV potential is still insufficient to achieve the NZEB goal for the entire city. Nevertheless, utilizing both roof and facade PV can enable many areas to achieve NZEL. Cost-benefit analysis reveals that rooftop PV systems are more economically viable than facades, with a payback period of 7 years and a net-benefit of $71.98 billion over the 25-year life cycle. This research provides scientific decision support for urban PV planning and NZEB policy formulation.
UAV-based 3D reconstruction has proven to be an effective technique for assessing the current state of architectures during the maintenance stage. Many buildings are designed with irregular shapes to fulfill aesthetic requirements. These buildings have lots of occlusion regions and edge regions, which makes it difficult for most commercially available flight planners to perform high-quality 3D reconstruction. To address this problem, we propose a novel UAV viewpoint generation strategy of 3D flight planning. Firstly, we detect the occlusion regions and edge regions through mesh analysis, devoid of any dependence on prior knowledge or deep-learning techniques. This makes our algorithm applicable to various scenes. Besides, we formulate the viewpoints optimization problem based on triangles in the mesh model with the consideration of two regions and MVS influence factors. Then we propose the SSA-variant algorithm to directly optimize the viewpoints' location and orientation within a continuous flyable space. We validate our view planning strategy by assessing both the optimization convergence and reconstruction quality in a synthetic scene. The result shows that our SSA-variant algorithm performs higher convergency speed and accuracy than other optimization algorithms, such as GA and PSO. Compared to off-the-shelf flight planners and state-of-art 3D UAV path planning methods, images captured by our strategy can generate finer 3D reconstruction model for irregular-shaped architecture.
Wearable sensing technologies (WSTs) are valuable in monitoring status and behaviour of construction workers, providing insights into their response under varying conditions and potentially improving their performance. Despite their importance, a comprehensive review of WSTs for evaluating construction worker behaviour and status is lacking. This paper conducted a quantitative and qualitative review of relevant studies. A bibliometric analysis revealed the selected 200 publications between 2011 and 2023 focusing on musculoskeletal disorders, worker activity, worker status, construction safety and occupational risks. Accordingly, a knowledge framework was proposed for evaluating workers' status and behaviour, compassing data collection, artifact removal, analysis, worker evaluation, and applications. Following a qualitative review, six future research directions were identified: sensor selection and placement, experiment validity, end-to-end data analysis, data fusion, human-technology interaction, and modelling worker status. This review provides the current research state and future trends, aiding the practical implementations of wearable technologies on construction sites.
This paper presents an innovative and fully automatic solution of generating as-built computer-aided design (CAD) drawings for landscape architecture (LA) with three dimensional (3D) reality data scanned via drone, camera, and LiDAR. To start with the full pipeline, 2D feature images of ortho-image and elevation-map are converted from the reality data. A deep learning-based light convolutional encoder-decoder was developed, and compared with U-Net (a binary segmentation model), for image pixelwise segmentation to realize automatic site surface classification, object detection, and ground control point identification. Then, the proposed elevation clustering and segmentation algorithms can automatically extract contours for each instance from each surface or object category. Experimental results showed that the developed light model achieved comparable results with U-Net in landing pad segmentation with average intersection over union (IoU) of 0.900 versus 0.969. With the proposed data augmentation strategy, the light model had a testing pixel accuracy of 0.9764 and mean IoU of 0.8922 in the six-class segmentation testing task. Additionally, for surfaces with continuous elevation changes (i.e., ground), the developed algorithm created contours only have an average elevation difference of 1.68 cm compared to dense point clouds using drones and image-based reality data. For objects with discrete elevation changes (i.e., stair treads), the generated contours accurately represent objects' elevations with zero difference using light detection and ranging (LiDAR) data. The contribution of this research is to develop algorithms that automatically transfer the scanned LA scenes to contours with real-world coordinates to create as-built computer-aided design (CAD) drawings, which can further assist building information modeling (BIM) model creation and inspect the scanned LA scenes with augmented reality. The optimized parameters for the developed algorithms are analyzed and recommended for future applications.
Construction project governance (CPG) acts as a ‘steering wheel’ that keeps a construction project steady in a challenging external environment. CPG frameworks can be divided into two types: control-based hierarchical governance and trust-based relational governance. To increase cooperation and networking between the participants in a project, effective CPG depends on well-combined control and trust mechanisms. However, the current CPG framework, which is control-based, is biased toward hierarchical governance and can lead to a lack of trust in the construction project. These issues have resulted in several studies that have focused on blockchain technology (BT) as a potential enabler for addressing the low levels of trust in CPG. However, existing research in this area has failed to address the nature of the relationship between trust as a key CPG challenge and BT. Hence, this paper aims to identify trust as a major CPG challenge and potential BT capabilities to reveal the relationship, through a state-of-the-art review. The findings show that the effective use of three decentralized BT capabilities (decentralized data storage, validity, and access) can positively influence the three relational norms (mutuality, flexibility, and solidarity), which are the functional tools of relational governance. Thus, a more flexible trust-based CPG framework is proposed by establishing relational governance with blockchain technology. Ultimately, this study is expected to enhance the trust between the key stakeholders in construction projects and enable more efficient collaboration and networking between them.
To reduce machine-related accidents on sites, automatically monitoring the full-body poses of operating heavy machines is crucial. Conventional pose estimation systems relying on homogeneous sensors are vulnerable to negative environmental impacts, leading to inaccurate and unstable estimation of machine states. Hence, a full-body pose estimation framework is proposed for excavators, with a data fusion strategy to utilize different types of onboard sensors for enhanced accuracy and robustness. Specifically, a non-invasive onboard visual-inertial sensor system is designed for data fusion. Then, through competitive and complementary data fusion, the keypoints describing the full-body poses of the excavator are tracked in 3D space. Especially, an EKF-based localization algorithm is developed for optimized multi-keypoint tracking, which is verified to improve the accuracy and robustness of pose estimation by a real-world excavator case study. The proposed sensor-fusion method can effectively improve operational safety, by accurately monitoring the motion of heavy machines operating on construction sites.
Sewer pipes are essential infrastructure for discharging wastewater. Regular pipe inspection is necessary to prevent malfunction of sewer systems, for which closed-circuit television (CCTV) crawlers are commonly used to capture images of the pipe interior. As manual assessment of pipe condition is labor-intensive and time-consuming, automated defect detection using computer vision and deep learning has been increasingly studied in recent years. However, deep learning approaches require large amount of annotated data for model training. Data collection in underground sewer pipes is expensive and difficult since they are inaccessible without the use of an inspection robot. Meanwhile, ground-truth annotation needs to be accurate and consistent, requiring massive time and expertise. This paper proposes a framework for synthetic data generation and augmentation to address the data shortage problem for sewer pipe defect detection. First, synthetic images of sewer pipes are generated by 3D modeling and simulation in virtual environment. The quality of the generated images is then enhanced using style transfer with reference to real inspection images. In addition, a contrastive learning module is developed to further improve the deep learning process for defect detection. Experiment results show that the average precision (AP) of the defect detection model is improved by 2.7% and 4.8% respectively after adding style-transferred synthetic images and applying the contrastive module. When both methods are applied, the AP of the model is boosted by 7.7%, from 22.22% to 23.92%, indicating the effectiveness of our proposed approaches. This study is expected to alleviate the burden on data collection and annotation for applying deep learning models in defect detection.
With the increase of service years, external walls of high-rise buildings tend to suffer from a variety of defects which impose great safety risks. Traditional methods for inspecting high-rise building external walls require inspectors to work at height and identify defects manually, which is dangerous and inefficient. In recent years, there has been an increasing trend of using unmanned aerial vehicles (UAV) for inspecting building external walls, but how to manage the information obtained by the UAV is still a problem. In addition, although building information modelling (BIM) with rich geometric and semantic information has been applied in the construction engineering industry, BIM models usually lack updated condition data of facilities. Therefore, this paper presents a method for managing the inspection results of building external walls by mapping defect data from UAV images to BIM models and modelling defects as BIM objects. First, images of building external walls obtained by UAV are processed and useful information such as coordinates are extracted. Considering the small scale of single buildings, a simplified coordinate transformation approach is developed to transform location of real-world defects to coordinates in the BIM model. Meanwhile, a deep learning-based instance segmentation model is developed to detect defects in the captured images and extract their features. In the end, the identified defects are modelled as new objects with detailed information and mapped to the corresponding location of the related BIM component. To validate the feasibility, the proposed method has been applied to a real office building, which successfully mapped and integrated the defects of external walls with the BIM model. This study is applicable to both buildings and infrastructure, and is expected to facilitate structure inspection and decision making in maintenance with integrated data of as-is condition and as-built BIM.