Quality inspection (QI) technology, as a preliminary method for physical assembly, is increasingly applied in the assembly of complex steel structures. Traditional methods using terrestrial laser scanning (TLS) and total stations face limitations in accuracy and efficiency. This study proposes an automated TLS-based quality inspection method for prefabricated ring truss steel structures. A multi-threshold KNN algorithm is introduced for the automatic identification of bolt hole point clouds, providing more accurate bolt hole boundaries and registration features. Precise fitting of assembly points is then conducted, using these as registration features for multi-feature constrained global registration. Finally, the effectiveness of this method is validated through practical application in ring truss assembly.
Walking direction recognition is vital for HCI, healthcare monitoring, and navigation, but real-world data are highly imbalanced: non-straight trajectories far outnumber straight ones, biasing classifiers and reducing accuracy. We collect synchronized inertial and plantar-pressure signals from wearables and deliberately construct an imbalanced dataset to mirror practice. We then propose a multimodal framework that couples CNN-based spatiotemporal feature extractors with a DRL Q-network; a reward-driven optimization promotes balanced decisions under skewed label distributions. Experiments on the constructed dataset show consistent gains over state-of-the-art methods, with the largest improvements under severe imbalance, demonstrating robustness and suitability for deployment in wearable gait analysis systems.
Inertial sensing technology brings new possibilities for emotion recognition by addressing occlusion and spatial constraints in vision-based solutions. Nevertheless, fully exploiting the interactive potential of inertial sensors remains a challenge. In this article, we propose a supervised contrastive learning with graph neural networks (GNNs) and transformer (SCL-GT) architecture to recognize emotions from upper-body gestures using inertial sensors. Specifically, SCL-GT leverages GNN and transformer branches to extract complementary emotional features from motion data in the local structure view and global interaction view, respectively, and introduces a supervised contrastive learning (SCL) task to ensure consistency across cross-view representations. For the GNN branch, a graph representation strategy is developed to model the relationships among sensors, and a dynamic adjacency matrix is integrated into graph convolutional networks (GCNs) to improve upper-body gesture emotion recognition. Furthermore, we construct the EmoUP dataset, which captures eight emotions from the upper-body gestures of sixteen participants using five inertial sensors. Experimental results show that SCL-GT outperforms the state-of-the-art (SOTA) methods in terms of overall performance across both positive and negative valences on the EmoUP dataset.
Early diagnosis of skin cancer is crucial for improving patient survival. However, achieving high-precision classification remains challenging due to the complexity of lesion characteristics and the inter-class differences are relatively subtle. In this study, a Segmentation-Guided Classification Network (SGC-Net) is proposed to facilitate accurate lesion segmentation and effective multi-class skin lesion classification. During the lesion segmentation phase, the classical U-Net architecture is utilized to precisely delineate the region of interest (ROI) associated with skin lesions. By mitigating the influence of surrounding skin textures and background artifacts, the discriminative quality of the extracted features is substantially improved. In the subsequent feature extraction phase, a dual-stream architecture combining DenseNet201 and ResNet101 is developed to capture complementary representations. Furthermore, the Convolutional Block Attention Module (CBAM) is utilized to refine feature representations by adaptively emphasizing informative lesion regions. In the last stage of classification, the CosFace (large-margin cosine loss) function is introduced to refine the decision boundaries. The discriminative performance of lesion features is substantially enhanced by increasing inter-class margins within the cosine embedding space. The proposed approach is assessed using the HAM10000 dataset, comprising seven categories of skin lesions. Experimental findings confirm that the proposed fusion framework maintains superior robustness when addressing class imbalance and the complexity of skin lesion classification, suggesting its applicability as a reliable computer-aided tool for clinical dermatological diagnosis.
Cooperative path planning of multi-agent systems in complex environments is a pivotal issue in the field of intelligent control, and its optimization is of great significance for improving task execution efficiency and reducing overall system energy consumption. Aiming at the cooperative task allocation problem of unmanned swarms under multi-base and time-window constraints, traditional heuristic algorithms are prone to falling into local optima and incur high computational overhead in large-scale scenarios. To this end, an advanced Probabilistic Hybrid Q-learning with Single-Pass 2-opt (HQ-2opt) algorithm was proposed in this paper to address the multi-base swarm cooperative path optimization problem under strict time window constraints. By reformulating the routing decision as a Markov Decision Process, the proposed framework successfully transitions from traditional randomized swarm mutations to sequential reinforcement learning exploration. The integration of a dynamic exploration decay mechanism and a probabilistically triggered Single-Pass 2-opt operator effectively balances global state exploration and precise local route exploitation while avoiding recursive computational overhead. Simulation experiments demonstrate that the HQ-2opt algorithm converges to a highly superior total routing distance of 403.59 km, achieving a significant reduction of 104.81 km compared to the CIAPSO baseline. Furthermore, the proposed algorithm exhibits exceptional operational stability across multiple independent trials and demonstrates robust scalability in large-scale task network scenarios.Future research will focus on integrating real-time environmental uncertainties and multi-objective optimization constraints, such as energy consumption minimization, into the reinforcement learning environment to further enhance the algorithm’s applicability in complex unstructured environments.
The swift advancement of artificial intelligence technology indicates that large language models provide significant promise for comprehension and creative endeavors. Point cloud data has emerged as an essential technology resource in urban building, providing extensive information that facilitates many operations. Current research has thoroughly examined the utilization of point cloud data in semantic segmentation and target detection; nevertheless, the perceptual outcomes of these methods are sometimes challenging to implement in scene design directly. This study presents a novel 3D-MELL network architecture designed to enhance the limits of large-scale language modelling in processing 3D data. The architecture employs elements from ancient architectural components as the data source, with each element assigned distinct ID attribute markers and spatial relationship markers. These markers accurately represent the characteristics of the items and their interactions within the 3D environment. The model can be optimized by a distributed training technique to accommodate diverse downstream jobs with particular commands during the fine-tuning phase, demonstrating favorable training metrics and fitting outcomes. This project created a front-end page utilizing HTML and CSS frameworks to represent the chat interface, offering a novel approach to transmitting and developing architectural historical knowledge. (c) 2025 Published by Elsevier Masson SAS.
Wooden architectural heritage, an irreplaceable repository of cultural value, is highly vulnerable to factors such as component aging, biological infestations, and various forms of damage. Most current preservation methods primarily focus on salvage and repair after damage occurs, lacking the ability to predict changes in the structural components in advance. This study addresses this gap by constructing a digital twin behavioral model through the design of a TLSA-PSO prediction network grounded in the broader digital twin framework. Using typical ancient wooden architectural heritage in China as a case study, the validity and accuracy of the behavioral prediction model are verified. Additionally, a digital twin behavioral model visualization system is developed to display the prediction results. Experimental outcomes demonstrate that the behavioral prediction model can accurately forecast the behavioral changes of architectural heritage, achieving a goodness of fit of 0.99. This makes preventive protection of architectural heritage feasible.
Virtual Trial Assembly (VTA) is increasingly employed in engineering projects to reduce costs and improve assembly efficiency. However, the complexity of construction sites and the high precision requirements for assembly make manual point cloud segmentation both time-consuming and ineffective in accurately extracting assembly features. In order to solve the problem, this paper proposes a VTA method based on semantic segmentation and 3D bounding box estimation. It enables efficient semantic segmentation of large-scale point clouds and calculates precise assembly feature points based on the estimated 3D bounding boxes, facilitating optimal assembly analysis of all components. The results show that this method enables automated VTA analysis of assembly components and guides physical assembly. This research contributes to improving the monitoring of the entire construction process, ensuring efficient and precise assembly of building components.
As a national treasure, architectural heritage carries multiple value dimensions such as history, technology, art, and culture. With the increasing demand for architectural heritage protection and utilization, the traditional static digital model of architectural heritage based on geometric expression can no longer meet the practical application of multi-stage and multi-level scenarios. To this end, this paper proposes a value-chain-driven multi-level digital twin model of architectural heritage. Based on the three-stage logic of protection, management, and dissemination of value-chain classification, it integrates four types of models: geometry, physics, rules, and behavior. Combined with different hierarchical application levels, the digital model of architectural heritage is refined into a VCLOD (Value-Chain-Driven Level of Detail) detail hierarchy system to achieve a unified expression from spatial form restoration to intelligent response. Through the empirical application of three typical scenarios: the full-area guided tour of the Forbidden City, the exhibition curation of the central axis and the preventive protection of the Meridian Gate, the model shows the following specific results: (1) the efficiency of tourist guidance is improved through real-time personalized path planning; (2) the exhibition planning and visitor experience are improved through dynamic monitoring and interactive management of the exhibition environment; (3) the predictive analysis and preventive protection measures of structural safety are realized, effectively ensuring the structural safety of the Meridian Gate. The research results provide a theoretical basis and practical support for the systematic expression and intelligent evolution of digital twins of architectural heritage.
High-precision road point cloud measurement using mobile LiDAR technology is essential digital infrastructure for various industries. Researchers focus primarily on developing high-precision automated semantic segmentation for road point clouds. Existing deep learning networks trained on uneven and sparse point clouds captured by self-developed Mobile LiDAR Systems (MLS) have low segmentation accuracy. This paper introduces a deep learning method that partitions data based on the spatial positions of road scene point clouds and considers the sampling radius of regional groups. We use a road point cloud dataset constructed with a self-developed MLS to train and test the semantic segmentation of road point clouds. Based on the linear characteristics of local road point clouds, Principal Component Analysis (PCA) and threshold filtering methods are applied to classify the point cloud into ground and non-ground points. Different sampling strategies are then employed for each class of points, which are subsequently fed into the network model for semantic segmentation. Experimental results show that the proposed method achieves an overall accuracy of 97.8
To address the degradation of visual simultaneous localization and mapping (VSLAM) pose estimation accuracy caused by extensive moving objects and the resulting contamination of feature points in indoor dynamic environments, this study proposes a dynamic feature point rejection strategy that integrates a You Only Look Once (YOLO) family detector, YOLO11, with the Lucas–Kanade (LK) optical flow method within the ORB_SLAM3 framework. First, YOLO11 is employed to detect potentially dynamic targets such as pedestrians in the image, and ORB features extracted within the detection bounding boxes are labeled as high-risk dynamic points. Subsequently, based on pyramidal LK optical flow, the motion consistency of feature points is analyzed across multiple frames, and a secondary discrimination is performed on features located at the edges and in the vicinity of the detection boxes. In this way, feature points that conform to a static motion model are maximally preserved, thereby alleviating the excessive removal and misclassification of static points caused by oversized detection boxes. On this basis, a new map is reconstructed using the filtered static feature points, and the pose estimation process is further refined. Experiments on multiple indoor dynamic scenarios demonstrate that, compared with the original ORB_SLAM3, the proposed method effectively suppresses dynamic interference, enhances the stability of feature points, and significantly improves pose estimation accuracy and trajectory continuity, thus providing a valuable reference for robust VSLAM applications in dynamic environments.
This study proposes an innovative method that integrates multi-source remote sensing technologies and artificial intelligence to meet the urgent needs of deformation monitoring and ecohydrological environment analysis in Great Wall heritage protection. By integrating synthetic aperture radar (InSAR) technology, low-altitude oblique photogrammetry models, and the three-dimensional Gaussian splatting model, an integrated air–space–ground system for monitoring and understanding the Great Wall is constructed. Low-altitude tilt photogrammetry combined with the Gaussian splatting model, through drone images and intelligent generation algorithms (e.g., generative adversarial networks), quickly constructs high-precision 3D models, significantly improving texture details and reconstruction efficiency. Based on the 3D Gaussian splatting model of the AHLLM-3D network, the integration of point cloud data and the large language model achieves multimodal semantic understanding and spatial analysis of the Great Wall’s architectural structure. The results show that the multi-source data fusion method can effectively identify high-risk deformation zones (with annual subsidence reaching −25 mm) and optimize modeling accuracy through intelligent algorithms (reducing detail error by 30%), providing accurate deformation warnings and repair bases for Great Wall protection. Future studies will further combine the concept of ecological water wisdom to explore heritage protection strategies under multi-hazard coupling, promoting the digital transformation of cultural heritage preservation.
A vast population of visually impaired individuals is currently facing intricate life challenges, particularly related to perceiving walking directions. Therefore, this paper proposes a novel deep learning method based on wearable sensors to address the problem of walking direction recognition. The information mining and fusion module, the multi-feature position information mining attention module, and the multi-feature content information mining attention module are proposed to comprehensively mine comprehensive information from walking data. To overcome the limitation of information gathered from a single type of sensor, this paper combines inertial sensors and pressure insoles for walking direction recognition. Experimental results demonstrate that compared to existing research methods, the proposed method in this paper achieves a higher recognition accuracy highlighting the superiority and effectiveness of this method.
With the increasing dependence of the logistics industry on plastics, the amount of plastic waste has gradually risen, leading to environmental problems and resource wastage. Due to the high absorption of visible light and near-infrared wavelengths by black plastics, most common optical technologies struggle to effectively sort black plastics, and traditional manual classification methods are inefficient and inaccurate. To address this issue, we propose an efficient algorithm for the identification of black plastics. First, Raman spectroscopy is used to analyze black plastic and obtain relevant data. Then, principal component analysis (PCA) is applied to reduce the dimensionality of the spectral data and extract effective features. Next, a classification algorithm combining Radial Basis Function Neural Network (RBFNN) and Neural Gas Network (NGN) clustering is designed. This algorithm integrates the adaptive learning ability and pattern recognition capability of neural networks, enabling efficient classification of black plastics. Experiments were conducted using a Raman dataset, and the proposed algorithm was compared with three methods: RBFNN based on MinMax, Fuzzy C-Means (FCM), and Hard C-Means Clustering (HCM). Additionally, experiments were also performed on the publicly available datasets of Glass, Abalone, MyHog, and Page. The experimental results show that the proposed algorithm demonstrates superior performance across all datasets.
Two-person interactions are prevalent in public places, and how to recognize the friendly and violent two-person interactions is crucial for ensuring public safety. However, there are three challenges in two-person interaction recognition: the difficulty in capturing interaction relationships, the interference of interaction-irrelevant actions, and the distortion of the spatial relationship caused by camera view variations. In this paper, we propose an end-to-end multi-stream feature fusion model integrating Convolutional Neural Network (CNN) and Long Short-Term Memory (LSTM) to capture interaction dynamics and perform fine-grained analysis of interactive actions. The model aims to overcome the limitations of insufficient interaction feature representation through hierarchical fusion of heterogeneous data streams. Meanwhile, to suppress interaction-irrelevant features, we propose a Mutual-Self Activity-level Evaluator (MSEr) to highlight interaction-related features that simultaneously represent individual actions and interaction semantics. Furthermore, to characterize view-invariant features, we employ the 2D Discrete Wavelet Transform (2D-DWT) to analyze action within the spatial-frequency domain. Experimental results demonstrate that the proposed method obtains competitive performance among the compared methods with accuracies of 98.3 +/- 0.2%, 94.6 +/- 0.1%, and 89.3 +/- 0.1% on the SBU-Kinect, NTU RGB+D 60, and NTU RGB+D 120 datasets respectively.
In recent years, walking direction recognition has demonstrated extensive application prospects in fields such as rehabilitation healthcare, sports assessment, fall prevention, and biometrics. Walking direction recognition is a technology that captures and analyzes human walking patterns to assess motion characteristics, health status, and behavior patterns. Therefore, in this paper, we propose a dual-stream CNN model based on the combination of lightweight AlexNet and SeBlock to recognise different walking directions and investigate the rationality of classifying different walking directions. The study conducted several experiments, including overlapping features of curved walking and deviation walking, overlapping features between left deviation and right deviation, and a four-classification experiment, which involved straight walking, left deviation, right deviation and curved walking. The results revealed significant feature overlap between curved walking and deviation walking, while left deviation and right deviation exhibited minimal overlap and could be merged into one class. This conclusion suggests that, in practical applications, classification strategies can be adapted according to specific conditions to achieve optimal resource utilization. Moreover, the highest accuracy of 89.5% was achieved in the four-classification experiment. The lightweight model proposed in this paper avoids the complex data processing and model optimization process, and the experimental results show the advantages and disadvantages of different classification methods in walking directions, and provide an effective reference for subsequent research.
Traditional visual SLAM methods are built on the strong assumption that the system operates in static environments, with limited consideration of moving objects. This assumption often leads to significant performance degradation when dynamic elements are present. To mitigate the impact of moving objects and enhance both localization accuracy and mapping quality, we propose a visual SLAM framework that explicitly removes dynamic object interference from the visual odometry and mapping modules. First, we refine the data association process in visual odometry by introducing motion consistency constraints, which reduce incorrect feature matches and thereby improve pose estimation accuracy. At the same time, depth information from RGB-D sensors is used to validate potentially dynamic feature points. Second, within the mapping module, we formulate keyframe selection as a vertex cover problem to ensure the local representativeness of keyframes. This approach not only reduces mapping artifacts but also enables the comprehensive detection and removal of dynamic objects. Finally, experiments conducted on the TUM RGB-D dataset demonstrate that our system achieves higher accuracy, robustness, and stability compared to baseline methods.
Traditional approaches to architectural heritage conservation often prioritize emergency maintenance, which lacks the capability to monitor and adapt to the dynamic, ongoing changes in structural integrity, thereby compromising long-term conservation objectives. This study introduces a digital twin five-dimensional model framework specifically tailored for architectural heritage conservation, designed to facilitate a shift from reactive to proactive conservation strategies. Incorporating dimensions of physical entity, digital twin, twin data, service, and connection relationships, this framework enables real-time situational awareness and hyper-realtime virtual extrapolation. Through the reconstruction of the historical building information model and the development of both a structural linkage rule model and a displacement trend model based on mechanical properties, the framework offers a comprehensive digital representation from geometric, physical, behavioral, and rule-based perspectives. Validated through a prototype system at the Yingxian Wooden Pagoda in China, a site susceptible to natural and anthropogenic hazards, the findings demonstrate that the digital twin effectively predicts and adapts to spatio-temporal dynamic changes in the pagoda's structure, enhancing resilience and promoting sustainable heritage conservation. This research supports disaster risk reduction strategies, contributing to vulnerability analysis and resilience in heritage conservation amid increasing threats from climate change and other emergent risks.