
Short-horizon vehicle state prediction has become a basic computational capability in intelligent transportation, where connected and automated vehicles, roadside sensing, and cooperative traffic management depend on forecasts that are accurate, dynamically consistent, and robust to operating-condition shift. Purely data-driven predictors can capture rich temporal structure, yet their rollouts may violate vehicle motion constraints; physics-based models remain interpretable, but their reliability deteriorates when tire–road interaction and environmental conditions depart from calibration. This paper presents DriveKAN, a Kolmogorov–Arnold (KA)-enhanced framework for multi-step vehicle state prediction. DriveKAN combines a structured encoder–decoder built on learnable univariate basis functions, a differentiable single-track consistency residual with bounded learnable physical parameters that jointly regularizes force–moment balance and ground-frame kinematics, and an unsupervised cross-domain adaptation branch that integrates adversarial alignment, kernel matching, and physics-guided pseudo-label consistency. Under the controlled simulation protocol coupling Simulation of Urban MObility (SUMO) with the Traffic Control Interface (TraCI), DriveKAN reduces average displacement error (ADE) by 6.4% relative to the strongest purely data-driven baseline and reduces the cross-domain transfer gap from 112.16% for the source-only sequence model to 28.14%. Independent hardware-in-the-loop (HIL) reference derivatives and an evaluation-only dynamics model with load transfer and lateral aerodynamics provide separate physical evidence, showing reductions of 19.5% in lateral-acceleration error and 20.1% in extended force–moment defect relative to the strongest learned baseline. Under SUMO-to-HIL transfer, unlabeled adaptation reduces average cross-platform degradation from 62.5% to 29.7% and attains the lowest average HIL state error across nominal and perturbed operating conditions. The resulting framework brings expressive sequence modeling, prior-regularized dynamics, and cross-domain robustness into a single predictor for safety-relevant intelligent transportation scenarios.
Phase-based displacement measurement has attracted increasing attention in structural vibration monitoring because of its robustness to illumination variations and high subpixel accuracy. The Gabor wavelet is one of the earliest and most widely used phase extraction methods for phase-based displacement measurement. When Gabor wavelets are used for phase extraction, measurement performance is highly sensitive to the center frequency, while its selection in most previous studies relies on comparison with ground truth. To overcome this limitation, this study proposes a motion-region-based center frequency optimization method for Gabor wavelets in phase-based displacement measurement. Numerical experiments on numerical videos with different target motion characteristics show that the optimal center frequency varies systematically with these characteristics. In particular, the results reveal a clear relationship between the optimal center frequency and the target motion region, with larger motion regions generally requiring lower center frequencies for accurate displacement measurement. Based on this finding and the physical interpretation of the Gabor wavelet, a physics-inspired relationship between the optimal center frequency and the motion region is established, and an iterative center frequency optimization method is developed accordingly. The proposed optimization method is validated through a laboratory shaker test and an outdoor seismic response measurement of a cold-formed steel wall structure.
This paper presents a topology-adaptive physics-informed graph-to-design neural network (PI-GNN) framework for rapid nonlinear minimal-mass design of cable-strut structures. Although direct nonlinear optimization can yield rigorous solutions for individual feasible designs, repeated case-by-case optimization becomes computationally expensive in topology-level exploration, where numerous candidate structures must be evaluated and infeasible layouts or poor initial guesses often lead to failed or slow convergence. To address this challenge, the proposed framework reformulates repeated nonlinear optimization as a reusable graph-to-design learning problem that directly maps candidate structural graphs to preliminary prestress and member-level sectional designs. Variable-size graph modeling enables a unified network to accommodate different structural configurations. A feasibility screening network first evaluates candidate structures, after which the design network predicts prestress levels and member-level cross-sectional variables. An auxiliary displacement network estimates structural responses under multiple load cases, enabling a differentiable mechanics evaluator to quantify engineering-constraint violations. These violations are incorporated into augmented-Lagrangian physics losses and backpropagated to guide the design network toward lightweight and mechanically admissible solutions. Numerical studies on a planar photovoltaic cable truss and a spatial Levy cable dome demonstrate the accuracy, physical consistency, and computational efficiency of the proposed method for topology-level design exploration.
Engineering systems increasingly rely on data-driven optimization and closed-loop decision-making. Accordingly, computational models in engineering are no longer only tools for forward analysis, but also need to support parameter updating, design optimization, and feedback decision-making. In contrast, existing pavement mechanics computation still mainly serves response prediction under given parameter conditions and lacks a dedicated differentiable physics solver that directly provides stable gradient information. To address this gap, this study develops ∂Pave, a differentiable physics solver for pavement mechanical response computation. ∂Pave builds on a spectral element forward solution framework for layered pavement systems under moving loads. The solver connects load, material, geometric, and evaluation parameters to explicit stages of an end-to-end differentiable spectral element method solution path. Within this unified physical framework, ∂Pave returns selected mechanical responses at specified points in the layered pavement coordinate system together with their gradients with respect to specified input parameters. This study verifies the computed gradients through examples involving multiple response quantities. The results show that the automatic differentiation gradients from ∂Pave are generally consistent with finite difference results. In multi-parameter joint differentiation tasks, the efficiency advantage of ∂Pave over parameter-by-parameter finite differences increases with the number of differentiated parameters. This study further illustrates the use of the same response-gradient interface in two gradient-driven pavement engineering workflows, namely pavement structural design optimization and measurement-point layout optimization for pavement inspection. The two examples use task-specific variables, objectives, constraints, and decision procedures outside the physics solver. They demonstrate that ∂Pave can serve as a reusable physics-based response-gradient provider for different pavement mechanics task formulations.
Automated crack analysis remains largely reactive: most methods delineate damage only after it has been observed, whereas maintenance planning would benefit from evidence about how visible crack patterns may evolve between inspections. Forecasting future crack traces is difficult because crack pixels are sparse and thin, small discontinuities can change structural interpretation, and uncertainty increases with prediction horizon. We introduce the Crack Predictor benchmark and ASPECT (Adaptive Spatio-temporal Prediction of Evolving Crack Traces), a non-autoregressive framework that predicts four future crack masks directly from 16 observed RGB frames. ASPECT employs a compact StarNet spatial backbone augmented with Haar Wavelet Downsampling for frequency-aware encoding, ConvLSTM for temporal modeling on a compressed two-dimensional feature grid, and a Geometry-adaptive Decoding Head for reconstructing thin, irregular crack masks. On the validation split, ASPECT achieves 90.33% mIoU, 94.66% mF1, and 80.91% crack-class IoU with 0.32 M trainable parameters. It also ranks first in topology sensitivity and endpoint precision, recall, and F1, reaching an endpoint F1 of 66.56%. Across the four prediction horizons, ASPECT retains the highest Crack F1 among the evaluated methods. These results establish ASPECT as a compact, topology-aware framework for short-horizon crack forecasting on the controlled multi-temporal benchmark.
In structural health monitoring (SHM) of long-span bridges, distributed optical fiber sensing technology can provide strain information with high spatial resolution. However, due to the inherently low sampling frequency of the system, it is difficult to capture the high-frequency dynamic strain responses of bridges under complex operational environments. To address this hardware bottleneck, a hybrid-driven method based on dynamic and static responses for dynamic girder strain reconstruction of long-span suspension bridges is proposed. First, a joint state estimation framework based on variational mode decomposition (VMD) and a Kalman filter is established. This framework utilizes high-frequency acceleration as a dynamic prior while treating low-frequency deflection measurements as absolute constraints, effectively resolving the persistent issue of low-frequency drift caused by uncertain initial conditions in double integration. Second, a multi-task BP neural network is introduced to decode the complex, non-linear spatiotemporal mapping between global girder deflection and localized strain fields, thereby mapping the reconstructed dynamic deflection field to high-frequency dynamic strains at multiple cross-sections of the entire bridge.. The proposed method is validated using numerical simulation and actual monitoring data of a long-span suspension bridge. Results demonstrate that compared to conventional identification methods relying solely on acceleration integration, the proposed approach reduces the maximum root-mean-square error (RMSE) and mean absolute error (MAE) of the reconstructed dynamic strains by up to 50% and 51%, respectively. Furthermore, under Gaussian white noise interference as high as 15%, the method still maintains highly consistent strain reconstruction trends, demonstrating excellent noise robustness.
For multi-support construction equipment operating in complex environments, load and configuration variations can induce significant whole-equipment CoM fluctuations that directly affect operational stability. Existing stability analysis methods are generally unable to simultaneously accommodate multiple constraints or explicitly account for CoM fluctuations, limiting their applicability to operation planning. To address these limitations, this paper proposes a multi-constraint three-dimensional quasi-static stability analysis and operation planning framework for multi-support construction equipment. The CoM fluctuation region induced by load and configuration variations is first characterized and regularized using a minimum enclosing sphere. A three-dimensional quasi-static stability boundary is then constructed by jointly considering force and moment equilibrium, kinematic constraints, actuation constraints, support conditions, and ground conditions. Together, these formulations transform equipment stability analysis into a feasibility problem of the CoM position and enable stability to be directly incorporated into operation planning. The proposed operation planning framework is applied to support-layout evaluation for static-support equipment, and to support-point feasibility evaluation, mobile-body target position and trajectory planning for the walking transition of mobile-support equipment. Simulations and prototype experiments on flat, uneven, and sloped terrain show that a quasi-static stability boundary and a trajectory between two target positions can be generated in approximately 2 s and 3.5 s, respectively, while achieving zero boundary violations under the adopted discrete verification. These results demonstrate the effectiveness of the proposed framework, which has the potential to further enhance the practicality, robustness, and autonomy of multi-support construction equipment operating in complex environments.
Reliable three-dimensional (3D) scene understanding is essential for robotic systems operating in complex, unstructured environments, such as construction sites. Existing methods, however, struggle with limited training data and rarely quantify epistemic uncertainty, often producing overconfident yet erroneous predictions that compromise operational safety. This paper introduces Probabilistic Gaussian Grouping (PGG), an uncertainty-aware two-dimensional-to-3D lifting segmentation approach that embeds probabilistic identity features directly within 3D Gaussian Splatting primitives. The approach leverages multi-view consistency alignment and Kullback–Leibler-regularized distributions via Monte Carlo sampling to explicitly capture ambiguity in the feature space. Experiments across four construction scenes characterized by uneven illumination and adverse weather conditions demonstrate strong performance, with PGG achieving up to 87.27% mean Intersection over Union (mIoU) and 81.09% mean boundary IoU (mBIoU), exceeding state-of-the-art baselines by 14% mIoU. Complementary uncertainty analyses further indicate superior calibration, with Mutual Information and Adaptive Calibration Error reaching 0.019 and 0.0140, respectively. These results establish a calibrated and trustworthy perceptual basis for robotic applications, including risk-aware motion planning and reliable spatial interaction within construction environments.
Ground penetrating radar (GPR) is widely used for subsurface inspection, yet its interpretability is often degraded by clutter induced by diverse electromagnetic interference conditions. Existing deep learning-based methods are typically tailored to specific clutter types and show limited robustness across different inspection scenarios. To address this challenge, this paper proposes MoA-RefineNet, an adaptive GPR data processing framework that integrates a shared RefineNet backbone with a Mixture-of-Adapters (MoA) mechanism to enable scattering-aware feature refinement within a unified network. Experiments on three representative interference conditions demonstrate improved reconstruction accuracy and structural fidelity compared with baseline configurations. Validation on field GPR data further shows enhanced target visibility and improved interpretability of subsurface structures and imaging results. These results indicate that MoA-RefineNet provides an adaptive and interpretable solution for GPR data enhancement under diverse electromagnetic interference conditions, with potential for practical infrastructure inspection.
Compaction quality control is essential for the safety of rockfill dams, yet traditional contact-type inspections suffer from invasiveness, spatiotemporal sparsity, and labor intensity. This study develops a non-intrusive visual framework for assessing rockfill compaction states under field conditions. The Rockfill Surface Compaction Condition (RSCC) dataset is established, systematically categorizing the visual characteristics of complex surface states throughout the construction lifecycle to provide a benchmark dataset for this task. Building on RSCC, a Transformer-based visual detection model, RockViT, is proposed to achieve precise classification of rockfill surface states. To address the challenge of long-tailed class distribution, a generative data augmentation strategy is employed using a Stable Diffusion (SD) model fine-tuned with Low-Rank Adaptation (LoRA). This process is optimized via Distribution-Alignment Parameter Optimization (DAPO) to improve the utility of the generated samples for minority-class learning. Experimental results demonstrate that RockViT outperforms the state-of-the-art (SOTA) learning models, achieving an overall accuracy of 96.24 % on the baseline dataset. Grad-CAM analysis indicates that RockViT captures both global context and subtle discriminative surface cues. The proposed framework offers a fast, accurate and end-to-end solution for identifying rockfill compaction states, particularly for defect identification. By linking visual recognition with physical validation, it supports targeted engineering actions and helps reduce detection difficulty and workload.
In practical engineering, high-resolution (HR) imaging devices have become increasingly utilized for capturing structural surface crack images. However, the effectiveness of current deep learning (DL) segmentation models in accurately predicting refined masks for HR crack images is hindered by the discrete sampling methods inherent in traditional DL architectures and the limited computational resources of GPUs. To tackle this issue, this investigation incorporates the point-based rendering methodology originating from computer graphics disciplines into the encoding-decoding framework, introducing an innovative Crack Boundary Point Rendering Network (CBPRN). The CBPRN endeavors to accomplish elaborate delineation of crack visual samples possessing resolutions surpassing 4K. Initially, an edge feature extractor integrated with a super-resolution encoder is devised to guide rendering heads in efficiently focusing computational power on ambiguous edge regions. Subsequently, a rendering-based prediction head is introduced with the function of efficiently sampling rendering points for the training and inference phases, respectively. Furthermore, a tailored composite objective function is deployed to enhance the learning procedure, enabling the architecture to equilibrium substantial disparities in pixel counts among positive and negative instances within crack visual data. Ultimately, to substantiate the practical applicability of the CBPRN, an on-site crack identification investigation was executed on an actual bridge structure located in Changsha utilizing an unmanned aerial vehicle (UAV). The CBPRN demonstrated remarkable effectiveness on 4K-resolution visual samples acquired by the unmanned aerial vehicle, attaining achieving overlap ratio (Intersection over Union, IoU), average boundary precision (mean Boundary Accuracy, mBA), and Dice similarity index metrics of 85.46%, 86.00%, and 92.16%, correspondingly. This outstanding effectiveness improves both the operational security and processing efficiency of unmanned aerial vehicle-assisted crack assessment procedures, offering enhanced flexibility in choosing flight trajectories for the inspection workflow.
Safeguarding tunnel boring machine (TBM) cutterheads against deep-seated, small-scale rock hazards remains challenging for low-frequency ground-penetrating radar (GPR) due to attenuation, clutter, and reflection-dominated wavefields that obscure weak diffractions. This study proposes a diffraction-centric inversion framework that integrates multi-velocity migration, edge-guided initial model construction, and an optimized diffraction wavefront-driven inversion (Optimized-DWI) to enhance the contribution of localized discontinuities to permittivity reconstruction. Extensive synthetic benchmarks, including complex spires and Marmousi models with multi-scale sensitivity analysis, and field validation at nuclear/thermal power plant tunnel sites confirm that the method delivers superior resolution and boundary delineation compared to conventional approaches. Recovered rock hazard positions and depth intervals align closely with borehole logs, demonstrating robust performance in wavefield illumination and inversion resolution under real-world interference. This paradigm offers a diffraction-centric utilization of wavefields and advances GPR-based detailed geotechnical investigation for hazard warning and TBM steering guidance.
Real near-fault pulse-like ground-motion records are scarce and unevenly distributed over magnitude, rupture distance, and site conditions, limiting the use of generative artificial intelligence for pulse-like time-history generation. This study proposes a two-stage generative framework for near-fault pulse-like velocity histories. Complete velocity records are decomposed into pulse and residual components and modeled separately in the wavelet-packet domain. In Stage 1, pulse center time, dominant period, local energy, and pulse peak velocity are predicted from magnitude, rupture distance, and , and a conditional variational autoencoder generates the local pulse wavelet-packet patch. In Stage 2, a conditional diffusion model learns residual wavelet-packet coefficients, with non-pulse residual pretraining followed by pulse-residual fine-tuning to improve stability under limited samples. Complete velocity histories are then reconstructed through pulse-arrival alignment, pulse–residual energy-ratio matching, and screening based on pulse-identification criteria. Experiments using 201 near-fault pulse-like velocity records from the NGA-West2 database show that the proposed method reproduces pulse parameters and outperforms traditional simulation methods in amplitude attenuation, energy accumulation, response spectra, and Fourier spectral characteristics. Under both testing-set mrv conditions and randomly sampled mrv conditions, the generated motions satisfy pulse-identification criteria and exhibit reasonable nonstationary acceleration, velocity, and displacement histories. These results indicate that the proposed framework provides an effective data-driven approach for generating near-fault pulse-like ground motions when recorded pulse-like data are limited.
The determination of tunnel locations is a critical aspect of railway planning and alignment design, especially in mountainous regions whose complex terrain challenges traditional methods. To address this issue, a tunnel location determination method based on detailed terrain analysis is proposed for mountain railway alignment design that integrates spectral feature analysis with multi-constraint optimization. First, terrain data are parameterized as spatial waves using a spatial Fourier transform, which abstracts the topographic information into frequency-domain representations. Low-pass filtering is then applied to retain macro-scale terrain features while suppressing high-frequency insensitive topographic data, thereby highlighting key topographic structures relevant to long tunnel locations. Subsequent steps involve further valley identification via local ternary patterns and density-based clustering to refine candidate tunnel portal regions. Afterward, feasible tunnel portal combinations are screened and connected through a customized shortest path search algorithm considering multiple constraints, with generated paths further optimized into smooth railway alignments. The proposed method is validated through a complex mountain railway case. The sensitivity analysis finds that the best low-pass filter cutoff frequency for identifying major tunnel locations is 1/2,000 m⁻¹ in this example. Then, compared with the best manually-designed solution, the proposed method can find an alternative that reduces the overall engineering cost by 5.71%, demonstrating its effectiveness in solving real-world problems.
UAV-based road damage detection is essential for intelligent infrastructure maintenance, as it achieves continuous monitoring and timely assessment of pavement conditions. Existing approaches commonly formulate this task as a object detection problem, while overlooking a critical structural prior: road damage inherently occurs within road regions. Consequently, object-detection-based predictions often suffer from limited spatial credibility, particularly in complex scenes containing shadows, lane markings, and other damage-like visual patterns. To address these limitations, this paper proposes CoMT, a coupled multi-task learning algorithm that simultaneously performs road damage detection and road region segmentation in an end-to-end manner. Specifically, a coupled multi-task branch is designed to promote cross-task collaboration through bidirectional feature sharing and interaction. In this design, road-region structural cues serve to constrain damage localization, while damage-related semantic features provide complementary contextual information for region segmentation. Moreover, an orthogonal attention module is introduced to strengthen direction-sensitive damage representations by modeling feature consistency and discrepancy along horizontal and vertical orientations. Furthermore, a spatial-distance fusion loss is developed by integrating Complete IoU and Normalized Wasserstein Distance, enabling the joint optimization of spatial overlap and distributional similarity for more robust bounding-box regression. Experimental results demonstrate that CoMT achieves excellent performance, with 72.3% and 82.0% mAP@50 on UAV-PDD2023 and BDD100K datasets, respectively. The proposed framework also achieves real-time inference and competitive segmentation performance, highlighting its effectiveness and practical potential for UAV-based intelligent road infrastructure inspection and maintenance.
Masonry structures form an important part of the housing stock and of transport and waterway infrastructure worldwide. The assessment of these structures relies heavily on visual surveying of geometric features and on monitoring displacements. Over the last decades, automated algorithms have been developed for these tasks. However, their quantitative assessment is limited by the scarcity of data with consistent task-specific labels and dense displacement references. To address this limitation, we present an automated, label-preserving synthetic data pipeline that couples masonry geometry generation, discrete-element mechanical simulation, Blender-based reconstruction, and virtual sensing. The pipeline creates extruded masonry wall models, generates mechanically informed crack and displacement fields, explicitly represents constituent interfaces, and propagates semantic, instance, and synthetic deformation labels to rendered images and point clouds. This paper presents the computational implementation and evaluates two downstream uses of the generated data. First, synthetic image–mask pairs are used to assess image–label consistency and to improve SAM-based segmentation on held-out real masonry images through synthetic-to-real mask-decoder fine-tuning. Second, synthetic point clouds with controlled noise, surface roughness, and known synthetic displacement fields are used as a controlled benchmark for point-cloud full-field displacement monitoring using the PICCA algorithm. These experiments indicate that, within the tested rubble-masonry segmentation and controlled synthetic point-cloud settings, the proposed generator can provide transferable image supervision and diagnostically useful monitoring test cases. Our code and data are available at https://github.com/Yilong-Yang/masonry-synthetic-pipeline.git.
Structural optimization in engineering traditionally focused on minimizing material weight to enhance efficiency and performance. However, due to fabrication, transportation, and assembly requirements, minimizing material consumption does not necessarily minimize total construction cost. Lightweight designs may involve increased nodal connectivity, a higher diversity of cross-sections, or the use of oversized or overweight members, thereby raising production effort, installation complexity, and logistical costs. To address this issue, three dimensionless indexes are proposed in this study, namely typological complexity, nodal complexity, and logistical complexity. These indexes enable a quantitative assessment of structural complexity to be performed alongside material weight. When embedded into a multi-objective optimization framework, they allow the structural design to be evaluated in terms of material efficiency, manufacturing standardization, topological simplicity, and logistical convenience. In this way, trade-offs among these competing design aspects can be systematically explored. The practical significance is demonstrated through three case studies: a planar cantilever truss, a three-dimensional geodesic dome, and an industrial steel building. It is demonstrated that, when the proposed complexity indexes are incorporated within a multi-objective optimization framework, the trade-offs between material efficiency and structural complexity can be systematically explored. Consequently, the framework supports more informed decision-making during early design stages.
Corrosion assessment in steel bridges requires not only estimating the spatial extent of visible deterioration but also distinguishing visual corrosion appearances that may support inspection interpretation. Existing vision-based corrosion-segmentation studies predominantly formulate the task as binary corrosion-versus-background mapping, which supports damage localization but provides limited information regarding corrosion-pattern distribution. This study proposes a two-level semantic segmentation framework for visual corrosion assessment of steel bridge components. Level 1 performs binary pixel-wise segmentation to quantify visible corrosion coverage, whereas Level 2 assigns corrosion pixels to four inspection-oriented visual categories: Uniform, Crevice, Underfilm, and Other Localized corrosion. Field images were annotated using an operational visual taxonomy by the first author, and category-label reproducibility was subsequently evaluated through a blinded second-annotator audit of the categorized validation polygons. Transformer-based and CNN-based models were evaluated using common image subsets, input resolution, and evaluation metrics, while model-specific implementation and optimization settings were documented explicitly. Performance was assessed using region-overlap metrics, class-wise precision, recall, and F1-score, and Boundary-F1 to quantify edge agreement along irregular corrosion fronts. The highest-performing binary configuration achieved an mIoU of 0.843. The adopted formula-derived Level 2 configuration achieved a five-class mIoU of 0.605 and a mean corrosion-category F1-score of 0.678; a manual-weighting variant produced a marginally higher five-class mIoU of 0.607 but was not adopted as the primary formula-derived SegFormer configuration. Binary-merged evaluation showed that Level 2 retained substantial corrosion-coverage capability, although dedicated binary models remained superior for foreground–background segmentation. The framework provides complementary representations of corrosion coverage, visual-category distribution, and boundary quality. It is intended to support, rather than replace, expert inspection, field measurement, and engineering evaluation.
Rockfill dams exhibit pronounced stage-dependent deformation during construction, impoundment, and long-term operation, requiring monitoring-updated deformation models for lifecycle prediction and digital twin-based safety assessment. However, existing updating methods are often deterministic and stage-isolated, while conventional probabilistic approaches insufficiently capture dependence among mechanically related constitutive parameters. For multi-stage lifecycle analysis, these methods also lack cross-stage adaptability and often require repeated finite-element simulations or full surrogate retraining, leading to high computational cost. This study proposes a dependence-aware sequential probabilistic framework with posterior-guided continual surrogate learning for dynamic updating of rockfill dam deformation models. The framework progressively assimilates monitoring data to track lifecycle parameter evolution, preserves empirical parameter dependence via Copula-enhanced prior modeling within a sequential Bayesian updating scheme, and adaptively refines the surrogate by concentrating new training samples in posterior-relevant regions while retaining historical deformation knowledge. Application to the 303 m Lianghekou core-wall rockfill dam shows that the proposed method reduces the global relative settlement prediction error from 37.4% to 9.7% and reduces the dominant computational cost of lifecycle surrogate updating by 45.8% compared with full surrogate retraining. The updated model further enables uncertainty-informed settlement prediction, spatiotemporal deformation reconstruction, and posterior-informed diagnosis of safety-related deformation characteristics, providing an interpretable and efficient approach for lifecycle deformation analysis and digital twin-based dam safety management.
The reliable assessment of concrete surface damage is important for the maintenance and safety of civil infrastructure. In practice, visual inspection is often time consuming and subjective, while supervised learning methods require large amounts of labeled data that are often difficult to obtain for infrastructure inspection. Beyond this annotation challenge, crack analysis also benefits from representing gradual morphological variation, including orientation, width, and continuity. This motivates unsupervised methods that not only separate crack-related patterns from intact surfaces but also organize fine-grained crack variability in a structured way. In this paper, we introduce TopoCLR, an unsupervised framework for concrete crack analysis that combines contrastive learning with a Self-Organizing Map (SOM) in a unified training process. A convolutional encoder is trained on augmented image pairs using a contrastive objective and a topology-preserving constraint induced by the SOM. While the contrastive loss promotes discriminative and invariant features, the SOM guides the encoder toward a structured two-dimensional topology that preserves latent-space neighborhood relations. TopoCLR operates without manual annotations and provides a structured map-based organization, in which visually related samples are arranged in neighboring regions without relying on rigid predefined classes. Experiments on three public concrete crack datasets demonstrate purity values above 0.93 for the binary separation between crack and intact samples. To evaluate topology preservation under controlled gradual variation, we additionally introduce CCIC-O, an orientation benchmark derived from real cracks of one of these datasets. Under controlled rotations, TopoCLR achieves a purity of 0.835 and a label rank correlation of 0.972, the highest among the compared map-based methods. On the natural configuration, which consists entirely of real, un-rotated cracks with measured, continuous orientation labels, TopoCLR shows a similar trend at lower absolute values, as expected for the unbalanced, continuous labels, validating the method on fully real data. The learned map captures orientation-related crack variation while preserving the ordinal structure of related crack appearances.