Urban Air Mobility (UAM) is advancing rapidly, and its services are inherently linked with ground transport networks, making their impacts on traffic flow and efficiency a critical concern for urban mobility planning and policy decision-making. However, research on the implications of UAM for ground transportation remains limited and contested. Existing studies show divergent findings at the system-wide level and often overlook spatial heterogeneity and localized effects. This study develops an extended SUMO-TraCI framework that (i) incorporates realistic car-based first-/last-mile access and dispatch mechanisms, (ii) enables dual-scale analysis by capturing both system-wide traffic efficiency and localized congestion around vertiports, and (iii) supports multi-scenario testing, providing a broadly applicable framework for UAM evaluation. Based on five scenario experiments varying vertiport numbers (10, 30, 60, 80, and 100) relative to a baseline, UAM demonstrates modest yet positive citywide impacts on traffic flow reducing VKT and TTT modestly under the baseline eligibility setting, while sensitivity tests further reveal a non-linear and saturating benefit pattern around the 60–80 vertiport range as infrastructure density increases. Network-level spatial patterns indicate slightly greater relief on peripheral corridors and more frequent minor increases in central areas, with localized changes more noticeable on high-capacity roads than on local streets. This study introduces a scalable, multi-scenario evaluation framework for understanding and managing the multimodal impacts of UAM, providing actionable insights to guide evidence-based UAM planning and support sustainable city development.
Understanding urban environment change is essential for sustainable development. However, current approaches, particularly remote sensing change detection, often rely on rigid, single-modal analysis. To overcome these limitations, we propose MMUEChange, a multi-modal agent framework that flexibly integrates heterogeneous urban data via a modular toolkit and a core module, Modality Controller for cross- and intra-modal alignment, enabling robust analysis of complex urban change scenarios. Case studies include: a shift toward small, community-focused parks in New York, reflecting local green space efforts; the spread of concentrated water pollution across districts in Hong Kong, pointing to coordinated water management; and a notable decline in open dumpsites in Shenzhen, with contrasting links between nighttime economic activity and waste types, indicating differing urban pressures behind domestic and construction waste. Compared to the best-performing baseline, the MMUEChange agent achieves a 46.7 % improvement in task success rate and effectively mitigates hallucination, demonstrating its capacity to support complex urban change analysis tasks with real-world policy implications.
Evaluating the hyper-local effectiveness of transport policies like the Ultra Low Emission Zone (ULEZ) on commuter air pollution exposure is critical for urban health yet remains challenging. Existing studies often lack high-resolution pollution data integrated with realistic pedestrian pathways, which constrains precise exposure evaluation. This study proposes an innovative framework that combines machine learning-based spatial-temporal street-level air pollution (SLAP) estimation with a network-based metric, the Public Transit-oriented Exposure Value (PTEV), to quantify last-mile cumulative PM2.5 exposure within a station's walkable service areas. Applying this framework to pre-and post-ULEZ periods in London, we find that while the policy reduced overall exposure, a central transit hub persisted as a high-risk hotspot, revealing significant spatial heterogeneity in policy impact. Our research contributes a replicable, fine-scale methodology for assessing transport-environment interactions, providing critical evidence for targeted planning to mitigate commuter health risks within complex urban transport systems.
Urban air pollution severely impacts pedestrians on sidewalks, yet traditional fixed-site monitoring lacks the spatial coverage needed for fine-scale exposure assessment due to high costs. This study proposes a novel machine-learning framework to predict short-term sidewalk PM2.5 and PM1 concentrations using multimodal audiovisual features extracted from self-collected street-view videos, alongside meteorological and background pollution data. Based on a mobile monitoring campaign in Shenzhen, China, we evaluated multiple models (linear regression, XGBoost, and LightGBM) across different temporal resolutions (10 s and 1 min) and validation strategies. LightGBM achieved the best performance, yielding R2 values of 0.64-0.65 for 10 s predictions and 0.80 for 1 min predictions under random cross-validation. Under rigorous spatial cross-validation, the model maintained moderate generalizability, with R2 reaching 0.41-0.48 at the 1 min resolution. Furthermore, developing a hybrid model that incorporated static geospatial context further improved the overall predictive accuracy. Variable interpretation revealed that while background PM and meteorology were dominant predictors, dynamic audio-derived features and visual indicators provided substantial additional predictive power. These findings demonstrate that integrating multimodal audiovisual sensing with ancillary data enables scalable, high-resolution estimation of street-level PM, effectively complementing conventional monitoring for urban air-quality management.
Protecting pedestrians from air pollution requires understanding how exposure varies across age groups, yet most regional-scale pedestrian exposure studies lack demographic stratification. This challenge is compounded for PM1, which remains underexplored despite evidence of heightened health risks compared to larger particles. In this study, we develop an age-stratified, spatially explicit exposure framework integrating mobile-monitored PM1 concentrations with deep learning-derived age-stratified pedestrian volume from street-view imagery. A transfer learning-based model automatically classifies pedestrians into three age groups (children, adults, elderly), enabling large-scale exposure estimation across urban sidewalks. Tree-based machine learning models achieve R2 values of 80.9% for PM1 concentrations and 64.38%-79.26% for age-stratified exposure prediction. Variable interpretation analysis reveals distinct determinants: ambient PM1 is primarily driven by urban morphology and meteorology, whereas pedestrian exposure is governed by points of interest distributions reflecting age-specific destination patterns. Age-specific exposure analysis demonstrates that high-pollution zones do not necessarily coincide with high-exposure locations for vulnerable populations, indicating that pollution reduction alone is insufficient without considering demographic-specific activity patterns. Furthermore, the differential importance of points of interest across age groups provides compelling evidence of distinct activity spaces and destination preferences throughout the life course. The proposed framework enables identification of critical intervention zones for each age group, supporting evidence-based urban planning strategies for equitable exposure mitigation.
Whether nighttime is the better half of life varies across cities depending on how safe residents perceive their neighborhood environments to be. Although street view imagery (SVI) is instrumental in auditing perceived safety, city-wide nighttime SVI does not exist. Consequently, the extent to which day-night safety perceptions diverge remains unclear. Although emerging studies have used day-to-night (D2N) translations to generate nighttime SVIs, the model convergence and transferability are unknown. Using paired day-night SVIs from U.S. and Chinese cities, we confirmed the converging sample size (similar to 2000 pairs) and demonstrated both cross-density and cross-cultural transferability of the D2N model. Medium-density SVIs achieved the best cross-density transferability, while models trained on U.S. samples outperformed those trained on Chinese samples, indicating moderate concerns about the adequacy of cross-cultural training data for applying D2N models across regions. Moreover, interpretable machine learning reveals a pronounced day-night asymmetry in perceptual mechanisms: daytime safety perception is primarily shaped by pedestrian-oriented configurations, whereas nighttime safety perception relies more on visibility-related cues. Meanwhile, brightness emerges as a key positive predictor in both periods. Additionally, urban greening intensifies day-night divergence in perceived safety disproportionately rather than uniformly enhancing perceived safety. This pattern illustrates the urban greening paradox: trees and plants increase daytime perceived safety, yet are linked to lower perceived safety after dark. Lastly, city-scale mapping in Boston reveals salient spatial heterogeneity: central and northwest corridors retain relatively high safety perception after dark, while southern areas experience the steepest nighttime declines. Our scalable GenAI framework extends urban scene studies to the nighttime domain and enables city-wide mapping of perceived safety after dark, informing more inclusive nighttime urban planning.
The pursuit of autonomous agents with predictive cognitive world models is hindered by a fundamental flaw in current vision-language models (VLMs): they lack cognitive inertia. Operating on isolated snapshots, these models cannot form a temporally coherent world view, leading to erratic decision jitter and a failure to execute complex, multi-step maneuvers. To remedy this, we introduce CogDriver, a framework designed to build a coherent world model by instilling this crucial cognitive property. Our work makes two key contributions: (1) We present CogDriver-Data, a large-scale vision-language-action dataset whose narrative annotations provide the supervisory signal for learning the temporal dynamics of a world model. (2) We develop the CogDriver-Agent, an architecture featuring a sparse temporalmemory to maintain a stable internal state, the foundation of a world model. This is enabled by a spatiotemporal knowledge distillation approach that explicitly teaches decision coherence. Comprehensive experiments validate our paradigm: CogDriver-Agent achieves a 22\% increase in the closed-loop Driving Score on Bench2Drive and a 21\% reduction in mean L2 error on nuScenes, establishing a new state-of-the-art. These significant gains in both long-term decision-making and imitation accuracy provide strong evidence that our agent is developing a more stable internal world model.
Humanoid robots are expected to traverse complex terrains, where the plantar support may vary dramatically due to foot placement errors, ground properties, and transient dynamics. To achieve robust locomotion, the robots are required to adapt to uneven terrain and uncertain foot–ground interactions. Existing locomotion policies rely primarily on proprioception or exteroceptive terrain perception, where the former provides only indirect evidence of plantar support, while the latter predicts contact conditions before touchdown but cannot observe the actual support in real-time. Although some studies incorporate plantar contacts as an auxiliary perception, they rely mainly on summary statistics, overlooking the spatial topology of plantar pressure, which provides a more direct characterization of the realized contact state. To bridge this gap, we present Tac4Loco, a tactile-perceptive framework that incorporates multi-array plantar pressure as direct feedback for humanoid locomotion. We formulate a topology-preserving ordinal representation to map simulated and physical sensor signals into a shared observation space, with a dual-branch encoder for extracting their spatial and temporal representations. Subsequently, the learned spatiotemporal features are integrated with augmented proprioception including terrain estimation cues, and provided to an asymmetric actor-critic architecture for policy learning. Extensive simulation and real-world experiments demonstrate improved tracking performance and support adaptation on terrains with inclined, partial, asymmetric, and changing support. We further demonstrate its zero-shot deployment on unseen compliant and unstructured terrains, including a foam platform and a gravel road. All code and experimental configurations will be released as open-source to facilitate reproducibility.
Accurate pedestrian estimation is essential for developing sustainable, livable cities and supporting human-centered urban planning. However, existing studies often rely on coarse, point-based estimates that lack the spatial resolution needed to capture fine-scale pedestrian dynamics. This study proposes R-PIN (Regional Pedestrian Inpainting Network), a novel deep learning model that reformulates regional pedestrian estimation as an image inpainting task to generate high-resolution, spatially continuous pedestrian distributions. R-PIN integrates multi-source urban features through a dual-branch encoder and attention-based fusion block to capture both local details and global spatial dependencies. Applied to a real-world case of New York City, R-PIN exhibits robust performance across diverse urban landscapes, especially in Manhattan and Brooklyn with stable and homogeneous pedestrian patterns. Compared with representative deep learning baselines, R-PIN achieves 57.29% and 79.96% reductions in MAE (Mean Absolute Error) and MSE (Mean Squared Error), respectively. By explicitly leveraging the surrounding urban context, R-PIN captures spatial continuity and spillover effects, reducing MAE and MSE by 55.23% and 76.97%. Feature importance analysis highlights the key roles of streetscape design, road network, points of interest, and land-use diversity in shaping pedestrian distribution. Overall, this framework provides fine-grained insights into pedestrian dynamics and offers a robust analytical tool for street-level management and pedestrian-oriented urban analytics.
The accurate and efficient calculation of building areas is a critical aspect of the AEC industry, yet it remains prone to human errors and inefficiencies. Current methodologies rely excessively on manual corrections due to three persistent bottlenecks: the inability to automate height-dependent area coefficients, ineffective extraction of building boundaries, and the misclassification of vertical spaces. To address these gaps, this paper proposes ArchAreaSynth, an integrative framework, combining three key algorithms: the VarHeight Solver for dynamic height analysis, the GeoMax Contour Detection for precise external wall boundary extraction, and Shaft Detection algorithm to distinguish vertical shafts from atriums. Validation experiments on 100 diverse floor plans demonstrate that the framework achieves over 99% accuracy in complex scenarios. It significantly outperforms commercial tools (e.g., Revit built-in tool) and manual measurements, improving mean time efficiency by over 30 times. This research offers a promising solution to streamline BIM workflows and enhance productivity in the AEC industry.
Rapid urbanization is a primary driver of terrestrial carbon loss, with most studies emphasizing external regulations as essential for mitigation. Yet whether urban systems can stabilize ecologically through self-regulated processes remains unclear. This study examines Dongguan, a global manufacturing hub in China, to assess the potential for endogenous stabilization. We employ an integrated land use-carbon simulation method to project trajectories under a “natural evolution” scenario, combined with a Geodetector analysis to identify the driving factors. The results show that impervious surface expansion decelerates sharply, with its compound annual growth rate declining from 5.85% (2000-2005) to 0.64% (2015-2020), and projected to nearly cease by 2030 (0.04%). Carbon loss follows a parallel slowdown, stabilizing around 23.24 million. GeoDetector results reveal a shift in urban growth driving forces: natural factors such as soil and topography dominate, while GDP plays only marginal influence. These coupled trajectories provide empirical evidence that the trade-off between urban expansion and carbon preservation can narrow intrinsically, through intrinsic stabilization rather than top-down regulation. The findings challenge the regulation-dependent paradigm and offer a new perspective on urban resilience and long-term ecological sustainability.
Comprehensive urban waste management systems, addressing municipal waste collection and construction waste disposal, are essential for maintaining livable and sustainable cities. Understanding the spatial distribution patterns of controlled and uncontrolled waste, along with their underlying environmental and socioeconomic determinants, is essential for developing more effective urban waste management strategies. However, comprehensive analysis of street-level waste distribution and its relationship with socioeconomic factors remains limited, particularly regarding environmental justice implications. This study developed a computer vision approach to detect controlled and uncontrolled waste in New York City and analyzed their spatial distribution patterns to examine associations with urban environmental and socioeconomic characteristics. We employed Swin Transformer architecture for automated waste detection from street-view imagery. Spatial analysis, logistic regression, and interpretable machine learning using SHAP (SHapley Additive exPlanations) were applied to analyze 43 variables across environmental and socioeconomic factors. Results revealed contrasting distribution profiles where controlled waste concentrated in high-density, well-developed areas with higher education levels, while uncontrolled waste exhibited dual marginalization in urban peripheries and socioeconomically disadvantaged communities. Hispanic populations showed 14.9 % higher odds of uncontrolled waste exposure (OR = 1.149, p = 0.032), confirming environmental justice concerns. This research provides the first comprehensive quantitative evidence of street-level waste inequality, revealing significant spatial and social disparities that support targeted policy interventions through data-driven hotspot identification for vulnerable communities in urban waste management systems.
Urban and transportation research has long sought to uncover statistically meaningful relationships between key variables and societal outcomes such as road safety, aiming to generate actionable insights that guide the planning, development, and renewal of urban mobility systems. However, traditional workflows face several key challenges: (1) reliance on human experts to propose hypotheses, which can be time-consuming and prone to confirmation bias; (2) limited interpretability, particularly in deep learning approaches; and (3) underutilization of unstructured data that encodes critical urban context. To address these limitations, we propose a Multimodal Large Language Model (MLLM)-based approach for interpretable hypothesis inference, enabling the automated generation, assessment, and refinement of hypotheses concerning urban form and transportation safety. Specifically, we leverage MLLMs to generate road safety-relevant questions and automatically answer them based on street view images (SVIs) through visual question answering (VQA). These responses are used to construct interpretable embeddings for each SVI, which are then incorporated into linear statistical models for transparent and explainable regression analysis. URBANX supports iterative hypothesis testing and refinement guided by statistical evidence, such as coefficient significance, thereby enabling rigorous, transparent scientific discovery of previously overlooked correlations between urban design and transportation risk. We evaluate our framework on Manhattan street segments and demonstrate that it outperforms pretrained deep learning baselines while offering full interpretability. We demonstrate that UR-BANX matches or exceeds the explanatory power of existing expert-curated built environment variables, validating its potential to replace labor-intensive feature engineering with automated, scalable discovery of potential safety-related factors. Beyond road safety, URBANX can serve as a general-purpose foundation for hypothesis-driven urban mobility analysis, extracting structured insights from unstructured data across diverse socioeconomic and environmental outcomes. This approach establishes a scalable and trustworthy pathway for interpretable, data-driven scientific discovery in urban and transportation systems using foundation models such as MLLMs.
Drivable areas and curbs are critical traffic elements for autonomous driving, forming essential components of the vehicle visual perception system and ensuring driving safety. Deep neural networks (DNNs) have significantly improved perception performance for drivable area and curb detection, but most DNN-based methods rely on large manually labeled datasets, which are costly, time-consuming, and expert-dependent, limiting their real-world application. Thus, we developed an automated training data generation module. Our previous work generated training labels using single-frame LiDAR and RGB data, suffering from occlusion and distant point cloud sparsity. In this paper, we propose a novel map-based automatic data labeler (MADL) module, combining LiDAR mapping/localization with curb detection to automatically generate training data for both tasks. MADL avoids occlusion and point cloud sparsity issues via LiDAR mapping, creating accurate large-scale datasets for DNN training. In addition, we construct a data review agent to filter the data generated by the MADL module, eliminating low-quality samples. Experiments on the KITTI, KITTI-CARLA and 3D-Curb datasets show that MADL achieves impressive performance compared to manual labeling, and outperforms traditional and state-of-the-art self-supervised methods in robustness and accuracy.
High-resolution (HR) imaging devices are crucial for ensuring the safety and efficiency of unmanned aerial vehicles (UAVs) during bridge crack detection tasks. However, due to the limitations of executing sampling discretely in traditional deep learning (DL) architectures and the constraints of GPU computing resources, it is challenging to perform fine-grained segmentation for HR crack images. To effectively address the challenge, the authors drew inspiration from the fine-grained rendering technology in the field of computer graphics (CG) and proposed the Crack Boundary Refinement Transformer (CBRFormer). Through three customized improvements, this architecture fully leverages the advantages of the rendering head in the refined representation of HR crack images. Firstly, a lightweight Transformer-based encoding architecture is designed, enabling the network to accurately capture crack backbone features from complex backgrounds. Subsequently, a boundary-guided branch based on super-resolution reconstruction technology is introduced to assist the network in capturing deep semantic information about crack boundary details. Additionally, two types of refined rendering point sampling methods are tailored for hard example areas during training and inference stages, ensuring that the prediction head used for refined rendering effectively focuses on ambiguous crack boundaries and tiny crack regions. Finally, the effectiveness of each component in the CBRFormer and the network's practicality are demonstrated through ablation and the field experiment. Compared to the current advanced HR segmentation architectures like CascadePSP and Segfix, the CBRFormer achieved average performance improvements of 2.16% in mean Intersection over Union (IoU), 7.80% in mean Edge Accuracy (mEA), and 2.46% in Dice coefficient, respectively. The utilization of the CBRFormer enables precise segmentation of HR crack images, providing inspectors with more comprehensive and accurate structural crack information, thereby offering technical support for structural safety assessment and maintenance decision-making.
The process of querying building codes has long been time-consuming and labor-intensive, requiring extensive manual effort to repeatedly consult and confirm regulations throughout design and review phases. Existing regulation query systems are limited to simple searches, often resulting in errors and failing to provide intelligent responses to users. Although the emergence of large language models (LLMs) offers potential solutions due to their natural language processing abilities, they face challenges such as insufficient domain knowledge, semantic misalignment, and difficulties in handling complex tabular data. To address these limitations, we propose a novel system, LLMs Query Building Codes (LLM-QueryBC), which integrates LLMs with a semantic network-enhanced retrieval-augmented generation (SN-RAG) for text-based queries and a specialized agent called Agent for Tables in Building Codes (TaBCe) for tabular queries. TaBCe leverages the Reasoning and Acting (ReAct) framework, think-by-structure planning method, RAG, computational tools, and a memory module to enhance its functionality. The system autonomously determines internal logic, invokes external tools, and mitigates hallucination issues associated with LLMs in the building code domain. We evaluated our system using fire safety regulations-a critical domain due to its impact on life safety, property protection, and legal compliance. Experimental results demonstrated significant improvements: textual query accuracy increased by 24%, and tabular query accuracy rose by 25%. Additionally, we conducted a case study involving fire safety code queries for a real-world design of a mixed-use high-rise building, further validating the system's practical applicability. These advancements offer greater value to users and promote broader adoption of intelligent regulation query systems.
In practical engineering, high-resolution (HR) imaging devices have become increasingly utilized for capturing structural surface crack images. However, the effectiveness of current deep learning (DL) segmentation models in accurately predicting refined masks for HR crack images is hindered by the discrete sampling methods inherent in traditional DL architectures and the limited computational resources of GPUs. To tackle this issue, this investigation incorporates the point-based rendering methodology originating from computer graphics disciplines into the encoding-decoding framework, introducing an innovative Crack Boundary Point Rendering Network (CBPRN). The CBPRN endeavors to accomplish elaborate delineation of crack visual samples possessing resolutions surpassing 4K. Initially, an edge feature extractor integrated with a super-resolution encoder is devised to guide rendering heads in efficiently focusing computational power on ambiguous edge regions. Subsequently, a rendering-based prediction head is introduced with the function of efficiently sampling rendering points for the training and inference phases, respectively. Furthermore, a tailored composite objective function is deployed to enhance the learning procedure, enabling the architecture to equilibrium substantial disparities in pixel counts among positive and negative instances within crack visual data. Ultimately, to substantiate the practical applicability of the CBPRN, an on-site crack identification investigation was executed on an actual bridge structure located in Changsha utilizing an unmanned aerial vehicle (UAV). The CBPRN demonstrated remarkable effectiveness on 4K-resolution visual samples acquired by the unmanned aerial vehicle, attaining achieving overlap ratio (Intersection over Union, IoU), average boundary precision (mean Boundary Accuracy, mBA), and Dice similarity index metrics of 85.46%, 86.00%, and 92.16%, correspondingly. This outstanding effectiveness improves both the operational security and processing efficiency of unmanned aerial vehicle-assisted crack assessment procedures, offering enhanced flexibility in choosing flight trajectories for the inspection workflow.
Large Language Models (LLMs) offer transformative potential for transportation, yet most of the existing review focus on vehicle-level techniques, leaving a gap in understanding the potential of LLM in city-scale and long-term transportation planning. Our survey addresses this gap by systematically analysing 61 studies across problem formulation, solution generation and evaluation - including transport analysis, modelling, system design, and decision-making. We examine existing applications strategies, identifying underexplored areas for leveraging LLMs in large-scale transportation planning. While LLMs excel in transportation analysis and optimisation, system design and decision-making remains untapped. Current applications focus on data processing over reasoning and the lack of multimodal LLM, causal reasoning, and real-time adaptability restricts their role in more complex transportation challenges. Based on these findings, we propose a Cognitive-Adaptive-Interactive framework, offering a structured approach to future LLM applications in transportation and urban systems. Our framework enhances adaptability, efficiency, and decision-making in modern transportation systems.
Crowdsourced street-view imagery from social media provides valuable real-time visual evidence of urban flooding and other crisis events, yet it often lacks reliable geographic metadata for emergency response. Existing image geo-localization approaches, also known as Visual Place Recognition (VPR) models, exhibit substantial performance degradation when applied to such imagery due to visual distortions and domain shifts inherent in cross-source scenarios. This paper presents VPR-AttLLM, a model-agnostic framework that integrates the semantic reasoning and geospatial knowledge of Large Language Models (LLMs) into established VPR pipelines through attention-guided descriptor enhancement. By leveraging LLMs to identify location-informative regions within the city context and suppress transient visual noise, VPR-AttLLM improves retrieval performance without requiring model retraining or additional data. To evaluate the framework, we conduct comprehensive testing across two morphologically distinct urban environments: San Francisco and Hong Kong. The evaluation utilizes established query sets, synthetic flooding scenarios, and real social media flood images integrated into the San Francisco benchmark, alongside a newly curated Hong Kong dataset. Integrating VPR-AttLLM with three state-of-the-art VPR models-CosPlace, EigenPlaces, and SALAD-consistently improves recall performance, yielding relative gains typically between 1%-3% and reaching up to 8% on the most challenging real flood imagery. Beyond measurable gains in retrieval accuracy, this study demonstrates a robust pipeline for LLM-guided multimodal fusion in visual retrieval systems. By embedding principles from urban perception theory into attention mechanisms, VPR-AttLLM bridges human-like spatial reasoning with modern VPR architectures. Its plug-and-play design, strong cross-source robustness, and interpretability highlight its potential for scalable urban monitoring and rapid geo-localization of crowdsourced crisis imagery in heavily urbanized environments.
Manual constructability checking of complex mechanical, electrical, and plumbing (MEP) designs is inefficient and error-prone. However, automating this process is hindered by the complexity of MEP drawings, the variety of design specifications, and the insufficient integration of checking algorithms with fragmented checking processes. Addressing these issues, this paper proposes an automated checking system that integrates large language models (LLMs) and graph technologies within a collaborative multi-agent framework. The system comprises three components: the MEP graph agent that constructs structured scene graphs from drawings; the MEP rule agent that converts textual rules into knowledge graphs; and the MEP checking agent that identifies violation via graph matching and reasoning. Validated on three real-world projects, the system achieved an average accuracy of 92% in violation detection. By offering an intuitive interface for issue identification, this approach significantly enhances coordination efficiency and reduces manual effort. This framework establishes a robust foundation for future AI-driven automated compliance checking in the construction domain.