Multi-camera Visual Simultaneous Localization and Mapping (V-SLAM) increases spatial coverage through multi-view image streams, improving localization accuracy and reducing data acquisition time. Despite its speed and generally robustness, V-SLAM often struggles to achieve precise camera poses necessary for accurate 3D reconstruction, especially in complex environments. This study introduces two novel multi-camera optimization methods to enhance pose accuracy, reduce drift, and ensure loop closures. These methods refine multi-camera V-SLAM outputs within existing frameworks and are evaluated in two configurations: (1) multiple independent stereo V-SLAM instances operating on separate camera pairs; and (2) multi-view odometry processing all camera streams simultaneously. The proposed optimizations include (1) a multi-view feature-based optimization that integrates V-SLAM poses with rigid inter-camera constraints and bundle adjustment; and (2) a multi-camera pose graph optimization that fuses multiple trajectories using relative pose constraints and robust noise models. Validation is conducted through two complex 3D surveys using the ATOM-ANT3D multi-camera fisheye mobile mapping system. Results demonstrate survey-grade accuracy comparable to traditional photogrammetry, with reduced computational time, advancing toward near real-time 3D mapping of challenging environments.
Abstract. High-resolution architectural documentation goes beyond geometry—it requires a deep understanding of the building’s structure, materials, and historical layers. This often means interpreting hidden construction logic and identifying even the smallest components, such as individual stones or bricks, to produce meaningful data for conservation, analysis, and interpretation. Identifying and describing all the individual components that constitute the building, such as the type, arrangement, and state of preservation of stones, bricks, mortars, or decorative materials embedded in the walls, is a real challenge due to the large quantity and the complex spatial distribution of each element. Recent advances in AI, particularly foundational models and zero-shot models, offer potential solutions to speed up the documentation process. Taking the gothic complex of Milan Cathedral as the monument object of study, the research hereby presented implements a SAM2 (Segment Anything Model) based stone-by-stone segmentation, leveraging object detector for semantic interpretation. The proposed framework integrates 2D stone block segmentation with photogrammetric 3D reconstruction, enabling accurate projection of semantic labels and geometric data from images to 3D point cloud, allowing a detailed 3D segmentation in all the components of the structure.
The capture of 3D reality has demonstrated increased efficiency and consistently accurate outcomes in architectural digitisation. Nevertheless, despite advancements in data collection, 3D reality-based modelling still lacks full automation, especially in the post-processing and modelling phase. Artificial intelligence (AI) has been a significant focus, especially in computer vision, and tasks such as image classification and object recognition might be beneficial for the digitisation process and its subsequent utilisation. This study aims to examine the potential outcomes of integrating AI technology into the field of 3D reality-based modelling, with a particular focus on its use in architecture and cultural-heritage scenarios. The main methods used for data collection are laser scanning (static or mobile) and photogrammetry. As a result, image data, including RGB-D data (files containing both RGB colours and depth information) and point clouds, have become the most common raw datasets available for object mapping. This study comprehensively analyses the current use of 2D and 3D deep learning techniques in documentation tasks, particularly downstream applications. It also highlights the ongoing research efforts in developing real-time applications with the ultimate objective of achieving generalisation and improved accuracy.
This paper presents a comprehensive scoping review of the application of 3D digital technologies in the documentation, conservation, and management of historic gardens and related cultural heritage. By analyzing a curated selection of literature, this study assessed the current state of research, highlighting trends in publications, the geographic distribution of contributors, and the key technologies employed. Using bibliometric methods and visualization tools, followed by a case study review, this review identified significant research hotspots and technical methodologies, particularly focusing on advanced techniques such as mobile laser scanning, UAV photogrammetry, and point cloud processing and their relationships with end users. The findings emphasize the importance of integrating multiple technologies to capture the diverse elements of historic gardens, including architectural features, vegetation, and topography. This review also underscores the significance of dynamic landscapes facing challenges posed by environmental degradation and urban development pressures. Moreover, it discusses the limitations of existing research and outlines future opportunities, such as the development of 4D documentation systems and the incorporation of AI for improving heritage management. This paper concludes by recommending interdisciplinary collaboration and public engagement to enhance the accessibility, understanding, and sustainable management of historic gardens through innovative technological applications.
This study demonstrates the applicability of air sampling for the detection of SARS-CoV-2 in a hospital by means of active bioaerosol samplers following a specifically designed air sampling strategy based on digital mapping of the architectural layout of the ward to minimize disruptions of health care activities and reducing operator risks. Prior to the experimental study, some model tests were conducted using the air sampler with a tunable flow rate to determine the most suitable real time polymerase chain reaction (RT-PCR) based detection method. Preliminary results showed the need to perform intensive extraction protocols combined with Real-time reverse transcription PCR (rRT-PCR), rather than conventional, to enhance sensitivity. The experimental study was conducted within the general medicine ward of Spedali Civili Hospital in Brescia during the winter of 2021/2022, a period marked by a high prevalence of COVID-19 cases using three active air sampling devices: Coriolis Compact (R), Coriolis Micro (R), and BioSpot GEM (R). Environmental parameters, such as room size, occupancy, ventilation rates, and activities per-formed during sampling, and patients' conditions were documented to contextualize the findings. The virus was detected in a few rooms with concentrations ranging from 1171 to 2225 copies/m3. These findings support the integration of routinary air sampling as tools for control and assess-ment of transmission risks, not only for SARS-CoV-2 but generalized to all airborne pathogens, supporting patient management and infection control in health care settings.
The Municipality of Venice, through Insula srl (Insula, 2024), started the RAMSES (Rilievo Altimetrico, Modellazione Spaziale E Scansione 3D) project in 2005 with the unprecedented intention of conducting a static laser scanner survey of an entire city. The authors of this paper, who have been involved in the project on behalf of the client, including drafting general contract technical specifications, wished to revisit the survey’s findings nearly two decades later. This contribution illustrates the procedures implemented to guarantee the future accessibility of the surveyed data. It is interesting to highlight how the detailed technical specifications outlined in the general technical contract section have facilitated the retrieval of the historical three-dimensional laser scanner measurements archive. Tests have been conducted to determine how the existing mobile mapping technologies may be utilised to update the three-dimensional historical data obtained in the Ramses project efficiently. Furthermore, the paper describes the surveying approach that has never been adequately described in the literature. The surveying and geo-referencing methodologies continue to have several interesting and relevant aspects, especially regarding how the topographic network has been implemented. The Ramses three-dimensional model represents an extraordinary, valuable digital archive containing portraits of the city’s conditions at the time of the mapping. Ramses 3D model, when enriched with field activities conducted using more updated technologies, can provide interesting and unique evaluations of the evolution of Venice’s landscape.
Digital technology provides methods to record and preserve cultural heritage, support conservation and restoration efforts, and share our collective past with a worldwide audience. Between 2011 and 2017, the 3D Survey Group from Politecnico di Milano operated an annual workshop in the medieval village of Ghesc in which photogrammetry and laser-scanner surveys were carried out. The point cloud data acquired in these activities has become “time slices” documenting different stages of the preservation interventions in Ghesc and the evolution of advanced survey techniques. The main objective of this research is to streamline the workflow of delivering immersive and interactive experiences for complex heritage by directly utilising the 3D survey point cloud data, whether derived from a photogrammetric survey, static laser scanner, or mobile mapping.A point cloud-based multiplatform application is designed and delivered with versatile functions. It runs on PC and VR devices to provide virtual access to the village and narrate its revitalisation story. Additionally, it operates on mobile devices with an AR feature that brings vibrancy to the on-site experience. This application integrates high-fidelity point cloud models, detailed information on vernacular architecture in the Ossola Valley, and information on the preservation project with gamified learning experiences. The unconventional approach of using points as rendering primitives in virtual applications offers a practical solution for visualising complex heritage, enabling an efficient transition from the data collection stage to the data sharing stage without the need for 3D reconstruction and intricate BIM modelling.
This paper presents an effective and low-cost approach for the digitisation of a Roman bronze head exhibited at the Santa Giulia Museum in Brescia using close-range photogrammetry. The artefact posed significant challenges due to its dark, reflective exterior surface and its hollow interior, accessible only through an 8 cm neck opening. The digitisation had two main objectives: (1) to support restoration activities by accurately measuring the thickness of the bronze cast, and (2) to enhance the artefact's dissemination by producing a web-optimised digital replica.Special emphasis was placed on the survey of the internal surface, achieved using a custom-designed camera probe equipped with a 5 megapixels RGB global shutter camera and LED ring lights. With this setup it was possible to effectively captured the inner geometry using low-cost tools. The internal dataset was connected with the exterior dataset, acquired by a regular DSLR and turntable setup supported by cross-polarization.The processing workflow involved generating a high-resolution mesh model (~80M faces) and computing the cast's thickness using two methods: the M3C2 algorithm and the Shrinking Sphere algorithm, which provided consistent results. Additionally, an optimised low-poly version of the model was created for web sharing. This study highlights the potential of low-cost tailored photogrammetric solutions/workflows to address the complexities of digitising intricate museum artefacts.
Multi-camera devices are increasingly popular in various metrological applications, including cultural heritage digitalisation, where these devices are adopted as low-cost alternatives to more traditional methods or mobile mapping systems. They can be of two types: panoramic and non-panoramic configurations, with the former usually more compact and ready-made off-the-shelves and the latter usually custom-developed for metrological applications. In the paper, we compare the accuracy and reliability performance of two types of multi-camera: the spherical camera INSTA 360 Pro2 and the custom multi-camera rig Ant3D. The case study is a challenging spiral staircase environment, typical in many cultural heritage survey projects. The processed image datasets were evaluated in the most common constrain scenario (GCPs at both ends of the staircase) and the worst-case scenario (open-ended path, GCPs at the start). The datasets were processed with precalibrated IO and various degrees of multi-camera constraints up to precalibrated relative orientations. The results highlight that the nominal scale 1:50 can be achieved, e.g. an accuracy of <2 cm plus complete and precise point clouds and mesh results.
Photogrammetric applications nowadays envisage the use of more and more low-cost cameras such as those equipped on commercial UAV platforms. Typically, these low-grade cameras suffer from extreme radial distortion and strong vignetting among other defects. This, initiated a trend among the low-cost cameras’ manufacturers to try to hide the camera defects by applying software pre-corrections to the images. These Built-In Correction Profiles gets applied to both the JPG files, directly in-camera, and usually to the raw files as well, through the opcode functions of the DNG standard. In this paper we rise this issue that is still under-reported in the literature and further assess the accuracy implication of applying or discarding the Built-In Correction Profile in the scenario of UAV mapping. We tested the commercial UAV DJI Phantom 4 Pro v2 in a calibration environment and a field test to compare the performance of pre-corrected versus uncorrected images. In our tests, processing the original uncorrected images led to improved IO calibration and reduced bowing effect in the field test.
Image classification and object detection techniques have been widely discussed and developed in recent years; they are the basis of various prosperous applications, for example, real-time mapping. Promising as it is, the practical test in the cultural heritage field encountered multiple problems. In this paper, the authors attempt to share the research experimentations and the empirical knowledge focusing on the classification and detection of architectural pathology. The tests are built on elaborated training sets annotated with analysed and in-advance defined categories. The trained models were examined from the perspective of evaluation sets, model explanation and unseen datasets. The outcomes indicated the mistakes and confusions behind things and stuff in the object detection efforts, to which cultural heritage and architectural field are closely related. The model also reveals specific visual patterns for recognition from thousands of instances in the training set. By digging into different aspects of model performance, the potential and limitations of these techniques in practical applications can be better understood.
Traditional 3D surveying methods often fall short in complex spaces due to lack of mobility, time constraints and high risk. For this reason there is an actual demand for 3D data acquisition tools and methods, particularly suitable for complex and narrow environments, due to their capacity for efficiently capturing detailed and accurate spatial information, maybe also automatically. This study presents a novel approach for fusing 3D spatial data collected by two separate and independent mobile mapping systems: (1) ATOM-ANT3D and (2) MandEye. We propose an innovative fusion technique that combines visual and LiDAR data from asynchronous acquisitions, reducing the need for strict temporal and spatial synchronizations between the two systems. We compare the outputs of both systems before and after fusion, studying the individual limitations and highlighting the complementary benefits achieved by the proposed fusion framework. Results demonstrate improved accuracy of global alignment and spatial completeness of the final point clouds, proving the efficiency and flexibility of the proposed approach.
Historical buildings and monuments are typically subject to degradation over time due to the passage of time and constant exposure to external agents. The use of artificial intelligence (AI) to support the work of conservation and restoration specialists in identifying surface decay is a research topic of considerable interest at present. This study presents two approaches: ChatGPT and an object detection architecture (YOLOv5). Specifically, this investigation sought to evaluate the ChatGPT’s ability to identify and describe surface degradation pathologies by exploiting its pre-trained models for image analysis. The ICOMOS-ISCS: Illustrated Glossary on Stone Deterioration Patterns (2008) was provided as a reference to guide the use of specific terminology. In the first test phase, to verify the accuracy of the ChatGPT results, benchmark images (depicting different types of damage) extracted from the UNI 11182 (2006) standard referring to the definition of degradation types were used. Only later were images from literature studies and other photographic datasets also used. In general, the results of the analysis were validated with the conclusions of professionals and with the conclusions of other AI techniques, as well as with the descriptions provided by reference manuals in the literature. In particular, the decay annotations predicted by the pre-trained object detection model were compared with those made by human experts. The capabilities and limitations of both approaches as tools for identifying deterioration pathologies are illustrated.
Large archaeological areas or archaeological parks located within urbanized areas suffer more than other sites from anthropic pressure (urbanization and tourism), to which are added all the problems related to their conservation, safeguarding, fruition, management, maintenance, etc. Effective management strategies are essential. The use of Web Platforms in the management of 3D data is emerging as an effective solution for balancing conservation and development within large archaeological sites. This paper presents a case study of the Naxos Archaeological Park, Sicily, Italy. This case study highlights the role of advanced survey techniques, such as Mobile Mapping Systems and UAV-based photogrammetry, in generating extensive and comprehensive 3D datasets. It explores the possible contribution given by commercial web platforms for 3D data management in supporting design and development projects with reality capture data. Web platforms offer significant advantages by providing a shared environment for visualizing, interacting with, and processing large datasets. These platforms enable real-time collaboration among professionals with varying levels of expertise, fostering data sharing and reducing the need for local computing resources. This article provides a comparison methodology by testing and evaluating four commercial platforms: FlyVast, ATIS.cloud, Flai, and Cintoo, testing their functionalities and performances. It is then presented the adopted platform, Cintoo, and explained the reasons of the choice and its suitability for the project goal. Finally, it is provided an overview of the principal reasons that support the adoption of web platforms for sharing and collaboration supported by accurate 3D data, focusing on the potential cost reduction and efficiency.
The increasing request for digitized data among several fields, including the built environment and cultural heritage (CH), highlights the need for proficient ways to access, archive, and share 3D data and related information among users. The sector of reality capture produces accurate and reliable products that can support building management and CH maintenance, at the price of heavy and resource-demanding data. An emerging solution to this problem is represented by the web platforms for 3D data management, that promise to relieve users from the costs of archive and hardware, providing effective visualization, access and sharing tools. The panorama of commercial web platforms is analyzed according to the Software-as-a-Service business model, and the features of some representative platforms are exposed. The paper discusses the main advantages of diffused access and collaboration and the potential issues concerning long-term archival and data persistence. It provides a general overview of the main available platforms and describes their main features, comparing their specific pros and cons according to their category. The future perspectives of the web platform sector are promising as, according to the current development path, they may be able to empower built environments and the CH sector with a diffused, systematic, and conscious use of 3D data.
Nowadays, fisheye image has become commonly used in the 3D reality capturing field. Although AI integration for image recognition has become mature with normal images, providing available annotated dataset and pre-trained models, its application for fisheye images is rarely seen. While the object detection models have generalization ability, dealing with barrel distortion requires specific data for fine-tuning. This paper seeks to acquire prior knowledge from normal image and transfer it to the application that deal with fisheye images. This research is devoted to test the annotation shape that could possibly improve the accuracy when representing the shape of objects. It also seeks a way to prove that the annotation can be converted to fisheye images, resulted into a pre-process, which will facilitate the data preparation process. The tests involve annotations with standard box and quadrilateral polygon, the later turned out to be preserving most of the wanted image content after the conversion. The test result shows that the model trained on converted annotations using quadrilateral polygons, compared to detection model trained on non-converted ones, improves the mean average precision by 8%.
The advent of mobile mapping systems (MMSs) and computer vision algorithms has enriched a wide range of navigation and mapping tasks such as localisation, 3D motion estimation and 3D mapping. This study focuses on Visual Simultaneous Localisation and Mapping (V-SLAM) in the context of two in-houses MMSs: Ant3D, a patented five-fisheye multi-camera rig and GeoRizon, a high-resolution stereo fisheye rig. The aim is to leverage V-SLAM to enhance the systems performance in near-real-time and non-real-time 3D reconstruction applications. The research investigates both Monocular and Stereo V-SLAM applied to both MMSs and tackles the challenge of combining the V-SLAM estimated trajectory of one or a pair of cameras with known multi-camera relative orientation. We propose a state-of-the-art code that serves as a flexible and extensible platform for MMSs image acquisition and processing, along with an adapted version of the well-established ORB-SLAM3.0. Evaluation is performed in a cultural heritage challenging setup: the Minguzzi spiral staircase in the Duomo di Milano Cathedral. Performed tests highlight that introducing V-SLAM trajectories as well as pre-calibrated interior orientation and multi-camera constraints improve speed, applicability and accuracy of 3D surveys.