In situations with a limited number of posed images, choosing the most suitable viewpoints becomes crucial for accurate Neural Radiance Fields (NeRF) modeling. Current approaches for view selection often rely on heuristic methods or are computationally intensive. To address these challenges, we introduce a new framework, OptiViewNeRF, which leverages scene uncertainty to guide the view selection process. Initially, an uncertainty estimation model of the entire scene is developed based on a preliminary NeRF model. This model then informs the selection of new perception viewpoints using a batch view selection strategy, allowing the entire process to be completed in a single iteration. By selecting viewpoints that provide informative data, this approach improves novel view synthesis results and accurately reconstructs 3D scenes. Experimental results on two selected datasets show that the proposed method effectively identifies informative viewpoints, resulting in more accurate scene reconstructions compared to baseline and state-of-the-art methods.
Accurate localization in GPS-denied environments has always been a core issue in computer vision and robotics research. In indoor environments, vision-based localization methods are susceptible to changes in lighting conditions, viewing angles, and environmental factors, resulting in localization failures or limited generalization capabilities. In this paper, we propose the TransCNNLoc framework, which consists of an encoding–decoding network designed to learn more robust image features for camera pose estimation. In the image feature encoding stage, CNN and Swin Transformer are integrated to construct the image feature encoding module, enabling the network to fully extract global context and local features from images. In the decoding stage, multi-level image features are decoded through cross-layer connections while computing per-pixel feature weight maps. To enhance the framework’s robustness to dynamic objects, a dynamic object recognition network is introduced to optimize the feature weights. Finally, a multi-level iterative optimization from coarse to fine levels is performed to recover six degrees of freedom camera pose. Experiments were conducted on the publicly available 7scenes dataset as well as a dataset collected under changing lighting conditions and dynamic scenes for accuracy validation and analysis. The experimental results demonstrate that the proposed TransCNNLoc framework exhibits superior adaptability to dynamic scenes and lighting changes. In the context of static environments within publicly available datasets, the localization technique introduced in this study attains a maximal precision of up to 5 centimeters, consistently achieving superior outcomes across a majority of the scenarios. Under the conditions of dynamic scenes and fluctuating illumination, this approach demonstrates an enhanced precision capability, reaching up to 3 centimeters. This represents a substantial refinement from the decimeter scale to a centimeter scale in precision, marking a significant advancement over the existing state-of-the-art (SOTA) algorithms. The open-source repository for the method proposed in this paper can be found at the following URL: github.com/Geelooo/TransCNNloc.
Trees are an important part of the cityscape,and 3D models of trees are indispensable for real-time 3D design,construction of vir-tual geographic environments,and construction of digital twin cities.Current 3D models of trees are reconstructed based on images or model libraries.The former show cluttered triangular network clusters,and the latter are vastly different from the real situation in terms of geomet-ric expression and realism,which makes directly using the reconstructed tree models in the practical applications of smart cities difficult.Therefore,in this paper,a bionic reconstruction method for 3D tree models is proposed based on high-precision laser scanning point cloud data for building realistic scenes in virtual geographic environments,which enables the automated reconstruction of 3D tree models at mul-tiple levels of detail while preserving morphological features. First,a skeleton-based parametric tree model reconstruction method that extracts branch geometry by generalized cylinder fitting and extracts the trunk,main branches,models of fine branches,and crown elements in a hierarchical manner according to the growth parameters of the tree is proposed.Second,the refinement requirements of modeling distinct parts of trees are considered,and a refined tree geometry reconstruction method by integrating the conformal Poisson network and parametric fitting is presented.Finally,the texture mapping method is applied to map the texture of multilevel tree branches automatically to achieve a detailed 3D reconstruction of tree models by considering the texture extension of the tree structure.Based on the laser point cloud acquired with a backpack or station,this method can produce a re-fined 3D tree model with high accuracy of morphological features. The overall geometric error of the model is better than 10 cm,and the geometric error of the trunk model is better than 3 cm.Under the same data conditions,the method has the highest degree of reproduction of 3D tree morphology and real texture compared with various mainstream tree modeling methods.Based on the results of this paper,the method can further advance the extraction of tree structure infor-mation and the calculation of 3D green volume for the realistic 3D China and national strategies such as green low-carbon development,which have great practical value. This paper proposes a 3D bionic reconstruction method for constructing high-fidelity scenes in virtual geographic environments to achieve highly accurate geometric reconstruction and texture mapping of individual tree roots,trunks,branches,and leaves.The core of the method is to consider the requirements of distinct parts of the tree reconstruction at multiple levels of detail and integrate Poisson mesh and parameter fitting to complete the 3D reconstruction of the tree with high accuracy.The experimental results show the proposed tree 3D re-construction method provides a highly accurate reconstruction of the tree geometry and texture.The research results are used for the accurate extraction of tree parameters,which can provide an important basis for tree structure information extraction,3D green volume calculation,and realistic modeling and simulation of virtual geographic environments.
This paper addresses the challenge of indoor space segmentation from 3D point clouds, which is essential for understanding interior layouts, reconstructing 3D structures, and developing indoor navigation maps. While current deep learning-based methods rely on projecting 3D point clouds into 2D for instance extraction, they often fail to capture the local and global 3D features necessary for effectively segmenting complex indoor spaces, such as multi-ring nested structures. These methods also struggle with generalization across different scenes. In response, this paper proposes an efficient indoor space segmentation method that integrates both 2D and 3D geometric constraints. By leveraging the distribution characteristics of point clouds in 2D and the local and global features in 3D, the method achieves reliable extraction of vertical structural information in complex indoor environments. To address under-segmentation in small spaces due to varying scales, the paper introduces an adaptive extraction method for space partition anchors, guided by local features. During instance-level space segmentation, a hierarchical contour tree structure is employed to precisely partition complex indoor spaces, effectively handling circular and composite structures. The proposed approach was tested on 96 RGB-D scans from the Beike dataset and 6 large-scale indoor scenes from the S3DIS dataset, covering a range of complexities, sizes, and structures. The experimental section includes ablation studies and thorough comparisons with existing state-of-the-art spatial partitioning algorithms based on morphology and deep learning. Results demonstrate that the proposed method significantly outperforms existing approaches in terms of accuracy, robustness, and generalization ability, providing a solid foundation for indoor space modeling and robotic navigation. The source code and datasets will be made publicly available via the “EISPGeo” link.
Large datasets are required to develop Artificial Intelligence (AI) models in AI powered smart farming for reducing farmers' routine workload, this paper contributes the first large lion-head goose dataset GooseDetectlion, which consists of 2,660 images and 98,111 bounding box annotations. The dataset was collected with 6 cameras deployed in a goose farm in Chenghai district of Shantou city, Guangdong province, China. Images sampled from videos collected during July 9 -10 in 2022 were fully annotated by a team of fifty volunteers. Compared with another 6 well known animal datasets in literature, our dataset has higher capacity and density, which provides a challenging detection benchmark for main stream object detectors. Six state-of-the-art object detectors have been selected to be evaluated on the GooseDetectlion, which includes one two-stage anchor-based detector, three one-stage anchor-based detectors, as well as two one-stage anchor-free detectors. The results suggest that the one-stage anchor-based detector You Only Look Once version 5 (YOLO v5) achieves the best overall performance in terms of detection precision, model size and inference efficiency.
This paper presents a novel fully automated approach for generating structured 3D synthetic tree models, addressing the limitations of existing datasets used in applications like digital twin construction, carbon stock calculation, and environmental assessments. The method allows for the automated creation of a large-scale dataset containing 13,000 tree models of ten common species, each featuring a detailed 3D point cloud with hierarchical structures, precise parameters, and separate branch and leaf information. The dataset includes both original and noise-added point clouds to enhance method testing and evaluation. It stands out by providing comprehensive structural data, including branch numbering and detailed tree skeleton information with node hierarchies and radius data. Furthermore, this study introduces randomly distributed batches of tree models within specific terrains. It provides results from airborne laser scanning simulations, which facilitate the individualized segmentation of these tree models. This first-of-its-kind, extensive synthetic dataset is designed for accurate algorithm evaluation in tasks such as branch-leaf separation, 3D reconstruction, individual tree segmentation and carbon stock estimation. The paper validates the dataset’s utility by applying state-of-the-art algorithms to demonstrate its effectiveness in various applications, marking a significant advancement in 3D tree modeling research. The datasets are publicly available, accessible via the ‘‘TreeNet3D Dataset’’ link.
With the rapid development of autonomous driving and SLAM technology, the perception system of a vehicle heavily relies on laser and image sensors to capture the real-world scenario and avoid obstacles autonomously. To achieve accurate and robust multi-sensor fusion computation, high-precision extrinsic calibration of camera and laser scanner is a necessary requirement. Traditional multi-sensor calibration methods based on manual features rely on specific scenarios and may not provide feature information over long distances. In this paper, we present a novel approach for robustly calibrating the extrinsic parameters of a solid-state(SS) lidar-camera system in a natural environment. Our proposed method begins with obtaining robust line feature information. we first innovatively employ a super-voxel clustering method to extract global 3D line features from the complete point cloud and then back-project these 3D line features into 2D space. Afterward, a transformer-based edge detection network, EDTER, is used to detect the edge features and estimate the probability pixel-by-pixel. To consider the uncertainty of two-dimensional line features and the inconsistency of residuals at different distances, we construct a line feature weight model for line feature residual calculation. Finally, we minimize the residual errors using least squares optimization to recover the relative pose of the camera and the lidar sensor. We conducted a performance study to compare our proposed method against existing targetless calibration methods on various natural scenarios. The experimental results demonstrate that our proposed method achieves higher robustness, accuracy, and consistency, making it suitable for real-world applications.
The dense urban metro network plays an important role in urban transportation, and it is becoming increasingly important to improve the operation and management of metro facilities. The monitoring and management of passenger flow is a main concern in metro operation, and the reliable analysis of passenger flow can greatly improve the operational efficiency and safety of a metro station. Therefore, using the long-time in-and-out smart card data of Shenzhen metro stations, this paper proposes a Coarse-to-Fine passenger flow analysis method for the characterization of passenger flow on multiple time scales. This method, which is proposed from a new perspective based on the time series clustering of metro stations, precisely defines the peak travel hours and then extracts features for the evaluation of the crowdedness and disorderliness at each station. Finally, the metro stations are classified into nine levels according to those features. The stations that need urgent attention in terms of their passenger flow, including Shenzhen North Station, Buji, and Grand Theater, are identified for the reference of city managers.
针对现有通过检测窗户角点实现窗户检测方法中存在窗户误检的问题,该文在窗角点分组阶段,以建筑物立面窗户的分布规律及其自身的几何结构特征为依据,提出一种参数自适应的窗角点分组方法.该方法是在使用深度学习方法获取窗户4个角点坐标的基础上,结合窗户角点及其连线的空间位置关系、平行垂直关系,建立窗角点分组判别依据,实现对窗角点检测结果的准确划分,进而得到有效窗户检测结果.为验证该方法的有效性,选用4个公开数据集进行窗户检测实验,结果表明:该方法可有效支持多类图像数据、实现全自动化运行,且与现有方法相比,具有更高的检测精度.
The present work analyzes the application of deep learning in the context of digital twins (DTs) to promote the development of smart cities. According to the theoretical basis of DTs and the smart city construction, the five-dimensional DTs model is discussed to propose the conceptual framework of the DTs city. Then, edge computing technology is introduced to build an intelligent traffic perception system based on edge computing combined with DTs. Moreover, to improve the traffic scene recognition accuracy, the Single Shot MultiBox Detector (SSD) algorithm is optimized by the residual network, form the SSD-ResNet50 algorithm, and the DarkNet-53 is also improved. Finally, experiments are conducted to verify the effects of the improved algorithms and the data enhancement method. The experimental results indicate that the SSD-ResNet50 and the improved DarkNet-53 algorithm show fast training speed, high recognition accuracy, and favorable training effect. Compared with the original algorithms, the recognition time of the SSD-ResNet50 algorithm and the improved DarkNet-53 algorithm is reduced by 6.37ms and 4.25ms, respectively. The data enhancement method used in the present work is not only suitable for the algorithms reported here, but also has a good influence on other deep learning algorithms. Moreover, SSD-ResNet50 and improved DarkNet-53 algorithms have significant applicable advantages in the research of traffic sign target recognition. The rigorous research with appropriate methods and comprehensive results can offer effective reference for subsequent research on DTs cities.
High-quality 3-D point cloud maps are essential for precise indoor environment modeling. However, constructing such maps in multistory indoor environments is challenging due to the presence of narrow nonstructural spaces, such as staircases, corners, and corridors with similar textures. Simultaneous localization and mapping (SLAM) in these scenes is particularly difficult, as cumulative errors can lead to incorrect loop closures and drastic degradation in map quality. To address these challenges, this article proposed an SLAM method based on multiple ground constraints pose optimization (MGCPO), which uses a backpack light detection and ranging (LiDAR) system. The proposed method includes two novel modules. First, a regression analysis-based scenarios recognition (RASR) module provides a reference for the construction of ground constraints. Second, based on different scene detection results, the MGCPO module constrains the sensor pose using the floor plane to reduce localization errors and effectively decrease loop closure detection errors. Qualitative experiments demonstrate that our proposed method outperforms state-of-the-art methods in challenging scenarios. Quantitative experiments show that our method achieves an error rate of just 1.06% using only LiDAR sensors.
Objectives: To address the problem of internal inconsistency of classification targets in existing three dimensional(3D) point cloud data segmentation and classification methods. we propose a high-precision classification method for indoor point cloud jointly optimized by super voxel random forest and long short-term memory(LSTM) neural network. Methods: The method takes into account that the super voxel structure has the characteristics of internal feature consistency, divides the original point cloud into super voxels, and uses super voxels as the basic unit for multivariate feature calculation to build a super voxel random forest classification model for indoor point cloud to achieve coarse classification of point cloud data. On this basis, LSTM is introduced to train and predict the neural network model for the hyper voxel neighborhood connectivity of coarse classification to achieve the optimization of hyper voxel coarse classification results. The validity and accuracy of the proposed classification method are verified based on the open dataset.Results: The results show that the classification accuracy of the proposed classification method can reach 83.2% for 13 types of elements in the open dataset. The training data of the LSTM optimization network proposed in this paper used only the label information of region 1 for model training, while other deep learning frameworks used regions 1-5 for model training, so from the perspective of training data requirements, the point cloud data classification framework proposed in this paper can achieve a relatively better prediction result with a small portion of the training data set. The super voxel-based LSTM optimization method approach has high classification accuracy on objects with obvious set features such as ceiling, floor and wall,however, it is inferior to the deep learning algorithm RandLA-Net in classifying objects with complex structures such as chair, sofa and bookcase. Conclusions: In this paper, we consider the association characteristics between different types of elements embedded in the connection relations among super voxels, and introduce LSTM to train and predict the model for the coarse classification of super voxel neighborhood connection relations to achieve the optimization of coarse classification results of super voxels. The proposed method can achieve better classification accuracy when trained with small samples compared with the classical deep learning framework.
Obtaining building instance models directly through photogrammetry and other means results in high polygon counts and structural detail levels. Therefore, it is an important challenge to preserve the features of instant building models and generate lightweight building models from them. In this paper, we propose an improved lightweight reconstruction method for 3D building models based on planar primitives extraction, topology correction, and optimal planes selection. Improvements due to our method arise from three aspects: (1) After plane segmentation based on a simple region growing method, a second plane segmentation is performed on the building model based on the similarity of normal vectors. (2) According to the building structural characteristics, the initial plane segmentation is checked to generate new plane regions within the necessary connection regions. (3) The intersection of candidate planes is improved by enlarging the region of candidate faces, and the topological connection is improved. Furthermore, the topological relations of all planar primitives that have been optimized are recorded in an undirected graph. Finally, a watertight and manifold lightweight model is extracted from the faces of a candidate set by energy minimization. Experiments on different data sets show that the improved method is more reasonable in plane segmentation, and has superior performance in algorithm robustness and necessary structure recovery of buildings. Even when dealing with imperfect data, a watertight model is still obtained.
Despite offshore aquaculture brings great economic benefits, it also has destructive effects on the ecological environment of coastal regions. Therefore, the accurate monitoring of offshore aquaculture areas is vital. Existing methods for extracting aquaculture ponds are still limited by the lack of high-resolution hyperspectral imagery and inadequate feature extraction. In this paper, we proposed an unsupervised aquaculture ponds extraction method based on hyperspectral imagery super-resolution, feature fusion and stepwise extraction strategy. First, the resolution of the original hyperspectral imagery is enhanced via a deep learning-based super-resolution method. Then we introduce a feature fusion method by combining multi-dimensional features including elevation, reflectance, biochemistry, principal component analysis, and fast-Fourier transform features to enhance the feature sensitivity to the aquaculture ponds. The aquaculture ponds are finally extracted via a step-wise extraction strategy. Experiments on ZH-1 and GF-7 imageries are conducted, and the experimental results show that our method achieves an OA of 97.9% on the aquaculture pond extraction and that the method can be generalized for the object extraction on multispectral datasets. Ablation experiments show that all three of our innovative modules can effectively improve the extraction accuracy and confirm the generalization of the method.
The limited amount of high-quality training data available in indoor understanding with deep learning is a major problem. A possible solution to this problem is to use synthetic data to improve network training. In this study, a fully automatic method to generate synthetic noisy point clouds from as-built building information modeling (BIM) models is presented and it assesses the potential of these synthetic point clouds to improve deep neural network training. Based on a skeleton-guided strategy, all hypothetical scanning sites are located along the central axis of the buildings, which are obtained through equidistant sampling. Then, the synthetic labeled point cloud is generated station-by-station, and data augmentation is achieved using a random combination of data from different stations. The proposed approach involves generating over 44 sets of synthetic noisy point clouds based on BIM models. The performance of state-of-the-art (SOTA) deep learning methods in understanding indoor scenes enhanced by the synthetic point clouds is thoroughly assessed, and the effectiveness of various combinations of real and synthetic datasets is investigated. The experimental results demonstrate that leveraging synthetic point clouds generated from BIM models leads to a remarkable 5%–10% improvement in 3D semantic segmentation accuracy. The research signifies the value of synthetic point clouds as an effective tool for improving deep neural network training. All simulation datasets are publicly available, including original BIM models, full synthetic point clouds, and point clouds after IHPR processing, accessible via the BIMSyn Dataset link. In future research, an exploration of how synthetic point clouds will be further improved by considering specific characteristics of objects such as color, material reflectance, and illumination.
In order to reconstruct high-quality 3D tree models, trunks and crowns could be reconstructed using appropriate methods separately and merged together. During this process, gaps will appear after tree models are spliced, which will affect the models' topological connectivity and visual effects. In this paper, a gap-repair algorithm for tree mesh models based on boundary restriction and coordinates projection is proposed. The algorithm first extracts all the holes in the tree trunk meshes and crown meshes and matched them according to their relative positions. Then, the gaps in the tree meshes are identified. After that, based on the projection of the hole vertices from 3D space to 2D space, the relative positions of the hole vertices were determined, and vertex connection and surface addition were carried out to generate a 2-manifold patch. Finally, the patch was refined to keep its density of vertices similar to the surrounding meshes. Real 3D Tree meshes of different species were used to conduct experiments, and the quality of the repaired tree models are evaluated qualitatively and quantitatively. Experimental results revealed that the proposed algorithm could fill gaps in tree meshes quickly and effectively, so as to achieve the topology connection and smooth transition between the trunk and crown meshes.
This paper introduces a novel framework, Tree-GPT, which incorporates Large Language Models (LLMs) into the forestry remote sensing data workflow, thereby enhancing the efficiency of data analysis. Currently, LLMs are unable to extract or comprehend information from images and may generate inaccurate text due to a lack of domain knowledge, limiting their use in forestry data analysis. To address this issue, we propose a modular LLM expert system, Tree-GPT, that integrates image understanding modules, domain knowledge bases, and toolchains. This empowers LLMs with the ability to comprehend images, acquire accurate knowledge, generate code, and perform data analysis in a local environment. Specifically, the image understanding module extracts structured information from forest remote sensing images by utilizing automatic or interactive generation of prompts to guide the Segment Anything Model (SAM) in generating and selecting optimal tree segmentation results. The system then calculates tree structural parameters based on these results and stores them in a database. Upon receiving a specific natural language instruction, the LLM generates code based on a thought chain to accomplish the analysis task. The code is then executed by an LLM agent in a local environment and . For ecological parameter calculations, the system retrieves the corresponding knowledge from the knowledge base and inputs it into the LLM to guide the generation of accurate code. We tested this system on several tasks, including Search, Visualization, and Machine Learning Analysis. The prototype system performed well, demonstrating the potential for dynamic usage of LLMs in forestry research and environmental sciences.
Accurate individual tree reconstruction based on laser point clouds is vital for precise biomass estimation, virtual geographic environment modeling, and simulation. However, mobile laser 3D scanning systems often capture tree point clouds obscured by leaves, resulting in missing or incomplete branch data, posing significant challenges to detailed tree reconstruction. In this paper, a novel method for fine-grained 3D tree model reconstruction from incomplete point clouds is proposed. This approach addresses the structural reconstruction of trees, tackling two main problems: reconstructing the main trunk and branches based on data completeness. Initially, a 3D morphological algorithm is employed to separate the main trunk and branches in the point cloud. Next, branch point clouds are node-aggregated using clustering, and the trunk’s skeleton point is computed using a multi-scale curve fitting method. To account for incomplete branch point clouds, the Alpha shape is calculated and used as a constraint for the growth model of the L-system, enabling the automatic generation of branch structures. Finally, a morphologically constrained multi-level trunk fusion method is utilized to achieve complete tree model reconstruction. To validate the effectiveness of this method, five trees with varying structures and levels of complexity were selected for structural reconstruction and compared with state-of-the-art (SOTA) methods. Experimental results evince that the method delineated in this study exhibits an aptitude for adeptly approximating tree nodes, even in scenarios characterized by incomplete point cloud data. This approach transcends the reconstruction accuracy manifested by SOTA algorithms specific to three-dimensional tree modeling, simultaneously demonstrating a pronounced resilience to noise and enhanced robustness.
Street tree extraction based on the 3-D mobile mapping point cloud plays an important role in building smart cities and creating highly accurate urban street maps. Existing methods are often over- or under-segmented when segmenting overlapping street tree canopies and extracting geometrically complex trees. To address this problem, we propose a method based on improved 3-D morphological analysis for extracting street trees from mobile laser scanner (MLS) point clouds. First, the 3-D semantic point cloud segmentation framework based on deep learning is used for preclassification of the original point cloud to obtain the vegetation point cloud in the scene. Considering the influence of terrain unevenness, the vegetation point cloud is deterraformed and slice point cloud containing tree trunks is obtained through spatial filtering on height. On this basis, a voxel-based region growing method constrained with the changing rate of convex area is used to locate the stree trees. Then we propose a progressive tree crown segmentation method, which first completed the preliminary individual segmentation of the tree crown point cloud based on the voxel-based region growth constrained by the minimum increment rule, and then optimizes the crown edges by “valley” structure-based clustering. In this article, the proposed method is validated and the accuracy is evaluated using three sets of MLS datasets collected from different scenarios. The experimental results show that the method can effectively identify and localize street trees with different geometries and has a good segmentation effect for street trees with large adhesion between canopies. The accuracy and recall of tree localization are higher than 96.08% and 95.83%, respectively, and the average precision and recall of instance segmentation in three datasets are higher than 93.23% and 95.41%, respectively.
With the development of deep learning technology, a large number of indoor spatial applications, such as robotics and indoor navigation, have raised higher data requirements for indoor semantic model. However, creating deep learning classifiers requires a large number of labeled datasets, and the collection of such datasets requires a lot of manually labeling proces, which is labor-intensive and time-consuming. In this paper, we propose a method to automatically create 3D point clouds datasets with indoor semantic labels based on parametric BIM model. First, a automatic BIM generation method is proposed through simulating the structure of interior space Secondly, we use a viewpoint-guided labeled point cloud generation method to generate synthetic 3D point clouds with different labels, color information. Especially, noise are also simulated with a gaussian model. As shown in the experiments, the point cloud data with labels can be quickly obtained from existing BIM models, which will largely reduce the complexity of data labeling and improve efficiency. These simulated data can be used in the deep learning training process and improve the semantic segmentation accuracy.