Three-dimensional imaging technology has significant application value in endoscopic medicine. Accurate 3D depth information can assist both surgical procedures and routine diagnosis. This paper proposes a narrow-baseline binocular endoscopic system and a corresponding 3D reconstruction method. After completing binocular calibration and epipolar rectification, Foundation Stereo is introduced to achieve dense disparity estimation. Combined with chessboard calibration information for scale correction, 3D reconstruction results with real-world metric scale are obtained. The proposed method is validated on both real endoscopic images and the Middlebury 2014 benchmark dataset, and is compared with the traditional Semi-Global Block Matching (SGBM) method. Experimental results show that in real endoscopic scenarios, the proposed approach can generate structurally continuous 3D reconstructions with reasonable geometric shapes under complex imaging conditions. The results provide a feasible technical route and experimental basis for narrow-baseline binocular endoscopic 3D imaging.
Grid-laser structured light is an effective technique for industrial three-dimensional (3D) inspection, and diffractive optical elements (DOEs) offer a compact, low-cost, and low-power means of generating dense grid patterns. However, DOE projection introduces non-uniform energy distribution, distance-dependent stripe-width variation, diffraction-induced distortion, intersection coupling, and local stripe discontinuities. These effects cause centerline drift, false responses, broken observations, and unstable stereo correspondences, ultimately degrading reconstruction accuracy.To address these challenges, this paper proposes a binocular DOE grid-laser 3D reconstruction framework that integrates grid-intersection detection, topology-guided stereo matching, continuity-enhanced sub-pixel centerline extraction, and segment-constrained triangulation. Grid intersections are detected as structural anchors and assigned consistent lattice coordinates across the stereo pair. Laser center points are then extracted through multi-scale derivative responses, adaptive thresholding, quality-gated zero-crossing localization, intensity-centroid correction, and geometric segment regularization. The recovered grid topology further restricts cross-view laser correspondence to one-dimensional segment corridors, reducing matching ambiguity in repetitive grid patterns. Experiments covering representative objects, exposure variation, noise perturbation, and standard reference geometries show that the proposed method improves centerline continuity and reconstruction stability. Geometric validation on plane and sphere references demonstrates sub-millimeter precision, with residual standard deviations below 0.17mm. These results confirm that tightly coupling image-domain centerline extraction with grid topology and calibrated stereo geometry provides a reliable solution for high-precision DOE grid-laser 3D reconstruction.
Telecentric Fringe Projection Profilometry (TFPP) is widely used in high-precision industrial 3D measurement due to its constant magnification and elimination of parallax. However, micron-level accuracy is often limited by the numerical instability of implicit reconstruction algorithms and systematic distortions caused by parameter coupling in conventional 2D calibration. In this paper, we propose a high-precision telecentric 3D measurement framework. An explicit closed-form analytical reconstruction model is derived for the hybrid telecentric–pinhole system, which avoids matrix inversion and significantly improves numerical stability and computational efficiency. To further suppress coupling-induced distortions, a global parameter optimization strategy based on standard gauge blocks is introduced to refine system parameters through embedded geometric constraints. Experimental results show that the proposed method outperforms conventional telecentric reconstruction pipelines, reducing the RMSE of planarity reconstruction by 59.3
In structured-light (SL) 3D measurement systems, underexposed coded fringe images (UCFIs) often fail to decode and reconstruct sparse point clouds due to low visibility. Unfortunately, existing low-light enhancement algorithms primarily target natural non-coded images, cannot faithfully preserve the encoded information of the fringe patterns, and are prone to introducing undesirable visual artifacts, noise, and reconstruction distortion. To this end, we first propose a novel bidirectional coded content similarity (B2CS) rule, specifically designed to enhance UCFIs captured by SL systems. Based on this rule, it reveals that the failure reason of existing methods occurs by breaking coded content similarity between under- and well-exposed fringes, and accordingly we construct a hybrid L2-L1-L0 norm-constrained decomposition model with illumination-reflectance-noise priors. Considering the non-convexity caused by the joint priors, we design an effective variational decomposition framework and employ the alternating direction method of multipliers (ADMM) for efficient optimization. Based on the reliable decomposed solution, the enhanced UCFIs can assist the structured light algorithm in recovering structurally rich and high-density point clouds. Extensive experiments on real-world SL datasets demonstrate that our approach outperforms state-of-the-art methods in both subjective and objective evaluation, and significantly improves the completeness of the 3D reconstructed point cloud (by ≥ 20 %).
Phase measurement deflectometry (PMD) is widely recognized for its high precision and full-field capability in measuring the three-dimensional (3D) shapes of specular surfaces. By analyzing the deformation of structured light patterns, PMD accurately reconstructs detailed surface geometry. However, monocular PMD (Mono-PMD) systems often struggle to decouple height from gradient information, which reduces the accuracy of gradient integration during reconstruction. To overcome this limitation, we propose an iterative reconstruction method augmented by the integration of a line laser into the Mono-PMD framework. The proposed method begins with the use of a calibrated line laser model to precisely acquire the 3D coordinates of multiple points on the specular surface. The continuous laser line, whose contour closely follows the object’s morphology, yields a spatially distributed set of seed points rich in geometric information. Leveraging the spatial distribution and height data of these seed points, the method iteratively refines the surface model by statistically estimating the global height offset and applying a linear correction at each iteration. This process progressively enhances the morphological accuracy until convergence, resulting in a high-fidelity reconstruction. Extensive experimental validation demonstrates the efficacy and enhanced precision of the proposed method, confirming its potential for robust and accurate 3D surface reconstruction.
In oral healthcare and digital dentistry, structured-light-based 3D reconstruction is widely used for tooth scanning; however, the lack of distinctive textures and the high similarity among tooth morphologies make point-cloud registration particularly challenging, often leading to tracking loss. We present a robust and efficient global relocalization framework designed for tracking-loss recovery. It first enhances surface texture with adaptive histogram equalization so that ORB can reliably extract rich 2D keypoints; these keypoints are then lifted to produce a sparse, high-quality set of 3D candidate points. Building on this sampling strategy, we compute FPFH descriptors guided by ORB salient features and replace costly k-d tree queries with a grid-based indexing scheme and depth-consistency checks. This reduces neighborhood queries for M ORB-lifted 3D keypoints from O(M logN) to O(M), where M ≪ N. Finally, we employ TurboReg, a linear-complexity graph-based solver, to obtain deterministic, millisecond-level backend pose estimates. Experiments on dental datasets simulating severe tracking loss demonstrate an over 11× acceleration of the full pipeline while preserving high local accuracy, with a 0.020 mm inlier RMSE and up to 83.8% inlier ratio under large viewpoint variations.
Acquiring complete and high-fidelity 3D facial geometry is essential for applications such as digital avatar generation and biometric analysis. However, existing methods struggle to balance high-resolution data acquisition with efficient processing, often suffering from heavy computational costs and cumulative drift during rapid head movements. To address these challenges, we present a complete high-speed 3D facial reconstruction system that tightly integrates a custom monocular structured light scanner with an efficient, globally consistent registration pipeline. First, our hardware system acquires 1.77-megapixel facial geometry at 35Hz. To handle this massive data stream, we implement a spatiotemporal keyframe selection strategy based on spherical grid sampling to effectively filter redundant frames. For global registration, we introduce an angular-guided pipeline. By utilizing a spherical projection mechanism, we bypass exhaustive pairwise feature matching, thereby significantly reducing the graph construction complexity. Furthermore, by replacing the heavy non-linear Pose Graph Optimization (PGO) backend with cascaded matrix multiplications along a Minimum Spanning Tree (MST) structure, we completely bypass the time-consuming optimization phase. Experimental results demonstrate that our entire system achieves a $\mathbf{1 4. 8} \times$ overall speedup compared to standard PGO baselines, while maintaining robust sub-millimeter geometric precision (Inlier RMSE of 0.339 mm) under large pose variations.
Due to the lack of large-scale, accurately annotated 3D facial datasets, most current 3D facial landmark detection algorithms rely on 2D texture assistance or non-real digital 3D faces. The performance of these algorithms is limited by the accuracy of texture mapping and the difference between digital faces and actual faces. To tackle these challenges, we built a large-scale, high-precision, and accurately annotated 3D facial database using a structured light system. Building upon this foundation, we proposed a novel point cloud sampling method and 3D facial landmark detection algorithm. This method utilizes a curvature-fused graph attention network to directly predict landmark coordinates from 3D point clouds. Firstly, we extracted a simplifying point set carrying curvature information from the original 3D facial point cloud via geometric point sampling. Then, we incorporated curvature-encoded positional information as the learning component of the attention module and employed it as a feature extractor to construct the network. We evaluated the performance of the method on BU-3DFE, FaceScape and our custom dataset. Compared to existing facial landmark detection algorithms, our model achieved higher accuracy. To facilitate future research on face related applications, we have made the database available on GitHub at https://github.com/CCtwelve/Face-dataset/tree/main
With the rapid advancement of industrial and intelligent manufacturing, robotic automation has become a key approach for improving efficiency, reducing costs, and enabling flexible production. Precision robotic assembly, as a critical stage in the manufacturing process, directly influences the performance, lifespan, and reliability of products. This study addresses the challenge of insufficient localization accuracy in complex assembly tasks caused by missing features or a limited field of view (FOV). We propose a multi-stage target localization method that integrates a large FOV 2D camera, deep learning, robotic manipulation, and 3D scanning reconstruction, specifically designed for robotic precision assembly scenarios where conventional 3D cameras suffer from FOV limitations. Our method economically overcomes the single 3D camera's FOV constraints by constructing an integrated system capable of fully capturing the assembly area. This enhances the accuracy and automation of industrial assembly tasks and improves the precision positioning of robots in complex assembly operations. Experimental results on a self-constructed HDMI dataset, which was collected with a 3D structured light camera in a robotic assembly environment, demonstrate that the proposed pose estimation method achieves an accuracy of 91.12%. Further testing under USB and VGA interfaces demonstrates that this method maintains good generalization and robustness under various input conditions.
The precision of vision-based 3-D scanning techniques, such as structured light (SL) means is usually dependent on the scanning area and resolution of the imaging sensors. However, large objects typically require scanning from multiple views, followed by the stitching of point clouds. This process becomes particularly challenging when dealing with featureless surfaces. For instance, existing vehicle inspection methods struggle to handle automobile components with large, smooth, and featureless surfaces. Generally, these methods involve either manually placing markers during the scanning process or performing iterative fine registration after robotic scanning. The manual marking method is cumbersome and prone to introduce contact errors, while the smooth surfaces of the components make it difficult to extract the distinct features needed for fine registration. Both methods have limitations when it comes to completely measuring large and featureless surfaces. In response, this article introduces a novel, high-accuracy point cloud stitching approach for large, featureless objects, employing robotic scanning alongside projected grid lasers. We unveil a corner extraction method for projected grid features that capitalizes on skeleton lines and rotational symmetry. To enhance the accuracy, we leverage a shape-based feature selector and a registration method based on the square root cubature Kalman filter (SRCKF) for effective noise suppression. In addition, a global search strategy is employed for final registration, by balancing costs weighted by overlapping areas and registration errors across multiple views. Our comprehensive experiments confirm the accuracy and reliability of the proposed method for on-site inspection of large surfaces devoid of distinct features.
In photogrammetry, developing specialized measurement equipment tailored to diverse scenarios and requirements often results in heightened operational complexity and increased costs. To mitigate these challenges, this study introduces a Versatile Handheld Tracking Target (VHT-T) utilizing infrared markers. The VHT-T offers seamless integration with various terminal devices, ensuring a cost-effective and efficient solution for measurement needs. The VHT-T adopts a multi-marker staggered planar constraint structure, and a marker detection and tracking algorithm has been developed to enhance its robustness and adaptability. Experimental results demonstrate that the system achieves stable matching within the range of -40 degrees to +40 degrees when rotating around the X-axis and Y-axis, while maintaining robust matching across the full 360 degrees range when rotating around the Z-axis. This paper further explores the integration of the VHT-T with two typical terminal devices: (1) when combined with a measurement probe, forming Combined Structure A (CS-A), representing contact-based measurement. For this configuration, a self-calibration algorithm based on rotational spherical constraints and a multi-station tracking method are proposed, achieving handheld coordinate measurement functionality; (2) when combined with a structured light camera, forming Combined Structure B (CS-B), representing non-contact measurement. For this configuration, a self-calibration algorithm based on third-party corner features and a multi-station stitching method are introduced to facilitate handheld surface measurement. These two integration methods provide a feasible reference framework for combining the VHT-T with other terminal devices. Experimental results demonstrate that the VHT-T system can meet tracking requirements in various scenarios for both coordinate and surface measurements, achieving efficient and stable performance.
Hand-eye coordination is one of the important ways to enhance the flexibility of industrial robots. The mid-joint configured camera is an emerging configuration mode that combines the operational flexibility of the eye-in-hand mode with the imaging stability of the eye-to-hand mode. However, the mismatch between the camera and the robot’s degrees of freedom (DOF) prevents traditional visual servo methods from fully playing their guiding role. This article studies a vision-guided positioning method for a typical cooperative system with mismatched DOF between the hand and eye. First, the kinematic model and Jacobian matrix of the hand-eye system are established. The control difference caused by the mismatching of DOF is clarified. A depth estimation method is introduced that can calculate depth values without hand-eye parameters. On this basis, a complementary visual servo algorithm is developed. Independent controllers are designed to control the manipulator in different DOF, and algorithm fusion is carried out. The Lyapunov method is adopted to prove the system’s stability. An adaptive adjustment algorithm has been designed to accelerate system convergence and overcome the accuracy degradation caused by system friction. Experiments are carried out to verify the high positioning accuracy and good control performance of the proposed algorithm.
Facial paralysis is a prevalent disorder affecting the facial nerve. In clinical settings, the severity of facial paralysis is typically assessed by physicians based on their subjective experience, evaluating the range of facial muscle movements and facial symmetry. To address these limitations, this paper proposes a method for the quantifiable evaluation of facial paralysis. We have developed a multi-view real-time facial acquisition system utilizing three infrared structured light units, enabling high-precision dynamic capture. Through non-rigid registration of the collected point cloud sequences, we generated a unified topological mesh sequence. Subsequently, we employed a novel facial asymmetry operator to quantitatively assess facial paralysis. Extensive experimental results demonstrate that the proposed method is both effective and accurate.
In the field of structured light technique, it is a developing trend to use MEMS mirror instead of Digital Light Processing projector as an active light-emitting device with its advantages such as smaller size, reduced power consumption and faster scanning speeds. Given that MEMS mirrors operate as lens-free optical emitters and are limited to projecting uni-axial patterns, the classic pinhole model and projected-checkerboard-based calibration method are infeasible for the MEMS-based structured light system. Thus, the modeling and calibration of the MEMS-based structured light system present a formidable challenge. In this paper, we propose an improved pinhole model by introducing a linear transition function to model MEMS mirrors, along with a novel two-step calibration method to calibrate the system parameters. The proposed model, containing only 32 parameters, simplifies the calibration procedure without compromising accuracy. With no need of the projected checkerboard, the proposed two-step calibration method can acquire the system parameters accurately. In experimental parts, the ablation study on the proposed model and the convergence of the two-step calibration method are discussed firstly. The calibration and reconstruction accuracy based on the proposed model and calibration method are validated by 3D reconstruction of standard and free-from objects quantitatively and qualitatively.
To achieve accurate and robust object detection in the real-world scenario, various forms of images are incorporated, such as color, thermal, and depth. However, multimodal data often suffer from the position shift problem, i.e., the image pair is not strictly aligned, making one object has different positions in different modalities. For the deep learning method, this problem makes it difficult to fuse multimodal features and puzzles the convolutional neural network (CNN) training. In this article, we propose a general multimodal detector named aligned region CNN (AR-CNN) to tackle the position shift problem. First, a region feature (RF) alignment module with adjacent similarity constraint is designed to consistently predict the position shift between two modalities and adaptively align the cross-modal RFs. Second, we propose a novel region of interest (RoI) jitter strategy to improve the robustness to unexpected shift patterns. Third, we present a new multimodal feature fusion method that selects the more reliable feature and suppresses the less useful one via feature reweighting. In addition, by locating bounding boxes in both modalities and building their relationships, we provide novel multimodal labeling named KAIST-Paired. Extensive experiments on 2-D and 3-D object detection, RGB-T, and RGB-D datasets demonstrate the effectiveness and robustness of our method.
With the rapid advancement of marine technology, underwater 3D structured light reconstruction has emerged as a pivotal tool for ocean exploration, marine robotics, and underwater assembly. However, the aquatic environment typically contains numerous suspended particles that scatter and absorb signal light from the target, degrading underwater image quality and leading to unsatisfactory results with traditional image-based 3D reconstruction methods. To address this challenge, we have developed a temporal-multiplexing structured light system with polarization imaging to obtain dense and accurate target geometry information. Subsequently, we propose an underwater turbidity removal network based on this system. This network mitigates the impact of underwater turbidity on imaging accuracy and quality, significantly enhancing the precision and effectiveness of 3D reconstruction. Initially, we built a polarization system to capture fringe images of underwater objects at four different polarization angles. By calculating the second Stokes parameter and the degree of polarization, and then combining Canny features as input for the global contour capture block, we effectively extract the object's contour information, restoring the fringe boundaries along its edges. Additionally, we introduce a fringe capture block to refine the boundaries of the fringe patterns, resulting in clearer internal reconstructions of the objects. By leveraging both the global contour capture block and the fringe capture block, we achieve high-precision underwater structured light reconstruction. Extensive experimental results demonstrate that our model achieves outstanding performance in terms of PSNR, SSIM, and LPIPS, with scores of 21.226, 0.969, and 0.050, respectively-significantly outperforming the reconstruction accuracy of state-of-the-art methods.
With the rapid development of artificial intelligence and computer vision, numerous technologies have been introduced to automate manufacturing in the industry. Typical metal workpieces in the industry often have highly reflective surfaces, come in various sizes, and are positioned irregularly. The motor rotor presented in this paper is one such representative workpiece. Traditional grasping methods for workpiece loading and unloading are pre-programmed and often struggle to cope with complex and disordered situations. In this paper, we introduce a structured light (SL) sensor as the visual guide for the triaxial robot. Furthermore, we propose a high-precision hand-eye calibration method for the non-orthogonal coordinate system of the triaxial robot. Additionally, a motor rotor center localization method based on U-Net image segmentation is proposed. By combining the high-precision hand-eye calibration and localization, we can accurately and automatically locate and grasp the rotor. We have conducted sufficient experiments to verify the effectiveness and accuracy of our system.
With the rapid development of automatic driving field, automatic parking has become increasingly concerned commercially. SLAM (Simultaneous Localization and Mapping) as an important technology is applied to autonomous valet parking (AVP) in recent years. However, it is difficult to take full advantage of the data from vision sensors owing to abundant similar elements in the parking lot. In this paper, a semantic SLAM based on slot number in parking lot is proposed. The system recognizes the slot number as semantic markers by a CNN (Convolution Neural Network) and classifies the slot according to the semantic markers. Then a semantic ICP (Iterative Closest Point) algorithm is used for mapping and localization modules. The results of the experiments on our simulation dataset show that this method performs well and has a better accuracy than traditional ICP. This work may have a promotion for autonomous valet parking and be applied to commercial products one day.