Sphere-based calibration for camera-projector pairs achieves high accuracy with simple setups but suffers under projector illumination, which distorts sphere contours and complicates conic extraction. This paper presents a robust single-sphere calibration method that leverages the properties of active illumination to overcome this challenge. We first analyze the geometric structure of sphere boundaries under projector lighting and show that they originate from distinct apparent contours in the camera and projector views. Exploiting dual epipolar geometry, we segment these boundaries and fit separate image conics for each view. Based on this, we derive two novel orthogonality constraints between the conics and the image of the absolute conic, enabling complete calibration from a single sphere. To enhance precision, we further propose a nonlinear optimization strategy that jointly minimizes sphere reconstruction and point reprojection errors, refining both intrinsic and extrinsic parameters as well as lens distortions. Experimental results demonstrate that our method significantly enhances conic fitting accuracy under self-shadowing conditions and provides a flexible, accurate calibration solution with a single sphere.
Color-encoded single-shot fringe projection methods offer clear advantages in dynamic 3D measurement, but often suffer from distortion when measuring objects with complex surface textures, such as print circuit boards and camouflage-coated components. To address this issue, we propose a robust single-shot fringe projection profilometry approach that integrates enhanced morphological component analysis with fast Fourier transform-based fringe extraction. The method decomposes each RGB channel of the captured fringe image to suppress texture interference and employs a frequency-domain filtering strategy to isolate fringe information via a tailored fringe mask. By exploiting the energy concentration characteristics of fringe components in the frequency domain, fringe signals are directly extracted from the tunable Q-factor wavelet transform coefficients, thereby eliminating the need for iterative low-rank approximation or prior estimation of statistical parameters. This approach significantly improves computational efficiency and measurement accuracy while enhancing robustness to texture-induced artifacts. Additionally, it removes the need for pre-calibration of color channel compensation, improving adaptability to real-world conditions. Experimental results demonstrate that the proposed approach achieves accurate 3D reconstruction with superior performance in both processing speed and resistance to texture-induced artifacts.
The paper introduces a novel image stitching method driven by the projection model, utilizing epipolar displacement field (EDF) to ensure geometric consistency in panoramic images. By leveraging the principles of epipolar geometry and infinite homography, this method ensures alignment accuracy and maintains global projectivity across stitched images. The process begins by establishing a pixel warping rule in epipolar geometry through the infinite homography. Then, the epipolar displacement field, which quantifies the displacement of pixels along the epipolar lines, is constructed using thin-plate splines derived from the principle of local elastic deformation. The final panoramic image is produced by inversely warping pixels according to the epipolar displacement field. This method integrates epipolar constraints into the warping rules, ensuring high-quality alignment and preserving projection accuracy. Comparative experiments, both qualitative and quantitative, show that the method effectively reduces parallax artifacts and maintains geometric fidelity, achieving an average SSIM of 0.928 and PSNR of 29.900 across 12 scenes, outperforming existing techniques.
In structured light systems, the accuracy of measurement notably diminishes when assessing complex texture objects, especially encountering boundaries between various colors. To address this challenge, this paper meticulously analyzes and establishes an error model, elaborating the correlation between phase errors and the gradients of phase and gray-scale. Based on this analysis, a novel high-precision method is proposed for measuring complex texture objects via bidirectional fringe projection. This approach firstly leverages horizontal and vertical fringe projections to derive bidirectional phase information and calculates the angles between the tangent of the texture edges and the phase gradient. Subsequently, a refined temporal phase correction algorithm is formulated based on the epipolar matching algorithm and the devised error model, effectively mitigating numerical instability issues within the algorithm and significantly reducing errors of bidirectional phases. Ultimately, corrected point clouds are calculated based on bidirectional phases, and the obtained point clouds are merged to further diminish phase errors. Comparison experiments indicate that this method can reduce Mean Absolute Error (MAE) and Root Mean Square Error (RMSE) by 65.74% and 67.75%, respectively. Compared to existing methods, it improves performance by 27.29% and 33.74%, respectively, demonstrating superior performance.
In fringe projection profilometry (FPP), telecentric lens is usually used to narrow the measurement field of view (FOV) when measuring small objects. The imaging process of the telecentric system differs from the classical pinhole imaging principle, necessitating a new calibration method. Presently, calibration techniques used for telecentric systems persist in following the conventional practice of separating intrinsic and extrinsic parameters, leading to imprecise system parameters. To tackle this issue, we present a novel joint calibration approach founded on the homography matrix. Our approach concurrently optimizes both the intrinsic and extrinsic parameters, effectively mitigating issues associated with invalid extrinsic parameters and fluctuating intrinsic parameters observed in existing methods. In addition, the sign ambiguity concerning extrinsic parameters is successfully resolved, even in the absence of a precise positioning stage. This approach necessitates only a single adjustment in the pose of the planar calibration target, without any additional equipment requirements. Consequently, it can be flexibly applied across various domains where precise data acquisition of small objects is essential. This proposed method is evaluated in terms of calibration and reconstruction through contrast experiments, which demonstrate its feasibility and accuracy.
In structured light systems, measurement accuracy tends to decline significantly when evaluating complex textured surfaces, particularly at boundaries between different colors. To address this issue, this paper conducts a detailed analysis to develop an error model that illustrates the relationship between phase error and image characteristics, specifically the blur level, grayscale value, and grayscale gradient. Based on this model, a high-precision approach for measuring complex textured targets is introduced, employing a multiple filtering approach. This approach first applies a sequence of filters to vary the blur level of the captured patterns, allowing calculation of phase differences under different blur conditions. Then, these phase differences are used in the constructed error model to identify the critical parameter causing phase errors. Finally, phase recovery is performed using the calibrated parameter, effectively reducing errors caused by complex textures. Experimental comparisons exhibit that this method reduces the Mean Absolute Error (MAE) and Root Mean Square Error (RMSE) by 40.31% and 40.78%, respectively. In multiple experiments, its performance generally surpassed that of existing methods, demonstrating improved accuracy and robustness.
Spheres are commonly employed for both camera calibration and light calibration. The apparent contours of spheres are utilized to determine the camera’s intrinsic parameters, while the highlights on the sphere aid in estimating the light direction. This paper introduces a novel calibration method using spheres to simultaneously calibrate both the camera and light sources. The reflected irradiance distribution of the Lambertian sphere is analyzed under illumination from a near-field point light source. Consequently, points with the same luminance form coaxial circles. Their imaging equations are formulated under a general projection model, offering orthogonal constraints on the image of the absolute conic for fully calibrating the camera. Next, the geometric determination of the sphere centers and point light source positions is implemented. Finally, a photometric constraint is constructed to minimize the discrepancy between the synthetic and real images of the sphere, aiming to optimize the entire parameters of the system. The experiments are executed on synthetic and real datasets, with results indicating the effectiveness and precision of the proposed method.
Spectral image reconstruction is an important task in snapshot compressed imaging. This paper aims to propose a new end-to-end framework with iterative capabilities similar to a deep unfolding network to improve reconstruction accuracy, independent of optimization conditions, and to reduce the number of parameters. A novel framework called the reversible-prior-based method is proposed. Inspired by the reversibility of the optical path, the reversible-prior-based framework projects the reconstructions back into the measurement space, and then the residuals between the projected data and the real measurements are fed into the network for iteration. The reconstruction subnet in the network then learns the mapping of the residuals to the true values to improve reconstruction accuracy. In addition, a novel spectral-spatial transformer is proposed to account for the global correlation of spectral data in both spatial and spectral dimensions while balancing network depth and computational complexity, in response to the shortcomings of existing transformer-based denoising modules that ignore spatial texture features or learn local spatial features at the expense of global spatial features. Extensive experiments show that our SST-ReversibleNet significantly outperforms state-of-the-art methods on simulated and real HSI datasets, while requiring lower computational and storage costs. https://github.com/caizeyu1992/SST
In order to accurately estimate the position and pose of an object in the camera coordinate sys-tem in challenging scenes with severe occlusion and scarce texture,while also enhancing network efficien-cy and simplifying the network architecture,this paper proposed a 6-DoF pose estimation method using auxiliary learning based on RGB-D data.The network took the target object image patch,corresponding depth map,and CAD model as inputs.First,a dual-branch point cloud registration network was used to obtain predicted point clouds in both the model space and the camera space.Then,for the auxiliary learn-ing network,the target object image patch and the Depth-XYZ obtained from the depth map were input to the multi-modal feature extraction and fusion module,followed by coarse-to-fine pose estimation.The es-timated results were used as priors for optimizing the loss calculation.Finally,during the performance evaluation stage,the auxiliary learning branch was discarded and only the outputs of the dual-branch point cloud registration network are used for 6-DoF pose estimation using point pair feature matching.Experi-mental results indicate that the proposed method achieves AUC of 95.9%and ADD-S<2 cm of 99.0%in the YCB-Video dataset;ADD(-S)result of 99.4%in the LineMOD dataset;and ADD(-S)result of 71.3%in the LM-O dataset.Compared with existing 6-DoF pose estimation methods,our method using auxiliary learning has advantages in terms of model performance and significantly improves pose estimation accuracy.
3D reconstruction is a fundamental task in robotics and AI, providing a prerequisite for many related applications. Fringe projection profilometry is an efficient and non-contact method for generating 3D point clouds out of 2D images. However, during the actual measurement, it is inevitable to experiment with translucent objects, such as skin, marble, and fruit. Indirect illumination from these objects has substantially compromised the precision of 3D reconstruction via the contamination of 2D images. This paper presents a fast and accurate approach to correct for indirect illumination. The essential idea is to design a highly suitable network architecture founded on a precise error model that facilitates accurate error rectification. Initially, our method transforms the error generated by indirect illumination into a sine series. Based on this error model, the multilayer perceptron is more effective in error correction than traditional methods and convolutional neural networks. Our network was trained solely on simulated data but was tested on authentic images. Three sets of experiments, including two sets of comparison experiments, indicate that the designed network can efficiently rectify the error induced by indirect illumination.
This article presents a method for unwrapping the phase using only the geometric constraints and photometric information of the structured light system, which can cope with objects in a large depth range without the need for additional image acquisition or other cameras. Following the minimum phase method, this article also establishes an artificial plane and generates an initial unwrapped phase map from the absolute phase map of the plane. Starting from the reference plane, we divide the space into several $2\pi $ intervals pixel by pixel. The boundaries at phase discontinuities in the initial unwrapped phase map are then used to segment the object into independent regions. Regarding the projector as a point light source, synthetic images of each region corresponding to different spatial $2\pi $ intervals are generated sequentially according to Lambert’s cosine law. The interval in which the region corresponds is identified by comparing the similarity of these synthetic images with the modulation image, leading to the acquisition of an entire absolute phase map of the object. Experiments are conducted to evaluate the performance of the proposed method.
Objective A video stitching method based on dense viewpoint interpolation is proposed to solve the problem of artifacts and defects caused by parallax when stitching under wide baseline scenes. Video stitching technology can facilitate access to a broader field of view and plays a vital role in security surveillance, intelligent driving, virtual reality, and video conferencing. One of the biggest challenges of the stitching task is the parallax. When the cameras' optical centers perfectly coincide, they are unaffected by parallax and can easily synthesize perfect images. However, achieving the complete coincidence of camera optical centers in practical applications is not easy. The cameras are also scattered in some scenes, such as vehicle-mounted panoramic systems and wide field security surveillance systems. Therefore, it is important to study the problem of stitching in wide baseline scenes. A standard method uses a global homography matrix for alignment, but it has no parallax processing capability, which results in obvious flaws in wide baseline and large parallax scenes. In order to solve the above problems, many researchers have proposed corresponding solutions from the perspectives of multiple homography and mesh optimization. However, the mesh deformation may have significant shape distortion. Some deep learning methods combine vision tasks of optical flow, semantic alignment, image fusion, and image reconstruction to help deal with the stitching problem. However, the parameter information of cameras is not fully utilized, so the stitching results sometimes still show defects. Therefore, we wish to make full use of the parameter information of cameras and synthesize the smooth interpolated view by supplementing intermediate viewpoints between cameras to achieve better visual perception. Methods The present study proposes a real-time video stitching method based on dense viewpoint interpolation. The method focuses on the overlapping regions of stitching and synthesizes the smooth interpolated view by supplementing dense intermediate viewpoints on the baseline of cameras, which can better align multiple inputs. In the first place, binocular camera calibration is performed to obtain internal parameters and the transformation matrix of the cameras. The original views acquired by cameras are de-distortioned and adjusted to the same horizontal plane for stitching in the horizontal direction. The maximum possible overlap regions are separated and adjusted to coplanarity and row alignment by stereo correction so that the image data can be processed in only one dimension. Subsequently, pixel-level displacement fields sampled in the original views for the overlapping regions are predicted by using the cost volume in stereo matching. Without the ground truth of the interpolated view, the network is guided to learn view generation rules by using spatial transformation relationships between viewpoints. Through the pixel-level displacement fields generated by the network, two images are sampled in the input views respectively and fused by linear weights to generate the interpolated view of the overlapping regions. Finally, the generated interpolated view is combined with non-overlapping regions of two views. The cylinder projection is performed to align the fusion boundaries of three regions and obtain the final stitching result. Results and Discussions In this paper, the stitching results of the proposed method are compared with mainstream stitching methods. Multiband blending may show artifacts under the influence of parallax, while the method based on multiple homography and mesh optimization may have significant shape distortion in non-overlapping regions after mesh deformation. The proposed method can eliminate artifacts and smoothly align the inputs with little shape distortion, resulting in better visual perception (Fig. 9 and Fig. 10). Furthermore, we evaluate the alignment quality of the overlapping regions. The traditional methods only deal with stitching from the perspective of image features, and the alignment quality is relatively low in the case of large parallax variations. The proposed method combines camera calibration information for preprocessing and deals explicitly with the parallax problem to obtain better alignment quality (Table 1). Regarding model size and speed, the proposed method has advantages because it can initially align images after camera calibration and uses a lightweight construction method of cost volume. The processing frame rate of 720 p video can reach more than 30 fps to meet the demand for online video stitching (Table 2). In the analysis of the variation of baseline width, the proposed method can align well under different baseline widths (Fig. 12). In addition, all of them can obtain a high improvement of indicators (Table 3), which is robust to the variation of the baseline width. In conclusion, the proposed method can improve the visual perception after stitching, eliminate artifacts, and smoothly align the inputs. It has high alignment quality, little shape distortion, and great application value because of its lightweight design and fast processing speed. Conclusions Applying the proposed video stitching method based on dense viewpoint interpolation can effectively deal with the problem of stitching in wide baseline and large parallax scenes. The interpolated view with the smooth transition is synthesized for the overlapping regions of stitching by supplementing dense intermediate viewpoints on the baseline of the left and right cameras. A network for generating the interpolated view is proposed, which is divided into modules of feature extraction, correlation calculation, and high-resolution optimization to predict the sampling locations in the original views. The generated interpolated view is combined with the non-overlapping regions to obtain the stitching result. Moreover, the proposed method calculates the three-dimensional information at the original viewpoint in the virtual environment without the ground truth of the interpolated view. The corresponding spatial region of the interpolated viewpoint is searched by dichotomization. The interpolated view is transformed into the original viewpoint under the constructed loss function, which guides the network to learn the view generation rules. Various experiments have proved that the proposed method can improve the visual perception of video frames after stitching. It is adaptive for different baseline widths, has great generalization ability, and achieves real-time performance to meet the online stitching requirements in practical applications.
Large parallax image stitching is a challenging task. Existing methods often struggle to maintain both the local and global structures of the image while reducing alignment artifacts and warping distortions. In this paper, we propose a novel approach that utilizes epipolar geometry to establish a warping technique based on the epipolar displacement field. Initially, the warping rule for pixels in the epipolar geometry is established through the infinite homography. Subsequently, Subsequently, the epipolar displacement field, which represents the sliding distance of the warped pixel along the epipolar line, is formulated by thin plate splines based on the principle of local elastic deformation. The stitching result can be generated by inversely warping the pixels according to the epipolar displacement field. This method incorporates the epipolar constraints in the warping rule, which ensures high-quality alignment and maintains the projectivity of the panorama. Qualitative and quantitative comparative experiments demonstrate the competitiveness of the proposed method in stitching images large parallax.
In this paper, a convolutional neural network is proposed to obtain high quality absolute phase from single frame composite images. The composite image used in the proposed method is the fringe image embedded with speckle. The convolutional neural network consists of two sub-networks, which use the fringe mode component and the speckle mode component in the composite image to solve and unfold the wrapping phase. In the process of phase unwrapping, the proposed method uses the pre-photographed composite image and its fringe order as auxiliary information to ensure the accuracy of phase unwrapping. Experimental results show that the proposed method can minimize the number of projected images by using single-frame composite images and obtain high precision absolute phase, which provides a feasible solution for 3D measurement in high precision dynamic scenes.
A newly developed calibration algorithm for camera-projector system using spheres is presented in this paper. Previous studies have exploited image conics of sphere to calibrate the camera, whereas this approach can be strengthened to apply in the projector and ultimately achieve the overall calibration for single or multiple pairs of camera-projector. Following the concept of taking the projector as an inverse camera, we retrieve the image conic of the sphere on the projector plane based on a pole-polar relationship we found. At least 3 image conics on the image plane of each device are required to calculate the intrinsic parameters of the device. The extrinsic parameters for all devices in the system are determined by the position of sphere centers in each coordinates frame of the device. Based on the isotropy of the calibration object (sphere), this work is mainly interested in accomplishing the entire calibration for multiple camera-projector systems in which sensors surround a central observation volume. Experiments are conducted on both synthetic and real datasets to evaluate its performance.
A deep learning-based method is proposed to recover the absolute phase value from a single fringe pattern. We propose a deep neural network architecture that includes two subnetworks used for wrapping phase calculation and phase unwrapping, respectively. The training set is generated with the absolute phase obtained by the combination of phase shifting and gray coding. In addition, a reference plane is adopted to provide periodic range information for phase unwrapping. Then according to the output of the well-trained network, a high-quality absolute phase is obtained through only a single fringe pattern of the measured object. Experiments on the test set verify that high accuracy for complex texture objects is acquired using the proposed method, which indicates its potential in high-speed measurement. (C) 2021 Society of PhotoOptical Instrumentation Engineers (SPIE)
This paper presents a calibration parameters refinement specifically for a fringe projection profilometry system to assure the final accuracy, even when using an imperfect calibration target. Unlike existing camera-projector calibration methods, we arrange a refinement in subsequent process of target measurement. Following the trend of additionally estimating the target's geometry, a novel formulation is presented that allows point reconstruction from the infinity homography, thereby introducing the target geometry from the scene. The final objective function is built on the reprojection error and implemented under an equality constraint of the fundamental matrix. In our approach, the fundamental matrix is estimated from the reliable feature correspondences exclusively provided by the structured light system. Experiments are conducted on both synthetic and real datasets to evaluate the performance of our proposed approach.
A newly developed flexible calibration algorithm for fringe projection profilometry system is presented in this paper. Previous studies have exploited images of spheres to calibrate the camera. It is shown in this paper that this approach can be improved to suit for the projector and ultimately achieve the overall calibration of FPP (Fringe Projection Profilometry) system. Taking the projector as a virtual camera, the images of sphere contour on the projectors plane can also be obtained through the phase information. The derivation and acquisition of intrinsic parameters for projector are just the same way used in the camera. In our algorithm, at least 3 images of sphere contour on both camera and projector are obtained to calculate the homography between these two views. Then the image of the sphere and its shadow on an induced plane settled in the back of the sphere are added to recover the epipolar geometry for the FPP system. Experimental results on real data are presented, which demonstrate the feasibility and accuracy achieved by our proposed algorithm.
Point cloud has achieved great attention in 3D object classification, segmentation and indoor scene semantic parsing. In terms of face recognition, although image-based algorithm become more accurate and faster, open world face recognition still suffers from the influences i.e. illumination, occlusion, pose, etc. 3D face recognition based on point cloud containing both shape and texture information can compensate these shortcomings. However training a network to extract discriminative 3D feature is model complex and time inefficient due to the lack of large training dataset. To address these problems, we propose a novel 3D face recognition network(FPCNet) using modified PointNet++ and a 3D augmentation technique. Face-based loss and multi-label loss are used to train the FPCNet to enhance the learned features more discriminative. Moreover, a 3D face data augmentation method is proposed to synthesize more identity-variance and expression-variance 3D faces from limited data. Our proposed method shows excellent recognition results on CASIA-3D, Bosphorus and FRGC2.0 datasets and generalizes well for other datasets.