Spectral Domain Optical Coherence Tomography(SD-OCT) systems have been extensively studied for biological tissue measurement, featuring non-invasiveness, micron-level resolution, and real-time detection capabilities. To expand their application scope to the measurement of industrial products with multilayer structures, it is necessary to optimize the performance of SD-OCT systems. Firstly, this study optimizes the parameters of the sample arm field lens using ZEMAX. Aberrations are reduced by adjusting the lens curvature, focal length, and optical path layout, and a matched scanning lens is designed to expand the detection range. Secondly, the spectrometer parameters are optimized: appropriate lenses and gratings are selected, and the optical path layout and device tilt angle are adjusted to improve spectral resolution and imaging depth. Finally, MATLAB is used to establish a simulation model for the system's axial resolution, which supports tomographic imaging simulation experiments of multilayer sample structures. The results show that the comprehensive performance of the optimized SD-OCT system is significantly improved. The lateral resolution is enhanced to 20 & micro;m, while a large scanning range of 34 mm & times; 34 mm is achieved. The imaging depth is increased to 2. 44 mm, enabling effective penetration of multilayer industrial structures. The axial resolution is verified by simulation to be 6. 40 & micro;m, with the simulation results highly consistent with the theoretical calculation values. Additionally, the tomographic imaging simulation of multilayer sample structures is successfully realized. The proposed optimization design method for the SD-OCT system provides valuable reference and promotion for its applications in other fields.
Structured light 3D measurement has been widely used in industrial inspection because of itsnon-contact, full-field, and high-accuracy characteristics. However, measuring complex highdynamic-range (HDR) surfaces with intricate structural features remains challenging due to fringe information loss induced by occlusions and surface reflections. To address these issues, this paper proposes a multi-camera structured-light reconstruction method with globally consistent calibration, self-correction, and reprojection-based multi-view fusion. Under the unified framework, a novel inverse camera calibration approach leveraging planar geometric constraints from main and auxiliary planes is first devised to generate globally consistent calibration parameters, thus suppressing the calibration residuals and the view-dependent stratification. Second, we propose a projector-domain pixel-wise phase-depth mapping model utilizing cross-ratio invariance to calculate 3D coordinates, and establish a data-driven epipolar-line function to eliminate abnormal phase pixels induced by reflections, occlusions or phase unwrapping failures. Third, a reprojection-based multi-view fusion method is developed to reconstruct the final point cloud by integrating pixel-wise 3D coordinates and phase information from all viewpoints, thereby eliminating the complicated point cloud registration step. Experiments demonstrate that the proposed method improves reconstruction accuracy and achieves complete 3D measurement for complex HDR objects.
Interferometric wavefront reconstruction is important for applications such as precision optical metrology. Traditional methods are hindered by challenges such as inaccurate phase extraction from a single-shot measurement and error accumulation in multi-frame approaches, while reported deep learning methods often lack physical constraints and introduce additional errors during the phase unwrapping process. To address these issues, a physics-driven hybrid feature calibration network (HFC-Net) is proposed and optimized for extracting a high-precision wavefront from a single-shot interferogram. Simulation and experimental results demonstrate that the HFC-Net achieves root mean square (RMS) errors as low as 0.00031λ on simulated data and 0.0021λ on real data, with a processing time of 87.14 ms. Its efficiencies are validated by comparing it with the traditional Fourier transform method and reported deep learning models. The presented method is also valuable for phase object or dynamic aberration measurement.
ABSTRACT Deep learning is becoming a core driving force for advances in topological photonics. This work provides how deep learning helps address key challenges in topological photonics and discusses future research directions. We summarize and discuss recent AI‐assisted methods for topological photonic crystals (TPCs) and their representative application pathways. The main TPC architectures and recent progress on representative topological states in two‐ and three‐dimensional photonic crystals are re‐examined, with an emphasis on how deep learning can improve design efficiency, enable robustness evaluation, and help bridge the gap between numerical models and as‐fabricated structures. We analyze the key bottlenecks that currently limit practical deployment, such as fabrication tolerance, loss, thermal drift, and topology‐aware diagnosis. We discuss emerging opportunities for learning‐ enabled topological photonic platforms and outline how AI may accelerate the transition of TPCs from proof‐of‐concept physics to deployable photonic technologies.
In the single-pixel imaging (SPI) process, the illumination patterns need to be projected onto the object surface to modulate and sample the object information. If the illumination pattern is defocused, it will cause blurring in the SPI results. This paper proposes a method using spatially multiplexed phase modulation to generate pseudo non-diffracting structured light pattern, such as random binary patterns, and applies it to SPI, thereby extending the axial imaging range of SPI. Experimental results show that the generated non-diffracting structured light can maintain its pattern stable and unchanged within the designed nondiffracting distance, achieving high-quality SPI imaging even at low sampling rates. This method can obtain clear images without requiring precise focusing, enable SPI for multiple objects along the axial direction simultaneously, and can also be applied to SPI for axially moving objects.
As a fundamental vision task, tracking gains fewer benefits from recent flourishing vision-language (VL) learning. Most existing VL trackers treat language as auxiliary information for vision, ignoring the interaction learning and complementarity between them. In this paper, a novel VL tracking framework named DynaMix vision-language tracker (DVLT) is developed to learn joint feature extraction and interaction based on the transformer backbone. First, a dynamic multi-head attention (DMHA) mechanism is designed to alleviate the low-rank bottleneck of traditional multi-head attention by employing a collaboration function among heads at the level of the attention matrix. Then, a trainable shared composition map is proposed for dynamically synthesizing new attention heads to maintain computational speed and simultaneously improve the attention matrix rank. In addition, to further enhance the ability of model representation, a modal mixing strategy is utilized to make the extracted features highly target-aware by assigning weights to the visual tokens under language guidance. Finally, a semantic location head is introduced that utilizes reference information to reduce visual ambiguity and mines target features by introducing high-level semantic information based on a similarity function. Extensive experiments are conducted on multiple language-assisted tracking datasets, including OTB99-LANG, TNL2K, LaSOT and LaSOText, to demonstrate the superiority of the proposed tracker compared with existing VL trackers.
Objective In structured light-based three-dimensional measurement, global illumination effects pose significant challenges that compromise measurement accuracy, such as inter-reflections and subsurface scattering. Traditional Gray code phase-shifting methods often suffer from stripe edge blurring induced by these effects, leading to decoding failures. Although techniques like complementary Gray code can alleviate phase jump errors, they remain ineffective in suppressing global illumination. High-frequency illumination strategies can separate direct and global components but usually require additional projected patterns, reducing system efficiency. This study introduces a novel coding strategy that integrates parity-check and exclusive OR (XOR) operations to simultaneously mitigate global illumination influences and correct phase jump errors, enabling high-precision 3D reconstruction of objects with complex surface properties. Methods This study presents a robust 3D reconstruction method based on a parity-check pattern (PCP) and XOR coding strategy. First, a unique auxiliary fringe pattern (PCP) is generated based on the parity of Gray code words. The PCP is then XORed with the original Gray codes to create high-frequency coded patterns, enhancing resistance to global illumination's low-pass filtering effect. Furthermore, a confidence map is computed from the direct and global components separated using the inherent high-frequency characteristics of phase-shifting patterns, dynamically filtering unreliable pixels. Finally, during decoding, an error detection and correction mechanism based on parity check is introduced to actively rectify phase jump errors, ensuring the continuity of the unwrapped phase. Results and Discussions Experiments on objects with complex optical properties demonstrate the effectiveness of the proposed method, including a ceramic bowl under strong inter-reflection and a semi-transparent nylon sphere affected by subsurface scattering. In the case of the ceramic bowl, where global illumination leads to overexposure and reconstruction artifacts, the proposed method recovers 2854839 valid points, outperforming the complementary Gray code method (2570544 points) and the method in reference [18] (2788165 points), corresponding to increases of 11.06% and 2.39%, respectively, while producing a more complete and smoother reconstruction. For the translucent nylon sphere, the proposed method achieves an average root mean square error (RMSE) of 0.166 mm in sphere fitting, significantly lower than the complementary Gray code method (0.278 mm) and also lower than reference [18] (0.201 mm), confirming higher accuracy under subsurface scattering. Comparative evaluations further confirm the competitiveness of the PCP-XOR strategy in handling complex optical interactions. Despite a slight increase in the number of projected patterns, the method maintains computational efficiency through confidence-based pixel filtering and localized error correction. Conclusions The proposed method effectively addresses the critical issues of global illumination and phase jump errors in structured light 3D measurement. By incorporating a high-frequency encoding strategy and an intelligent error correction mechanism, the method improves both the robustness and accuracy of reconstructions under challenging illumination conditions. Experimental results demonstrate its capability to achieve high-quality 3D reconstruction for highly reflective and scattering surfaces, highlighting its practical value in industrial and scientific applications.
Objective As semiconductor manufacturing processes move towards high integration, surface defects on wafers directly determine chip yield, especially scratches and other defects generated during chemical mechanical polishing (CMP) processes that can easily lead to chip functional failure. The geometric discontinuity and complex background interference of defects make precise detection a necessity in the industry. Traditional optical detection relies on manually designed features and thresholds, resulting in high false detection rates in the recognition of small defects and complex texture backgrounds. The statistical analysis and image segmentation methods based on manual features are difficult to cope with multi form defects due to poor adaptability. Although deep learning provides a new path for detection, models such as the YOLO series have been applied due to their real-time advantages. However, existing algorithms still face bottlenecks such as insufficient accuracy in detecting multi-scale weak defects and inaccurate positioning of detection boxes, making it difficult to meet the strict requirements of defect detection in nanoscale manufacturing. Therefore, this article optimizes the algorithm based on YOLOv8, aiming to break through the dilemma of identifying and locating multi-scale weak defects, improve the accuracy and robustness of defect detection in complex backgrounds, provide efficient and reliable quality control technology support for semiconductor manufacturing, and ensure the production yield of highly integrated chips. Methods This study proposes a wafer defect detection algorithm based on SLDC-YOLOv8 to address the issues of insufficient detection accuracy and inaccurate detection boxes in the detection of multi-scale weak defects. First, the designed SLD-C2f module introduces Sobel algorithm to actively capture edge features and complete the weak defect contour information lost by ordinary convolution. By concatenating the outputs of Sobel and convolution branches, a comprehensive feature representation containing key edge information and detailed spatial features is constructed, and LDConv is used to maintain efficient feature extraction. Second, introducing coordinate attention (CA) attention mechanism to fully focus on channel and spatial features enhances the model's ability to extract defect features, and only increases a small amount of model parameters and computational complexity. Third, a detection box merging algorithm (BMA) is designed to replace the traditional non-maximum suppression (NMS) algorithm. This algorithm effectively reduces false positives and false negatives for long strip scratches, crack defects, and scattered oxidation point defects, improving the overall performance of object detection. Results and Discussions The comparative experiments on attention mechanisms verify the effectiveness of the introduced CA by comparing it with various attention mechanisms. The model incorporating the CA attention mechanism achieves the highest improvements in both mAP@0.5 (mean average precision when intersection-over-union threshold is 0.5) and mAP@0.5 0.95(mean average precision computed over intersection-over-union thresholds ranging from 0.5 to 0.95 with a step of 0.05), while only increasing the model parameters and computational complexity slightly. The comparative experiments on different kernel sampling point counts of LDConv analyze the impact of different parameters on model performance. A sampling point count of 3 ensures detection accuracy while maintaining the model's lightweight property. The ablation experiment results verify the function of each improved module. The SLD-C2f module captures more information, and the CA attention module captures long-distance dependencies, so that the model can effectively avoid false detection of the detected object. On the Wafer-Scratch dataset, the improved model shows increases of 15.2 percentage points in precision (P), 9.8 percentage points in recall (R), 11.8 percentage points in mAP@0.5, and 7.8 percentage points in mAP@0.5 0.95. On the Dunnan60mil dataset, it achieves improvements of 8.9 percentage points in P, 5.4 percentage points in R, 9.3 percentage points in mAP@0.5, and 5.4 percentage points in mAP@0.5: 0.95, demonstrating its effectiveness in wafer multi-scale defect detection. The performance comparison of different network models indicates that SLDC-YOLOv8 maintained high accuracy while having relatively low model parameters and computational complexity. The visualization results show that the algorithm reduces missed detections and false detections, with improved detection performance. Experimental results on the NEU-DET steel defect dataset reveal that compared with the baseline model, SLDC-YOLOv8 achieves increases of 1.9 percentage points in P, 3.8 percentage points in R, 3.2 percentage points in mAP@0.5, and 0.5 percentage point in mAP@0.5 0.95. Additionally, its mAP@0.5 and mAP@0.5 0.95 outperform all other comparative models, with advantages in parameters and computational complexity, proving the model's generalization capability. Conclusions To address issues such as insufficient detection accuracy and false positives in the detection of multi-scale weak defects on wafer surfaces, this paper proposes an improved defect detection algorithm based on the YOLOv8 network, named SLDC-YOLOv8. By designing the feature extraction module SLD-C2f, the model's ability to perceive detailed information, such as edges, is enhanced, while the computational load of the model is reduced. The CA attention mechanism is employed to capture both channel and spatial information, strengthening the model's capability to extract defect features. Additionally, the designed BMA enables the model to obtain more accurate detection box results. Experimental results across multiple datasets demonstrate that the improved algorithm enhances the detection accuracy of multi-scale industrial defects, achieving high detection performance and showing promising prospects for industrial applications. However, the method proposed in this paper still has some limitations. Subsequent research can consider carrying out more complex background of weak defect detection on different workpiece surfaces and lightweight design of models, and combining the characteristics of different defect types to match the optimal model parameters for various defects, to narrow the gap between model generalization performance and industrial scenario requirements, and further improve the universality and efficiency of models.
The multi-functionalities of metasurface structure usually need to increase a high system complexity and a low efficiency, which can be mathematically expressed by the transfer matrix. This study presents an all-dielectric diatomic metasurface design approach, deduced from the transfer matrix to achieve high degrees of freedom and multifunctionalities in multimodal polarization imaging analysis. Here a compact and efficient all-dielectric diatomic metasurface platform is proposed, which can effectively achieve the selectivity of arbitrary orthogonal polarization. It is proved that the phase and amplitude of two orthogonal circular polarizations can be controlled independently by designing the rotation angle relationship between the diatoms. Furthermore, we propose an optical polarization imaging encryption that generates four near-field amplitude and far-field diffraction holograms, which can only be observed when the incident light is at a specific polarization state, whereas no image can be discerned for the orthogonal polarization incidence case, indicating the realization of incidence-polarization secured meta-image. Through data fusion, we are able to combine multiple polarization imaging modes to improve the accuracy and security of information. We envision the metasurface to enhance the security of information transmission and open new possibilities for creating compact multifunctional optical devices for optical data storage and security.
The nonlinear response characteristics of stripe projection of structured light systems often lead to stripe gray-level distortion,adversely affecting both phase and measurement precision.To address the nonlinearity within structured light systems and the differential phase error caused by various degrees of defocused stripes,a novel phase gradient-based regional phase error self-correction methodology is introduced.Initially,the phase is segmented into distinct regions by leveraging data modulation of the stripe image and the gradient information of the wrapped phase.A phase error function is then constructed utilizing the wrapped phase across multiple frequencies,and the phase values from each region are subsequently input into this function for iterative correction.Experimental outcomes indicate a 79.07%reduction in phase error after applying this method.In addition,this approach eliminates the need for stripe pre-encoding,directly addresses errors within the unwrapped phase map,and demonstrates impressive resilience to stripe defocus.
Graphene-based photodetectors can exhibit excellent performance within a wide spectral range, and thus have attracted extensive attention. In the field of infrared detection, graphene-based silicon carbide (SiC) photodetectors have not been extensively and thoroughly studied, and their potential value remains to be further explored and developed. In this paper, we present a graphene-based SiC photodetector that can synergize with HgTe colloidal quantum dots and orthogonally arranged metal nanostructures. This designed allows the photodetector to achieve perfect absorption of infrared light and reach a high responsivity of 8.58 A/W at the wavelength of 10.63 mu m. It can be flexibly tuned to detect wavelengths from 10.3 mu m to 10.9 mu m through adjusting the vertically oriented strips and the Fermi level of graphene. Furthermore, this photodetector also has sensitivity to incident light polarization and dual-band detection performance. Our structure opens up a new path for high-performance and multifunctional infrared photodetectors.
Phase-shifting interferometry is a widely used technique in precision optical metrology.However,conventional PSI methods typically require three or more phase-shifted interferograms,which limits their applicability in dynamic measurements and vibration-sensitive environments.To address this issue,this paper proposes a deep learning-based framework named PSI-IPENet.The framework adopts a Two-to-One structure,where two phase-shifted interferograms are used as dual-channel input,corresponding to a single phase map as the supervisory signal,and a dedicated dataset is constructed for training.PSI-IPENet leverages the feature extraction capability of IPENet and incorporates the physical characteristics of interferometric imaging,thereby enhancing the robustness and noise resistance of phase retrieval.Experimental results demonstrate that the proposed method maintains high-precision phase recovery performance even with a low number of input frames.Compared with the traditional four-step phase-shifting method,it shows significant advantages in terms of signal-to-noise ratio and phase error metrics.
We propose a double-layer metasurface to realize the polarization switching of fractional perfect composite vortex beams (FPCVBs). The metasurface is designed by the Jones matrix formalism and optimized through the “Random Forest” machine learning algorithm, achieving high polarization switching efficiency. The grafted-FPCVBs are presented with featuring multivariate topological charges, and double-ring FPCVBs are achieved with independent on/off modulation and shaping transformation of inner and outer ring intensities. These capabilities unlock orbital angular momentum density modulation and light field reconfiguration in composited light architectures. Moreover, we design an optical encryption protocol inspired by the hierarchical structure of chinese characters, where semantic radicals are mapped to polarization-encoded FPCVBs. These innovations present significant potential for applications in optical information security, particle manipulation, and next-generation photonic communications.
Perfect vector vortex beams, distinguished by their unique polarization states and vortex characteristics, have recently become a central focus in advanced photonics research. However, the generation of these vector beams has been limited by static topological charge (TC) configurations and restricted information encryption capabilities. Herein, we present a metasurface-based method employing trigonometric-function topological engineering to generate grafted perfect vector vortex beams (GPVVBs). Through computational analysis, we establish the relationship between the rotation angle and continuously tunable TCs. By dynamically adjusting fractional-order beams through polarizer rotation, we achieve customizable field distributions. Furthermore, we demonstrate the application of GPVVBs in multichannel optical encryption, where the grafted continuous TCs considerably enhance information encryption capacity. This innovative method introduces unprecedented flexibility and holds transformative potential for both optical encryption and high-density communications.
Visual anomaly detection has become a significant solution in industrial production due to its remarkable effectiveness and efficiency. However, traditional generative model-based methods are constrained in overall performance due to limitations in reconstruction quality and speed. Recently, generative approaches based on diffusion models have brought new opportunities to visual tasks, owing to their high-quality and diverse generation capabilities. Therefore, this paper proposes a novel anomaly detection framework based on diffusion models, termed F-ADMD, which is primarily composed of two parts: image reconstruction and defect segmentation. In the image reconstruction portion, we introduce a multi-scale guided rapid reconstruction method that is not only significantly faster than traditional diffusion models but also effectively handles the reconstruction of various anomaly-type regions. In the defect segmentation section, we integrate a Transformer attention mechanism that can capture both local and global context information at different feature levels, aiding the model in focusing more effectively on key regions of the image and thus improving the accuracy of identifying the boundaries of anomalous areas and enhancing segmentation precision. F-ADMD achieves state-of-the-art performance in image-level detection and anomaly localization on the challenging and widely used MVTec dataset, demonstrating the framework's effectiveness and broad applicability. This new method not only enhances detection accuracy but also significantly boosts processing speed, providing a novel comprehensive solution for visual anomaly detection in industrial production.
The voice coil deformation mirror, due to its large modulation and lack of hysteresis, is widely used as an adaptive secondary mirror in astronomical telescopes. This paper proposes an optimization method for the voice coil driver used in a 37-element large modulation voice coil deformation mirror, followed by simulation and validation. A dynamic magnetic structure is utilized to derive the theoretical formula for the electromagnetic force of the driver and the relationship between the electromagnetic force and mirror surface strain is briefly discussed. Detailed optimization of the structure is conducted using finite element software. Ultimately, simulation analysis and aberration correction validation of the 37-element voice coil deformation mirror are performed based on the optimization results. The designed voice coil driver has an efficiency of 0. 362 N/W-1/2, a response speed of 0. 3 ms, linearity of 1, and allows for a maximum deformation mirror stroke of 50 mu m. Moreover, the designed driver excels in simplicity, miniaturization, low cost, high response speed, linear response, and high efficiency, indicating promising applications in atmospheric turbulence aberration correction, laser weapons, laser communication, and fundus imaging.
The single-pixel imaging (SPI) technique requires multiple samplings during the imaging process, which leads to significant motion blur when capturing moving objects due to the temporal integration of dynamic scenes. This limitation hinders the application of SPI in practical and complex scenarios involving high-speed motion. In existing SPI methods for moving objects, research typically focuses on a single moving object. This work addresses the scenario where two objects with distinct motion trajectories coexist in a scene. First, the mixed radon spectrum of both objects is acquired using SPI. Subsequently, a series of operations-including separation, calibration, and reconstruction-is performed. Finally, blur-free images of the moving objects are obtained. We evaluate the proposed method by testing various scenarios involving images of diverse objects, distinct motion trajectories, and varying velocities. The test results demonstrate the feasibility and robustness of the approach.
Binocular structured light 3D measurement systems have been widely used in the industrial field due to their non-contact and high-precision characteristics. However, when it comes to addressing the issue of stripe information loss caused by occlusion and surface reflection, current prevalent methods are inclined to exhibit accuracy degradation and incomplete reconstruction results. Some intricate pre-processing procedures can mitigate this situation, but they simultaneously reduce system efficiency. To address this issue, this paper proposes a binocular structured light reconstruction method based on disparity fusion, which could enhance the accuracy and completeness of reconstruction results by fusing disparities from multiple views. This work makes four major contributions as follows: (1) An optimized stereo matching algorithm is proposed by leveraging the phase unwrapping order to narrow the search range of matching points, thus addressing the time-consuming issue in disparity calculation. (2) A novel pixel-wise calibration method is developed to solve the problem of incomplete reconstruction. This method realizes single-view disparity calculation by constructing a phase-disparity mapping function, and then fuses disparity information from dual views to significantly enhance the reconstruction completeness. (3) The construction of the phase-disparity mapping function allows disparity calculation to bypass the time-consuming stereo matching algorithm, thus greatly improving the reconstruction efficiency. (4) To further enhance reconstruction accuracy, the proposed method proves that there exists a planar relationship between the disparity of an ideal plane and its corresponding pixel positions, and accordingly corrects the local errors in the disparity map. Experimental results demonstrate that this method not only enhances the accuracy and completeness of the reconstruction results but also exhibits high flexibility and efficiency.
Phase imaging technology holds significant research value in fields such as biomedicine, materials science, and precision measurement. Single-pixel imaging (SPI) for phase objects has also attracted growing attention. In existing SPI-based phase imaging approaches, phase information is typically retrieved by combining holographic techniques or complex iterative optimization algorithms. To address these limitations, we propose a self-supervised, physics-driven neural network model. By incorporating Fresnel and SPI layers into an autoencoder architecture, the network directly reconstructs the phase object from the measured intensity sequence in SPI, with network parameters updated in an end-to-end manner. Both simulation and experimental results demonstrate the effectiveness and robustness of the proposed method, showing that high-quality phase images can be reconstructed even at low sampling rates. Notably, the same SPI system can be used for both phase and amplitude objects without hardware modifications-only minor adjustments to the network structure are required. The proposed framework is also extendable to complex amplitude imaging and shows promise for applications in biomedical imaging and adaptive optics.