Industrial anomaly detection faces critical challenges of sample imbalance and complex defect patterns in real-world manufacturing. This paper proposes a novel DMGLAD (Dynamic Multi-scale distillation with Global-Local Decoupling for Anomaly Detection) framework that synergizes multi-scale structural distillation and global-local anomaly reasoning for fabric defect detection. Our framework integrates: (1) A Teacher-Student architecture with multi-scale decoupled distillation that captures local structural anomalies through patch-based feature matching, combined with an autoencoder for global anomaly reasoning; (2) A Dynamic Focusing Feature Loss (DFFL) that adaptively reweights hard samples through curriculum annealing, addressing extreme class imbalance in feature space; (3) A Differentiable Anomaly Map Fusion (DAMF) mechanism that probabilistically combines local and global anomaly scores through learnable Bayesian fusion. Extensive experiments demonstrate state-of-the-art performance on MVTec-AD, VisA and a new RealFabric dataset containing subtle textile defects, where our method achieves 93.26% AU-ROC with 11.97% improvement over baseline through effective fusion of multi-scale features. Qualitative analysis shows our method effectively detects fine-grained defects (e.g., texture variations <10% contrast) and handles long-range anomalies in large images (1000 & times; 1000 resolution). These results highlight the practicality of our framework for real-world industrial inspection systems.
Optical 3-D measurement is a key noncontact method for capturing the surface geometry of objects. In fringe projection technology, the fringe pattern period significantly affects the accuracy of point cloud reconstruction. This study analyzed reconstruction results obtained with fringe patterns ranging from 16 to 128 pixels per period and identified a systematic rotational error across all periods. This error manifests as a global, rotation-structured phase deviation around a stable center rather than random noise or local fluctuations, and it appears in both opaque and translucent objects, with more pronounced distortions on translucent surfaces. To investigate the characteristics of this error, multiple real-world experiments were conducted to systematically analyze the origins and properties of the rotational error and propose a mathematical model. Furthermore, based on this mathematical model, the epipolar line-guided rotation (EGR) error correction algorithm is proposed, which utilizes the characteristics of rotational error, combining prior error models with polar line constraints to correct phase information and rectify error points in anomalous regions around the polar line. This approach requires no extra hardware, as it directly calculates correction values through well-defined constraints, thereby avoiding cumulative errors from multiparameter calibration and significantly simplifying the correction process. The experimental results demonstrate the effectiveness of the algorithm in correcting rotational errors, providing novel insights and methodologies for future error correction research in structured light 3-D measurement.
Computational spectral imaging based on coded-aperture snapshot spectral imaging (CASSI) aims to reconstruct hyperspectral images (HSIs) from compressed measurements and is inherently an ill-posed inverse problem. Although recent deep learning methods have achieved remarkable progress, two challenges remain. First, RGB images provide complementary spatial-spectral priors, yet effectively incorporating such multimodal information into iterative CASSI reconstruction remains challenging. Second, existing loss functions primarily focus on pixel-wise reconstruction errors and often neglect the geometric characteristics of spectral curves across wavelengths. To address these issues, We propose a deep unfolding framework for dual-input hyperspectral reconstruction, where RGB observations provide complementary spatial constraints for CASSI measurements. The proposed framework jointly optimizes CASSI fidelity and RGB-guided constraints within an iterative deep unfolding framework. Furthermore, motivated by the continuity of spectral signatures in natural scenes, we introduce a higher-order derivative loss that constrains spectral curve geometry and promotes consistency of spectral variations across wavelengths. Extensive experiments on benchmark datasets demonstrate the effectiveness of the proposed framework. The proposed network achieves 37.92 dB and 44.83 dB PSNR in single-camera and dual-camera reconstruction settings, respectively. Furthermore, the proposed higher-order derivative loss improves PSNR by an average of 0.5 dB across multiple reconstruction methods while producing spectral curves that better match the ground truth. (https://github.com/xintangjin/MaskFree.git).
Surface defect detection for synthetic fiber bobbins is crucial to intelligent textile manufacturing. Existing vision-based methods are limited by the lack of task-specific datasets, the inefficient adaptation of detectors to this application, and the difficulty of deployment on edge devices. To address these challenges, this study proposes a lightweight detection framework. First, an image acquisition system equipped with a bobbin pose-adjustment device is developed, and a new industry-oriented dataset, FiberBobbin-40K, containing 40,497 high-quality images, is established to fill the gap in fiber bobbin defect detection. Subsequently, a computationally efficient detector adaptation strategy is proposed. Finally, a compression framework integrating a redesigned detection head, layer pruning, and channel pruning is developed to enable efficient inference on edge devices. Experimental results show that, compared with the conventional YOLOv8, the optimized model reduces the number of parameters by approximately 71%, while incurring only 2.03% and 2.49% decreases in mAP and F1 score, respectively. In addition, it achieves 94.73% accuracy and 26.37 FPS on a CPU. The optimized model has only 2.47 million parameters and a model size of 4.86 MB, which is substantially smaller than Faster R-CNN and YOLOv7 while maintaining competitive detection performance.
3D multi-object tracking (MOT) for autonomous driving remains challenging due to frequent identity switches in crowded scenes, trajectory fragmentation during occlusions, and the difficulty of adapting association strategies to varying scene complexities. While existing methods rely on fixed geometric or appearance-based associations, they struggle to handle ambiguous cases and detection failures. We present an adaptive multi-level 3D MOT framework that achieves robust tracking through three key innovations: (1) multi-granularity temporal modeling that captures both fine-grained short-term motion and coarse long-term trends via dual-scale spatio-temporal attention, enabling accurate motion prediction across different object dynamics; (2) Transformer-based Appearance Association that employs cross-attention to model global inter-object relationships, resolving ambiguous associations in crowded scenarios where geometric cues alone fail; and (3) scene-adaptive learned thresholds that automatically adjust association strictness based on object density, motion complexity, and occlusion levels, avoiding the one-size-fits-all limitations of fixed thresholds. Our hierarchical four-level tracking strategy progressively handles cases from easy geometric matching (Level 1) to complex interval-frame recovery (Level 4), with SOT-based virtual detection generation bridging detector failures. Extensive experiments on the nuScenes benchmark demonstrate state-of-the-art performance.
6D pose estimation is a key technology in computer vision and robotic manipulation. However, many methods remain heavily dependent on CAD models that are difficult to obtain. Object-level 3D reconstruction provides an alternative route, and 3D Gaussian Splatting (3DGS) shows convincing potential owing to its training and rendering efficiency. Nevertheless, under sparse reference views, 3DGS is prone to floating artifacts and appearance overfitting, which weakens the stability of pose estimation. We present PoseGaussian, a method for sparse-view 6D pose estimation for unseen object that builds on improved 3DGS. First, we use sparse RGB-D views to inject a depth structure prior into the 3DGS initialization for stable structure, and we adopt adaptive density control, view-warping augmentation, and joint photometric–depth supervision to reduce floaters and appearance overfitting under sparse reference views. Next, in the pose estimation stage, we apply a two-stage learning-guided ICP initializer that exploits geometric features to obtain a stable initial pose. Finally, we introduce a 3DGS-based iterative pose refiner that aligns rendered and query images in both appearance and geometry, further improving pose estimation accuracy. Experiments on LINEMOD, GenMOP, and our real-world datasets show that PoseGaussian achieves significant improvements over baseline methods under model-free and sparse-view settings, demonstrating strong generalization to unseen objects and robustness to view sparsity.
Interval Type-2 (IT2) fuzzy systems have gained significant attention due to their strong capability in handling system uncertainties. This paper investigates the robust stability analysis of conventional Takagi-Sugeno (TS) IT2 fuzzy systems under an observer-based control framework. A non-uniform piecewise linear approximation method is introduced to more accurately capture the boundary characteristics of IT2 membership functions (MFs), allowing key variation information of MFs to be effectively exploited. Subsequently, an error model transformation strategy is proposed to reconstruct approximation-induced errors into an auxiliary fuzzy model, enabling richer error-related and MF information to be explicitly incorporated into the stability conditions and thereby reducing conservatism. By leveraging Lyapunov stability theory and a scaling approach, sufficient stability criteria are derived in terms of linear matrix inequalities (LMIs), which can be efficiently solved using standard convex optimization tools. Simulation results and comparative studies demonstrate that the proposed method achieves less conservative stability conditions and enhanced robustness compared with existing approaches.
Sphere-based calibration for camera-projector pairs achieves high accuracy with simple setups but suffers under projector illumination, which distorts sphere contours and complicates conic extraction. This paper presents a robust single-sphere calibration method that leverages the properties of active illumination to overcome this challenge. We first analyze the geometric structure of sphere boundaries under projector lighting and show that they originate from distinct apparent contours in the camera and projector views. Exploiting dual epipolar geometry, we segment these boundaries and fit separate image conics for each view. Based on this, we derive two novel orthogonality constraints between the conics and the image of the absolute conic, enabling complete calibration from a single sphere. To enhance precision, we further propose a nonlinear optimization strategy that jointly minimizes sphere reconstruction and point reprojection errors, refining both intrinsic and extrinsic parameters as well as lens distortions. Experimental results demonstrate that our method significantly enhances conic fitting accuracy under self-shadowing conditions and provides a flexible, accurate calibration solution with a single sphere.
Current research indicates that camera defocus can adversely affect phase accuracy in regions with abrupt reflectivity changes, particularly for objects with complex textures. To address this challenge, this paper constructs an error model, revealing a trigonometric relationship between phase errors and the tangent angle of the texture. By investigating the variation patterns of this error model, a one-dimensional phase-weighted average and parameter fitting model is proposed to correct phase errors. The proposed method requires only the projection of unidirectional fringe patterns for phase error correction, offering faster processing and higher stability compared to methods relying on bidirectional fringe patterns. The experimental results demonstrate that the algorithm proposed in this paper can effectively reduce systematic errors and improve the accuracy of the final point clouds reconstruction.
Effective suppression of pseudo-changes remains a critical challenge in optical remote sensing image change detection. To address this issue, a correlation-assisted and similarity-guided discriminative perception network, abbreviated as CSDPNet, is proposed. It leverages the synergy between temporal-channel interdependencies and feature-space similarities. The architecture of CSDPNet centers on four key components: a global-local information distillation encoder, a channel correlation difference-aware module (CCDAM) that integrates dual cross-weight fusion convolution (DCWFC), a multivariate difference feature fusion module (MDFFM), and a similarity-guided cross-layer dynamic fusion module (SCDFM). Specifically, the global-local information distillation encoder progressively distills hierarchical features through its information distillation mechanism. Then, the CCDAM learns discriminative features from multiple dimensions, including channel correlations, to extract robust global-local difference features and thereby suppress pseudo-changes. Subsequently, the MDFFM adaptively integrates global and local difference features through a hybrid gating and cross-attention mechanism. Finally, the SCDFM performs cross-layer feature fusion guided by cosine similarity, which enables the module to conduct layer-wise decoding and progressively refine the change features. Evaluations demonstrate that the proposed CSDPNet achieves the best performance on the LEVIR-CD, WHU-CD, and SYSU-CD datasets, with F1/IoU scores of 91.59%/84.48%, 94.96%/90.41%, and 83.24%/71.29%, respectively. This is underscored by particularly significant gains on the WHU-CD and SYSU-CD datasets, where F1/IoU improvements reach 0.53/0.96, and 0.92/1.34 percentage points over the prior state-of-the-art, respectively.
In visual servoing micromanipulation, the hysteresis nonlinearity of piezoelectric stages and image transmission delay significantly degrade positioning accuracy. To address these issues, this paper proposes a dual-layer control strategy based on an improved Extended Kalman Filter (B-W-EKF). First, image block matching combined with a Gaussian kernel interpolation algorithm is employed to obtain high-precision displacement measurements from microscopic image sequences, from which the voltage-displacement hysteresis loop is constructed. Then, the EKF is integrated with the Bouc-Wen (B-W) model, incorporating hysteresis nonlinearity into the state observation equations. Based on this model, a dual-layer control architecture that combines upper-layer Model Predictive Control (MPC) with lower-layer Sliding Mode Control (SMC) is designed: the upper-layer MPC performs global optimization, while the lower-layer SMC regulates position and velocity, thereby improving tracking accuracy. Experimental results show that the RMSE values for SMC, MPC, SMPC, iMPC, and the proposed dual-layer MPC-SMC are 0.1292 µm, 0.1366 µm, 0.0635 µm, 0.0827 µm, and 0.0372 µm, respectively, under triangular wave reference input, demonstrating the effectiveness of the control strategy in enhancing tracking precision.
Color-encoded single-shot fringe projection methods offer clear advantages in dynamic 3D measurement, but often suffer from distortion when measuring objects with complex surface textures, such as print circuit boards and camouflage-coated components. To address this issue, we propose a robust single-shot fringe projection profilometry approach that integrates enhanced morphological component analysis with fast Fourier transform-based fringe extraction. The method decomposes each RGB channel of the captured fringe image to suppress texture interference and employs a frequency-domain filtering strategy to isolate fringe information via a tailored fringe mask. By exploiting the energy concentration characteristics of fringe components in the frequency domain, fringe signals are directly extracted from the tunable Q-factor wavelet transform coefficients, thereby eliminating the need for iterative low-rank approximation or prior estimation of statistical parameters. This approach significantly improves computational efficiency and measurement accuracy while enhancing robustness to texture-induced artifacts. Additionally, it removes the need for pre-calibration of color channel compensation, improving adaptability to real-world conditions. Experimental results demonstrate that the proposed approach achieves accurate 3D reconstruction with superior performance in both processing speed and resistance to texture-induced artifacts.
Hybrid monocular visual odometry, which combines the advantages of learning-based and geometry-based methods, has gained widespread attention for its superior robustness and accuracy. Recent works use networks to predict optical flow to establish pixel correspondences between images while retaining a geometry-based nonlinear optimization backend. However, existing methods primarily rely on local features to refine optical flow. A critical limitation of these methods is that the loss of local features due to occlusion can introduce significant uncertainty, degrading the performance of visual odometry. To enhance the system's robustness, we propose an occlusion-aware monocular visual odometry that aggregates both spatial and temporal features, effectively leveraging global information to reduce the impact of occlusion. Our method consists mainly of a Spatial Feature Aggregation (SFA) module and a Temporal Feature Aggregation (TFA) module. SFA models image self-similarity, utilizing visible regions to guide optical flow estimation in occluded regions. Meanwhile, TFA captures dynamic variations across consecutive frames, enhancing the model's understanding of motion trends. Experiments demonstrate that our method achieves SOTA performance on multiple benchmarks. Notably, compared to the baseline, it reduces the average Absolute Trajectory Error (ATE) on the TartanAir test split by 25.71%.
To address the problem of robust stability analysis for interval type-2 fuzzy systems (IT2FSs), this article proposes an innovative analysis approach based on model transformation. First, a classical piecewise linear approximation method is utilized to process the upper boundary membership functions and lower boundary membership functions of the footprint of uncertainty in IT2FSs, resulting in linear boundary membership functions (MFs) that are more convenient for analysis, along with the corresponding approximation error functions. Subsequently, a novel error model transformation method is introduced to handle these error terms. By constructing new fuzzy rules and MFs, the boundary error terms are converted into a new fuzzy model, thereby incorporating more information about the error functions into the stability analysis. Based on this model, a robust stability condition in the form of linear matrix inequalities are derived, achieving improved robustness. Finally, the effectiveness of the proposed method is validated through simulations on real-world systems, and its superiority is demonstrated by comparison with existing methods.
Multi-object tracking (MOT) aims to associate objects of the same identity across video frames, with robust similarity measurement being crucial for maintaining tracking performance. However, the current inefficient integration of motion and appearance cues often leads to tracking failures in challenging scenarios, such as occlusions and missed detections. In this paper, we introduce LV2DMOT, a tracker that employs a novel paradigm for integrating motion and appearance cues through language and visual multi-modal feature learning, thereby generating more distinctive data association similarities. We propose three key techniques: I) A text matching task between tracking trajectories and candidate detections. This method uses text encoding of detection geometric information combined with a temporal model, Mamba, to extract temporal motion features of trajectories, enabling more accurate motion similarity calculations. II) A multi-modal, multi-level feature fusion model that integrates motion and appearance features via cross modal learning mechanism, resulting in more robust fused similarities. III) A learnable temporal attention model for trajectory appearance feature updates, which effectively aggregates historical visual features to improve the representational ability of trajectory appearance features, employing k-medoids for feature selection. Extensive experiments on the MOT17 and MOT20 datasets demonstrate that our method achieves state-of-the-art tracking performance.
Monocular depth estimation is an important task in computer vision, which aims to predict pixel-wise depth maps from input images. Previous works neglect the relationship between temporal and single image information, depth network and pose network, introducing inferior optimization efficiency and performance. In order to better utilize the temporal information from image sequences, an attention based Multi Single attention (MSA) module is introduced to guide the fusion of image features and cost volume. In addition, towards the purpose of increasing the performance, we propose a dual mask scheme to the well-established photometric and smoothness loss to prevent large variations of depth maps from harming the training process. Our experiments demonstrate that our method outperforms state-of-art depth estimation models on KITTI dataset.
In structured light 3D measurement systems, defocusing of the camera is inevitable. Under its influence, the complex texture on the object's surface causes significant errors to the wrapped phase, affecting measurement accuracy. To address this issue, this paper analyzes and establishes an error model for them, highlighting their relationship with the texture direction, and proposes a correction method for complex texture errors based on bidirectional fringe projection point cloud fitting. Theoretically, the point clouds obtained in both directions should be identical. Thus, the method corrects the phase by minimizing the distance between the corresponding points in the two point clouds, ultimately yielding the corrected point cloud. To eliminate overall point cloud displacement caused by calibration parameter errors, a pre-correction process is applied. Comparative experiments show that the proposed method can reconstruct objects with complex textures with higher accuracy. Compared to traditional methods, the mean absolute error (MAE) and root mean square error (RMSE) of the proposed method can be reduced by up to 33.6% and 39.1%, respectively.