Sequential segmentation is fundamental for analyzing dynamic behaviors in complex sequence data. While most existing methods rely on single-view representations and impose temporal constraints simply on the self-representation matrix, which limits their ability to exploit complementary multi-view information and temporal continuity. To address these issues, we propose Implicit Multi-view Temporal Subspace Clustering (IMTSC) for sequence segmentation, which leverages two regularization terms to propagate temporal information across multi-view sequence data. First, temporal constraints are imposed on the latent representations derived from multi-view sequence data to preserve the intrinsic temporal continuity between adjacent frames. Second, recognizing that clustering results depend on the construction of the self-representation matrix, a weighted temporal consistency constraint is applied to propagate temporal information across latent representations. This dual-layer modeling strategy, referred to as the dual-level temporal consistency structure, enforces temporal constraints on the intrinsic properties of the data while facilitating information propagation. It enables the acquisition of more accurate temporal information and effectively enhances the stability of temporal data clustering. The proposed model can be efficiently solved using the Alternating Direction Method of Multipliers (ADMM). Experimental results on multiple benchmark datasets demonstrate that, IMTSC achieves higher segmentation accuracy on image-based datasets and exhibits superior robustness.
Existing image dehazing methods often rely on handcrafted priors or heuristic rules for atmospheric light estimation, which tend to fail in the presence of specular highlights, reflective surfaces, and spatially nonuniform haze. Such inaccuracies usually propagate to transmission map estimation and subsequently degrade the visual quality of restored images. To address these issues, we propose a variational image dehazing framework based on joint color-brightness decomposition. Instead of estimating atmospheric light solely from extreme pixel intensities, the proposed method decomposes the hazy image into a Color Component, a Structural Component, and a Highlight-Suppression Term, and then estimates atmospheric light by jointly integrating chromatic and brightness cues. This design improves robustness in highlight-contaminated regions and reduces the risk of atmospheric light misestimation. Building on this decomposition, we further develop a semi-decoupled variational dehazing model in which transmission map estimation and scene radiance recovery are formulated as two alternating but interrelated subproblems. This strategy reduces the optimization complexity of fully coupled formulations while preserving important structural details in challenging regions. In addition, an adaptive brightness enhancement scheme is introduced to refine the restored image according to the mean intensity and the estimated transmission map, thereby alleviating over-enhancement and preserving fine details. Extensive subjective and objective experiments on both synthetic and real-world hazy images demonstrate that the proposed method performs favorably against several representative dehazing approaches, particularly in scenes containing large highlight regions, complex illumination, and uneven haze distribution.
Pedestrian detection plays a vital role in autonomous driving and intelligent surveillance systems. However, traditional pedestrian detection methods suffer from significant performance degradation in nighttime scenes due to challenging factors such as low illumination, poor contrast, and non-uniform lighting distributions. This paper presents an innovative nighttime pedestrian detection method by integrating traditional Retinex image enhancement theory with deep learning-based detection networks to achieve illumination-invariant pedestrian feature representation. Specifically, we design a Light Decomposition Module (LD-Module) that decomposes in put nighttime images into reflectance and illumination components. By eliminating the illumination component, we extract intrinsic pedestrian features that remain robust against lighting variations, effectively mitigating the adverse impact of illumination on nighttime pedestrian detection. Additionally, we introduce a Context-aware Pedestrian-Background Attention (CPBA) mechanism featuring a novel triple-branch architecture that efficiently captures contextual information across both spatial and channel dimensions, while adaptively fusing illumination-invariant reflectance features with semantic features. This mechanism establishes comprehensive global relationships between pedestrians and their nighttime background, enhancing discriminative capability while suppressing background noise interference. Extensive experimental results demonstrate the effectiveness of our proposed RIFL-Net, which achieves a remarkable MR-2 of 6.10% on the NightOwls Reasonable sub set, significantly outperforming state-of-the-art methods. The proposed network demonstrates strong robustness against diverse lighting variations, offering a highly reliable solution for pedestrian detection in complex low-light environments.
Multi-model fitting aims to recover multiple geometric structures from data contaminated by gross outliers and pseudo-outliers. Recent learning-based methods improve sampling efficiency, but most still rely on loosely coupled pipelines in which feature extraction, weight prediction, and hypothesis sampling are optimized separately, lacking a reliable iterative refinement process to progressively stabilize predictions before sampling. As a result, hypothesis generation is often based on unreliable one-shot predictions, leading to low-purity minimal subsets in cluttered multi-structure scenes.To address this issue, we propose a Structure-Aware Gated Parallel Fitting (SGPF) framework. SGPF employs a Dynamic Graph Convolutional Neural Network backbone to encode local neighborhood structure, a Joint Iterative Refinement module to progressively refine features and predictions in a closed loop, and a Feature-Guided Pivot Sampling strategy to generate structure-consistent minimal subsets by conditioning subset selection on feature similarity. Together, these components form a unified framework that couples local structural encoding, iterative refinement, and structureaware sampling, thereby improving hypothesis purity and robust multi-model fitting in complex scenes. Experiments on vanishing point, homography, and fundamental matrix estimation benchmarks demonstrate that SGPF achieves strong and consistent performance across diverse multi-model fitting tasks, with particularly notable gains on multi-homography fitting and cross-domain vanishing point estimation, while maintaining a favorable accuracy–efficiency trade-off.
Repetitive action counting is a key task in human-centric video analysis, supporting fitness monitoring and medical rehabilitation. Most current systems require dense frame-level labels, which are expensive to obtain. Furthermore, real-world human motion is non-stationary, often featuring tempo drifts and intensity changes that degrade counting accuracy. To address these issues, we propose TWCRAC, a Time-Window-Cycle network based on periodic representation learning. Our framework introduces two technical innovations. First, we develop an intra-video Temporal Cycle Consistency (TCC) objective for self-supervision. This mechanism exploits the intrinsic self-similarity of motion within a single sequence. It allows the model to learn structured cyclic features using only video-level labels. Second, we design an adaptive counting algorithm that leverages dynamically computed statistical thresholds from local features. By calculating these thresholds from the periodicity score curves optimized by our TCC loss, this design overcomes the limitations of fixed thresholds under complex motion, increases robustness, and supports reliable performance in non-stationary pattern recognition scenarios. We evaluated TWCRAC on three benchmarks: RepCount, Countix, and UCFRep. Experimental results show that our method achieves state-of-the-art accuracy while significantly reducing annotation costs. The system is also computationally efficient, reaching 43+ FPS on a single GPU, making it suitable for real-time autonomous video systems.
Introduction:To address the lack of integrated and clinically applicable motion capture systems for hand function assessment, we developed a wearable device capable of simultaneously recording finger curvature and surface electromyography (sEMG) signals from both healthy individuals and patients with motor impairments. Methods:The dataset comprises 900 measurements of six predefined gestures collected from 15 participants using a six-channel sEMG motion-capture glove. Data were obtained through hospital-based field acquisition, ensuring clinical relevance and independence of the hardware-database framework. The recorded signals were processed using a Savitzky-Golay filter, followed by Short-Time Fourier Transform (STFT) for spectrogram generation. Multiple machine learning models, including SVM, LightGBM, and MLP, were employed for gesture classification. Results:Most models achieved over 90% precision on both cross-validation and test sets, demonstrating robust classification performance across different gesture types and subject conditions. Discussion:These results confirm that the proposed system maintains high recognition accuracy even in severely impaired subjects. The dataset presented here offers substantial value for gesture recognition research, rehabilitation assessment, and neuromuscular signal analysis.
Nuclear segmentation and classification play a crucial role in pathological image analysis. However, it is frequently challenged by blurred nuclear boundaries and complex structures in digital pathology slides, due to factors such as staining techniques and imaging methods, posing a significant challenge for accurate segmentation and classification. To this end, we propose a novel and efficient approach for nuclear identification, termed Information Propagation with Multi-Granularity Morphology-Guided Network (IPMMG). Specifically, IPMMG progressively captures edge morphology information from different network layers while simultaneously incorporating structural morphology features at multiple granularities. By explicitly propagating features related to both the edge and the structure, our approach constrains semantic features to focus on contours of the region of interest in the nuclear segmentation task, thus mitigating the challenge of blurred morphology. Experiments on public datasets demonstrate that IPMMG achieves state-of-the-art (SOTA) performance in segmentation, as measured by Dice and IoU scores, while also attaining competitive results in classification with DQ, SQ, and PQ metrics. In particular, our proposal IPMMG excels in handling nuclei with blurred edges and complex structures.
Repetitive action counting from human pose data is a fundamental yet challenging task, often hindered by viewpoint variation and occlusion. To overcome these challenges, we propose the Contrastive-Enhanced and Partitioned Multi-Scale Network (CEPMS-Net), an end-to-end pose-level framework that enhances feature representations for robust and accurate counting. The proposed architecture integrates four key components, highlighted by two novel modules designed to address feature robustness. Specifically, the Pose Contrast Enhancement Module (PCE-Module) systematically augments training data through diverse geometric and occlusion-based transformations. Combined with contrastive learning, it aligns feature representations across augmented data, enhancing the model's invariance to viewpoint variation and occlusion. Furthermore, the Partitioned Multi-Scale Convolution Module (PMSC-Module) hierarchically models local and global action features by exploiting the human body's topological structure. This architecture facilitates the extraction of fine-grained limb dynamics and global body coordination, ensuring comprehensive action representation. Additionally, the Bidirectional Cross-Attention Module (BCA-Module) enables effective multi-scale feature fusion, while the Projection Head and Action Counting Module (PHAC-Module) performs the final stage of action counting. Experimental results show that CEPMS-Net sets a new benchmark for robustness and counting accuracy, consistently outperforming existing methods across three major benchmarks: the pose-annotated versions of the Repetition Counting (RepCount), University of Central Florida Repetition (UCFRep), and Countix-Fitness datasets (namely RepCount-pose, UCFRep-pose, and Countix-Fitness-pose). Notably, on the challenging RepCount-pose dataset, CEPMS-Net achieves a Mean Absolute Error (MAE) of 0.225 and an Off-By-One (OBO) accuracy of 0.599, providing compelling evidence of its superior performance in handling complex real-world scenarios.
Nighttime imaging is degraded by multiple scattering from artificial light sources, resulting in prominent glow artifacts. Existing methods often fail to suppress brightness scattering and exhibit limited robustness. To address these limitations, we propose the Nighttime Glow Suppression via Structural Cascade Decomposition (NGSCD) model, which decomposes images into hidden and salient structure layers and a glow layer via inverse multiple scattering derivation. Furthermore, a robust camera response model based on Retinex theory utilizes a local maxima channel to enhance brightness and contrast. To address multicolor illumination, the model incorporates the grayscale world hypothesis, thereby minimizing color distortion. The proposed method demonstrates robust performance across diverse nighttime scenes and outperforms state-of-the-art methods in both glow suppression and visibility enhancement, as validated by experimental results.
To enable more natural motion mapping between the human arm and a robotic counterpart while reducing control complexity, this paper presents a novel seven-degree-of-freedom (7-DoF) bionic robotic arm with hybrid pneumatic–electric actuation in an antagonistic configuration inspired by the skeletal structure and muscular actuation of the human upper limb. The design combines the high power density and intrinsic compliance of pneumatic artificial muscles with the precision and stability of electric motors, improving motion adaptability and payload-to-weight performance. Kinematic feasibility and motion smoothness for human-like waving are validated via forward kinematics and redundancy-resolved inverse kinematics, together with trajectory simulations. To quantitatively evaluate dexterity and operational range, Monte Carlo sampling is used to generate reachable postures across the workspace, producing a wrist activity map that characterizes attainable orientations and maneuverability. A prototype testbed is built to verify physical performance. Joint-angle tracking experiments for the wrist and elbow, as well as whole-arm coordinated-motion tests, demonstrate accurate trajectory tracking, smooth transitions, and stable motion. These results confirm the mechanical soundness and effectiveness of the proposed hybrid antagonistic actuation scheme. This work provides a practical basis for advanced control development and offers insights into hybrid actuation design for bionic robotic systems.
Nighttime pedestrian detection suffers from feature contamination under extreme illumination, while traditional enhancement methods often destroy semantic structures. This paper proposes RMFNet, a novel Retinex-based multiplicative fusion network that achieves robust nighttime pedestrian detection through illumination decomposition and hierarchical multiplicative fusion. First, an Illumination Separation Module (IS-Module) is introduced to physically decouple illumination from reflectance, obtaining lighting-invariant features while preserving semantic structures. Second, a Hierarchical Multiplicative Fusion Module (HMF-Module) is designed to amplify pedestrian features and suppress background noise via element-wise multiplication. Experiments on two challenging benchmarks demonstrate that RMFNet achieves a superior 6.0% miss rate on the NightOwls Reasonable subset, providing a robust solution with strong generalization for all-condition detection.
The accuracy and stability of pathology image segmentation have become critical factors in clinical applications such as cancer screening and tumor grading. However, the presence of complex local structures, uncertain regions, and subtle morphological variations in pathological images continues to pose significant challenges. Most existing feature fusion approaches rely on the simplistic aggregation of extracted features, neglecting the unique characteristics and relative importance of distinct feature representations, which ultimately limits their potential to enhance model performance. To address these issues, we propose a Gestalt-Inspired Feature Integration Network (GeNet), a novel architecture inspired by Gestalt theory that mirrors the human visual system's ability to derive holistic understanding from partial information. Embracing the principle that ‘the whole is greater than the sum of its parts,’ GeNet introduces a mechanism to synergistically leverage multi-scale information, which assesses the similarity between features to achieve a more meaningful fusion of global context and local detail. Given the variability in target appearance within pathological images, we use information entropy to quantify feature uncertainty, allowing the model to prioritize uncertain regions and reduce the occurrence of ambiguous results. To explicitly eliminate multi-feature redundancy and misalignment, the refinement block utilizes parallel convolutional recalibration to fully leverage the advantages of various features. Extensive experiments on multiple pathological image segmentation datasets, including GlaS, GCaSeg, and EBHI-Seg, demonstrate that GeNet achieves high accuracy and strong robustness, offering a new perspective for joint modeling of global and local features in medical image analysis.
Estimating the 6D pose of unseen objects is a fundamental challenge in robotics and industrial automation, where reliance on textured 3D models, multi-view sequences, or Structure-from-Motion (SfM) often limits scalability due to the vast diversity of real-world objects. To address these limitations, we propose an open-vocabulary framework that leverages vision-language models (VLMs) to estimate 6D poses of unseen objects without requiring prior object-specific training data or CAD models. Unlike existing approaches that rely on relative pose estimation without explicit feature alignment, our method introduces the Feature Alignment and Correspondence Module (FACM), a hierarchical architecture that integrates reference and query features via self-attention and cross-attention, while preserving geometric consistency through layer-wise pose encoding. In addition, we introduce a customized feature guidance mechanism and spatial and class-level aggregation strategy to enhance discriminability and robustness under occlusion and viewpoint variations. Through extensive ablation studies, we demonstrate the critical role of textual cues in boosting pose estimation accuracy and validate the contributions of each component in our framework. Experiments on REAL275 and Toyota-Light datasets show that our method outperforms state-of-the-art techniques, achieving superior generalization to unseen objects. Our work advances the field of 6D pose estimation and provides a scalable solution for real-world applications where 3D models are unavailable. The code will be available on: \url{https://github.com/Metwalli/OVZS6D}
Surface defect detection of printed circuit boards (PCBs) is critical for quality control in electronic manufacturing. Existing public datasets often rely on synthetic images or are collected under controlled laboratory conditions, and may not fully capture the complex illumination variations and large-scale defect diversity present in real production lines. Here, this Data Descriptor presents PCB-IND, a large-scale real-world dataset directly acquired from an industrial automated optical inspection (AOI) system. The dataset contains 4,789 images covering eight typical surface defect categories common in industrial inspection, with a total of 5,932 annotated defect instances. Compared to existing benchmarks, PCB-IND preserves non-uniform illumination, high-contrast imaging characteristics typical of etching inspection, and cross-scale defect patterns found in industrial environments. It also differs from prior PCB datasets by combining inline AOI acquisition, natural long-tail defect distribution, and standardized multi-format release within a single industrial dataset resource. To evaluate the usability and benchmark stability of the dataset, we conducted experiments using several representative object detection models. The results indicate that PCB-IND supports the training and evaluation of various models under real industrial conditions, providing a useful real-world data resource for research in industrial defect detection, early-stage process monitoring, and defect repair decision-making.
Multi-model fitting is fundamental for robust geometric estimation in computer vision. However, recent deep learning methods enable parallel model detection but rely on simple architectures that inadequately model spatial relationships. Moreover, current methods typically generate hypotheses only through minimal solvers on randomly sampled points, thus failing to explore the full diversity of the solution space. To address these limitations, we propose a novel Jacobian-based Gaussian uncertainty modeling framework, which analytically propagates covariance through geometric transformations and enables efficient expansion of the hypothesis space with strong theoretical guarantees. We further introduce a Gaussian Hypothesis Generation Network (GHG-Net) to learn global parameter distributions, enabling the generation of diverse and geometrically valid hypotheses. Additionally, our network captures spatial relationships among observations by employing a dynamic graph neural network with a multi-head attention mechanism. This yields more accurate sample and inlier weights, significantly improving the quality of hypothesis generation. Extensive experiments on three representative geometric estimation tasks (i.e. vanishing point detection, fundamental matrix estimation, and homography estimation) demonstrate that our method achieves new state-of-the-art accuracy and stability, while maintaining high computational efficiency.
Single-image dehazing remains a challenging low-level vision task because haze degradation is inherently depth-dependent and spatially non-uniform. To address this problem, we propose DMSH-Net, a Depth-Aware Multi-Scale Hybrid Vision Network specifically designed for robust single-image dehazing. DMSH-Net is designed to implicitly capture haze variations through hierarchical feature recalibration, nonlinear residual refinement, and multi-scale contextual aggregation. Specifically, we introduce a redesigned convolutional squeeze-and-excitation attention (CSEA) module, which replaces fully connected transformations with convolutional operations and global average pooling to jointly model channel dependencies and spatial context. Building on CSEA, a nonlinear CSEA-coupled residual block (NCCRB) is developed to enhance local feature representation and improve adaptability to haze with varying densities. Furthermore, a multi-scale dilated convolution bottleneck is incorporated to enlarge the receptive field and aggregate haze-aware contextual information across multiple spatial scales, thereby improving the restoration of regions with varying scene depths. Extensive experiments on standard benchmarks demonstrate that DMSH-Net consistently achieves superior quantitative performance across full-reference and no-reference evaluations, thereby validating its robustness in complex real-world dehazing scenarios.
With the rapid development of deep object detection, there is an increasing need for more efficient and accurate day-night domain adaptation methods. Many domain-adaptation-based detectors perform well in night-time scenes, but pixel-level appearance alignment can lead to excessive focus on irrelevant details, hindering task-related feature learning. To address this, we propose a method based on structural consistency constraints. Specifically, we first design an illumination-invariant structural representation learning module (ISM) to extract structural information as pseudo-references, and construct a bridge coder (BCoder) to map pseudo-references to the shallow feature space. Then, we reuse the deep detection network to improve pseudo-reference detection compatibility, and introduce a flexible structural consistency constraint. Experimental results show that, compared with DAINet, DoSC-Net achieves a 4.53% improvement in mAP on the DarkFace dataset and a 1.84% improvement on the ExDark dataset, demonstrating its effectiveness and robustness in day-night domain adaptation detection.
Multi-exposure image fusion (MEF) is a fundamental technique for High Dynamic Range (HDR) image generation. However, existing MEF approaches commonly suffer from detail loss and motion-induced artifacts, which severely degrade the visual quality of the fused results. To tackle these issues, this paper presents a robust MEF framework based on brightness adaptation and dual-weight decomposition. Specifically, an adaptive brightness weighting strategy is developed that jointly accounts for the global and local luminance characteristics of cross-exposure inputs and dynamically adjusts the corresponding weight distribution. This enables an effective trade-off between local contrast enhancement and global exposure consistency. In parallel, by introducing multi-channel vector gradient weights, the proposed framework can accurately detect edge regions that exhibit similar brightness yet pronounced color differences, thereby providing more precise guidance for structural information extraction. Furthermore, a joint constraint mechanism that integrates structural similarity and brightness consistency is constructed. Through the generation of consistency maps and a reference-frame-guided motion detection and artifact suppression scheme, ghosting is effectively mitigated and the overall fusion quality is substantially improved. Extensive experiments on public datasets demonstrate that the proposed method delivers superior performance in both static and dynamic scenes.
Existing multi-exposure fusion methods face two primary challenges: loss of detail and ghosting artifacts. To address these challenges, we propose a novel ghost-free multi-exposure fusion method. A joint constraint mechanism, integrating structural similarity index and luminance consistency constraints, is introduced to effectively detect and suppress ghosting artifacts. Furthermore, by decoupling luminance and chrominance components in the YUV color space, we introduce a dual-weight fusion model to address initial detail loss. During the fusion process, we employ a weight-guided pyramidbased strategy to adaptively integrate image information across multiple scales, thereby enhancing detail preservation. Extensive subjective and objective evaluations confirm that the proposed method consistently outperforms existing methods in terms of visual quality and quantitative metrics.
Human motion data has found widespread application across multiple domains. However, raw motion data acquisition often suffers from information loss or distortion caused by inevitable environmental occlusions and hardware limitations of capture devices. In recent years, Low-Rank Matrix Completion (LRMC)-based motion recovery methods have gained growing research attention. Most of these methods treat all human motion parts uniformly, inadvertently overlooking the inherent correlations within the same body parts. In this paper, we have developed a multi-level fine-grained operator for motion data that allows us to treat data from the same body part as a cohesive unit. This approach focuses the recovery effort on strongly correlated data. To integrate the characteristics of data segmented from different fine-grained levels, we fuse the recovery results obtained from these various levels. Given that data from the same body part exhibit more pronounced low-rank characteristics, compared to methods that treat all motion parts equally, our proposed Multi-level Fine-grained Fusion (MFF) model can fully exploit the local similarity and low-rank properties of human motion data. Furthermore, we have developed an efficient Alternating Direction Method of Multipliers (ADMM) algorithm to solve the proposed model. Experiments conducted on multiple motion sequences demonstrate that the recovery performance of our proposed method surpasses that of other state-of-the-art approaches.