Photometric stereo is a technique aimed at determining surface normals through the utilization of shading cues derived from images taken under different lighting conditions. However, existing learning-based approaches often fail to accurately capture features at multiple stages and do not adequately promote interaction between these features. Consequently, these models tend to extract redundant features, especially in areas with intricate details such as wrinkles and edges. To tackle these issues, we propose MSF-Net, a novel framework for extracting information at multiple stages, paired with selective update strategy, aiming to extract high-quality feature information, which is critical for accurate normal construction. Additionally, we have developed a feature fusion module to improve the interplay among different features. Experimental results on the DiLiGenT benchmark show that our proposed MSF-Net significantly surpasses previous state-of-the-art methods in the accuracy of surface normal estimation.
Accurate underwater surface normal estimation is essential for high-precision underwater 3D measurement, yet remains challenging due to complex optical effects. Photometric stereo (PS) can recover fine normal details but is sensitive to motion and non-Lambertian effects, while stereo matching provides disparity-based normal but loses high-frequency details. Moreover, existing benchmarks mainly focus on ideal in-air scenes and are unsuitable for challenging underwater illumination. To address these limitations, we develop a portable underwater measurement system for synchronized stereo and multi-illumination image acquisition, and establish the first synthetic underwater benchmark for surface normal estimation. We further propose a dual-branch frequency-domain fusion network that exploits the complementary characteristics of binocular vision and Multispectral Photometric Stereo (MPS). By performing frequency information fusion via the Fast Fourier Transform (FFT), the proposed network adaptively integrates reliable global geometry and fine local surface details. Experiments demonstrate that our method achieves accurate and robust surface normal estimation in challenging underwater environments.
Hyperspectral image super-resolution is essential for enhancing the spatial fidelity of HSI data, yet existing deep learning methods often struggle with substantial spectral redundancy and the limited non-linear modeling capacity of standard feed-forward networks (FFNs). To address these challenges, we propose Spectral Dynamic Attention Network (SDANet), a framework designed to adaptively suppress redundant spectral interactions. SDANet integrates two key components: 1) Dynamic Channel Sparse Attention (DCSA) module that computes channel-wise correlations and selectively preserves the most informative attention responses through dynamic and data-dependent sparsification. 2) Frequency-Enhanced Feed-Forward Network (FE-FFN) that jointly models spatial and frequency-domain representations to enhance non-linear expressiveness. Extensive experiments on two benchmark datasets demonstrate that SDANet achieves state-of-the-art HISR performance while maintaining competitive efficiency. The code will be made publicly available at https://github.com/oucailab/SDANet.
Stereo image super-resolution aims to generate high-resolution images by leveraging complementary information from binocular systems. Although previous studies have achieved impressive results, the potential of intra-view and cross-view information has not been fully exploited. To address this issue, we propose a novel multi-scale interaction network for stereo image super-resolution. Specifically, we design a Multi-scale Spatial-Channel Attention Module that utilizes multi-scale large separable kernel attention and simple channel attention to improve intra-view feature extraction. Additionally, we propose a Dual-View Epipolar Attention Module, utilizing an optimal transport algorithm to achieve more accurate matching along the epipolar line. Extensive experimental and ablation studies show that our method achieves competitive results that outperform most SOTA methods.
Semantic segmentation of point clouds plays a crucial role in computer vision, with diverse applications in urban modelling, autonomous driving, and virtual reality. Despite its significance, many existing methods face challenges when dealing with large-scale datasets, such as (1) unclear or incomplete boundary segmentation and (2) poor performance on sparse objects. These limitations stem from inadequate local context extraction and insufficient handling of density variations, which hinder the accuracy and robustness of segmentation. To address these challenges, we propose TFNet, an end-to-end deep neural network specifically designed to enhance local geometric feature extraction and improve performance on density variations. TFNet introduces three key components: (1) Rotation-Invariant and Geometric Feature Extractor (RIGFE), which independently captures rotation-invariant and geometric features; (2) Annularly Convolutional Attention Pooling (ACAP), which leverages annular convolution for effective relational feature extraction in both feature and geometric spaces; and (3) Subgraph Vector of Locally Aggregated Descriptors (SGVLAD), which learns position- and scale-invariant point set features. Experimental evaluations on benchmark datasets, including S3DIS, Toronto-3D, and Nanning Power Grid, demonstrate that TFNet outperforms existing methods by effectively addressing these challenges. The results highlight its ability to deliver superior segmentation accuracy and robustness in diverse scenarios.
Hyperspectral image (HSI) and light detection and ranging (LiDAR) data joint classification is a challenging task. Existing multisource remote sensing data classification methods often rely on human-designed frameworks for feature extraction, which heavily depend on expert knowledge. To address these limitations, we propose a novel dynamic cross-modal feature interaction network (DCMNet), the first framework leveraging a dynamic routing mechanism for HSI and LiDAR classification. Specifically, our approach introduces three feature interaction blocks: bilinear spatial attention block (BSAB), bilinear channel attention block (BCAB), and integration convolutional block (ICB). These blocks are designed to effectively enhance spatial, spectral, and discriminative feature interactions. A multilayer routing space with routing gates is designed to determine optimal computational paths, enabling data-dependent feature fusion. Additionally, bilinear attention mechanisms are employed to enhance feature interactions in spatial and channel representations. Extensive experiments on three public HSI and LiDAR datasets demonstrate the superiority of DCMNet over the state-of-the-art methods. Our codes are available at https://github.com/oucailab/DCMNet.
Camouflaged object detection (COD) from a single image is a challenging task due to the high similarity between objects and their surroundings. Existing fully supervised methods require labor-intensive pixel-level annotations, making weakly supervised methods a viable compromise that balances accuracy and annotation efficiency. However, weakly supervised methods often experience performance degradation due to the use of coarse annotations. In this paper, we introduce a new weakly supervised approach for camouflaged object detection to overcome these limitations. Specifically, we propose a novel network, MGNet, which tackles edge ambiguity and missed detections by utilizing initial masks generated by our custom-designed Cascaded Mask Decoder (CMD) to guide the segmentation process and enhance edge predictions. We introduce a Context Enhancement Module (CEM) to reduce the missing detection, and a Mask-guided Feature Aggregation Module (MFAM) for effective feature aggregation. For the weak supervision challenge, we propose BoxSAM, which leverages the Segment Anything Model (SAM) with bounding-box prompts to generate pseudo-labels. By employing a redundant processing strategy, high quality pixel-level pseudo-labels are provided for training MGNet. Extensive experiments demonstrate that our method delivers competitive performance against current state-of-the-art methods.
In this paper, a video inpainting framework that combines Local Flow Propagation with the Global Multi-scale Dilated Transformer, referred to as LFP-GMDT, is proposed. First, optical flow is utilized to guide the bidirectional propagation of features between adjacent frames for local inpainting. With the introduction of deformable convolutions, optical flow errors are corrected, substantially enhancing the accuracy of both local inpainting and frame alignment. Following the local inpainting stage, a multi-scale dilated Transformer module is designed for global inpainting. This module integrates multi-scale feature representations with an attention mechanism, introducing a multi-scale dilated attention mechanism that balances the modeling capabilities of local details and global structures while reducing computational complexity. Experimental results show that, compared to existing models, LFP-GMDT performs exceptionally well in detail restoration and structural integrity, particularly excelling in the recovery of edge structures, leading to an overall enhancement in visual quality.
We propose a method that is able to use the monocular zoom technology for real-scale 3-D reconstruction of the scene. To reconstruct the scene, we take a sequence of zoomed-in and zoomed-out figures. First, we can estimate zoomed-in camera parameters using the known zoomed-out camera parameters, which avoids calibrating the camera parameters twice. Then, we use the structure from motion (SfM) method (COLMAP) to reconstruct free-scale translations among these figures. After that, as we have pairs of zoom frames in the same scene, we can calculate the true scale of the scene by comparing the ratio between the free-scale translation of a pair of zoom frames and the difference in zoomed-out and the zoomed-in focal length. Finally, we use RAFT-stereo to compute the depth of the scene. In detail, we select two adjacent figures taken at the same focal length, make a stereo correction for them, and remove the nonco-vision area of the corrected images. This way, we obtain a more accurate matching of these images and then get a dense real-scale 3-D reconstruction. Experimental results have demonstrated that our method achieves good performance on monocular 3-D reconstruction with the real scale.
Private set intersection (PSI) allows Sender holding a set X and Receiver holding a set Y to compute only the intersection X∩ Y for Receiver. We focus on a variant of PSI, called fuzzy PSI (FPSI), where Receiver only gets points in X that are at a distance not greater than a threshold from some points in Y . Most current FPSI approaches first pick out pairs of points that are potentially close and then determine whether the distance of each selected pair is indeed small enough to yield FPSI result. Their complexity bottlenecks stem from the excessive number of point pairs selected by the first picking process. Regarding this process, we consider a more general notion, called fuzzy mapping (Fmap), which can map each point of two parties to a set of identifiers, with closely located points having a same identifier, which forms the selected point pairs. We initiate the formal study on Fmap and show novel Fmap instances for Hamming and L_∞ distances to reduce the number of selected pairs. We demonstrate the powerful capability of Fmap with some superior properties in constructing FPSI variants and provide a generic construction from Fmap to FPSI. Our new Fmap instances lead to the fastest semi-honest secure FPSI protocols in high-dimensional space to date, for both Hamming and general L_𝗉∈ [1, ∞ ] distances. For Hamming distance, our protocol is the first one that achieves strict linear complexity with input sizes. For L_𝗉∈ [1, ∞ ] distance, our protocol is the first one that achieves linear complexity with input sizes, dimension, and threshold.
Photometric stereo (PS) endeavors to ascertain surface normals using shading clues from photometric images under various illuminations. Recent deep learning-based PS methods often overlook the complexity of object surfaces. These neural network models, which exclusively rely on photometric images for training, often produce blurred results in high-frequency regions characterized by local discontinuities, such as wrinkles and edges with significant gradient changes. To address this, we propose the Image Gradient-Aided Photometric Stereo Network (IGA-PSN), a dual-branch framework extracting features from both photometric images and their gradients. Furthermore, we incorporate an hourglass regression network along with supervision to regularize normal regression. Experiments on DiLiGenT benchmarks show that IGA-PSN outperforms previous methods in surface normal estimation, achieving a mean angular error of 6.46 while preserving textures and geometric shapes in complex regions.
Sketch-based image retrieval is an important research topic in the field of image processing. Hand-drawn sketches consist only of contour lines, and lack detailed information such as color and textons. As a result, they differ significantly from color images in terms of image feature distribution, making sketch-based image retrieval a typical cross-domain retrieval problem. To solve this problem, we constructed a perceptual space consistent with both textures and sketches, and using perceptual similarity for sketch-based texture retrieval. To implement this approach, we first conduct a set of psychological experiments to analyze the similarity of visual perception of the textures, then we create a dataset of over a thousand hand-drawn sketches according to the textures. We proposed a layer-wise perceptual similarity learning method that integrates perceptual similarity, with which we trained a similarity prediction network to learn the perceptual similarity between hand-drawn sketches and natural texture images. The trained network can be used for perceptual similarity prediction and efficient retrieval. Our experimental results demonstrate the effectiveness of sketch-based texture retrieval using perceptual similarity.
The imbalanced data classification has gained popularity in machine learning research domain due to its prevalence in numerous applications and its difficulty. However, the majority of contemporary work primarily focuses on addressing between-class imbalance issues. Previous researches have shown that combined with other elements, such as within-class imbalance, small sample size and the presence of small disjuncts, the imbalanced data significantly increase the difficulties for the traditional classifiers to learn. Therefore, we propose a novel MeanShift-guided oversampling with self-adaptive sizes for imbalanced data classification. The proposed MeanShift-guided oversampling technique can simultaneously consider the distribution of minority class and majority class within the sphere with the current minority instance as its center, which can favor addressing small sample size and avoiding overlapping issues often caused by the nearest neighbor (NN)-based oversampling techniques. The incorporation of random vector and flexible cut-off mechanism for vector length can enhance the diversity among the generated synthetic minority instances and avoid overlapping, which makes it suitable for small sample size and small disjuncts problems. To address between-class and within-class imbalance issues, we also introduce a self-adaptive sizes assignment strategy for each minority instance to be oversampled, where the assigned size is inversely proportional to its density and its distance from the majority class. In addition to eliminating within-class imbalance, the strategy can ensure that the informative border minority instances have more opportunities to be oversampled, thus improving classification performance. Extensive experimental results on some datasets with different distributions and imbalance ratios show the proposed algorithm outperforms other compared ones with significant difference.
Hyperspectral image (HSI) contains abundant spatial and spectral information, making it highly valuable for unmixing. In this paper, we propose a Dual-Stream Attention Network (DSANet) for HSI unmixing. The endmembers and abundance of a pixel in HSI have high correlations with its adjacent pixels. Therefore, we adopt a “many to one” strategy to estimate the abundance of the central pixel. In addition, we adopt multiview spectral method, dividing spectral bands into multiple partitions with low correlations to estimate abundances. To aggregate the estimated abundances for complementary from the two branches, we design a cross-fusion attention network to enhance valuable information. Extensive experiments have been conducted on two real datasets, which demonstrate the effectiveness of our DSANet.
The conventional perceptual hashing algorithms are constrained to a singular global feature extraction algorithm and lack efficient scalability adaptation. To address this problem, an image-perceptual hashing algorithm based on convolutional neural networks is proposed in this paper. First of all, the entire image is convolved by the backbone network to obtain a feature map. The Region Proposal Network (RPN) is employed to generate multiple-sized proposal frames at each location by using sliding windows. Considering the complexity and diversity of the object, proposal boxes of various sizes and shapes are formulated, and the local features are comprehensively exploited in an image, thereby, generating a perceptual hash code that can represent the semantic features of an image strongly. Moreover, The Mean Square Error (MSE) loss is incorporated into the optimization process to evaluate the coincidence between the proposal frame and the actual frame, generating more representative hash codes. Finally, an image perceptual hash code with high intuitive features can be formulated through iterative training of the proposed convolutional neural networks. Extensive experimental results demonstrate that the proposed image perceptual hashing algorithm based on a convolutional neural network surpasses other state-of-the-art methods.
Multi-source remote sensing data classification has emerged as a prominent research topic with the advancement of various sensors. Existing multi-source data classification methods are susceptible to irrelevant information interference during multi-source feature extraction and fusion. To solve this issue, we propose a sparse focus network for multi-source data classification. Sparse attention is employed in Transformer block for HSI and SAR/LiDAR feature extraction, thereby the most useful self-attention values are maintained for better feature aggregation. Furthermore, cross-attention is used to enhance multi-source feature interactions, and further improves the efficiency of cross-modal feature fusion. Experimental results on the Berlin and Houston2018 datasets highlight the effectiveness of SF-Net, outperforming existing state-of-the-art methods.
Camouflaged object detection (COD) presents a persistent challenge in accurately identifying objects that seamlessly blend into their surroundings. However, most existing COD models overlook the fact that visual systems operate within a genuine 3D environment. The scene depth inherent in a single 2D image provides rich spatial clues that can assist in the detection of camouflaged objects. Therefore, we propose a novel depth-perception attention fusion network that leverages the depth map as an auxiliary input to enhance the network's ability to perceive 3D information, which is typically challenging for the human eye to discern from 2D images. The network uses a trident-branch encoder to extract chromatic and depth information and their communications. Recognizing that certain regions of a depth map may not effectively highlight the camouflaged object, we introduce a depth-weighted cross-attention fusion module to dynamically adjust the fusion weights on depth and RGB feature maps. To keep the model simple without compromising effectiveness, we design a straightforward feature aggregation decoder that adaptively fuses the enhanced aggregated features. Experiments demonstrate the significant superiority of our proposed method over other states of the arts, which further validates the contribution of depth information in camouflaged object detection. The code will be available at https://github.com/xinran-liu00/DAF-Net.
In this work, we concentrate on exciting the intrinsic local consistency of stereo matching through the incorporation of superpixel soft constraints, with the objective of mitigating inaccuracies at the boundaries of predicted disparity maps. Our approach capitalizes on the observation that neighboring pixels are predisposed to belong to the same object and exhibit closely similar intensities within the probability volume of superpixels. By incorporating this insight, our method encourages the network to generate consistent probability distributions of disparity within each superpixel, aiming to improve the overall accuracy and coherence of predicted disparity maps. Experimental evalua tions on widely-used datasets validate the efficacy of our proposed approach, demonstrating its ability to assist cost volume-based matching networks in restoring competitive performance.
Cross-Domain Few-Shot Learning has witnessed great stride with the development of meta-learning. However, most existing methods pay more attention to learning domain-adaptive inductive bias (meta-knowledge) through feature-wise manipulation or task diversity improvement while neglecting the phenomenon that deep networks tend to rely more on high-frequency cues to make the classification decision, which thus degenerates the robustness of learned inductive bias since high-frequency information is vulnerable and easy to be disturbed by noisy information. Hence in this paper, we make one of the first attempts to propose a Frequency-Aware Prompting method with mutual attention for Cross-Domain Few-Shot classification, which can let networks simulate the human visual perception of selecting different frequency cues when facing new recognition tasks. Specifically, a frequency-aware prompting mechanism is first proposed, in which high-frequency components of the decomposed source image are switched either with normal distribution sampling or zeroing to get frequency-aware augment samples. Then, a mutual attention module is designed to learn generalizable inductive bias under CD-FSL settings. More importantly, the proposed method is a plug-and-play module that can be directly applied to most off-the-shelf CD-FLS methods. Experimental results on CD-FSL benchmarks demonstrate the effectiveness of our proposed method as well as robustly improve the performance of existing CD-FLS methods. Resources at https://github.com/tinkez/FAP_CDFSC.
Mike Chantler合作论文数School of Mathematical & Computer Sciences;Heriot-Watt University8