With the rapid development of the Underwater Internet of Things (UIoT), multi-view image stitching in large-scale underwater visual monitoring has become a key technology for marine sensing networks to achieve deep-sea exploration and environmental monitoring. In dark deep-sea environments, imaging heavily relies on artificial light sources equipped on underwater carriers. The color cast induced by underwater light attenuation and the non-uniform illumination caused by artificial lighting consequently result in noticeable stitching seams. Therefore, this paper proposes a two-stage stitching-oriented enhancement framework. In the first stage, single-image restoration is conducted to mitigate underwater color cast, low contrast, and blurred details. In the second stage, pseudo-overlapping image pairs are constructed to simulate illumination imbalance in overlapping regions, and cross-view consistency losses are introduced to constrain color and brightness variations across views. Experimental results on UWIS show that our method reduces the mean color difference in overlapping regions by 37.1% and achieves an inlier ratio of 90.81%. These results verify that the proposed method effectively balances single-frame visibility and cross-view consistency for underwater image stitching.
Detector-free methods have achieved notable progress in recent years, but the limited capacity of existing models to leverage multi-frequency features continues to constrain matching performance. To address this challenge, we propose a novel feature matching approach based on a dual-frequency Transformer model, which effectively exploits multi-level image information. The proposed architecture employs dual attention branches, specifically designed to capture high-frequency details and low-frequency structural features. The high-frequency attention branch incorporates a feature enhancement module to accentuate edge visual features, which play a pivotal role in matching tasks. In addition, a frequency-based loss function is designed to constrain the consistency and integrity of features in the frequency domain during the feature extraction process, effectively mitigating frequency feature distortion. The proposed method not only enhances the model’s ability to represent contextual features across different frequency components but also improves selective attention to reliable feature details. Experimental results demonstrate the proposed method achieves superior performance in multiple feature matching tasks.
Person search is a crucial area of image retrieval and has diverse practical applications in real-world surveillance applications such as intelligent traffic, smart cities and so on. The objective is to locate and identify a specific pedestrian within a gallery of scene images, based on the query image. The person search task involves the localization and re-identification of target pedestrians, and can be divided into two subtasks: pedestrian detection and person re-identification. With the rapid evolution of technology, various deep learning-based person search methods have been proposed in recent years. Despite these advancements, there remains a lack of comprehensive overview over the research developments in person search. Therefore, we present an in-depth analysis within the domain. The survey begins with an exploration of fully supervised person search methods, covering the advances in framework and loss function. From the perspective of task processing, we examine the challenges such as accuracy of detection, discriminative re-identification features, relationship between detection and re-identification, background influence, and discuss how existing methods address these issues. Furthermore, we investigate weakly supervised person search, where limited annotations pose challenges. We discuss the annotation settings used in recent methods, which determine how the model will be trained. We also explore the role of clustering techniques and contextual information in addressing the annotation quality issue. The survey covers the evaluation metrics, benchmark datasets, and a performance evaluation of person search methods. Finally, we discuss the potential future directions for research in the field of person search.
Multisource remote sensing (MRS) image matching can provide more accurate data support for a variety of remote sensing tasks, but the disparities in imaging characteristics among different sensors bring significant challenges for effective image matching. In this article, we propose an image matching method based on the contrastive network with similarity statistics weighting for MRS registration. First, we proposed a sampling strategy that selects hard negatives based on similarity statistics measures. These hard samples with more valuable information can promote the abundance of multimodal consistency information. Subsequently, the samples are encoded with a self-attention contrastive network for semantic feature embedding, enabling the learning of effective representations in the latent space. In addition, a similarity-weighted loss function is designed to guide the training process, aiming to improve the discrimination ability of the network. Experimental results demonstrate that our proposed method can derive accurate and robust results in MRS matching, exhibiting superior matching performance with respect to other methods.
Image enhancement technology plays an important role in the practical application of detecting underwater tunnels. By improving image quality, it enhances the accuracy and reliability of detection results. However, in deep-sea scenes, the problem of image color distortion and detail loss is common. To address this issue, a real-time underwater image enhancement algorithm based on weighted multi-branch correction has been proposed in this paper. The algorithm uses a combination of multi-scale and multi-branch techniques to process the image. The multi-branch correction module has been improved by increasing the weighting and downsampling strategy, which not only improves the processing speed but also ensures the processing effect. In the air-frequency domain interactive processing module, zero padding and window functions have been introduced to reduce the ringing effect and block effect. In addition, the smart grid system provides stable power support for underwater tunnel search and rescue, which ensures the safety of the search and rescue environment and also provides a good experimental lighting environment for the image enhancement algorithm. At the same time, the model has been trained and validated using self-constructed real deep-sea datasets. The existing algorithms and the algorithm proposed in this paper have been compared and analyzed comprehensively. The experimental results show that the algorithm proposed in this paper is better than the existing algorithms in terms of processing effect, and the processing speed is improved by about 12.7%.
One of the successful strategies developed for studying the gene function of aphids is to silence aphid gene expression by plant-mediated RNA interference (RNAi). In this study, we analyzed the expression patterns and the biological functions of genes related to chitin metabolism ApCht7 and ApCht10 by using gene-specific plant-mediated RNAi in the green pea aphid, Acyrthosiphon pisum (Ap). The RT-qPCR results demonstrated that the RNAi-mediated silencing of these ApChts suppressed the expression of most genes involved in the chitin degradation pathway, but enhanced the expressions of ApHK, ApGFAT, ApPGM and ApCHS, indicating that the RNAi of ApChts could affect the expression of genes related to chitin metabolism and regulate the chitin metabolism. Furthermore, it resulted in a decrease in aphid body weight of aphids, with mortality rates ranging from 3.3 to 26.1
The convolution neural network is still the main tool for extracting the image features and the motion features for most of the optical flow models. The convolution neural networks cannot model the long-range dependencies, and more details are lost in deeper layers. All the deficiencies in the extracted features affect the estimated flow. Therefore, in this work, we concentrated on optimizing the convolution neural network in both the encoder and decoder parts to improve the image and motion features. To enhance the image features, we utilize the involution to provide rich features and model the long-range dependencies. In addition, we propose a Multi-Scale-Interaction module which utilizes the self-attention to make an interaction between the feature scales to avoid detail loss. Additionally, we propose a Motion-Features-Optimization block that utilizes the deformable convolution to enhance the motion features. Our model achieves the state-of-the-art performance on Sintel and KITTI 2015 benchmarks.
Spinal microglia and astrocytes are both involved in neuropathic and inflammatory pain, which may display sexual dimorphism. Here, we demonstrate that the sustained activation of spinal astrocytes and astrocyte-derived interleukin (IL)-17A promotes the progression of mouse bone cancer pain without sex differences. Chemogenetic or pharmacological inhibition of spinal astrocytes effectively ameliorates bone cancer-induced pain-like behaviors. In contrast, chemogenetic or optogenetic activation of spinal astrocytes triggers pain hypersensitivity, implying that bone cancer-induced astrocytic activation is involved in the development of bone cancer pain. IL-17A expression predominantly in spinal astrocytes, whereas its receptor IL-17 receptor A (IL-17RA) was mainly detected in neurons expressing VGLUT2 and PAX2, and a few in astrocytes expressing GFAP. Specific knockdown of IL-17A in spinal astrocytes blocked and delayed the development of bone cancer pain. IL-17A overexpression in spinal astrocytes directly induced thermal hyperalgesia and mechanical allodynia, which could be rescued by CaMKIIα inhibitor. Moreover, selective knockdown IL-17RA in spinal Vglut2 + or Vgat +neurons, but not in astrocytes, significantly blocked the bone cancer-induced hyperalgesia. Together, our findings provide evidence for the crucial role of sex-independent astrocytic signaling in bone cancer pain. Targeting spinal astrocytes and IL-17A/IL-17RA-CaMKIIα signaling may offer new gender-inclusive therapeutic strategies for managing bone cancer pain.
Feature representation is a crucial issue in multimodal image registration. The handcrafted features extracted by traditional methods are highly sensitive to nonlinear radiation differences, while supervised learning methods are limited by deficient labeled samples in the remote sensing field. Therefore, this article proposes a consistency feature representation learning method for multimodal image registration, which involves mapping data into a common feature space to realize the accurate alignment of remote sensing images. First, a contrastive network with a spatial attention mechanism is driven to enhance the capability to highlight high-level features of images. Second, a positive sample augmentation strategy is implemented with contrastive loss, which helps the model learn the inherent features better, and imposes constraints on the sample similarity to optimize the feature projection. Finally, a multimodal image registration framework is proposed to enhance the stability of feature matching. The proposed framework achieves accurate feature extraction and consistency feature description for multimodal images, ensuring robustness against nonlinear radiometric differences. Experimental results demonstrate that the proposed method obtains more reliable registration results on the SEN1-2 dataset. Furthermore, the proposed algorithm achieves superior performance on data from other modalities, indicating strong generalization ability.
Person search is a computer vision task that aims to locate and re-identify specific pedestrians in images captured by non-overlapping cameras. However, the identity annotation in person search is labor-intensive, especially as the amount of data increases. Therefore, more and more studies consider training person search models using weakly supervised learning with only location annotations. The context information is useful to improve feature representations in the absence of pedestrian identity as supervision. Existing weakly supervised person search methods focus on logic-driven contexts while ignoring feature contexts. In this paper, we propose a hybrid deep network for weakly supervised person search. The hybrid architecture consists of a Transformer-based feature extraction network and a fully convolution-based region recognition head network. The purpose is to enable the model to learn feature contexts at different levels. In our network, hierarchical vision Transformers are used to extract features in order to obtain discriminative representations of scene images. The context-enhanced head network is designed to integrate different features for candidate pedestrians. In addition, a pedestrian proposal network is proposed to improve the quality of predicted proposals. Experiments are conducted on the CUHK-SYSU and the PRW benchmarks to evaluate the effectiveness of the proposed method.
The continuous mutation of the novel coronavirus necessitates effective urban public health safety measures. In public spaces, enforcing social distancing is crucial to prevent droplet transmission of the virus. Existing video detection methods are limited by blind spots, especially in densely populated areas. This paper introduces an innovative millimeter wave radar method for effective social distancing control. Unlike traditional radar technology, our approach employs energy accumulation and Doppler compensation techniques to reduce preprocessing interference notably and accurately differentiate between human and non-human entities. We have fine-tuned the sensor configuration through distance feedback, thereby enhancing detection resolution. Multiple micro-Doppler images at various spatial scales are acquired and concatenated to construct a three-dimensional data cube. This serves as input to our network model to achieve higher detection accuracy. Our experiments demonstrate that this method surpasses mainstream video and radar detection solutions, offering broader coverage, increased stability, and the ability to discern distances to humans with a resolution as acceptable as 10 centimeters.
The comprehensive utilization of multi-sensor data can provide comprehensive support for Internet of Things tech-nology. Modality difference is one of the difficulties in multi-sensor image registration tasks. As a commonly used image-to-image translation approach, the generative adversarial network can effectively alleviate the problem of heterogeneous image registration. However, the quality of the generated image greatly limits the accuracy of the registration. In this paper, we proposed a structure consistency generative adversarial network(SCGAN), which enhances the similarity of structures by constraining the consistency of information in the neighborhood. Introducing a coordinate attention mechanism improves the spatial accuracy of location information, and multiscale frequency domain con-sistency loss helps capture structure features essential for image registration. Compared to various metrics, experimental results verify the proposed method's effectiveness in image quality and registration accuracy.
Common feature representation in optical and synthetic aperture radar (SAR) image registration is one of the most challenging tasks due to the significant geometric and radiometric differences. This letter proposed a Prominent structure-guided feature (PSGF) representation for SAR and optical image registration. Firstly, the prominent structure of the image is highlighted based on windowed inherent variations, which is conducive to identifying more accurate and reliable corresponding points. Secondly, the maximum response index filter banks are proposed to extract structure features with multi-orientation convolution results. Then the structure feature-guided representation generated from this filtering map is quantized in histograms. Finally, the descriptor with radiation invariance is employed for feature matching, enabling automatic image registration with high accuracy. Comparative analysis with state-of-the-art methods on diverse terrain data demonstrates the superiority of the proposed PSGF method for SAR and optical image registration.
In this paper, we investigate deep-learning-based image inpainting techniques for emergency remote sensing mapping. Image inpainting can generate fabricated targets to conceal real-world private structures and ensure informational privacy. However, casual inpainting outputs may seem incongruous within original contexts. In addition, the residuals of original targets may persist in the hiding results. A Residual Attention Target-Hiding (RATH) model has been proposed to address these limitations for remote sensing target hiding. The RATH model introduces the residual attention mechanism to replace gated convolutions, thereby reducing parameters, mitigating gradient issues, and learning the distribution of targets present in the original images. Furthermore, this paper modifies the fusion module in the contextual attention layer to enlarge the fusion patch size. We extend the edge-guided function to preserve the original target information and confound viewers. Ablation studies on an open dataset proved the efficiency of RATH for image inpainting and target hiding. RATH had the highest similarity, with a 90.44% structural similarity index metric (SSIM), for edge-guided target hiding. The training parameters had 1M fewer values than gated convolution (Gated Conv). Finally, we present two automated target-hiding techniques that integrate semantic segmentation with direct target hiding or edge-guided synthesis for remote sensing mapping applications.
Multi-source image registration has often suffered from great radiation and geometric differences. Specifically, grayscale and texture from similar landforms in different source images often show significantly different visual features, and these differences disturb the corresponding point extraction in the following image registration process. Considering that edges between heterogeneous images can provide homogeneous information and more consistent features can be extracted based on image edges, an edge consistency radiation-change insensitive feature transform (EC-RIFT) method is proposed in this paper. Firstly, the noise and texture interference are reduced by preprocessing according to the image characteristics. Secondly, image edges are extracted based on phase congruency, and an orthogonal Log-Gabor filter is performed to replace the global algorithm. Finally, the descriptors are built with logarithmic partition of the feature point neighborhood, which improves the robustness of the descriptors. Comparative experiments on datasets containing multi-source remote sensing image pairs show that the proposed EC-RIFT method outperforms other registration methods in terms of precision and effectiveness.
To clarify the effects of different temperatures and parental reproductive ages on the offspring color morph differentiation,the newly-born first-instar nymphs of the red and green pea aphid(Acyrthosiphon pisum)on the 2nd,7th,and 12th parental reproductive age from F0-F5 were continuously raised for five generations at three tem-peratures(15,22,and 26 ℃).Results showed that temperature and parental reproductive age had a significant effect on the offspring color morph differentiation.Red color morph A.pisum tended to produce green color morph offsprings at 15 ℃,while green color morph A.pisum tended to produce red color morph offsprings at 26 ℃.The proportion of non-parental color morph offspring produced by red and green parents at 22 ℃ was significantly lower than that at 15 and 26 ℃(P<0.05).Furthermore,different parental reproductive age also had a certain effect on the offspring color morph differentiation.The parental reproductive age affected the proportion of color differentiation and the time of color differentiation in the offsprings of pea aphid(P<0.05),indicating that the onset of color dif-ferentiation in the offsprings of red and green pea aphid was related to different parental reproductive ages at the same temperature.
Wolbachia and Rickettsia are bacterial endosymbionts that can induce a number of reproductive abnormalities in their arthropod hosts. We screened and established the co-infection of Wolbachia and Rickettsia in Bemisia tabaci and compared the spatial and temporal distribution of Wolbachia and Rickettsia in eggs (3–120 h after spawning), nymphs, and adults of B. tabaci by qPCR quantification and fluorescent in situ hybridization (FISH). The results show that the titer of Wolbachia and Rickettsia in the 3–120 h old eggs showed a “w” patterned fluctuation, while the titers of Wolbachia and Rickettsia had a “descending–ascending descending–ascending” change process. The titers of Rickettsia and Wolbachia nymphal and the adult life stages of Asia II1 B. tabaci generally increased with the development of whiteflies. However, the location of Wolbachia and Rickettsia in the egg changed from egg stalk to egg base, and then from egg base to egg posterior, and finally back to the middle of the egg. These results will provide basic information on the quantity and localization of Wolbachia and Rickettsia within different life stages of B. tabaci. These findings help to understand the dynamics of the vertical transmission of symbiotic bacteria.
The utilization of LiDAR point cloud in various applications, such as autonomous driving and intelligent transportation, has gained significant attention and become a research hotspot. However, existing 3D object detection methods face two main challenges. (1) random or farthest point sampling can lead to a loss of foreground points and missed detections, particularly for small objects. (2) the sparsity of LiDAR point clouds causes the object point clouds to be non-uniformly distributed. In this paper, we present a novel method for 3D object detection. Our approach is based on the PV-RCNN network architecture and incorporates Centroid-aware (CA) Sampling and Local Attention Feature Encoding techniques. The voxel branch in feature extraction uses the Focals Conv module to predict the importance of voxel features at different locations and enhances the representation of local spatial structure by attention convolution. The key point branch uses CA Sampling to obtain key points. We calculate mask weights based on the position relationship between each point and the 3D bounding boxes, and the K highest scoring points are retained as key points. We further improve the local feature expression capability by aggregating each key point with the original point cloud and voxel features in the neighborhood using the PointNet++ network. Finally, we combine the collocated features and feed them into the classification and regression network of PV-RCNN to complete the detection process. Our simulation experiments show that the detection performance of our proposed method outperforms the PV-RCNN on the KITTI dataset.