Objective Synthetic aperture radar (SAR), an active imaging sensor, is pivotal in military reconnaissance, maritime surveillance, and disaster assessment due to its all-weather and all-day imaging capability. However, unlike optical images, SAR imagery is inherently affected by severe speckle noise and exhibits significant variations in target scale, orientation, and background complexity. These factors pose major challenges to accurate target detection. Traditional approaches, such as threshold-based statistical detectors and template-matching methods, heavily rely on prior knowledge and handcrafted features, resulting in limited robustness and adaptability in complex environments. Although deep learning-based detectors, including both two-stage and single-stage frameworks, have achieved remarkable success in optical image detection, their performance degrades when applied directly to SAR imagery. This is primarily due to the lack of rich texture and color cues and the presence of strong noise interference. Recent studies indicate that frequency-domain features contain abundant structural and scale-dependent information, which can effectively complement spatial-domain representations by highlighting the distinct energy distribution between targets and background. Nevertheless, most existing frequency-domain methods rely on fixed-frequency component extraction, overlooking the spectral diversity among targets of different scales and structures. To address these limitations, this study aims to develop a frequency-domain enhanced SAR target detection framework that adaptively models spectral characteristics and achieves robust multi-scale feature fusion, thereby improving detection accuracy and generalization under complex SAR imaging conditions. Methods This paper proposes a dynamic frequency-aware feature enhancement network for SAR target detection. The core objective is to fully exploit the complementary advantages of spatial and frequency-domain features to enhance the robustness and accuracy of SAR target detection. First, a frequency-aware feature modulation (FAFM) module is introduced to replace conventional convolutional operations that rely solely on spatial-domain modeling. Specifically, the FAFM module transforms input feature maps into the frequency domain, decomposing them into distinct spectral components to separately model low-frequency global structure and high-frequency local details. A region-guided weighting mechanism is then applied to adaptively adjust the response intensity of different frequency bands, establishing a correspondence between spatial and spectral domains. This enables the network to suppress high-frequency noise while enhancing edge and texture representations. The modulated spectral features are subsequently remapped back to the spatial domain, yielding a more stable and structure-aware representation for downstream multi-scale fusion. In the multi-level feature fusion stage, to address large target scale variations and the loss of fine details for small objects, a selective bidirectional aggregation (SBA) network is proposed. This network constructs two complementary information flows: a bottom-up semantic flow and a top-down boundary flow. These flows facilitate bidirectional interaction between shallow features, which provide spatial structure and boundary cues, and deep features, which offer high-level semantic context. Through an attention-guided recalibration mechanism, the network dynamically balances the importance of features at different levels, mitigating redundancy and alignment inconsistencies during fusion. Consequently, the SBA module achieves fine-grained feature integration, ensuring the fused representation retains detailed boundary information while maintaining strong semantic expressiveness. Results and Discussions Experimental results demonstrate that the proposed method significantly outperforms existing detectors in terms of both accuracy and robustness. As shown in Table 3, on the RSAR dataset, our method achieves the highest precision (84.10 % ), recall (77.20 % ), and mAP@0.5 (79.60 % ), outperforming YOLOv11-OBB by 3.4 percentage points in mAP@0.5. The F1-score reaches 80.52 degrees o, indicating a well-balanced trade-off. On the RSDD-SAR dataset, which features more complex backgrounds and greater scale variations, our method maintains a clear advantage, with an mAP@0.5 of 95.30 % and an mAP of 55.90%, surpassing YOLOv11-OBB by 2.4 percentage points in mAP. Favorable results are also obtained on the SSDD+ dataset. These results collectively confirm the strong generalization capability of the proposed network across diverse SAR scenarios. Qualitatively, visualization results reveal that the FAFM module effectively enhances target contours and suppresses background clutter, while the SBA network enables more precise bounding box localization and reduces false detections. Ablation studies further validate the complementary effects of the two modules: FAFM contributes to noise suppression and texture recovery, whereas SBA facilitates multi-scale semantic alignment and structural consistency. Their combination effectively addresses key challenges in SAR target detection. Conclusions This paper addresses key challenges in SAR image target detection, such as severe speckle noise interference, significant target scale variations, and insufficient feature representation, by proposing a detection method based on frequency-domain dynamic perception and feature enhancement. The designed FAFM effectively utilizes the complementary nature of frequency-domain and spatial features, improving the model's perception of target edges and texture details. Concurrently, the SBA network enables cross-level bidirectional feature fusion, enhancing the model's ability to represent and discriminate multi-scale targets. Experimental results on the publicly available RSAR, RSDD-SAR, and SSDD+ datasets demonstrate that the proposed method outperforms existing mainstream approaches across multiple metrics, including precision, recall, mAP@0.5, and mAP. It exhibits particularly strong robustness and generalization capabilities in complex backgrounds and multi-scale target scenarios, fully validating its effectiveness and advancement for SAR target detection tasks. Future work will involve testing the model in larger-scale and more diverse real-world application scenarios to further verify its practicality. Additionally, exploring integration with Transformer architectures or generative modeling methods could be pursued to more fully leverage multi-level information from both frequency and spatial domains.
In recent years, the increase of multimodal image data has offered a broader prospect for multimodal semantic segmentation. However, the data heterogeneity between different modalities make it difficult to leverage complementary information and create semantic understanding deviations, which limits the fusion quality and segmentation accuracy. To overcome these challenges, we propose a hybrid attention driven CNN-Mamba multimodal fusion network (HACMNet) for semantic segmentation. It aims to fully exploit the strengths of optical images in texture and semantic representation, along with the complementary structural and elevation information from the digital surface model (DSM). This enables the effective extraction and combination of global and local complementary information to achieve higher accuracy and robustness in semantic segmentation. Specifically, we propose a progressive cross-modal feature interaction (PCMFI) mechanism in the encoder. It integrates the fine-grained textures and semantic information of optical images with the structural boundaries and spatial information of DSM, thereby facilitating more precise cross-modal feature interaction. Second, we design an adaptive dual-stream Mamba cross-modal fusion (ADMCF) module, which leverages a learnable variable mechanism to deeply represent global semantic and spatial structural information. This enhances deep semantic feature interaction and improves the ability of the model to distinguish complex land cover categories. Together, these modules progressively refine cross-modal cues and strengthen semantic interactions, enabling more coherent and discriminative multimodal fusion. Finally, we introduce a global-local feature decoder to effectively integrate the global and local information from the fused multimodal features. It preserves the structural integrity of target objects while enhancing edge detail representation, thus enhancing segmentation results. Through rigorous testing on standard datasets like ISPRS Vaihingen and Potsdam, the proposed HACMNet demonstrates advantages over prevailing methods in multimodal remote sensing analysis, particularly on challenging object classes.
3D Gaussian Splatting (3DGS) has achieved significant progress in the field of novel view synthesis. However, there are challenges associated with using spherical harmonics to learn scenes that involve specular reflections, due to the presence of high-frequency details in such scenes. To address above problem, we propose an image-based view-dependent appearance model to jointly extracts both high-and low-frequency information from the scene, to more efficiently represent the appearance field of 3D Gaussians. Specifically, by statistically assessing the dot product between the view direction and the normal at the respective Gaussian within the image, we develop a view-dependent appearance module that calculates the variances of these dot products; the module is able to adaptively assign weights to both specular and diffuse reflection colors. We propose a normal-guided specular reflections module to extract view-dependent high-frequency information, which effectively filters out specular colors by using a threshold on the variance of the dot product between the view direction and the normal. In addition, to extract low-frequency information, we design an image-based diffuse reflections module to compute the diffuse reflection colors and preserve full-frequency information. Experimental results show that our method outperforms the baseline in both quantitative and qualitative results, significantly enhancing the ability of 3DGS in processing specular reflection scenes.
We examined whether two variables-pattern recognition (i.e., predicting which chute a ball would come down), and exposure to repeated behaviors in the home environment (e.g., drinking from a cup or folding washing multiple times)-predicted children's later theory of mind (ToM). These two variables have been hypothesized to assist ToM development because they help children learn to recognize patterns in behavior, and therefore predict future behavior. Because this behavior is underpinned by mental states (e.g., people act in particular ways due to their desires and beliefs), repeated behaviors accompanied by pattern recognition should also help children to eventually acquire a ToM. We studied these questions using a longitudinal study of 56 children (primarily of European ethnicity, 22 girls, 34 boys) at four timepoints (21, 24, 27 and 30 months). Replicating previous research, we found that repeated behaviors were very frequent (M = 143.4 per hour). Crucially, we also found that both repeated behaviors and pattern recognition were unique predictors of children's subsequent ToM (along with their earlier language ability), consistent with the idea that pattern recognition and repeated behaviors facilitate children's subsequent ToM.
In recent years, contrastive learning has made significant progress in DeepFake detection. However, existing methods emphasize class granularity, and it is difficult to distinguish between the real instance and its forgery counterparts effectively. Furthermore, the diversity of forgery cues produced by different manipulation methods cannot be effectively clustered by class granularity alone. Thus, the model’s generalization capability is limited. To tackle the above problems, a Dual-Granularity Contrastive Learning (DGCL) for DeepFake detection is proposed in this paper. Specifically, Class Granularity Contrastive Learning (CGCL) and Instance Granularity Contrastive Learning (IGCL) are designed. Firstly, for semantic aggregation at the class level, CGCL incorporates the class prototype, which encourages anchor approaches to the prototype of the positive class, thereby pulling the intra-class features closer. Secondly, for distinguishing between real and fake instances, Real Instance Granularity Contrastive Learning (RIGCL) and Fake Instance Granularity Contrastive Learning (FIGCL) are proposed based on the instance characteristics. RIGCL endeavors to distinguish fake instances from original real instances by expanding the differentiation in the feature space. Meanwhile, FIGCL extracts consistent forgery features from various manipulation methods using cosine similarity constraints. Finally, the superiority and generalizability of DGCL are validated by the experimental results on CELEBDF, DFD, and DFDC datasets.
Open-vocabulary aerial object detection (OVAD) aims to detect objects outside the training sets, which typically involves distilling knowledge from pre-trained vision-language models, e.g., RemoteCLIP, to inherit its generalizable recognition ability and thereby enabling the models to detect novel categories. However, existing distillation methods lacks a customized perception incentive mechanism, leading to a disconnect between perception with discrimination knowledge and poor inductive generalization. To this end, we propose a Pi-Noise guided progressive prior accumulation student-teacher distillation network (P3AD-Net), which utilizes the structured perturbation of noise distribution to achieve tighter localization-classification synergy. Concretely, we design an adaptive guided noise generator to continuously accumulate universal patterns of object location distributions and refined semantic knowledge. These prior patterns not only provide implicit cues to simplify the detection task, but also establish bidirectional knowledge flow, forming a closed-loop optimization of localization and classification. Then, we introduce a semantic-aware dual-alignment reprojection module (SDAR), which employs cross-modal attention to achieve dual visual-semantic alignment and enhances feature discriminability. This module effectively mines the latent localization-aware information embedded in external teachers, thereby improving the model's recognition sensitivity to unseen categories. Furthermore, to enhance data diversity and increase the number of pseudo-samples, we propose a pixel-level adaptive CutMix strategy. This approach enrichs training scenarios by performing pixel thresholding after channel separation. Extensive experiments on multiple remote sensing object detection benchmarks demonstrate that the proposed method achieves highly competitive performance compared with recent state-of-the-art methods, while maintaining a simple training pipeline without relying on additional classification or caption datasets.
Nowadays, face forgery poses a significant threat to societal security, making the development of effective countermeasures imperative. Though most existing methods adopt neural networks to automatically extract discriminative features for forgery detection and have achieved promising results, significant challenges remain. Namely, when detecting forgery faces generated by unseen forgery methods, the detection performance degrades significantly, indicating poor generalization capability. To address such limitation, a novel deep supervised anomaly detection for generalized face forgery detection (DAGFD) is proposed in this paper. Specifically, the artifact map detector optimized by triplet focal loss and metric-softmax loss is first used to locate the forgery regions and obtain artifact maps. Next, forgery detection is reformulated from the supervised anomaly detection perspective, and the artifact map score is calculated to detect forgery videos. Furthermore, mean square error (MSE) loss is used to minimize the artifact map score of real samples while increase the one of forgery samples to generalize well to unseen forgery methods. Also, circle loss is used for auxiliary classifier to learn more discriminative artifact features. Finally, the experimental results demonstrate that the proposed method’s detection accuracy is better than other state-of-the-art methods.
Parent mental state talk (MST) is an important contributor for children's theory-of-mind development, with cognitive MST for Western children being more advanced, often being used with older children relative to talk about desires or emotions. We investigated the contexts in which cognitive MST is most prevalent, as well as the conditions that prompted mothers to provide scaffolding of MST with toddlers. We tested 89 mother-toddler dyads, examining mothers' MST and non-MST with familiar and unfamiliar children during a Picture Describing Task. Of particular interest was whether the age of the toddler, the toddler's preestablished mental state (MS) or non-MS vocabulary, spontaneous conversation dynamics, or the familiarity of the child influenced the mothers' MST. We found that mothers used significantly more cognitive talk in interactions with their own child than with an unfamiliar child. These findings persisted even after covarying out the child's age and MT vocabulary, as well as the mother's MST and non-MST during the picture task. We argue that mothers possessed a better understanding of their own child's MS vocabulary because of their rich relationship history, which resulted in greater use of more advanced cognitive MS discourse with the child. Further, our cross-lagged correlational analysis examining the first half and second half of the Picture Describing Task demonstrated that more child MS talk cued mothers to use more advanced cognitive talk. Overall, the findings of the present study underscore the importance of a relationship history between adults and children for the sensitive scaffolding of children's MT understanding. (PsycInfo Database Record (c) 2025 APA, all rights reserved).
In the past few years, feature fusion-based violence detection has made remarkable progress. However, existing detection methods primarily focus on temporal feature analysis, which may result in an insufficient representation of the subtle variations inherent in violent behaviors, ultimately compromising detection performance. To overcome this limitation, this study introduces a Wavelet-Based Time–Frequency Feature Fusion (WTFF) method. Specifically, the Wavelet-Dilated Separable Convolution Module (WDCM) and the Time–Frequency Feature Fusion (TFFF) Network are designed. Firstly, the input video data is utilized by the WDCM to extract and process frequency-domain features, enabling the model to capture fine-grained behavioral details often overlooked in temporal analysis. Secondly, the TFFF fuses the temporal and frequency-domain features, thereby improving the model’s ability to discriminate violent events. Ultimately, the effectiveness and superiority of the proposed approach are demonstrated by experimental results on UCF-Crime, XD-Violence, and ShanghaiTech datasets, achieving 85.87% (AUC), 84.77% (AP), and 97.91% (AUC), respectively.
Accurate detection of small objects plays an important role in the application of Autonomous aerial vehicles (AAV). However, current works mainly extract comprehensive features from unimodal images, which can obtain very limited distinguishable features for objects, especially those with small sizes. To address this issue, we propose a dynamic cascade cross-modal coassisted network, which integrates multimodal images fusion and fine-grained feature learning to generate powerful object semantic representations. Specifically, we design a multimodal high-order interaction module to achieve collaborative interaction of spatial details and channel dependencies between modalities, thereby enhancing object discrimination. To preserve multimodal fine-grained details, we devise a scale-adaptive dynamic feature prompt module, which dynamically motivates the backbone network to capture feature degradation clues. Meanwhile, to maintain the spatial correlation of multimodal cross-scale features and improve the quality of feature fusion, we derive a global collaborative enhancement module into the feature pyramid network for enhancing the detection accuracy across multiple scales. Extensive experimental results on multimodal datasets have shown that our method achieves favorable performance, surpassing other state-of-the-art methods.
Remote sensing change detection (RSCD) has become an essential tool in observing and analyzing geographical information. However, existing deep learning approaches dependent solely on visual modalities may encounter challenges in discerning subtle variations amidst noise interference. To overcome these issues, we propose a multiscale semantic-guided synergistic interaction network (MSSI-Net), which utilizes the advanced multimodal semantic representations for enhancing the capacity to perceive hierarchical changes. Specifically, we first devise a multiscale interaction module (MIM) that leverages a multiscale attention mechanism to guide the interaction between the coarse and fine stages of different visual features. The fine-grained visual features subsequently complement the semantic features through scale weight reassignment to enhance the discriminative capability of vision-language features. Furthermore, driven by the semantic-guided synergistic interaction mechanism, our developed cross-modal feature fusion module (CFFM) exploits both homogeneous and heterogeneous features among modalities. This ensures that the generated vision-language features are semantically representative. Finally, we formulate a manifold differential perception head (MDPH) to optimize the detection of changes by efficiently fusing diverse differential feature representations, achieving comprehensive performance enhancement. Extensive experiments conducted on four benchmark datasets (LEVIR-CD, CDD, SYSU-CD, and WHU-CD) indicate that the designed MSSI-Net achieves state-of-the-art performance compared with existing methods.
Previous research reveals that older adults have relatively intact well-being when excluded by others as compared to young adults. This observation can be attributed to two plausible explanations: Either older adults are unaware of their exclusion, thereby shielding their well-being from its impact, or they recognize the exclusion but respond to it rationally. We carried out two studies to compare young and older adults' awareness of and response to exclusion, and explored its potential mechanisms by assessing the explanatory roles of loneliness, general cognition, and rejection sensitivity. Study 1 measured young and older adults' loneliness, awareness of exclusion, and needs satisfaction after playing the Cyberball game, and Study 2 further examined other potential correlates including processing speed, working memory, and rejection sensitivity. Over the two studies, older adults were not worse at recognizing exclusion, and sometimes better than young adults. Older adults' awareness of exclusion predicted their responses to exclusion, whereas the same link was absent in younger adults. Despite older adults' relatively good performance, there were individual differences in recognizing exclusion; older adults with better general cognition and lower rejection sensitivity were particularly adept. In sum, older adults can be as aware of exclusion as young adults, but rather than reacting in an emotional way that is detached from reality, older adults are more likely to respond to it rationally based on the severity of exclusion they have perceived.
Sketch face synthesis aims to generate sketch images from photos. Recently, contrastive learning, which maps and aligns information across diverse modalities, has found extensive application in image translation. However, when applying traditional contrastive learning to sketch face synthesis, the random sampling strategy and the imbalance between positive and negative samples result in poor performance of synthesized sketch images regarding local details. To address the above challenges, we propose A Facial Structure Sampling Contrastive Learning Method for Sketch Facial Synthesis. Firstly, we propose a region-constrained sampling module that utilizes the distribution map of facial structure obtained by a dual-branch attention mechanism to segment the input photos into diverse regions, thereby providing regional constraints for sample selection. Subsequently, we propose a dynamic sampling strategy that dynamically adjusts the sampling frequency based on the feature density in the distribution map, thereby alleviating sample imbalance. Additionally, to diminish the background influence and enhance the delineation of character contours, we introduce the mask derived from the input photo as an additional input. Finally, to further enhance the quality of the synthesized sketch images, we introduce pixel-wise loss and perceptual loss. The CUFS dataset experiment demonstrates that our method generates high-quality sketch images, outperforming existing state-of-the-art methods in subjective and objective evaluations.
Recently, the rapid advancement of the unmanned aerial vehicle (UAV) remote sensing technology has positioned object detection in UAV imagery as a prominent research domain. However, object detection models designed for conventional imagery often fail to achieve satisfactory detection accuracy due to the challenges of varying object scales and a high proportion of dense small objects in UAV imagery. Based on this observation, we devise a multiscale frequency-aware and dual attention-guided feature fusion network (MFDAFF-Net) for UAV imagery object detection. MFDAFF-Net integrates spatial domain multiscale feature fusion with frequency domain information augmentation, effectively improving the detection accuracy of objects with varying scales in UAV imagery. Specifically, we construct a multiscale frequency-aware feature pyramid network as the neck of the model, which facilitates thorough top-down fusion of multiscale features through a meticulously designed feature fusion architecture. Then, we design a dual attention-guided adaptive feature fusion network (DAAFFN) as the specific feature fusion strategy. The DAAFFN effectively enhances and fully integrates multiscale features by leveraging spatial-channel collaborative attention and interscale feature interactions. Moreover, a wavelet-inspired frequency-aware module (WFM) is proposed to disentangle high-frequency object details from low-frequency backgrounds, eventually improving the detection performance for dense small objects. Comprehensive experimental evaluations conducted on the VisDrone2019 and UAVDT datasets demonstrate that MFDAFF-Net substantially outperforms existing state-of-the-art UAV imagery object detection methods.
Remote sensing change detection (RSCD) has achieved creditable success in recent years. However, the challenge of identifying changed objects with shape details persists in RSCD. In this letter, we proposed a cross-temporal interaction with difference refinement network (CTIDRNet) to solve interference-caused fake change and incomplete irregular change shape in RSCD tasks. Specifically, by combining cross-attention and self-attention to steer the temporal feature interaction of each input, we design a temporal feature attention (TFA) module to excavate the potential relation of change areas and suppress the unchanged object interference. Afterward, a deformable convolution is used to design a difference feature refinement (DFR) architecture to capture temporal difference information at diverse feature levels. At last, we proposed a multiscale-guided fusion (MGF) module to fuse pyramid features, thereby dealing with scaling changes. Experimental results on three datasets show that CTIDRNet can extract irregularly changed areas effectively, and the evaluation result outperforms other SOTA methods, with an improvement of 1.79%-19.82%, 2.9%-11.07%, and 0.97%-8.91% in terms of F1 for CDD, SYSU, and LEVIR datasets, respectively. The demo code of this work is publicly available at https://github.com/lucyjiong/CTIDR.
3D Gaussian Splatting (3DGS) achieves remarkable results in the field of surface reconstruction. However, when Gaussian normal vectors are aligned within the single-view projection plane, while the geometry appears reasonable in the current view, biases may emerge upon switching to nearby views. To address the distance and global matching challenges in multi-view scenes, we design multi-view normal and distance-guided Gaussian splatting. This method achieves geometric depth unification and high-accuracy reconstruction by constraining nearby depth maps and aligning 3D normals. Specifically, for the reconstruction of small indoor and outdoor scenes, we propose a multi-view distance reprojection regularization module that achieves multi-view Gaussian alignment by computing the distance loss between two nearby views and the same Gaussian surface. Additionally, we develop a multi-view normal enhancement module, which ensures consistency across views by matching the normals of pixel points in nearby views and calculating the loss. Extensive experimental results demonstrate that our method outperforms the baseline in both quantitative and qualitative evaluations, significantly enhancing the surface reconstruction capability of 3DGS.
Sketch face synthesis aims to create detailed and realistic sketch images from optical photos. Recently, diffusion models have effectively addressed challenges like over-smoothing and mode collapse by simulating the distribution of multi-channel input data, but due to limitations in capturing low-frequency information, the generated images lack authenticity in textural details. To address the concerns raised in the appeal, we propose the Conditional Denoising Diffusion Probability Model (DDPM) with The T-UNet. To improve the utilization of low-frequency information, we propose a T-UNet module. This module maintains UNet’s ability to capture high-frequency details while integrating the Transformer’s ability to process low-frequency information. As a result, it enhances the overall quality of the synthesized sketch images. Experiments on the CUHK dataset demonstrate that our approach generates sketches of excellent quality, surpassing existing methods in visual performance.
Objective High-resolution remote sensing imagery presents complex scene configurations,diverse semantic associations,and significant object scale variations,often resulting in overlapping feature distributions across categories in the latent space.These ambiguities hinder the model's ability to capture intrinsic associations between textual semantics and visual representations,reducing retrieval accuracy in image-text retrieval tasks.This study aims to address these challenges by investigating object-level attention mechanisms and cross-modal feature alignment strategies.By dynamically allocating attention weights to salient object features and optimizing image-text feature alignment,the proposed approach enables more precise extraction of semantic information and achieves high-quality cross-modal alignment,thereby improving retrieval accuracy in remote sensing image-text retrieval. Methods Building on the theoretical foundations above,this study proposes an Object Semantic and Dual-attention Perception Model(OSDPM)for remote sensing image-text retrieval.OSDPM first utilizes a pretrained CLIP model to extract features from remote sensing images and their associated textual descriptions.A Dual-Attention Perception Network(DAPN)is then developed to characterize both global contextual information and salient object regions in the imagery.DAPN adaptively enhances the representation of salient objects with large scale variations by dynamically attending to significant local regions and integrating attention across spatial and channel dimensions.To address cross-modal heterogeneity between image and text features,an Object Semantic-aware Feature Clustering Module(OSFCM)is introduced.OSFCM conducts statistical analysis of the frequency of semantic nouns associated with object categories in image-text pairs,extracting high-probability semantic priors for the corresponding images.These semantic cues are used to guide the clustering of image features that exhibit ambiguity in the cross-modal feature space,thereby reducing distribution overlap across object categories.This targeted clustering enables precise alignment between image and text features and improves retrieval performance in remote sensing image-text tasks. Results and Discussions The proposed OSDPM integrates spatial-channel attention and adaptive saliency mechanisms to capture multiscale object information within image features.It then leverages semantic priors from textual descriptions to guide cross-modal feature alignment,improving retrieval performance in remote sensing image-text tasks.Experiments on the RSICD and RSITMD benchmark datasets show that OSDPM outperforms state-of-the-art methods by 9.01%and 8.83%,respectively(Table 1,Table 2).Comparative results for image-to-text and text-to-image retrieval(Fig.6,Fig.7)further confirm the superior retrieval accuracy achieved by the proposed approach.Feature heatmap visualizations(Fig.5)indicate that the DAPN effectively captures both global contextual features and local salient object regions,maintaining spatial semantic consistency between visual and textual representations.In addition,t-SNE visualizations across training stages demonstrate that OSFCM mitigates feature distribution overlap among object categories,thereby improving feature alignment accuracy.Ablation studies(Table 3)confirm that each module in the proposed network contributes to retrieval performance gains. Conclusions This study proposes a remote sensing image-text retrieval method,OSDPM,to address challenges in object representation and cross-modal semantic alignment caused by complex scenes,diverse semantics,and scale variation in high-resolution remote sensing images.OSDPM first employs a pretrained CLIP model to extract global contextual features from both images and corresponding text descriptions.It then introduces the DAPN to capture salient object features by dynamically attending to significant local regions and adjusting attention across spatial and channel dimensions.Furthermore,the model incorporates an OSFCM,which extracts prior semantic information through frequency analysis of object category terms and uses these priors to guide the clustering of ambiguous image features in the embedding space.This strategy reduces semantic misalignment and facilitates accurate cross-modal mapping between image and text features.Experiments on the RSICD and RSITMD benchmark datasets confirm that OSDPM outperforms existing methods,demonstrating improved accuracy and robustness in remote sensing image-text retrieval.
The objective of sketch-based face identification is to match a target individual’s facial features from a collection of photographs using a sketched portrait as the search query. The method based on Contrastive Language-Image pre-training (CLIP) brings matched sketch and optical images closer together through learnable text tokens, therefore improving recognition performance. However, the existing sketch face datasets are relatively small. Training the CLIP with a large number of parameters on these datasets makes model overfit, leading to suboptimal performance. To address the above issue, we propose a sketch face recognition method based on Local-Gobal Adapter (LGAdapter). The method uses a transformer to capture facial details, and we segment the features through H windows, thus limiting the self-attention mechanism to localized windows, which better captures detailed information. In order to enhance the extraction of more comprehensive representational characteristics, we incorporate a streamlined Graph Convolutional Network (GCN) module subsequent to the local module. In addition, we propose a two-stage training strategy to fully utilize the LGAdapter to obtain more accurate visual features. In the first phase, only text tokens are optimized; in the second phase, only the LGAdapter is fine-tuned. The outcomes of our experiments reveal that our approach surpasses existing leading-edge methods across all three widely recognized datasets, namely UoM-SFGS, CUFSF, and PRIP-VSCG.
Semi-supervised object detection aims to enhance object detectors by utilizing a large number of unlabeled images, which has gained increasing attention in natural scenes. However, when these methods are directly applied to scenes with tiny objects, they face the challenge of selecting pseudo-labels with high localization quality due to the minuscule and blurred characteristics of these objects. To address this issue, we propose a novel method called semi-supervised tiny object detection (STOD). Firstly, to enhance the localization accuracy of pseudo-labels, we design a dense IoU-aware head that evaluates the quality of bounding box localization by incorporating additional predicted overlap values. Secondly, to mine more potential pseudo-labels, we propose a GMM-based multi-threshold pseudo-labels mining module that dynamically generates multiple thresholds using classification scores and overlap values to classify bounding boxes into strong positive and weak positive pseudo-labels. Lastly, we design the localization-aware weighting loss to incorporate the localization quality of both positive and negative samples in order to enhance the accuracy of pseudo-label localization. The experimental results show that STOD achieves comparable performance when compared to both fully and semi-supervised methods. Notably, on the VisDrone-Partial benchmark, STOD achieves outstanding results by outperforming our baseline model with improvements of 3.1 mAP, 5.7 AP_50 , and 3.0 AP_75 .