The widespread application of advanced multispectral detectors in surveillance and reconnaissance poses a serious threat to military equipment and personal safety. However, achieving comprehensive and effective camouflage across the visible (VIS) to infrared spectra and specific laser wavelengths, while maintaining efficient thermal management, remains a significant challenge. Herein, we propose a wavelength-selective thermal regulator (WSTR) that achieves excellent camouflage performance across a wide spectrum, including VIS, midwave infrared (MWIR, ε3-5μm = 0.21), long-wave infrared (LWIR, ε8-14μm = 0.17), and laser wavelengths (ε1.06μm = 0.99, ε1.55μm = 0.94, and ε10.6μm = 0.91). The regulator emits a radiative intensity of 725.2 W/m2 at high temperatures and 95.7 W/m2 at relatively low temperatures. It achieves efficient thermal regulation through radiative heat dissipation via two nonatmospheric windows (ε5-8μm = 0.60, ε14-20μm = 0.63). Additionally, by varying the thickness of the ZnS film, a range of structural colors suitable for camouflage against different backgrounds can be achieved. Besides, this regulator employs a multilayer structural design that maintains stable spectral emissivity across a wide range of incident angles, thereby further enhancing the stability of its infrared camouflage performance. This study provides valuable insights and feasible solutions for the development of multispectral camouflage technology.
Remote sensing image change detection (CD) aims to identify changes occurring in bitemporal remote sensing image of the same area at different times. Due to the presence of targets with various scales in the CD scenes, existing methods have designed multiscale encoders to extract change information at different scales. However, these methods do not directly enhance or enrich high-frequency features, resulting in persistent blurry change boundaries. To address this issue, we propose a multiscale enhanced CD method based on detail supplementation, dubbed MEDS-CD. It designs a multiscale edge-guided adaptive filter to extract high-frequency information and compensate for multiscale features. Specifically, a multiscale transformer encoder is employed to extract multiscale features. To capture high-frequency details lost during the encoding process, a multiscale edge-guided adaptive filter is introduced to extract high-frequency details. Furthermore, to fully exploit the restored high-frequency information for change extraction and edge enhancement, we propose a multiscale convolution modulation module and a fuzzy weighting strategy. Finally, the above process is supervised by a hybrid loss function. The proposed method MEDS-CD is validated on three public datasets, LEVIR-CD, WHU-CD, and SYSU-CD. The experimental results show that MEDS-CD achieves better segmentation performance.
With the rapid development of multi-band infrared detection, traditional static camouflage struggles to achieve coordinated concealment and thermal management across visible, laser, and infrared bands. To overcome the limitations of single-band regulation and costly lithography, this study proposes a lithography-free seven-layer film structure (Cr/ZnS/VO2/Ge/MgF2/Ge/ZnSe), optimized via Finite-Difference Time-Domain (FDTD). Leveraging the metal-insulator transition (MIT) of VO2 near 68 degrees C, the structure enables dynamic emissivity control: in the insulating state, average emissivity drops to 0.233 in the mid-wave infrared (MWIR, 3-5 & micro;m) and 0.172 in the long-wave infrared (LWIR, 8-14 & micro;m), suppressing infrared signatures; in the metallic state, emissivity rises to 0.754 in the non-atmospheric window (NAW, 5-8 & micro;m) for efficient radiative cooling. By tuning the top ZnSe layer thickness, visible multicolor camouflage is achieved while maintaining low reflectivity at key laser wavelengths (1.06, 1.55, 10.6 & micro;m). The design remains robust up to 60 degrees incidence and demonstrates effective thermal management across a broad temperature range, significantly reducing target-background temperature differences. This fabrication-friendly and lithography-free structure offers a scalable and practical solution for multispectral camouflage and adaptive thermal regulation.
Radiative cooling (RC) and solar heating (SH) are gaining attention as sustainable thermal management technologies. Existing material designs are often static and cannot adapt to dynamic weather conditions. This study introduces a dual-mode Janus structure for efficient RC and SH, integrated with a thermoelectric generator to form a compact thermoelectric device for effective thermal management. The structure comprises the cooling side of Al2O3-doped polydimethylsiloxane composite film, the heating side of Ti/SiO2/Ti/SiO2/Ti/ZnS nanoengineered metafilm, and a thermoelectric device in the middle. The cooling side has a solar reflectivity of up to 92% and an emissivity of 96.6% in the atmospheric window (AW), achieving an average cooling effect of 6.7 degrees C during the day. The heating side has a low AW emissivity of 11.6% and a high solar absorbance of 80.8%, enabling an average temperature increase of 18.7 degrees C. The system can harness the temperature difference for thermoelectric power generation. When the RC and SH ends are facing upwards, the average temperature differences between the two ends are 2.7 degrees C and 7 degrees C, with corresponding average output powers of 55 mW & sdot;m- 2 and 0.39 W & sdot;m- 2. This structure offers references for thermal management technology and provides a new approach to address global energy challenges.
Remote sensing spatio-temporal fusion (STF) aims to integrate remote sensing images with complementary temporal and spatial resolutions to enhance image quality with high spatio-temporal resolutions. Conditional diffusion models have shown promising performance in STF by learning data probability distribution to achieve accurate fusion results. However, existing conditional diffusion models for STF directly use the limited image exemplars as conditional signals, limiting their ability to extract coherent features across images with extended temporal gaps and significant resolution differences. To address this challenge, this paper presents a spatio-temporal Dual Priors guided Diffusion framework for STF, named DPDiff. The DPDiff has three key modules including a Prior Extraction Network (PENet), a residual diffusion framework (RDF) and a Difference-Aware Module (DAM). First, a lightweight CNN-based PENet is designed to learn the spatio-temporal dual priors as conditional signals to control the diffusion process. The dual priors include a spatial mapping prior (SMP) between the coarse and fine images at the same time, and a temporal change prior (TCP) between the coarse and fine images at different times. The SMP and the TCP are reciprocal to ensure precise control of the diffusion process. Then, an RDF is adapted to learn distribution differences between high temporal resolution and high spatial resolution images, rather than directly reconstructing the entire image. Third, a DAM is introduced to align the CNN-predicted space with the ground truth space. Experimental results on four benchmark datasets — CIA, LGC, WaterBank, and the global E-SMILE — demonstrate that DPDiff achieves superior generation fidelity and quantitative performance across diverse regions and platforms, outperforming state-of-the-art methods. The source code is available at https://github.com/paul623/DPDiff.
The development of tunable multilayer thin-film systems, leveraging phase-change materials as the functional core, represents a pivotal advancement for augmenting infrared stealth capabilities and facilitating adaptive thermal management across diverse operational scenarios. Such systems hold substantial promise for both engineering applications and scientific inquiry. In this study, we designed and fabricated a Ge/ZnS/GST/Ag multilayer camouflage architecture, wherein the phase-transition characteristics of Ge2Sb2Te5 (GST) serve as the principal mechanism for dynamically modulating thermal emissivity within the mid-wave (3-5 & micro;m) and longwave (8-14 & micro;m) infrared atmospheric windows. Specifically, when GST is maintained in its amorphous state (a-GST), the film stack demonstrates remarkably low emissivity-averaging 0.0254 in the mid-infrared and 0.1264 in the long-infrared regimes-thereby enabling effective infrared stealth by minimizing radiative signature. Conversely, upon transition to the crystalline phase (c-GST) induced by external stimuli, the system exhibits a substantial enhancement in emissivity, with prominent peaks at 3.77 & micro;m and 10.84 & micro;m, enabling a switching contrast in average emissivity of 0.65 in the mid-wave infrared (MWIR) and 0.58 in the long-wave infrared (LWIR), thus promoting efficient radiative cooling. Simulation-based infrared thermal imaging assays validated consistent camouflage performance under varying thermal gradients, underscoring the structure's robustness and versatility for multiband stealth applications in complex environments. These findings underscore the potential of non-volatile, phase-change-driven designs for next-generation adaptive thermal regulation technologies.
Remote sensing spatiotemporal fusion (STF) aims to integrate remote sensing images with complementary temporal and spatial resolutions to enhance image quality with high spatiotemporal resolutions. Conditional diffusion models have shown promising performance in STF by learning the data probability distribution to achieve accurate fusion results. However, existing conditional diffusion models for STF directly use the limited image exemplars as conditional signals, limiting their ability to extract coherent features across images with extended temporal gaps and significant resolution differences. To address this challenge, this article presents a spatiotemporal dual priors guided diffusion (DPDiff) framework for STF. The DPDiff has three key modules including a prior extraction network (PENet), a residual diffusion framework (RDF), and a difference-aware module (DAM). First, a lightweight convolutional neural network (CNN)-based PENet is designed to learn the spatiotemporal dual priors as conditional signals to control the diffusion process. The dual priors include a spatial mapping prior (SMP) between the coarse and fine images at the same time, and a temporal change prior (TCP) between the coarse and fine images at different times. The SMP and the TCP are reciprocal to ensure precise control of the diffusion process. Then, an RDF is adapted to learn distribution differences between high temporal resolution and high spatial resolution images, rather than directly reconstructing the entire image. Third, a DAM is introduced to align the CNN-predicted space with the ground-truth (GT) space. The experimental results on four benchmark datasets-CIA, LGC, WaterBank, and the global E-SMILE-demonstrate that DPDiff achieves superior generation fidelity and quantitative performance across diverse regions and platforms, outperforming state-of-the-art methods. The source code is available at https://github.com/paul623/DPDiff
Titanium alloy Ti-6Al-4V (TC4) is extensively utilized in aerospace and high-temperature structural components due to its excellent mechanical properties. However, surface oxidation of Ti-6Al-4V during high-temperature processing significantly affects its thermal radiation characteristics. This study establishes a prediction model for the spectral emissivity of Ti-6Al-4V by combining the Lorentz-Drude dielectric response related to temperature and heating rate with thin-film interference effects. The prediction results of the model are verified by measuring the spectral emissivity under different heating rates using a self-developed experimental platform, yielding a total emissivity standard error of less than 3.0%. Meanwhile, the research results of spectral emissivity indicate that the heating rate significantly influences the occurrence and intensity of the oxide film interference phenomena. By correlating the oxide layer thickness derived from the spectral emissivity with the non-isothermal oxidation model, the oxidation activation energy is quantitatively determined, revealing an increasing trend with the increase in heating rate. This study proves the feasibility of using spectral emissivity to characterize non-isothermal oxidation kinetics and provides a quantitative tool for optimizing titanium alloy high-temperature processes.
Remote sensing spatiotemporal fusion (STF) aims at fusing high temporal-resolution images with high spatial-resolution images to obtain both high temporal and spatial resolution images. Temporal uncertainty remains a critical challenge, as existing methods rely on implicit temporal modeling and thus struggle to handle land-cover or phenological changes over long intervals, leading to texture errors or overly smooth transitions. To overcome this challenge, this paper presents a novel land-cover change inpainting network (LCCINet) for STF. The LCCINet consists of two core stages: the enhancing feature fusion stage and the feature reconstruction stage, which are designed to reconstruct the unchanged and changed regions, respectively. First, to distinguish unchanged and changed regions, we design a change detection module to compute the difference between the prediction and the reference time, generating a change mask. In the fusion stage, we propose a feature mask modulator (FMM) that leverages the mask to modulate the features of the high spatial-resolution image. Then the feature fusion encoder employs multiscale convolutions and spatiotemporal attention to focus the model on fusing features from unchanged areas. Finally, in the reconstruction stage, the image coarse inpainting module (ICM) and image fine inpainting module (IFM) use the mask and rich contextual cues to reconstruct the features of changed regions from coarse to fine. Extensive experiments on three STF benchmarks (CIA, LGC, and E-SMILE) demonstrate that LCCINet consistently outperforms existing approaches, reducing mean absolute error, mean absolute error, and structural similarity assessment measure by 4-11%, and increasing SSIM by 0.003-0.008 over the strongest competing methods. These results confirm the effectiveness of the proposed method in preserving spatial textures and spectral values.
Multispectral thermometry has been widely utilized in various fields due to its advantages of non-contact operation, rapid response, and broad applicability. Currently, optimal functions and intelligent optimization algorithms are combined to solve multispectral thermometry. Nonetheless, the search range of emissivity or the parameters of optimization algorithms may be selected improperly, thus, the accuracy of temperature inversion will be affected significantly. In the research, a novel unconstrained optimization approach of multispectral thermometry is proposed, in which the emissivity trends and the discrepancy of emissivity error at each wavelength caused by temperature error are combined. The performance of the approach is verified by the simulation results for six different models and the experimental results for the rocket nozzle. The maximum relative error is 0.67%, the computation time is less than 0.041 s, and this can be obtained by analyzing the inversion results.
Existing leading Co-saliency Detection (CoD) framework aims to segment the co-salient objects by learning the consensus visual representation of the foreground objects. However, despite different categories, some distractors may have similar appearance to the co-salient objects, such as Apples vs. Bananas have similar color and textures. This makes it challenging to distinguish the distractors only through learning the co-salient object appearance representations. To address this issue, we propose a joint appearance and shape co-representation learner for CoD, dubbed as ASCoD. The ASCoD is composed of a Co-Appearance learning Module (CoAM) and a Co-Shape learning Module (CoSM). The CoAM first learns a co-salient object appearance embedding that encodes the global cross-image and spatial context information. Then, this embedding is set as a co-appearance prototype, which guides the model to enhance the features to highlight the co-salient object regions. Afterwards, we design the CoSM that is a cross-attention module, among which the key and the value encode the shape information from a set of salient tokens dynamically selected by a Co-Shape Prototype generation Module (CSPM). Finally, through jointly optimizing the cascaded CoAM and CoSM, the optimal appearance and shape co-representations are achieved, marrying the merits of both appearance and shape co-representations that are not only robust to co-salient objects appearance variations, but also can well discriminate the co-salient objects from the distractors with similar appearance. Extensive evaluations on three challenging benchmarks including CoCA, CoSOD3k and CoSal2015, demonstrate superiority of the ASCoD to a variety of state-of-the-art CoD methods.
Remote sensing change detection aims to identify changes on the Earth's surface from remote sensing images acquired at different times. However, the identification of changed areas is often hindered by pseudochanges in similar objects, leading to inaccurate identification of change boundaries. To address this issue, we propose a novel network named boundary-guided semantic context network (BSCNet), which decouples features to improve the feature representation ability for changing objects. Specifically, we design a selective context fusion module that selectively fuses semantically rich features by computing the similarity between features from adjacent stages of the backbone network, thereby preventing detailed features from being overwhelmed by contextual information. In addition, to enhance the ability to perceive changes, we design a context fast aggregation module that leverages a pyramid structure to help the model simultaneously extract and fuse detailed and semantic information at different scales, enabling more accurate change detection. Finally, we design a boundary-guided feature fusion module to aggregate edge-level, texture-level, and semantic-level information, which enables the network to represent change regions more comprehensively and precisely. Experimental results on the WHU-CD, LEVIR-CD, and SYSU-CD datasets show that BSCNet achieves F1 scores of 94.92%, 92.19%, and 82.55%, respectively.
Visible-infrared multispectral camouflage is essential for evading advanced detection systems. However, achieving simultaneous low emissivity in atmospheric windows, high emissivity in non-atmospheric windows for radiative cooling, and visible color matching remains a significant challenge. Here, we propose a fully planar multilayer structure (TiO2/Si/Al/Si/W) that enables synergistic control across 0.38–14 µm. Through a hybrid global–local optimization strategy, the structure achieves low average emissivities of 0.177/0.223 in the mid-wave infrared (MWIR, 3–5 µm) and long-wave infrared (LWIR, 8–14 µm) bands, alongside high emissivities of 0.749/0.742 in the non-atmospheric windows (2.5–3 µm and 5–8 µm)—suppressing infrared detection while enabling efficient radiative cooling. Furthermore, the top-layer TiO2 provides continuous visible reflection control for color matching without compromising infrared performance. This structure maintains stable spectral selectivity at elevated temperatures and wide observation angles. This study provides a simple, scalable, and fabrication-friendly approach for advanced multispectral camouflage with integrated thermal management.
Existing remote sensing change detection methods often struggle to accurately capture the contours of complex change targets and subtle textural differences. This makes it difficult to effectively distinguish between the boundaries of change targets and the background. To address this challenge, we propose a novel method called spatial-frequency decoupling alignment encoding (SDA-Encoding), which is designed to fully leverage information from both the spatial and frequency domains. Specifically, we first use a Transformer encoder to extract bi-temporal features. Next, we apply wavelet transform to decouple these features into low-frequency and high-frequency components. In the multi-scale high-frequency interaction (MHI) module, we combine local spatial enhancement using spatial pyramid pooling with cross-scale dependency supplementation via the dual-domain alignment fusion (DAF) module. Meanwhile, in the position-aware low-frequency enhancement (PLE) module, spatial position sensitivity is restored using coordinate attention, and region-level contextual dependencies are captured through the selective fusion attention (SFA) module. Finally, the two frequency-domain branches are complementarily fused within the spatial domain to achieve unified detection of both fine-grained and structural changes. Experimental results on three benchmark datasets demonstrate the significant performance improvements of SDA-Encoding.
Remote sensing spatiotemporal fusion (STF) remains a formidable challenge in scenarios with complex land-cover changes. Existing methods generally use masking mechanisms to cover changed regions, aiming to mitigate alignment biases induced by change information. However, such indiscriminate masking strategies ignore transferable structural cues embedded in pseudo changes, leading to texture collapse and boundary degradation. To address these issues, we propose a change-aware (CA) remote sensing spatiotemporal fusion method via frequency decoupling and symmetric gating, termed CA-STF. Specifically, a wavelet feature modulation module (WFMM) is first designed to perform frequency decoupling, which extracts reusable structural cues from pseudo changes under low-frequency guidance while suppressing spectral perturbations. Subsequently, a spatiotemporal synergistic enhancement module (SSEM) employs a symmetric gating mechanism to selectively fuse the cue-enhanced features, adaptively decoupling changed and unchanged regions. Finally, a mixture-of-experts based spatial restoration module (SRM) decodes the gated features to reconstruct fine-grained textures in changed areas. Experimental results on the public LGC and CIA datasets demonstrate that our CA-STF achieves superior fusion accuracy and visual quality in complex spatiotemporal dynamics.
Visual-Language Tracking (VLT) enhances traditional visual trackers by incorporating high-level semantic cues from natural language descriptions. However, existing methods often rely on full fine-tuning of the visual encoder and late-stage global alignment between visual and textual features, leading to high training costs and limited generalization ability. To address these issues, we propose CFPTrack, a generalizable VLT framework based on Cross-modality Fourier Prompt Tuning (CFPT), which efficiently aligns visual and textual tokens in the frequency domain using Fast Fourier Transform (FFT). CFPTrack consists of two core components: a CFPT module and a Dynamic Sparse Sinkhorn Attention (DSSA) module. Within CFPT, we introduce Visual-Language Fourier Prompt Tuning (VLFPT), which injects learnable frequency-aware prompts into both visual and language encoders. These prompts capture spatial and frequency-domain statistics and align cross-modal features using optimal transport, enabling the implicit modeling of structural mappings across modalities, thereby bridging the modality gap and improving generalization. The DSSA further enhances robustness by employing a sparse Sinkhorn attention mechanism to highlight informative target cues and suppress background noise. Guided by language semantics, DSSA ranks and prunes visual tokens through a differentiable Top-K sorting strategy, reducing redundancy and improving inference efficiency. Extensive experiments on five VLT benchmarks (TNL2K, LaSOT, LaSOText, OTB99-Lang, and VastTrack) and two UAV benchmarks (UAVDT and DTB70) demonstrate that CFPTrack achieves state-of-the-art real-time performance with significantly reduced training time while maintaining competitive accuracy.
Camouflaged Object Detection (COD) aims to delineate the boundaries between the Camouflaged Objects (COs) and their surrounding backgrounds explicitly. However, an excessive focus on the CO boundaries may make the features insufficient for the low-frequency information, thereby leading to the loss of semantics and incomplete segmentation. This issue can be naturally addressed by perceptual models of the human visual system, which first prioritize the CO within the central vision, and subsequently extend attention to the CO in the peripheral vision. Inspired by this, we propose a Partitioned Observation Network (PONet) for COD. The PONet comprises two branches that separately focus on the low-frequency semantics of the CO’s central region and the high-frequency textures of the CO’s peripheral region. To fully extract and reinforce the low-frequency semantics of the CO’s central region, a Global Semantic Enhancement (GSE) module is designed within the central region branch. Through stacking the GSE modules, the low-frequency information in the features is progressively refined and supplemented, achieving the rendering of the CO’s central region. To capture the high-frequency features of the CO’s peripheral region and to accentuate the contrast between the CO and its surrounding backgrounds, we design a Local Detail Extraction (LDE) module and a weighted peripheral region loss within the peripheral region branch. Guided by the semantics of the central region, the high-resolution details of the target are progressively restored, enhancing the discriminative high-frequency features of the peripheral region. Finally, to complement and integrate the low-frequency and high-frequency information from both regions, a Region Adaptive Fusion (RAF) module is proposed to achieve a complete “observation” of the CO. The PONet is evaluated on four challenging benchmarks, which achieves superior performance over 35 state-of-the-art methods under four widely-used evaluation metrics.
ABSTRACT Current co‐salient object detection (CoSOD) methods leverage common visual representations to identify recurring foreground objects across image groups. A significant challenge arises when distractors from different categories exhibit high visual similarity to the target objects—such as apples and bananas sharing comparable color and texture—making pure appearance‐based matching prone to failure. To address this issue, we introduce En‐ASCoD, an enhanced architecture that integrates both appearance and shape cues to form a more discriminative consensus representation. The model incorporates a Global Co‐appearance Module (GoAM) to extract group‐level appearance prototypes, a Local Co‐appearance Module (LoAM) that refines these features through contrastive learning within local contexts, and a Co‐shape Module (CoSM) that introduces structural constraints via cross‐attention operating on salient tokens selected through global average pooling. By jointly optimizing these modules in a cascaded manner, En‐ASCoD effectively suppresses visually similar distractors while maintaining robustness to intra‐group appearance variations. Extensive experiments on three challenging benchmarks including CoCA, CoSOD3k, and CoSal2015 show that our method consistently outperforms state‐of‐the‐art alternatives, achieving notable gains in both detection accuracy and generalization.
The goal of RGB-D Co-Salient Object Detection (CoSOD) is to integrate both depth and RGB information to identify common and salient objects across a set of images. However, the existing leading research paradigm often fuses raw depth and RGB cues indiscriminately via simple concatenation or weighted summation, neglecting a fundamental limitation: the low inter object contrast in depth maps often causes misalignment between RGB-salient regions and depth-highlighted areas, resulting in model performance degradation. To address this discrepancy, we propose a Semantic-level Multi-modal Alignment transFormer (SMAFormer) that learns to establish precise correspondence for the co-salient objects across RGB and depth modalities. The SMAFormer mainly incorporates two key components: a Learnable Depth-calibration Module (LDM) and a cross-modal transformer (CMT). In the LDM, a Semantic-level Cluster center Extractor (SCE) is first designed to generate semantic level masks, which are then converted to semantic-level cluster centers via the Segment Anything Model (SAM). Afterwards, a Recurrent Deep Clustering Module (RDCM) iteratively projects these semantic representations onto the depth features through a learnable clustering strategy, accurately calibrating the depth features with rich semantic alignment information between two modalities. Next, in CMT, cross-modal shared representations between RGB and calibrated depth features are extracted and mixed through interaction-attention mechanism. Extensive evaluations demonstrate that SMAFormer establishes new state-of the-art performance across three benchmarks including RGB-D CoSal150, RGB-D CoSeg183, and RGB-D CoSal1k.
Ming-Hsuan Yang合作论文数Vision and Learning Lab, University of California, Merced;Google DeepMind10