Cross-domain few-shot object detection (CD-FSOD) remains a challenging problem for existing object detectors and few-shot learning approaches, particularly when generalizing across distinct domains. As part of NTIRE 2026, we hosted the second CD-FSOD Challenge to systematically evaluate and promote progress in detecting objects in unseen target domains under limited annotation conditions. The challenge received strong community interest, with 128 registered participants and a total of 696 submissions. Among them, 31 teams actively participated, and 19 teams submitted valid final results. Participants explored a wide range of strategies, introducing innovative methods that push the performance frontier under both open-source and closed-source tracks. This report presents a detailed overview of the NTIRE 2026 CD-FSOD Challenge, including a summary of the submitted approaches and an analysis of the final results across all participating teams. Challenge Codes: https://github.com/ohMargin/NTIRE2026_CDFSOD.
Most existing lightweight image super-resolution (SR) methods rely on a static feature mapping paradigm, which cannot adaptively accommodate the spatial heterogeneity of texture complexity. Under tight parameter and computational budgets, this inherent static nature imposes a fundamental trade-off between restoration fidelity and inference efficiency. In this work, we propose LiquidSR, a lightweight SR framework that introduces an Ordinary Differential Equation (ODE)-driven adaptive feature update mechanism inspired by Liquid neural networks (LNNs) to overcome the limitations of static modeling. Specifically, we design a spatial-aware heterogeneity neuron (SSN) module that models feature evolution via ODEs, where learnable per-channel time constants τ assign distinct update rates to different feature channels. To handle the multi-time-step features generated by this continuous evolution, we propose an Adaptive Temporal Attention Weights (ATAW) module that enhances informative time steps and suppresses redundancy. Furthermore, we introduce a Differentiable Spatial Heterogeneity Self-Attention (DSSA) module that optimizes spatial heterogeneity representation via ODEs prior to self-attention, thereby synergizing continuous feature update with long-range dependency modeling. Extensive experiments on five standard benchmarks demonstrate that LiquidSR achieves competitive state-of-the-art performance among lightweight SR methods, with a well-balanced trade-off between reconstruction quality and inference efficiency. The code will be publicly available upon acceptance.
Accurate prediction of remaining useful life (RUL) for lithium-ion batteries remains a critical yet complex challenge due to highly non-linear degradation dynamics and profound data heterogeneity across varying operational profiles. While convolutional neural networks (CNNs) have shown promise in battery health management, traditional architectures struggle with gradient vanishing in deep feature spaces and lack the adaptive capacity to filter early-cycle noise under diverse degradation conditions. To improve robust RUL estimation across heterogeneous benchmark datasets, this paper proposes a deep learning framework that integrates residual connections with dual-attention mechanisms (ResCNN). Specifically, the residual structures effectively mitigate gradient degradation during the extraction of abstract degradation patterns. Concurrently, a synergistic Squeeze-and-Excitation (SE) and Multi-Head Attention module adaptively calibrates channel-wise feature importance and captures long-range temporal dependencies inherent in complex capacity fade processes. The proposed framework is evaluated under a wide spectrum of degradation conditions and distinct cathode systems (LFP and LCO) using both dataset-specific train/validation/test protocols and strict source-to-target cross-dataset transfer tests. Experimental results demonstrate that ResCNN achieves consistently lower prediction errors than baseline models across the evaluated datasets and maintains positive explanatory power on unseen target datasets without target-domain training. Ablation studies further validate the synergistic contribution of each architectural component toward capturing intrinsic battery aging phenomena.
The complex causes and varied visual manifestations of cotton diseases create significant obstacles for accurate disease identification in cotton plants. Under limited computational conditions, it is still challenging to construct lightweight cotton disease detection models that maintain high detection accuracy. To address this issue, we propose LMF-YOLO (Lightweight Multi-scale Feature-enhanced YOLO), aiming to achieve accurate cotton disease detection with reduced computational cost. Specifically, a lightweight multi-scale feature extraction block (LMFEBlock) is introduced to reconstruct the original C2f module, forming the C2f-LMFEB backbone, which enhances fine-grained disease representation while significantly reducing redundant computation. Meanwhile, a dynamic upsampling strategy, DySample, is incorporated into the neck network to adaptively optimize feature reconstruction across scales and alleviate spatial misalignment caused by fixed interpolation. Furthermore, a shared and refined detection head, LMFEB-Detect, is developed to improve feature consistency between classification and regression tasks, strengthen boundary localization, and suppress background interference with lower parameter overhead. Extensive experiments conducted on a publicly available cotton disease dataset demonstrate that LMF-YOLO achieves an mAP of 90.0%, outperforming YOLOv8n by 2.5%, while reducing computational complexity, parameter count, and model size by 50.0%, 40.1%, and 36.1\%, respectively. Additional evaluations on two independent public datasets further confirm its strong generalization capability. These results indicate that LMF-YOLO provides an effective balance between detection accuracy and computational efficiency, making it particularly suitable for real-time cotton disease monitoring in resource-constrained agricultural environments.
Vision-language tracking has gained increasing attention in many scenarios. This task simultaneously deals with visual and linguistic information to localize objects in videos. Despite its growing utility, the development of vision-language tracking methods remains in its early stage. Current vision-language trackers usually employ Transformer architectures for interactive integration of template, search, and text features. However, persistent challenges about low-semantic images including prevalent image blurriness, low resolution and so on, may compromise model performance through degraded cross-modal understanding. To solve this problem, language assistance is usually used to deal with the obstacles posed by low-semantic images. However, due to the existing gap between current textual and visual features, direct concatenation and fusion of these features may have limited effectiveness. To address these challenges, we introduce a pioneering Generative Language-AssisteD tracking model, GLAD, which utilizes diffusion models for the generative multi-modal fusion of text description and template image to bolster compatibility between language and image and enhance template image semantic information. Our approach demonstrates notable improvements over the existing fusion paradigms. Blurry and semantically ambiguous template images can be restored to improve multi-modal features in the generative fusion paradigm. Experiments show that our method establishes a new state-of-the-art on multiple benchmarks and achieves an impressive inference speed.
Video super-resolution is a fundamental task aimed at enhancing video quality through intricate modeling techniques. Recent advancements in diffusion models have significantly enhanced image super-resolution processing capabilities. However, their integration into video super-resolution workflows remains constrained due to the computational complexity of temporal fusion modules, demanding more computational resources compared to their image counterparts. To address this challenge, we propose a novel approach: a Frames-Shift Diffusion Model based on the image diffusion models. Compared to directly training diffusion-based video super-resolution models, redesigning the diffusion process of image models without introducing complex temporal modules requires minimal training consumption. We incorporate temporal information into the image super-resolution diffusion model by using optical flow and perform multi-frame fusion. This model adapts the diffusion process to smoothly transition from image super-resolution to video super-resolution diffusion without additional weight parameters. As a result, the Frames-Shift Diffusion Model efficiently processes videos frame by frame while maintaining computational efficiency and achieving superior performance. It enhances perceptual quality and achieves comparable performance to other state-of-the-art diffusion-based VSR methods in PSNR and SSIM. This approach optimizes video super-resolution by simplifying the integration of temporal data, thus addressing key challenges in the field.
To compare the blood vessel visualization with spiral MRA (MRAspiral) and compressed SENSE accelerated Cartesian MRA (MRACS) in moyamoya disease (MMD) patients, with digital subtraction angiography (DSA) as the reference standard. We prospectively collected MRAspiral with different acquisition windows (τ = 4, 6, 10 ms), MRACS and DSA in MMD patients. Contrast-to-noise ratio (CNR) was measured in the M1, M2, M3, and M4 segments of the middle cerebral artery (MCA) for each MRA sequence. Vessel visualization of the distal MCA, leptomeningeal artery (LMA) collaterals, distal external carotid artery (ECA), and internal carotid artery (ICA) steno-occlusion was qualitatively analyzed using 3- and 4-point likert scales compared to DSA. A linear fixed-effects model was used to determine differences among the four sequences. A total of 98 hemispheres from 55 MMD patients were included. CNR in the M2, M3 and M4 segments of the MCA was not significantly different between MRACS and MRAτ4 or MRAτ6, but it was significantly higher in MRACS than MRAτ10 (M2: P < 0.001, M3: P < 0.001, M4: P = 0.013). MRAspiral sequences provided better visualization of the distal MCA, LMA collaterals and distal ECA compared to MRACS (all P < 0.001). MRAspiral offers improved vessel visualization in distal arteries with adequate image quality for patients with MMD. Compared to MRACS, MRAspiral can reduce scan time by 32.31% when the τ value is set to 6 ms, while also providing superior image quality. Spiral MRA performs well in visualizing collateral vessels in moyamoya disease with shorter scan time.
Lane detection is to determine the precise location and shape of lanes on the road. Despite efforts made by current methods, it remains a challenging task due to the complexity of real-world scenarios. Existing approaches, whether proposal-based or keypoint-based, suffer from depicting lanes effectively and efficiently. Proposal-based methods detect lanes by distinguishing and regressing a collection of proposals in a streamlined top-down way, yet lack sufficient flexibility in lane representation. Keypoint-based methods, on the other hand, construct lanes flexibly from local descriptors, which typically entail complicated post-processing. In this paper, we present a "Sketch-and-Refine" paradigm that utilizes the merits of both keypoint-based and proposal-based methods. The motivation is that local directions of lanes are semantically simple and clear. At the "Sketch" stage, local directions of keypoints can be easily estimated by fast convolutional layers. Then we can build a set of lane proposals accordingly with moderate accuracy. At the "Refine" stage, we further optimize these proposals via a novel Lane Segment Association Module (LSAM), which allows adaptive lane segment adjustment. Last but not least, we propose multi-level feature integration to enrich lane feature representations more efficiently. Based on the proposed "Sketch and Refine" paradigm, we propose a fast yet effective lane detector dubbed "SRLane". Experiments show that our SRLane can run at a fast speed (i.e., 278 FPS) while yielding an F1 score of 78.9%. The source code is available at: https://github.com/passerer/SRLane.
High-order quantum coherence reveals the statistical correlation of quantum particles. Manipulation of quantum coherence of light in the temporal domain enables the production of the single-photon source, which has become one of the most important quantum resources. High-order quantum coherence in the spatial domain plays a crucial role in a variety of applications, such as quantum imaging, holography, and microscopy. However, the active control of second-order spatial quantum coherence remains a challenging task. Here we predict theoretically and demonstrate experimentally the first active manipulation of second-order spatial quantum coherence, which exhibits the capability of switching between bunching and anti-bunching, by mapping the entanglement of spatially structured photons. We also show that signal processing based on quantum coherence exhibits robust resistance to intensity disturbance. Our findings not only enhance existing applications but also pave the way for broader utilization of higher-order spatial quantum coherence.
Photonic orbital angular momentum (OAM) carried by phase-structured vortex light is an important and promising resource for the ever-increasing demand towards high-capacity data information due to its intrinsic unlimited dimensionality. Large superpositions of OAM are easy to be produced, but on-demand generation of arbitrary OAM spectra such as an OAM comb similar to a frequency comb is still a challenge; especially, the on-demand OAM comb and arbitrary multi-OAM modes have not yet been realized at the source. Here we report a versatile at-source strategy for developing a flexibly and dynamically switchable on-demand digital OAM comb laser for the first time, to our knowledge, by controlling the phase degree of freedom itself rather than any proxy. For this aim, we present a crucial design idea that a nested ring cavity configuration is composed of a degenerate cavity embedded into a stable ring cavity and a pair of conjugate two-fold symmetric multi-spiral-phase digital holographic mirrors loaded onto reflective phase-only spatial light modulators. In the nested ring cavity, the stable ring cavity and the degenerate cavity meet the requirements of high spatial coherence and supporting any transverse mode, respectively. The paired conjugate holographic mirrors located in mutual object and image planes circumvent the competing issue among different OAM modes and control the number and chirality of modes in OAM combs with ease. Our strategy has also universality as it has the ability of encoding OAM spectra with arbitrary distribution. The realization of a dynamic on-demand multi-OAM-mode laser is an important progress in the infancy of multi-OAM-mode sources. Our idea provides a promising solution for development of emerging high-dimensional technologies; in the future, there will be increasing opportunities in the fundamentals and applications of high-dimensional OAM modes, and beyond. Our strategy not only contributes to the development of new laser technology, but also provides a toolbox for both linear and nonlinear generation of the multiple OAM modes at the source. (c) 2024 Optica Publishing Group under the terms of the Optica Open Access Publishing Agreement
Visual Reinforcement Learning (RL) is a promising approach to achieve human-like intelligence. However, it currently faces challenges in learning efficiently within noisy environments. In contrast, humans can quickly identify task-relevant objects in distraction-filled surroundings by applying previously acquired common knowledge. Recently, foundational models in natural language processing and computer vision have achieved remarkable successes, and the common knowledge within these models can significantly benefit downstream task training. Inspired by these achievements, we aim to incorporate common knowledge from foundational models into visual RL. We propose a novel Focus-Then-Decide (FTD) framework, allowing the agent to make decisions based solely on task-relevant objects. To achieve this, we introduce an attention mechanism to select task-relevant objects from the object set returned by a foundational segmentation model, and only use the task-relevant objects for the subsequent training of the decision module. Additionally, we specifically employed two generic self-supervised objectives to facilitate the rapid learning of this attention mechanism. Experimental results on challenging tasks based on DeepMind Control Suite and Franka Emika Robotics demonstrate that our method can quickly and accurately pinpoint objects of interest in noisy environments. Consequently, it achieves a significant performance improvement over current state-of-the-art algorithms. Project Page: https://www.lamda.nju.edu.cn/chenc/FTD.html Code: https://github.com/LAMDA-RL/FTD
Reinforcement Learning (RL) has experienced rapid advancements in recent years. The widely studied RL algorithms mainly focus on a single input form, such as pixel-based image input or symbolic vector input. These two forms have different characteristics and, in many scenarios, will appear together, while few RL algorithms have studied the problems with mixed input types. Specifically, in the scenario where both pixel and symbolic inputs are available, symbolic input usually offers abstract features with specific semantics, which is more conducive to the agent's focus. Conversely, pixel input provides more comprehensive information, enabling the agent to make well-informed decisions. Tailoring the processing approach based on the properties of these two input types can contribute to solving the problem more effectively. To tackle the above issue, we propose an Internal Logical Induction (ILI) framework that integrates deep RL and rule learning into one system. ILI utilizes the deep RL algorithm to process the pixel input and the rule learning algorithm to induce propositional logic knowledge from symbolic input. To efficiently combine these two mechanisms, we further adopt a reward shaping technique by treating valuable knowledge as intrinsic rewards for the RL procedure. Experimental results demonstrate that the ILI framework outperforms baseline approaches in RL problems with pixel-symbolic input, and its inductive knowledge exhibits transferability advantages when pixel input semantics change.
Image super-resolution (SR) serves as a fundamental tool for the processing and transmission of multimedia data. Recently, Transformer-based models have achieved competitive performances in image SR. They divide images into fixed-size patches and apply self-attention on these patches to model long-range dependencies among pixels. However, this architecture design is originated for high-level vision tasks, which lacks design guideline from SR knowledge. In this paper, we aim to design a new attention block whose insights are from the interpretation of Local Attribution Map (LAM) for SR networks. Specifically, LAM presents a hierarchical importance map where the most important pixels are located in a fine area of a patch and some less important pixels are spread in a coarse area of the whole image. To access pixels in the coarse area, instead of using a very large patch size, we propose a lightweight Global Pixel Access (GPA) module that applies cross-attention with the most similar patch in an image. In the fine area, we use an Intra-Patch Self-Attention (IPSA) module to model long-range pixel dependencies in a local patch, and then a spatial convolution is applied to process the finest details. In addition, a Cascaded Patch Division (CPD) strategy is proposed to enhance perceptual quality of recovered images. Extensive experiments suggest that our method outperforms state-of-the-art lightweight SR methods by a large margin. Code is available at https://github.com/passerer/HPINet.
This paper reviews the NTIRE 2022 challenge on efficient single image super-resolution with focus on the proposed solutions and results. The task of the challenge was to super-resolve an input image with a magnification factor of ×4 based on pairs of low and corresponding high resolution images. The aim was to design a network for single image super-resolution that achieved improvement of efficiency measured according to several metrics including runtime, parameters, FLOPs, activations, and memory consumption while at least maintaining the PSNR of 29.00dB on DIV2K validation set. IMDN is set as the baseline for efficiency measurement. The challenge had 3 tracks including the main track (runtime), sub-track one (model complexity), and sub-track two (overall performance). In the main track, the practical runtime performance of the submissions was evaluated. The rank of the teams were determined directly by the absolute value of the average runtime on the validation set and test set. In sub-track one, the number of parameters and FLOPs were considered. And the individual rankings of the two metrics were summed up to determine a final ranking in this track. In sub-track two, all of the five metrics mentioned in the description of the challenge including runtime, parameter count, FLOPs, activations, and memory consumption were considered. Similar to sub-track one, the rankings of five metrics were summed up to determine a final ranking. The challenge had 303 registered participants, and 43 teams made valid submissions. They gauge the state-of-the-art in efficient single image super-resolution.
As an important degree of freedom, orbital angular momentum (OAM) plays a key role in the research of photonic quantum information. Combined OAM with other degrees of freedom of photons such as polarization, multi-degree-of-freedom photonic quantum information processing is possible. In addition, due to the property of natural discrete high dimensions, OAM is one of the optimal degrees of freedom for the research of high dimensional quantum information processing. Based on the spontaneous parametric down-conversion nonlinear optical process, the entangled source with OAM can be easily obtained. In recent years, quantum entanglement based on OAM of photons has attracted wide attention and many significant progresses have been made in many directions, such as multiple degrees of freedom, high dimension and multiple photons. However, there are still many key scientific issues that need to be further studied in this realm, including how to achieve efficient and high-quality OAM sorter, how to achieve higher-dimensional frequency conversion, how to improve the quality of multi-degree-of-freedom entangled sorter, how to obtain high-dimensional entangled states with more dimensions and more photons, and how to construct feasible high-dimensional quantum gates. Starting from the most basic two-dimensional manipulation of OAM of photons, the quantum state regulation of single photon with OAM and the entanglement manipulation of two photons and multiple photons with OAM are reviewed. Based on the characteristics of multiple degrees of freedom, large angular momentum and high dimension, quantum entanglement of OAM of photons is discussed systematically from the perspectives of generation, regulation, measurement and application. Meanwhile, the possible methods to overcome the challenges in this realm are explored.
Recently, lane detection has made great progress with the rapid development of deep neural networks and autonomous driving. However, there exist three mainly problems including characterizing lanes, modeling the structural relationship between scenes and lanes, and supporting more attributes (e.g., instance and type) of lanes. In this paper, we propose a novel structure guided framework to solve these problems simultaneously. In the framework, we first introduce a new lane representation to characterize each instance. Then a topdown vanishing point guided anchoring mechanism is proposed to produce intensive anchors, which efficiently capture various lanes. Next, multi-level structural constraints are used to improve the perception of lanes. In the process, pixel-level perception with binary segmentation is introduced to promote features around anchors and restore lane details from bottom up, a lane-level relation is put forward to model structures (i.e., parallel) around lanes, and an image-level attention is used to adaptively attend different regions of the image from the perspective of scenes. With the help of structural guidance, anchors are effectively classified and regressed to obtain precise locations and shapes. Extensive experiments on public benchmark datasets show that the proposed approach outperforms state-of-the-art methods with 117 FPS on a single GPU.