Large Reasoning Models (LRMs) have advanced automated multi-step reasoning, but their ability to generate complex Chain-of-Thought (CoT) trajectories introduces severe privacy risks, as sensitive information may be deeply embedded throughout the reasoning process. Existing Large Language Models (LLMs) unlearning approaches that typically focus on modifying only final answers are insufficient for LRMs, as they fail to remove sensitive content from intermediate steps, leading to persistent privacy leakage and degraded security. To address these challenges, we propose Sensitive Trajectory Regulation (STaR), a parameter-free, inference-time unlearning framework that achieves robust privacy protection throughout the reasoning process. Specifically, we first identify sensitive content via semantic-aware detection. Then, we inject global safety constraints through secure prompt prefix. Next, we perform trajectory-aware suppression to dynamically block sensitive content across the entire reasoning chain. Finally, we apply token-level adaptive filtering to prevent both exact and paraphrased sensitive tokens during generation. Furthermore, to overcome the inadequacies of existing evaluation protocols, we introduce two metrics: Multi-Decoding Consistency Assessment (MCS), which measures the consistency of unlearning across diverse decoding strategies, and Multi-Granularity Membership Inference Attack (MIA) Evaluation, which quantifies privacy protection at both answer and reasoning-chain levels. Experiments on the R-TOFU benchmark demonstrate that STaR achieves comprehensive and stable unlearning with minimal utility loss, setting a new standard for privacy-preserving reasoning in LRMs.
Multivariate time series forecasting plays a crucial role in various domains. Existing forecasting methodologies neglect the fact that one prediction task involves multiple forecasting horizons, each requiring distinct analytical perspectives. To address this issue, we propose the Temporal Decoupling Network (TDN), which explicitly models horizon disparities and facilitates horizon-specific predictions. To meet the feature analysis requirements across varying forecasting horizons, TDN systematically extracts dynamic features from recent historical inputs and inherent features by encoding the temporal properties of the datasets. Then, TDN employs a multi-head and multi-resolution mechanism to adaptively fuse the dynamic and inherent features according to diverse forecasting horizons, thereby yielding horizon-specific forecasts. The multi-resolution form provides horizon segmentation at varying granularities, enabling TDN to capture local dynamics and global regularities in predictions. Extensive experiments conducted across seven real-world datasets demonstrate the superiority of TDN. Code is available at https://github.com/dzysadasd/Temporal-Decoupling-Network.
Weakly-supervised Video Anomaly Detection (wVAD) aims to detect abnormal events using only binary labels, making it challenging to capture both the diversity of anomalies and their shared semantic cues. Existing methods either focus on a generic anomaly pattern, achieving strong generalization but weak discrimination, or rely on class-level diversity modeling, which ignores shared semantics and suffers from limited generalization. To overcome these limitations, we propose the Mixture of Memory Experts (MoME), a unified framework that jointly learns general and diverse patterns. Each expert in MoME possesses an internal memory for fine-grained specialization and shares an external memory for general knowledge aggregation. To enhance semantic diversity and improve generalization beyond coarse class-level supervision, we introduce an Anomaly Prototype Router that leverages large language models to construct generalized anomaly prototypes for semantically guided expert routing. Moreover, the regularization loss for APR ensures balanced routing, the distinctiveness loss for experts encourages diversity, and reconstruction together with memory tasks enhance pattern discriminability. Extensive experiments on UCF-Crime and XD-Violence demonstrate that our approach achieves state-of-the-art performance, validating the effectiveness of jointly modeling generality and diversity for robust anomaly detection under weak supervision.
Traffic forecasting is essential in city-level applications, where data-driven deep learning has become the most popular method. However, sufficient data in developing cities is not always accessible, posing a challenge for training effective models in scenarios with limited data. Recently, several works have promoted this issue through cross-city knowledge transfer and shown promising performances. However, existing methods can neither distinguish node divergence nor extract functional similarities between cities, which results in suboptimal performance. To overcome the limitations, we propose a Cross-city Correlation Learning (CCL) framework. Firstly, we construct a self-supervised learning model to infer accurate node-to-node and node-to-region cross-city correlations from multiple noisy labels without using any auxiliary information. Then, we achieve spatial knowledge transfer from a transfer-adaptive graph convolution network based on the learned correlations in two aspects: the learnable adjacency matrix and region-specific kernel parameters, which ensure the target models can transfer more and better utilize the knowledge from the source domain. The experiments are conducted on six real-world datasets and fully prove the effectiveness of the proposed framework.
Video Anomaly Detection (VAD) aims to identify abnormal frames from discrete events within video sequences. Existing VAD methods suffer from heavy annotation burdens in fully-supervised paradigm, insensitivity to subtle anomalies in semi-supervised paradigm, and vulnerability to noise in weakly-supervised paradigm. To address these limitations, we propose a novel paradigm: Single-Frame supervised VAD (SF-VAD), which uses a single annotated abnormal frame per abnormal video. SF-VAD ensures annotation efficiency while offering precise anomaly reference, facilitating robust anomaly modeling, and enhancing the detection of subtle anomalies in complex visual contexts. To validate its effectiveness, we construct three SF-VAD benchmarks by manually re-annotating the ShanghaiTech, UCF-Crime, and XD-Violence datasets in a practical procedure. Further, we devise Frame-guided Progressive Learning (FPL), to generalize sparse frame supervision to event-level anomaly understanding. FPL first leverages evidential learning to estimate anomaly relevance guided by annotated frames. Then it extends anomaly supervision by mining discrete abnormal events based on anomaly relevance and feature similarity. Meanwhile, FPL decouples normal patterns by isolating distinct normal frames outside abnormal events, reducing false alarms. Extensive experiments show SF-VAD achieves state-of-the-art detection results while offering a favorable trade-off between performance and annotation cost.
The hallucination of large multimodal models (LMMs), providing responses that appear correct but are actually incorrect, limits their reliability and applicability. This paper aims to study the hallucination problem of LMMs in video modality, which is dynamic and more challenging compared to static modalities like images and text. From this motivation, we first present a comprehensive benchmark termed HAVEN for evaluating hallucinations of LMMs in video understanding tasks. It is built upon three dimensions, i.e., hallucination causes, hallucination aspects, and question formats, resulting in 6K questions. Then, we quantitatively study 7 influential factors on hallucinations, e.g., duration time of videos, model sizes, and model reasoning, via experiments of 16 LMMs on the presented benchmark. In addition, inspired by recent thinking models like OpenAI o1, we propose a video-thinking model to mitigate the hallucinations of LMMs via supervised reasoning fine-tuning (SRFT) and direct preference optimization (TDPO)– where SRFT enhances reasoning capabilities while TDPO reduces hallucinations in the thinking process. Extensive experiments and analyses demonstrate the effectiveness. Remarkably, it improves the baseline by 7.65 hallucination evaluation and reduces the bias score by 4.5 are public at https://github.com/Hongcheng-Gao/HAVEN.
China is the largest producer of marine capture fisheries globally. Overfishing since the 1970s has led to a decline in fishery resources in Chinese coastal waters. After China’s reform and opening up, a series of management measures were implemented to alleviate marine fishing pressure and conserve the fisheries resources. We conducted a comprehensive assessment for multispecies fisheries in the South China Sea (SCS) to explore whether fisheries management has been effective in recovery of the resources. Indicators of the exploitation status of major commercial fish species were assessed using statistical catch data and survey data simultaneously. The results reveal a significant shift in bottom-trawl fishery, with its share of the total catch transitioning from an upward to a currently downward trend. The species composition of bottom-trawl fisheries has undergone substantial changes in the SCS over six decades. Stock assessment results based on catch data indicated some positive signals, with small pelagic fishes, such as herrings, anchovies, mackerel and scad recovering from overfished/overfishing to a healthy status. However, the exploitation status of high-trophic-level fish species, such as conger pike and groupers, were still in overfished status. Assessment based on length data was less optimistic. Our uncertainty analysis showed that the catch-based model is less sensitive to parameters compared with the two length-based models considered here. We advocate for more practical and precise fisheries management in China, such as category/species-based management, further optimization and improvement of the fishing structure, development of a scientific quota-based system, ecosystem management that incorporates climate factors, and establishment of marine protected areas for fish species that are severely overfished or have high ecological value.
With the continuous expansion of urban areas, accurate and effective traffic forecasting has become essential for intelligent urban traffic management. As traffic data inherently exhibits temporal dynamics, modeling its temporal patterns is critical to improve prediction performance. However, constrained by computational complexity, existing methods rely primarily on short-term historical data, which is typically noisy and limits the ability to capture global temporal patterns. To address this issue, we propose a novel Dual-Stream Transformer model (DSformer) that effectively captures global temporal patterns through a time-index model. To mitigate the impact of noise in short look-back windows, DSformer explicitly learns a temporal matrix that encodes structured temporal dependencies. Furthermore, we design a time-index loss that encourages similar representations for adjacent time indices, thereby reducing error propagation across time steps. In parallel, a historical-value stream is employed to model local information. Finally, a self-adaptive learning module is constructed to flexibly and accurately fuse global and local information. Extensive experiments on real-world traffic forecasting tasks across ten diverse scenarios demonstrate that our method consistently outperforms state-of-the-art baselines while maintaining competitive efficiency. The code is available at https://github.com/sky836/DSFormer.git.
The Beibu Gulf is a resource-rich bay in the northwestern South China Sea; however, its fishery resources have been in decline owing to overfishing and other factors. To rebuild the depleted fish stocks, China has implemented a series of fishery resources conservation and management measures; the most influential of these is possibly the summer fishing moratorium, in effect since 1999. This study used data from bottom-trawl surveys conducted in spring and autumn from 1998 to 2020 to analyze variations in the fish community of the Beibu Gulf following implementation of the fishery resources protection measures. The data analysis indicated improvements in fish species richness, diversity, and mean trophic level, but further declines in fish abundance and biomass since 1998. The composition of dominant species, except for the small-sized glowbelly Acropoma japonicum, changed obviously from 1998 to 2020, although still mainly composed of small-sized demersal and pelagic species. Abundance-biomass comparison curves indicate that the disturbance level affecting fish in the Beibu Gulf remained unchanged and that it is still in a heavily disturbed state. These findings suggest that the current fishery resource conservation policies are inadequate for reversing the trend of declining fishery resources. China's marine fishery resources management measures should be refined by incorporating more-precise conservation measures consistent with the characteristics of the local resources. In particular, relevant measures are needed to control the fishing output, such as the quota fishing system currently being piloted.
Video has emerged as a favored multimedia format on the internet. To better gain video contents, a new topic HIREST is presented, including video retrieval, moment retrieval, moment segmentation, and step-captioning. The pioneering work chooses the pre-trained CLIP-based model for video retrieval, and leverages it as a feature extractor for other three challenging tasks solved in a multi-task learning paradigm. Nevertheless, this work struggles to learn the comprehensive cognition of user-preferred content, due to disregarding the hierarchies and association relations across modalities. In this paper, guided by the shallow-to-deep principle, we propose a query-centric audio-visual cognition (QUAG) network to construct a reliable multi-modal representation for moment retrieval, segmentation and step-captioning. Specifically, we first design the modality-synergistic perception to obtain rich audio-visual content, by modeling global contrastive alignment and local fine-grained interaction between visual and audio modalities. Then, we devise the query-centric cognition that uses the deep-level query to perform the temporal-channel filtration on the shallow-level audio-visual representation. This can cognize user-preferred content and thus attain a query-centric audio-visual representation for three tasks. Extensive experiments show QUAG achieves the SOTA results on HIREST. Further, we test QUAG on the query-based video summarization task and verify its good generalization.
The development of parameter-efficient fine-tuning methods has effectively bridged the trade-off between model capability and optimization expenses for the adaptation of large models. LoRA-based methods have become particularly prominent in this domain, as they maintain baseline inference efficiency without requiring additional reasoning resources. Despite these advantages, such methods still exhibit a noticeable performance gap compared to full fine-tuning, or mitigate it at the expense of increased memory usage and extended training time. As an alternative, we introduce SVDLoRA, a parameter adjustment strategy that integrates Truncated Singular Value Decomposition (SVD) with orthogonal constraints on matrices. By projecting pre-trained weights into an optimal low-rank space via Truncated SVD, SVDLoRA significantly reduces resource demands while improving performance and training stability, and preserving the inference efficiency characteristic of standard LoRA. Experimental results show that SVDLoRA-tuned LLaMA and VL-BART models consistently outperform those using LoRA and DoRA across both commonsense reasoning and image-text understanding tasks.Code is available at github.
Change captioning aims to describe the semantic change between two similar images. In this process, as the most typical distractor, viewpoint change leads to the pseudo changes about appearance and position of objects, thereby overwhelming the real change. Besides, since the visual signal of change appears in a local region with weak feature, it is difficult for the model to directly translate the learned change features into the sentence. In this paper, we propose a syntax-calibrated multi-aspect relation transformer to learn effective change features under different scenes, and build reliable cross-modal alignment between the change features and linguistic words during caption generation. Specifically, a multi-aspect relation learning network is designed to 1) explore the fine-grained changes under irrelevant distractors ( e.g., viewpoint change) by embedding the relations of semantics and relative position into the features of each image; 2) learn two view-invariant image representations by strengthening their global contrastive alignment relation, so as to help capture a stable difference representation; 3) provide the model with the prior knowledge about whether and where the semantic change happened by measuring the relation between the representations of captured difference and the image pair. Through the above manner, the model can learn effective change features for caption generation. Further, we introduce the syntax knowledge of Part-of-Speech (POS) and devise a POS-based visual switch to calibrate the transformer decoder. The POS-based visual switch dynamically utilizes visual information during different word generation based on the POS of words. This enables the decoder to build reliable cross-modal alignment, so as to generate a high-level linguistic sentence about change. Extensive experiments show that the proposed method achieves the state-of-the-art performance on the three public datasets.
Weakly-supervised Video Anomaly Detection (wVAD) aims to detect frame-level anomalies using only video-level labels in training. Due to the limitation of coarse-grained labels, Multi-Instance Learning (MIL) is prevailing in wVAD. However, MIL suffers from insufficiency of binary supervision to model diverse abnormal patterns. Besides, the coupling between abnormality and its context hinders the learning of clear abnormal event boundary. In this paper, we propose prompt-enhanced MIL to detect various abnormal events while ensuring clear event boundaries. Concretely, we design the abnormal-aware prompts by using abnormal class annotations together with learnable prompt, which can incorporate semantic priors into video features dynamically. The detector can utilize the semantic-rich features to capture diverse abnormal patterns. In addition, normal context prompt is introduced to amplify the distinction between abnormality and its context, facilitating the generation of clear boundary. With the mutual enhancement of abnormal-aware and normal context prompt, the model can construct discriminative representations to detect divergent anomalies without ambiguous event boundaries. Extensive experiments demonstrate our method achieves SOTA performance on three public benchmarks. The code is available at https://github.com/Junxi-Chen/PE-MIL.
Diffusion models (DMs) have demonstrated remarkable proficiency in producing images based on textual prompts. Numerous methods have been proposed to ensure these models generate safe images. Early methods attempt to incorporate safety filters into models to mitigate the risk of generating harmful images but such external filters do not inherently detoxify the model and can be easily bypassed. Hence, model unlearning and data cleaning are the most essential methods for maintaining the safety of models, given their impact on model parameters. However, malicious fine-tuning can still make models prone to generating harmful or undesirable images even with these methods. Inspired by the phenomenon of catastrophic forgetting, we propose a training policy using contrastive learning to increase the latent space distance between clean and harmful data distribution, thereby protecting models from being fine-tuned to generate harmful images due to forgetting. The experimental results demonstrate that our methods not only maintain clean image generation capabilities before malicious fine-tuning but also effectively prevent DMs from producing harmful images after malicious fine-tuning. Our method can also be combined with other safety methods to maintain their safety against malicious fine-tuning further.
Neural tuning for visual words is essential for fluent reading across various scripts. This study investigated the emergence and development of N170 tuning for Chinese characters and its cognitive-linguistic correlates. Electroencephalogram data from 48 adult L2 learners and 23 native Chinese readers were collected using a color detection task. The N170 for real characters, pseudo-characters, false characters, stroke combinations and line drawings were recorded. We found beginner adult L2 learners showed larger N170 Chinese characters compared to stroke combinations (coarse neural tuning). The intermediate-level L2 Chinese learners demonstrated fine-tuning for Chinese orthographic regularities. Importantly, a clear shift from bilateral to left-lateralized coarse and fine-tuning for print was observed from beginner to intermediate L2 learners as their Chinese reading experience increased. Moreover, individual differences in neural print tuning moderately correlated with word-reading fluency, Chinese vocabulary knowledge and morphological awareness.
Change captioning aims to succinctly describe the semantic change between a pair of similar images, while being immune to distractors (illumination and viewpoint changes). Under these distractors, unchanged objects often appear pseudo changes about location and scale, and certain objects might overlap others, resulting in perturbational and discrimination-degraded features between two images. However, most existing methods directly capture the difference between them, which risk obtaining error-prone difference features. In this paper, we propose a distractors-immune representation learning network that correlates the corresponding channels of two image representations and decorrelates different ones in a self-supervised manner, thus attaining a pair of stable image representations under distractors. Then, the model can better interact them to capture the reliable difference features for caption generation. To yield words based on the most related difference features, we further design a cross-modal contrastive regularization, which regularizes the cross-modal alignment by maximizing the contrastive alignment between the attended difference features and generated words. Extensive experiments show that our method outperforms the state-of-the-art methods on four public datasets. The code is available at https://github.com/tuyunbin/DIRL.
The recently invented retina-inspired spike camera has shown great potential for capturing dynamic scenes. However, reconstructing high-quality images from the binary spike data remains a challenge due to the existence of noises in the camera. This paper proposes SpikeODE, a novel approach to reconstructing clear images by exploring temporal-spatial correlation to depress noises. The main idea of our method is to restore the continuous dynamic process of real scenes in a latent space and learn the temporal correlations in a fine-grained manner. Furthermore, to model the dynamic process more effectively, we design a conditional ODE where the latent state of each timestamp is conditioned on the observed spike data. Subsequently, forward and backward inferences are conducted through the ODE to investigate the correlations between the representation of the target timestamp and the information from both past and future contexts. Additionally, we incorporate a Unet structure with a pixel-wise attention mechanism at each level to learn spatial correlations. Experimental results demonstrate that our method outperforms state-of-the-art methods across several metrics.
Multi-change captioning aims to describe complex and coupled changes within an image pair in natural language. Compared with single-change captioning, this task requires the model to have higher-level cognition ability to reason an arbitrary number of changes. In this paper, we propose a novel context-aware difference distilling (CARD) network to capture all genuine changes for yielding sentences. Given an image pair, CARD first decouples context features that aggregate all similar/dissimilar semantics, termed common/difference context features. Then, the consistency and independence constraints are designed to guarantee the alignment/discrepancy of common/difference context features. Further, the common context features guide the model to mine locally unchanged features, which are subtracted from the pair to distill locally difference features. Next, the difference context features augment the locally difference features to ensure that all changes are distilled. In this way, we obtain an omni-representation of all changes, which is translated into linguistic sentences by a transformer decoder. Extensive experiments on three public datasets show CARD performs favourably against state-of-the-art methods.The code is available at https://github.com/tuyunbin/CARD.
Change captioning is to describe the semantic change between a pair of similar images in natural language. It is more challenging than general image captioning, because it requires capturing fine-grained change information while being immune to irrelevant viewpoint changes, and solving syntax ambiguity in change descriptions. In this paper, we propose a neighborhood contrastive transformer to improve the model's perceiving ability for various changes under different scenes and cognition ability for complex syntax structure. Concretely, we first design a neighboring feature aggregating to integrate neighboring context into each feature, which helps quickly locate the inconspicuous changes under the guidance of conspicuous referents. Then, we devise a common feature distilling to compare two images at neighborhood level and extract common properties from each image, so as to learn effective contrastive information between them. Finally, we introduce the explicit dependencies between words to calibrate the transformer decoder, which helps better understand complex syntax structure during training. Extensive experimental results demonstrate that the proposed method achieves the state-of-the-art performance on three public datasets with different change scenarios. The code is available at https://github.com/tuyunbin/NCT.