Open-set supervised anomaly detection (OSAD) aims to identify unseen anomalies using limited anomalous supervision. However, existing prototype-based methods typically model normal data via a unimodal Gaussian prior, failing to capture inherent multi-modality and resulting in blurred decision boundaries. To address this, we propose Mixture Prototype Flow Matching (MPFM), a framework that learns a continuous transformation from normal feature distributions to a structured Gaussian mixture prototype space. Departing from traditional flow-based approaches that rely on a single velocity vector, MPFM explicitly models the velocity field as a Gaussian mixture prior where each component corresponds to a distinct normal class. This design facilitates mode-aware and semantically coherent distribution transport. Furthermore, we introduce a Mutual Information Maximization Regularizer (MIMR) to prevent prototype collapse and maximize normal-anomaly separability. Extensive experiments demonstrate that MPFM achieves state-of-the-art performance across diverse benchmarks under both single- and multi-anomaly settings.
Few-shot continual learning (FSCL) has attracted increasing attention for real-world applications, where models must continuously adapt to new classes with only a few labeled samples while retaining prior knowledge. These abilities are essential in dynamic environments where data availability is often sparse and nonstationary. However, traditional FSCL methods are largely confined to closed data spaces, which limits their generalizability when diverse and evolving distributions are involved. Inspired by the paradigm of human lifelong learning, we propose a new self-adaptive evolution framework for FSCL that enables continuous interaction with and adaptation to external environments. To exploit latent knowledge in large-scale models, we use an adaptive diffusion-based generator that not only implicitly captures the distribution of new fewshot samples but also produces more high-quality samples. To mitigate the inevitable variability in generation quality, we also use a reinforced sample selection module, comprising a generated sample explorer and a selection evaluator, which explicitly guides the retained distributions toward alignment with the large-scale models. Integrated with the continual model, these components are optimized in an iterative self-adaptive evolution framework, ensuring stable knowledge retention while improving adaptability to newly emerging classes. We validate our approach through experiments on three benchmarks, revealing its effectiveness in exploiting external distributions and achieving notable performance improvements.
Identity-preserving text-to-video generation (IPT2V) empowers users to produce diverse and imaginative videos with consistent human facial identity. Despite recent progress, existing methods often suffer from significant identity distortion under large facial pose variations or facial occlusions. In this paper, we propose FaithfulFaces, a pose-faithful facial identity preservation learning framework to improve IPT2V in complex dynamic scenes. The key of FaithfulFaces is a pose-shared identity aligner that refines and aligns facial poses across distinct views via a pose-shared dictionary and a pose variation-identity invariance constraint. By mapping single-view inputs into a global facial pose representation with explicit Euler angle embeddings, FaithfulFaces provides a pose-faithful facial prior that guides generative foundations toward robust identity-preserving generation. In particular, we develop a specialized pipeline to curate a high-quality video dataset featuring substantial facial pose diversity. Extensive experiments demonstrate that FaithfulFaces achieves state-of-the-art performance, maintaining superior identity consistency and structural clarity even as pose changes and occlusions occur.
Graph pooling is crucial for enlarging the receptive field and reducing computational costs in deep graph representation learning. In this work, we propose a simple but effective graph probabilistic pooling (GP-Pool) framework to facilitate graph feature learning. Instead of either deterministic selection or random dropping, we design a probabilistic subgraph sampling to reach an expected distribution by deducing a variational bound. Accordingly, a Bernoulli graph pooling (BernPool) is first derived to sample nodes together with the local structures, for which a learnable reference set is introduced to encode nodes into a latent expressive probability space. Hereby, the resultant BernPool captures salient graph substructures while possessing much diversity on sampled nodes due to its nondeterministic manner. For more controllable pooling, we derive the Poisson-distributed version (aka PoissonPool) from BernPool to explicitly cut the node quantity with less variables in variational learning. Furthermore, considering the complementarity of node sampling and clustering, we propose a hybrid graph pooling (HGP) paradigm to combine a compact subgraph (via BernPool/PoissonPool) and a coarsening graph (via clustering), to retain both representative substructures and global topology. Extensive experiments on multiple public graph classification datasets demonstrate that our GP-Pool is superior to various graph pooling methods and achieves state-of-the-art performance.
Multi-view learning aims to integrate multi-source information for a comprehensive data representation, which has gained widespread attention in image processing. Each view contains view-specific noise and joint features associated with other views, and thus exploring the specificity and consistency among views is a typical solution to deal with multi-view data for learning discriminative representations. In this paper, we present a theory-induced model, termed Adversarial Distribution Alignment Network (ADAN), which learns view-invariant features and alleviate the negative impact of view-specific noise. We first demonstrate the necessity of suppressing view-specific noise and capturing view-invariant features inspired by the theory of view generalization, and then derive two collaborative modules: a feature disentangler and an adversarial alignment module. In detail, the feature disentanglement separates view-specific noise and view-invariant features by minimizing the mutual information between them. Following this, a negative entropy is proposed to suppress the negative impact of view-specific noise. Meanwhile, the adversarial module uses the adversarial technique that can fit more complex data conformed to different distributions to adaptively align cross-view features so that features encoded in different views converge. Substantial experiments are constructed on multi-view datasets, demonstrating that ADAN can achieve more promising performance compared to other superior methods. Code is available at https://github.com/huangsuj/ADANet.
Human image animation has witnessed significant advancements, yet generating high-fidelity hand motions remains a persistent challenge due to their high degrees of freedom and motion complexity. While reinforcement learning from human feedback, particularly direct preference optimization, offers a potential solution, it necessitates the construction of strict preference pairs. However, curating such pairs for dynamic hand regions is prohibitively expensive and often impractical due to frame-wise inconsistencies. In this paper, we propose Implicit Preference Alignment (IPA), a data-efficient post-training framework that eliminates the need for paired preference data. Theoretically grounded in implicit reward maximization, IPA aligns the model by maximizing the likelihood of self-generated high-quality samples while penalizing deviations from the pretrained prior. Furthermore, we introduce a Hand-Aware Local Optimization mechanism to explicitly steer the alignment process toward hand regions. Experiments demonstrate that our method achieves effective preference optimization to enhance hand generation quality, while significantly lowering the barrier for constructing preference data.
Synthesizing realistic and diverse anomalous samples from limited data is vital for robust model generalization. However, existing methods struggle to reconcile fidelity and diversity, often hampered by distribution misalignment and overfitting, respectively. To mitigate this, we introduce Anomaly Preference Optimization (APO), a novel paradigm that reformulates anomaly generation as a preference learning problem. Central to our approach is an implicit preference alignment mechanism that leverages real anomalies as positive references, deriving optimization signals directly from denoising trajectory deviations without requiring costly human annotation. Furthermore, we propose a Time-Aware Capacity Allocation module that dynamically distributes model capacity along the diffusion timeline— prioritizing structural diversity during highnoise phases while enhancing fine-grained fidelity in low-noise stages. During inference, a hierarchical sampling strategy modulates the coherencealignment trade-off, enabling precise control over generation. Extensive experiments demonstrate that significantly outperforms existing baselines, achieving state-of-the-art performance in both realism and diversity.
Movie posters are vital for captivating audiences, conveying themes, and driving market competition in the film industry. While traditional designs are laborious, intelligent generation technology offers efficiency gains and design enhancements. Despite exciting progress in image generation, current models often fall short in producing satisfactory poster results. The primary issue lies in the absence of specialized poster datasets for targeted model training. In this work, we propose a Movie Posters DataSet (MPDS), tailored for text-to-image generation models to revolutionize poster production. As dedicated to posters, MPDS stands out as the first image-text pair dataset to our knowledge, composing of 373k+ image-text pairs and 8k+ actor images (covering 4k+ actors). Detailed poster descriptions, such as movie titles, genres, casts, and synopses, are meticulously organized and standardized based on public movie synopsis, also named movie-synopsis prompt. To bolster poster descriptions as well as reduce differences from movie synopsis, further, we leverage a large-scale vision-language model to automatically produce vision-perceptive prompts for each poster, then perform manual rectification and integration with movie-synopsis prompt. In addition, we introduce a prompt of poster captions to exhibit text elements in posters like actor names and movie titles. For movie poster generation, we develop a multi-condition diffusion network that takes poster prompt, poster caption, and actor image (for personalization) as inputs, yielding excellent results through the learning of a diffusion model. Experiments demonstrate the valuable role of our proposed MPDS dataset in advancing personalized movie poster generation. MPDS is available at https://anonymous.4open.science/r/MPDS-373k-BD3B
Semi-supervised referring expression segmentation (SS-RES) aims to achieve precise pixel-level language grounding under limited annotation, yet suffers from limited supervision and unreliable pseudo-labels when exploiting unlabeled image–text pairs. In this work, we propose Learning to Label, a reinforced self-evolving framework (L2L) that casts pseudo-label construction as a learnable decision-making process. To build foundational understanding, we leverage a multimodal large language model to extract semantic–spatial priors, which are instantiated as initial soft segmentation proposals and elevated—together with textual cues—into learnable guidance signals that condition a hierarchical segmentation network. To ensure stable learning, a reinforced pseudo-label selection is further formulated as an exploratory decision process that adaptively rewards high-utility pixel-level supervision based on multimodal priors and model predictions. This reinforced self-evolving loop enables joint optimization of the segmentation model and pseudo-labels, progressively enhancing label reliability under sparse supervision. Extensive experiments on RefCOCO, RefCOCO+, and RefCOCOg datasets demonstrate improvements over existing methods, validating its effectiveness and generalization.
Federated Graph Learning (FGL) facilitates privacy-preserving collaborative training of graph neural networks, yet homophily heterogeneity across subgraphs triggers optimization conflicts that degrade model generalization. Most existing solutions rely on multi-channel architectures to mitigate such conflict, which increase the burden on edge devices and lack theoretical convergence guarantees. To overcome these limitations, we propose FedGCM, a novel FGL framework with Group-oriented Conflict Mitigation, which aligns inconsistent optimization objectives via a tailored gradient surgery scheme. Specifically, FedGCM first divides clients into distinct groups based on their homophily levels, a strategy that precludes exhaustive client-to-client conflict assessments. To resolve inter-group interference, we develop RPGrad, a gradient surgery mechanism based on residual projection, which integrates synergistic knowledge while filtering inter-group conflicts. The refined updates are then transmitted in a group-wise fashion, effectively alleviating optimization conflicts induced by homophily heterogeneity without augmenting the client-side burden. Furthermore, we provide a formal theoretical analysis establishing the convergence of FedGCM. Extensive experiments on both homophilous and heterophilous graphs demonstrate that FedGCM consistently achieves advanced performance.
Multimodal emotion recognition (MER) leverages heterogeneous cues to overcome unimodal limitations, yet real-world applications often suffer from modality absence that degrades multimodal fusion effectiveness. While existing recovery-based methods address missing modalities, they face challenges including uncontrollable restoration processes, low-fidelity outputs, and inter-modal inconsistencies. In this work, we propose an Incomplete Multimodal Probability Flow Recovery (IM-PFR) framework to boost modality-missed emotion recognition. By unifying previous approaches into an autoencoder-like paradigm, we derive a Reversible Probability Flow Transformation (RPFT) mode through non-stochastic ordinary differential equations (ODEs), which combines the advantages of controllable sampling and continuous-time modeling. This creates a controllable Available $\leftrightarrow$ Prior$\leftrightarrow$Missing paradigm supporting multiple recovery scenarios (one-to-many, many-to-one, many-to-many), enhanced by two key components: i) a time-dependent aligner coordinating multimodal generation processes, and ii) a cross-modal high-order ODE solver reducing error accumulation. Extending from our prior works, this work enables tractable prior learning while ensuring inherent reversibility without explicit constraints. Extensive experiments on various MER datasets demonstrate state-of-the-art performance, with both quantitative metrics and qualitative analyses confirming superior recovery fidelity across diverse missing-modality conditions.
Decomposition-based text-guided video editing paradigm aims to utilize the layered neural atlas model to decompose the input video into foreground and background parts and edit the video in a divide-and-conquer manner, which is meaningful and improves the controllability of editing. However, they may suffer from some limitations: 1) high computational cost of per-video training (i.e, 7 similar to 8 hours for training a single atlas model, 2) foreground object deformation is restricted by the foreground opacity value, and 3) restricted flexibility in manipulating multiple objects. In this paper, we propose TraFrCo, a Training-Free Controllable Text-guided Video Editing framework to mitigate these challenges. Instead of training complex atlas models, our method leverages pre-trained segmentation to rapidly decompose videos into foreground and background parts. This allows users to perform independent edits on foreground objects using existing video diffusion editing models without affecting the environment. To ensure visual consistency, we introduce a training-free mechanism that effectively propagates information across frames to fill missing background regions caused by the segmentation-derived foreground masks and reconstructs the scene behind moving objects. Finally, the edited components are seamlessly composited by re-predicting the new foreground masks. In contrast to prior works, TraFrCo enables efficient, fine-grained manipulation of video content without the burden of training. Experimental results verify that our TraFrCo consistently reduces the costs of decomposing video and achieves superior text-guided video editing performance. Codes and video demos will be released at https://github.com/mdswyz/TraFrCo
Objects in remote sensing images often exhibit arbitrary orientations and significant scale variations, posing substantial challenges for accurate detection. Mainstream approaches typically rely on regressing oriented bounding box (OBB) based on predefined anchor boxes. However, issues such as angle discontinuity during regression significantly degrade performance. To enhance flexibility, recent works have proposed using a set of adaptive points instead of traditional bounding boxes, particularly within the RepPoints framework, which models object geometry and pose via point-based representations. In this work, we revisit RepPoints-based methods and identify critical issues, including partial point usage problem and training-testing inconsistency problem. To address these limitations, we propose a novel concept of the deformable bounding box (DBB) and design an anchor-free detector, DBB-Det, built upon this representation. Furthermore, we introduce a Normalized Spatial Constraint Loss to ensure scale-invariant penalization of outlier points, effectively balancing detection across varied object sizes. Extensive experiments on three challenging remote sensing benchmarks including DOTA, DIOR-R, and HRSC2016, demonstrate that our approach achieves superior accuracy and robustness. Moreover, the proposed deformable bounding box design is easily transferable to other RepPoints-based detectors.
Active contour models (ACMs) offer a powerful approach for extracting object boundaries from remote sensing imagery. However, most existing ACMs focus solely on minimizing contour energy while overlooking the vast contour evolution space and the inherent challenges in determining whether and when convergence is achieved. In this work, we propose a physics-driven conservative contour evolution (PCCE) method to progressively enhance contour learning. By thoroughly reformulating the general optimization objective from a physical energy perspective, we construct an energy-conservative contour model comprising potential energy, kinetic energy, and damping energy components-each governed by the law of energy conservation. In this formulation, potential energy encodes the geometric shape of building contours, kinetic energy accelerates the evolution, and damping energy converges the model. To optimize this contour model, we derive a partial differential equation that captures the evolution dynamics, and further discretize it to obtain an explicit recursive inference equation over the image domain. This evolution process is seamlessly embedded into an end-to-end deep neural network, where both the model parameters and contour evolution dynamics are learnable through backpropagation. Moreover, to address the high memory demands of ultrahigh-resolution (UHR) remote sensing images, we design a memory-efficient evolution paradigm; evolution parameters are learned from image slices and then aggregated and applied to produce full-image predictions through the contour evolution mechanism.
Protein structure prediction from a single sequence has drawn increasing attention due to the high computational costs associated with obtaining homologous information. Here we propose a two-dimensional geometric template diffusion method, named TDFold, to generate high-quality pairwise geometries (including pairwise distances and orientations). These are subsequently used for accurate and highly efficient three-dimensional protein structure prediction. Given a protein sequence, TDFold infers three-dimensional structure via a network architecture consisting of two stages: two-dimensional geometric template generation and sequence-geometry collaborative learning. TDFold presents three key advantages compared with existing protein language models (for example, ESMFold and OmegaFold) and homology-based methods (for example, AlphaFold2, AlphaFold3 and RoseTTAFold): better single-sequence-based prediction performance, lower resource consumption and higher efficiency in inference. This work demonstrates the model effectiveness on homology-insufficient datasets such as Orphan and Orphan25 and popular CASP benchmarks, introducing an alternative solution for single-sequence protein structure prediction. It also accelerates protein-related research, particularly for resource-limited universities and academic institutions.
Cross-domain class-incremental learning (CD-CIL) requires models to continuously acquire new classes across shifting domains while retaining previously learned knowledge. Existing approaches often entangle what to update with how to update, resulting in unstable adaptation and severe forgetting under domain shifts. Inspired by the hippocampal learning mechanism that separates rapid adaptation from stable consolidation, we propose Parameter-Masked Decoupled Optimization (PMDO) that disentangles what knowledge is adapted from how learning proceeds in cross-domain class-incremental learning. Specifically, we introduce a domain-aware knowledge decoupler that selectively adapts domain-relevant shared parameters, constraining incremental updates while preserving prior representations. To regulate how learning proceeds, we further design a stability-aware trajectory regulation that guides optimization along transferable and stable optimization trajectories, thereby reducing interference across domain transitions. As a result, PMDO enables effective cross-domain adaptation while mitigating catastrophic forgetting and maintaining long-term learnability. Extensive experiments across multiple benchmarks demonstrate the effectiveness of PMDO and its superiority over state-of-the-art methods.
Unsupervised reinforcement learning (URL) enables agents to adapt efficiently to novel tasks but relies on robust environmental modeling. Traditional world models enhance exploration but are prone to single-step errors, leading to distributional shifts. To address this challenge, we introduce the Diffusion Dynamic Model (DDM), incorporating diffusion models into the environmental modeling component of unsupervised reinforcement learning. During pre-training, DDM generates future observations conditioned on current observations and actions. These generated observations are used to design an intrinsic reward function that incentivizes the agent to explore diverse and high-uncertainty states. In the fine-tuning phase, DDM leverages its generative capabilities to perform data augmentation. By producing diverse, high-quality synthetic data, DDM expands the training dataset, enabling the agent to generalize better and adapt more rapidly to task-specific environments. Our approach is validated across three domains and twelve downstream tasks in the URLB benchmark, demonstrating superior exploration and adaptability compared to existing methods.
Weakly supervised video anomaly detection (WsVAD) aims to localize anomalies at the snippet level using only video-level annotations for training. Existing multiple instance learning approaches often fail to capture the diversity and contextual dynamics of abnormal events, resulting in limited localization accuracy. To address this, we propose Gaussian-grounded Contextual Hierarchical Inference (GCHI), a novel framework that learns discriminative Gaussian-grounded feature representations via conditional normalizing flows, models long-range temporal dependencies through contextual aggregation, and performs joint coarse-to-fine anomaly inference by aligning visual features with textual semantics and anomaly prototypes. Extensive experiments on UCF-Crime and XD-Violence datasets demonstrate that GCHI achieves state-of-the-art performance, validating its effectiveness and robustness for WsVAD task.
The burgeoning presence of Large Language Models (LLM) is propelling the development of personalized recommender systems. Most existing LLM-based methods fail to sufficiently explore the multi-view graph structure correlations inherent in recommendation scenarios. To this end, we propose a novel framework, Hypergraph Enhanced LLM Learning for multimodal Recommendation (HeLLM), designed to equip LLMs with the capability to capture intricate higher-order semantic correlations by fusing graph-level contextual signals with sequence-level behavioral patterns. In the recommender pre-training phase, we design a user hypergraph to uncover shared interest preferences among users and an item hypergraph to capture correlations within multimodal similarities among items. The hypergraph convolution and synergistic contrastive learning mechanism are introduced to enhance the distinguishability of learned representations. In the LLM fine-tuning phase, we inject the learned graph-structured embeddings directly into the LLM's architecture and integrate sequential features capturing each user's chronological behavior. This process enables hypergraphs to leverage graph-structured information as global context, enhancing the LLM's ability to perceive complex relational patterns and integrate multimodal information, while also modeling local temporal dynamics. Extensive experiments demonstrate the superiority of our proposed method over state-of-the-art baselines, confirming the advantages of fusing hypergraph-based context with sequential user behavior in LLMs for recommendation.