Despite the remarkable process of talking-head-based avatar-creating solutions, directly generating anchor-style videos with full-body motions remains challenging. In this study, we propose Make-Your-Anchor+, a novel system necessitating only a one-minute video clip of an individual for training, subsequently enabling the automatic generation of anchor-style videos with precise torso and hand movements. Specifically, we finetune a proposed structure-guided diffusion model on input video to render 3D mesh conditions into human appearances. We adopt a two-stage training strategy for the diffusion model, effectively mapping movements with specific appearances to create digital avatars for online streamers, live shopping hosts, and other applications. To produce arbitrary long temporal video, we extract human motion information from video diffusion prior by adapting the frame-wise diffusion model to pretrained video diffusion weights with lower cost, and a simple yet effective batch-overlapped temporal denoising module is proposed to bypass the constraints on video length during inference. Finally, a novel identity-specific face enhancement module is introduced to improve the visual quality of facial regions in the output videos. Comparative experiments demonstrate the system’s effectiveness and superiority in visual quality, temporal coherence, and identity preservation, outperforming SOTA diffusion/non-diffusion methods.
Diffusion Transformers (DiTs) have exhibited robust capabilities in image generation tasks. However, accurate text-guided image editing for multimodal DiTs (MM-DiTs) still poses a significant challenge. Unlike UNet-based structures that could utilize self/cross-attention maps for semantic editing, MM-DiTs inherently lack support for explicit and consistent incorporated text guidance, resulting in semantic misalignment between the edited results and texts. In this study, we disclose the sensitivity of different attention heads to different image semantics within MM-DiTs and introduce HeadRouter, a training-free image editing framework that edits the source image by adaptively routing the text guidance to different attention heads in MM-DiTs. Furthermore, we propose a dual-token refinement module to refine text/image token representations for precise semantic guidance and accurate region expression. Experiments on multiple benchmarks demonstrate HeadRouter's performance in terms of editing fidelity and image quality. The code is available at https://github.com/ICTMCG/HeadRouter.
Unified image generation and editing models suffer from severe task interference in dense diffusion transformers architectures, where a shared parameter space must compromise between conflicting objectives (e.g., local editing v.s. subject-driven generation). While the sparse Mixture-of-Experts (MoE) paradigm is a promising solution, its gating networks remain task-agnostic, operating based on local features, unaware of global task intent. This task-agnostic nature prevents meaningful specialization and fails to resolve the underlying task interference. In this paper, we propose a novel framework to inject semantic intent into MoE routing. We introduce a Hierarchical Task Semantic Annotation scheme to create structured task descriptors (e.g., scope, type, preservation). We then design Predictive Alignment Regularization to align internal routing decisions with the task's high-level semantics. This regularization evolves the gating network from a task-agnostic executor to a dispatch center. Our model effectively mitigates task interference, outperforming dense baselines in fidelity and quality, and our analysis shows that experts naturally develop clear and semantically correlated specializations.
Multi-label image classification aims to recognize multiple object labels within an image. In the field of intelligent waste sorting, efficient classification can enhance the accuracy of robotic sorting. However, most existing waste classification tasks are single-label, and there is limited research on multi-label classification of urban kitchen waste, especially in complex backgrounds and diverse categories. To address this, we propose a Multi-Modal Semantic-Aware Graph (MM-SAG) framework, which includes a semantic-aware module designed for instance-level label relationship mining. The captured semantic features are then processed through graph convolution to generate a label correlation matrix, enhancing the efficiency and effectiveness of label correlation mining. To improve the integration of visual and linguistic modalities, we design an improved multi-head attention mechanism module. This module re-encodes and aligns visual and textual features, further enhancing feature extraction capabilities. Experimental results show that our proposed method achieves a mean Average Precision (mAP) of 83.1% on the MLKW dataset, delivering state-of-the-art performance. The method's strong generalization capability is also validated on public datasets VOC2007 and MS-COCO.
Segmentation of tubular structures in remote sensing imagery represents a domain of significant value for geographic information systems. However, achieving high-quality automated segmentation remains challenging due to morphological diversity and boundary ambiguity between tubular structures and their backgrounds. To address these challenges, we propose RLRAnet, an enhanced Unet-based segmentation architecture that incorporates boundary-aware mechanisms to improve tubular structure extraction. Specifically, we designed a reverse Attention Module that performs inverse calibration on low-level features within multi-scale skip connections, thereby enhancing boundary detail representation. Additionally, we introduced a reinforcement learning-based dynamic loss weight adjustment strategy that leverages a policy network to adaptively balance Soft Dice loss and boundary loss, achieving optimal reconciliation between global segmentation and local detail preservation. Extensive evaluations on two distinct remote sensing tubular structure datasets demonstrate that RLRAnet significantly outperforms multiple advanced segmentation architectures across IoU, Accuracy, Recall, and F1 Score, substantiating its superior performance in tubular structure segmentation tasks.
Both accuracy and timeliness are key factors in detecting fake news on social media. However, most existing methods encounter an accuracy-timeliness dilemma: Content-only methods guarantee timeliness but perform moderately because of limited available information, while social context-based ones generally perform better but inevitably lead to latency because of social context accumulation needs. To break such a dilemma, a feasible but not well-studied solution is to leverage social contexts (e.g., comments) from historical news for training a detection model and apply it to newly emerging news without social contexts. This requires the model to (1) sufficiently learn helpful knowledge from social contexts, and (2) be well compatible with situations that social contexts are available or not. To achieve this goal, we propose to absorb and parameterize useful knowledge from comments in historical news and then inject it into a content-only detection model. Specifically, we design the Comments ASsisted FakENews Detection method (CAS-FEND), which transfers useful knowledge from a comment-aware teacher model to a content-only student model and detects newly emerging news with the student model. Experiments show that the CAS-FEND student model outperforms all content-only methods and even comment-aware ones with 1/4 comments as inputs, demonstrating its superiority for early detection.
Magnetic particle imaging (MPI) is a promising technique for mapping magnetic nanoparticle (MNP) distributions within biological tissues. The system matrix (SM)-based reconstruction is a crucial research part in MPI. However, the time-consuming nature of SM measurements often requires repetition whenever there are changes in scan parameters, particle types, or environmental conditions. In this study, we proposed an iterative up-and-down sampling network based on pyramid pooling and attention mechanism for 3-D SM recovery (3-D-ISPAnet) to accelerate MPI calibration. Specifically, aiming at the smaller size and higher noise level of the MPI SM relative to the natural image, we firstly introduce a pyramid pooling module to better utilize global and local contextual information. Secondly, the deep relationship between low-resolution (LR) and high-resolution (HR) image pairs is efficiently captured by the convolutional block attention module (CBAM)-based iterative up-and-down sampling module. The experiments conducted on OpenMPI data demonstrate the excellent SM recovery capability of our proposed method. Average quantitative indicators of all test samples show that it has the best normalized root-mean-square error (nRMSE) and structure similarity index measure (SSIM), and the overall performance metrics of 3-D-ISPAnet were improved by approximately 2% to 10% over other methods. We believe that this research will enhance the practicality of MPI in biomedical applications and contribute to the future advancement of MPI technology.
Digital images of Chinas maps play a crucial role in map detection, particularly in ensuring national sovereignty, territorial integrity, and map compliance. However, there is currently no publicly available dataset specifically dedicated to problematic maps the CME dataset. Existing datasets primarily focus on general map data and are insufficient for effectively identifying complex issues such as national boundary misrepresentations, missing elements, and blurred boundaries. Therefore, this study creates a Problematic Map dataset that covers five key problem areas, aiming to provide diverse samples for problematic map detection technologies, support high-precision map compliance detection, and enhance map data quality and timeliness. This dataset not only provides essential resources for map compliance, national security monitoring, and map updates, but also fosters innovation and application of related technologies.
Organelle morphology and dynamics are closely linked to cellular function and fate, yet their relationships remain poorly defined across physiological and pathological contexts. Live-cell imaging enables the visualization of subcellular structures and dynamic processes but often requires extensive manual analysis, introducing variability and limiting reproducibility and throughput. Image segmentation partitions digital images into meaningful regions, facilitating the quantification of organelle morphology and molecular behavior for precise subcellular analysis. Herein, this review surveys recent advances in live-cell imaging segmentation algorithms across diverse organelles, from traditional thresholding-based methods to deep learning approaches that enhance accuracy and adaptability in complex biological environments. We discuss key challenges, including 3-dimensional imaging, multi-organelle segmentation, and generalization across diverse imaging modalities. We also highlight label-efficient strategies, synthetic data, and physics-guided modeling that reduce reliance on manual annotations and large annotated datasets. By advancing generalist models, these innovations improve quantitative cell biology, accelerate disease research, and drive therapeutic discovery, underscoring the transformative role of artificial intelligence in biomedical microscopy.
The unique nature of map data presents challenges for detecting key error areas in problematic maps, especially in terms of discontinuous boundary feature extraction and neglect of small target information. To address these issues, we propose a lightweight problematic map detection algorithm called YOLO-Map. First, to ensure the integrity of edge extraction and resistance to interference, we designed the Dual-branch Attention Convolution Module (DACM), which utilizes the synergistic effect of two branches to accurately identify national boundary regions. Next, the Multi-Path Feature Aggregation (MPFA) module adopts a bidirectional adaptive fusion strategy, enhancing recursive connections of multi-scale features and improving target localization accuracy. Additionally, we propose the Global Context Fusion Module (GCFM), which strengthens small target feature representation through a multi-branch collaborative attention mechanism. Experimental results show that YOLO-Map achieves an accuracy of 87.3% (mAP@.5) on the CME dataset, outperforming many larger models.
Nowadays, the proliferation of portraits or photographs containing human faces on the internet has created significant risks of illegal privacy collection and analysis by intelligent systems. Previous attempts to protect against unauthorized identification by face recognition models have primarily involved manipulating or adding adversarial perturbations to photos. However, it remains a challenge to balance privacy protection effectiveness and maintaining image visual quality. That is, to successfully attack real-world black-box face recognition models, significant manipulation is required for the source image, which will obviously damage the image visual quality. To address these issues, we propose an attribute-guided face identity protection (AG-FIP) approach that can protect facial privacy effectively without introducing meaningless or conspicuous artifacts into the source image. The proposed method involves mapping the images to latent space and subsequently implementing an adversarial attack through attribute editing. An attribute selection module followed by an attribute adversarially editing module is proposed to enhance the efficiency and effectiveness of adversarial attacks. Experimental results demonstrate that our approach outperforms SOTAs in terms of confusing black-box face recognition models, commercial face recognition APIs, and image visual quality.
This paper presents InvLatents, a novel framework for character animation that leverages latent inversion diffusion models to ensure consistent identity preservation across frames. Existing diffusion-based character animation methods often struggle with maintaining identity consistency due to the inherent randomness in the generation process. To address this issue, InvLatents introduces a latent inversion technique that incorporates target identity and pose guidance into the inference stage. By controlling different injection ratios in different branches, the method obtains richer identity information from the reference image. Additionally, a lightweight pose integration module is introduced to compensate for potential missing pose guidance. Experimental results on the TikTok dataset demonstrate that InvLatents achieves competitive performance compared to state-of-the-art approaches, effectively maintaining both identity and pose consistency without requiring additional training. The proposed method can be integrated as a plugin into other diffusion models, offering a promising solution for generating temporally coherent motion videos with consistent identity. Project page: https://github.com/SodaLee/InvLatents .
Personalized generation paradigms empower designers to customize visual intellectual property with the help of textual descriptions by adapting pre-trained text-to-image models on a few images. Recent studies focus on simultaneously customizing content and detailed visual style in images but often struggle with entangling the two. In this study, we reconsider the customization of content and style concepts from the perspective of parameter space construction. Unlike existing methods that utilize a shared parameter space for content and style learning, we propose a novel framework that separates the parameter space to facilitate individual learning of content and style by introducing "partly learnable projection" (PLP) matrices to separate the original adapters into divided sub-parameter spaces. A "break-for-make" customization learning pipeline based on PLP is proposed: we first break the original adapters into "up projection" and "down projection" for content and style concept under orthogonal prior and then make the entity parameter space by reconstructing the content and style PLP matrices by using Riemannian preconditioning to adaptively balance content and style learning. Experiments on various styles, including textures, materials, and artistic style, show that our method outperforms state-of-the-art single/multiple concept learning pipelines regarding content-style-prompt alignment. Code is available at https://github.com/ICTMCG/Break-for-make.
Personalized generation paradigms empower designers to customize visual intellectual properties with the help of textual descriptions by adapting pre-trained text-to-image models on a few images. Recent studies focus on simultaneously customizing content and detailed visual style in images but often struggle with entangling the two. In this study, we reconsider the customization of content and style concepts from the perspective of parameter space construction. Unlike existing methods that utilize a shared parameter space for content and style learning, we propose a novel framework that separates the parameter space to facilitate individual learning of content and style by introducing “partly learnable projection” ( PLP ) matrices to separate the original adapters into divided sub-parameter spaces. A “ break-for-make ” customization learning pipeline based on PLP is proposed: we first break the original adapters into “up projection” and “down projection” for content and style concept under orthogonal prior and then make the entity parameter space by reconstructing the content and style PLPs matrices by using Riemannian precondition to adaptively balance content and style learning. Experiments on various styles, including textures, materials, and artistic style, show that our method outperforms state-of-the-art single/multiple concept learning pipelines regarding content-style-prompt alignment. Code is available at: https://github.com/ICTMCG/Break-for-make.
Recent advances in diffusion models have enhanced multimodal-guided visual generation, enabling customized subject insertion that seamlessly "brushes" user-specified objects into a given image guided by textual prompts. However, existing methods often struggle to insert customized subjects with high fidelity and align results with the user's intent through textual prompts. In this work, we propose "In-Context Brush", a zero-shot framework for customized subject insertion by reformulating the task within the paradigm of in-context learning. Without loss of generality, we formulate the object image and the textual prompts as cross-modal demonstrations, and the target image with the masked region as the query. The goal is to inpaint the target image with the subject aligning textual prompts without model tuning. Building upon a pretrained MMDiT-based inpainting network, we perform test-time enhancement via dual-level latent space manipulation: intra-head "latent feature shifting" within each attention head that dynamically shifts attention outputs to reflect the desired subject semantics and inter-head "attention reweighting" across different heads that amplifies prompt controllability through differential attention prioritization. Extensive experiments and applications demonstrate that our approach achieves superior identity preservation, text alignment, and image quality compared to existing state-of-the-art methods, without requiring dedicated training or additional data collection.
Most fine-grained visual recognition methods endeavor to directly locate discriminative regions in intricate environments, but tend to overlook the object’s holistic structure, which may lead to misclassification due to overemphasizing incorrect areas. In this paper, we propose a coarse-to-fine paradigm, which prioritizes locating holistic structural regions of the target object, followed by a gradual search to locate discriminative areas. Specifically, we first design the “look into object” module to locate the areas encompassing the target’s holistic structure using prior information. Subsequently, without introducing additional parameters, we design a partial focus searching module to enhance feature representations of discriminative regions within the target’s structural composition. Ultimately, we segregate the foreground components from the original image, attaining a more precise characterization of the target. Furthermore, we demonstrate the practical application potential of our model in real-world industries through our self-constructed DHU-Fine-grained-6000 dataset. Comparative experiments on three public datasets indicate that the superiority of our approach over many recent methods and holds promising application potential in industrial production processes.
Box supervised instance segmentation (BSIS) aims to achieve an effective trade-off between annotation costs and model performance by solely relying on bounding box annotations during training process. However, we observe that BSIS model is bottlenecked by the intricate objective under limited guidance, and tends to sacrifice segmentation capability in order to effectively recognize multiple instances. To boost the BSIS model's perceptual ability for object shape and contour, we introduce MISA, that is, MIning Saliency-Aware semantic prior from a well-optimized box supervised semantic segmentation (BSSS) network, and incorporating cross-model guidance into the learning process of BSIS. Specifically, we first design a Frequency-Space Distillation (FSD) module to extract assorted salient prior knowledge from BSSS model, and perform cross-model alignment for transfering the prior to BSIS model. Furthermore, we introduce Semantic-Enhanced Pairwise Affinity (SEPA), which borrows the object perceptual ability of BSSS model to emphasize the contribution of salient objects for pairwise affinity, providing more accurate guidance for the BSIS network. Extensive experiments show that our proposed MISA consistently surpasses the existing state-of-the-art methods by a large margin in the BSIS scenario.
In this study, we revisit the fundamental setting of face-swapping models and reveal that only using implicit supervision for training leads to the difficulty of advanced methods to preserve the source identity. We propose a novel reverse pseudo-input generation approach to offer supplemental data for training face-swapping models, which addresses the aforementioned issue. Unlike the traditional pseudo-label-based training strategy, we assume that arbitrary real facial images could serve as the ground-truth outputs for the face-swapping network and try to generate corresponding input < source, target > pair data. Specifically, we involve a source-creating surrogate that alters the attributes of the real image while keeping the identity, and a target-creating surrogate intends to synthesize attribute-preserved target images with different identities. Our framework, which utilizes proxy-paired data as explicit supervision to direct the face-swapping training process, partially fulfills a credible and effective optimization direction to boost the identity-preserving capability. We design explicit and implicit adaption strategies to better approximate the explicit supervision for face swapping. Quantitative and qualitative experiments on FF++, FFHQ, and wild images show that our framework could improve the performance of various face-swapping pipelines in terms of visual fidelity and ID preserving. Furthermore, we display applications with our method on re-aging, swappable attribute customization, cross-domain, and video face swapping. Code is available under https://github.com/ICTMCG/CSCS.
Personalized generation paradigms empower designers to customize visual intellectual properties with the help of textual descriptions by tuning or adapting pre-trained text-to-image models on a few images. Recent works explore approaches for concurrently customizing both content and detailed visual style appearance. However, these existing approaches often generate images where the content and style are entangled. In this study, we reconsider the customization of content and style concepts from the perspective of parameter space construction. Unlike existing methods that utilize a shared parameter space for content and style, we propose a learning framework that separates the parameter space to facilitate individual learning of content and style, thereby enabling disentangled content and style. To achieve this goal, we introduce "partly learnable projection" (PLP) matrices to separate the original adapters into divided sub-parameter spaces. We propose "break-for-make" customization learning pipeline based on PLP, which is simple yet effective. We break the original adapters into "up projection" and "down projection", train content and style PLPs individually with the guidance of corresponding textual prompts in the separate adapters, and maintain generalization by employing a multi-correspondence projection learning strategy. Based on the adapters broken apart for separate training content and style, we then make the entity parameter space by reconstructing the content and style PLPs matrices, followed by fine-tuning the combined adapter to generate the target object with the desired appearance. Experiments on various styles, including textures, materials, and artistic style, show that our method outperforms state-of-the-art single/multiple concept learning pipelines in terms of content-style-prompt alignment.