Large vision-language models (LVLMs) have recently shown strong potential for industrial anomaly detection (IAD) by providing image-level anomaly judgments and interpretable defect reasoning. However, current LVLM-based IAD methods still struggle to produce precise pixel-level anomaly maps from generated language judgments. We aim to achieve precise pixel-level localization while using language as guidance rather than letting it dominate the visual response. Specifically, we propose OPD-IAD, an evidence-privileged dense on-policy self-distillation framework for LVLM-based IAD. OPD-IAD distills privileged defect evidence onto the model's own on-policy judgment trajectory, enabling the final generated judgment to be learned under dense supervision rather than treated only as a textual answer. The resulting judgment serves as a semantic condition for dense anomaly perception. To turn this condition into dense visual evidence, we introduce Language-guided Visual Anchoring, which uses a judgment reforward to re-encode the image and question under the final-judgment condition into semantic anchors and contrasts them with dense visual features through a contrastive heatmap head to generate anomaly maps. The language judgment therefore provides compact semantic guidance, while dense visual features remain the basis for pixel-level scoring, allowing language to guide anomaly localization without letting language quality directly dictate the pixel-level response. Extensive experiments show that OPD-IAD achieves the best overall performance among LVLM-based IAD methods, leading on most image-level, pixel-level, and QA metrics.
Existing generative models for unsupervised anomalous sound detection are limited by their inability to fully capture the complex feature distribution of normal sounds, while the potential of powerful diffusion models in this domain remains largely unexplored. To address this challenge, we propose a novel framework, TLDiffGAN, which consists of two complementary branches. One branch incorporates a latent diffusion model into the GAN generator for adversarial training, thereby making the discriminator's task more challenging and improving the quality of generated samples. The other branch leverages pretrained audio model encoders to extract features directly from raw audio waveforms for auxiliary discrimination. This framework effectively captures feature representations of normal sounds from both raw audio and Mel spectrograms. Moreover, we introduce a TMixup spectrogram augmentation technique to enhance sensitivity to subtle and localized temporal patterns that are often overlooked. Extensive experiments on the DCASE 2020 Challenge Task 2 dataset demonstrate the superior detection performance of TLDiffGAN, as well as its strong capability in anomalous time-frequency localization.
Mamba architecture has demonstrated excellent performance in various visual tasks. It outperforms many CNN and transformer based architectures due to its long-sequence modeling capability and linear complexity. However, when we aggregate cost volume for occluded regions, Mamba module fails to properly match features of the same category due to its token-independent parameter generation mechanism. To address this issue, we propose a novel aggregation scheme called Global-aware Mamba Cost volume Aggregation, which can enhance matching ability of pixels within the same category by introducing prompts with global semantic information into the discretization step size of state space model. Additionally, we introduce an uncertainty-guided loss to enhance the network’s self-optimization ability for poorly predicted regions during iteration. Experiments show that our design achieves outstanding results on the Sintel and KITTI datasets.
State-of-the-art 3D human pose estimation methods often rely on time-domain modeling of joint trajectories, yet struggle to distinguish smooth global motion from transient local jitter. To overcome this, we propose FreePose, a frequency-decoupled framework that separates motion into low-frequency and high-frequency components via discrete wavelet transform. We design two specialized modules: the Low-Frequency Coefficient Interaction Network (LFCI-net), which captures continuous and smooth motion patterns by applying temporal self-attention to low-frequency coefficients; and the High-Frequency Coefficient Interaction Network (HFCI-net), which extracts fine-grained local motion through multi-scale dilated convolutions and adaptively suppresses noise via learnable gated attention. Compared to time-domain baselines, FreePose reduces MACs by 50%, improves accuracy, and produces smoother motion. Experiments on Human3.6M and MPI-INF-3DHP validate the effectiveness of the proposed method under indoor and semi-controlled evaluation settings. Code is available at https://github.com/Lxg-233/FreePose.
Low-light image enhancement using RAW data holds significant promise due to its rich, unprocessed information. Nevertheless, existing RAW-based low-light enhancement methods are plagued by a critical trade-off: single-stage models introduce color artifacts, while multi-stage pipelines suffer from error propagation and detail loss. In this work, we argue that overcoming these challenges requires a holistic approach that jointly addresses spatial color fidelity and frequency-domain detail preservation. To this end, we propose RAWFormer, a new architecture featuring two key innovations. We first introduce a powerful color calibration module that leverages cascaded hybrid attention and a multi-scale adaptive feed-forward network to capture both local and global contexts and ensure accurate color rendition. Complementing this, we design a full-spectrum frequency loss that guides the network to restore high-frequency textures typically neglected by spatial losses. Our extensive experiments show that RAWFormer sets a new state-of-the-art on the SID and MCR benchmark, outperforming previous methods with PSNR gains of 0.21 dB and 0.32 dB on Sony and Fuji subsets of SID.
Motivation Structure-based drug design (SBDD) aims to generate ligand molecules that tightly bind to specific protein targets, a critical step in drug discovery. Diffusion models have shown promise for this task, yet existing methods struggle to effectively incorporate protein-ligand interaction priors during generation. Most approaches rely on protein-specific structural priors that remain fixed throughout generation, limiting molecular diversity and failing to capture the dynamic interplay between protein pockets and ligand atoms, which is essential for achieving high binding affinity.Results We propose DPDiff, a disentangled prior-conditioned diffusion model for protein-specific 3D molecular generation. DPDiff introduces two complementary interaction prior networks that capture geometry-based spatial interactions and sequence-based interactions robust to structural noise. During generation, the model dynamically extracts interaction priors using intermediate diffusion predictions and adaptively fuses them via a time-dependent adapter. A disentangled denoising network balances prior guidance with generative flexibility. Experiments on the CrossDocked2020 dataset demonstrate that DPDiff generates molecules with more realistic 3D structures and state-of-the-art binding affinities, achieving an average Vina Dock score of -8.58 and a high affinity ratio of 69.4%, outperforming existing methods while maintaining favorable drug-likeness and synthetic accessibility.Availability and implementation The source code of DPDiff is available at https://github.com/ZerinHwang03/DPDiff.
Despite remarkable advancements in multimodal large language models (MLLMs), their fine-grained visual understanding is constrained by reliance on pure textual supervision. To unify understanding and generation capabilities, unified autoregressive multimodal models introduce visual supervision; however, they impair multimodal understanding due to the effects of visual feature discretization and orthogonality between image-text loss gradients. In this paper, we observe that pixel-level image patches and textual tokens coexist in raw high-dimensional spaces with inherent input symmetry. Motivated by this insight, we propose UVU, a novel vision-language unified autoregressive framework that eschews vector quantization. It uniquely employs continuous visual encoding for lossless representation of visual inputs and proposes a large-scale iterative hierarchical clustering algorithm to construct a pixel-level visual codebook, thereby extending the vocabulary for unified supervision and enabling autoregressive generation of pixel-level image tokens alongside textual tokens. UVU effectively synergizes pixel-level visual perception with semantic-level visual understanding, internalizing visual generation capabilities and, for the first time, unlocking the facilitative role of visual supervision in enhancing understanding. Extensive experiments across multiple tasks demonstrate that MLLMs are capable of achieving superior multimodal understanding performance under the supervised learning paradigm of UVU.
BACKGROUND AND OBJECTIVE:Long QT syndrome (LQTS) has received increasing attention because of its association with sudden cardiac death. Traditional LQTS diagnosis relies on standard 12-lead electrocardiograms (ECGs) in hospital settings, which may limit timely detection outside clinical settings. With the proliferation of portable single-lead ECG devices, single-lead diagnostic solutions offer new potential for efficient arrhythmia screening. However, single-lead systems are inherently limited by information sparsity, and the complex, variable waveform morphology of LQTS further challenges accurate detection. METHODS:To address this, we propose a Single-Lead LQTS Detection (SLLD) framework based on knowledge distillation from a 12-lead teacher model. By leveraging an efficient knowledge transfer mechanism, the proposed framework enhances the representation learning capability of the single-lead network for LQTS-related ECG classification. Specifically, we design a feature fusion module to integrate local waveform morphology and temporal dependencies of ECG signals, and use a first-order temporal derivative augmentation strategy to supplement the raw signal with local dynamic variation information, thereby enabling the model to better characterize subtle waveform and temporal changes in LQTS-related ECG signals. RESULTS:Experimental results on public datasets demonstrate that the SLLD framework achieves improved performance over state-of-the-art methods. Specifically, SLLD achieved an AUC of 0.903 and an average F1-score of 0.706 on the Georgia dataset. Notably, our method improved the LQTS-specific F1-score by 4.9 percentage points over the strongest evaluated baseline, indicating the utility of the proposed framework. CONCLUSION:The SLLD framework mitigates the inherent information sparsity of single-lead signals by distilling diagnostic knowledge from multi-lead systems, demonstrating the potential of knowledge-distillation-based single-lead LQTS screening.
Accurate interpretation of electrocardiogram (ECG) remains challenging due to the scarcity of labeled data and the high cost of expert annotation. Self-supervised learning (SSL) offers a promising solution by enabling models to learn expressive representations from unlabeled signals. Existing ECG SSL methods typically rely on either contrastive learning or reconstructive learning. However, each approach in isolation provides limited supervisory signals and suffers from additional limitations, including non-physiological distortions introduced by naive augmentations and trivial correlations across multiple leads that models may exploit as shortcuts. In this work, we propose CoRe-ECG, a unified contrastive and reconstructive pretraining paradigm that establishes a synergistic interaction between global semantic modeling and local structural learning. CoRe-ECG aligns global representations during reconstruction, enabling instance-level discriminative signals to guide local waveform recovery. To further enhance pretraining, we introduce Frequency Dynamic Augmentation (FDA) to adaptively perturb ECG signals based on their frequency-domain importance, and Spatio-Temporal Dual Masking (STDM) to break linear dependencies across leads, increasing the difficulty of reconstructive tasks. Our method achieves state-of-the-art performance across multiple downstream ECG datasets. Ablation studies further demonstrate the necessity and complementarity of each component. This approach provides a robust and physiologically meaningful representation learning framework for ECG analysis.
Face super-resolution (FSR), which reconstructs high-quality faces from low-resolution (LR) inputs, has achieved impressive performance. However, the information available in the input images is severely limited, and most FSR methods encounter challenges in accurately reconstructing textures and preserving identity information. To this end, we propose a novel Code Diffusion paradigm, named Code-Diffusion FSR (CoDiFSR), exploiting global priors to guide super-resolution. We propose the Facial Prior Codebook (FPC), which quantizes and stores the high-quality Facial Prior Embeddings (FPEs) from high-resolution (HR) faces as prior knowledge. This codebook design effectively prevents over-fitting while ensuring that these codes remain expressive and beneficial for the FSR task. Then, through knowledge distillation training, we distill the HR priors into our student model. With LR inputs, we put forward a Code Diffusion model to refine the low-quality FPE towards its high-quality counterpart with few diffusion steps. This refined FPE eventually serves as the modulation signal, providing strong prior guidance for the FSR backbone. Extensive experiments demonstrate that our model outperforms advanced methods in both visual quality and identification accuracy.
Monocular 3D object detection has received considerable attention for its simplicity and low cost. Existing methods typically follow conventional 2D detection paradigms, first locating object centers and then predicting 3D attributes via neighboring features. However, these approaches mainly focus on local information, which may limit the model's global context awareness and result in missed detections, as the global context provides semantic and spatial dependencies essential for detecting small objects in cluttered or occluded environments. In addition, due to large variation in object scales across different scenes and depths, inaccurate receptive fields often lead to background noise and degraded feature representation. To address these issues, we introduce MonoASRH, a novel monocular 3D detection framework composed of Efficient Hybrid Feature Aggregation Module (EH-FAM) and Adaptive Scale-Aware 3D Regression Head (ASRH). Specifically, EH-FAM employs multi-head attention with a global receptive field to extract semantic features and leverages lightweight convolutional modules to efficiently aggregate visual features across different scales, enhancing small-scale object detection. The ASRH encodes 2D bounding box dimensions and then fuses scale features with the semantic features aggregated by EH-FAM through a scale-semantic feature fusion module. The scale-semantic feature fusion module guides ASRH in learning dynamic receptive field offsets, incorporating scale information into 3D position prediction for better scale-awareness. Extensive experiments on the KITTI and Waymo datasets demonstrate that MonoASRH achieves state-of-the-art performance. The code and model are released at https://github.com/WYFDUT/MonoASRH
LiDAR Simultaneous Localization and Mapping (SLAM) has been pivotal in various domains, including digital twins, geographic surveying, and autonomous mobile robotics. However, achieving an optimal balance between high precision and computational efficiency remains a significant challenge. In this letter, we propose MLS-SLAM, a multi-level submap guided 3D LiDAR SLAM method that incorporates three levels of submaps (local-frame submap, rotation-aware submap, and loop-optimized submap) and consists of four key modules: LiDAR Odometry, Local Frame Integration, Rotation-Aware Submap Integration, and Loop-Optimized Submap Integration. Firstly, the LiDAR Odometry module provides an initial pose for each LiDAR scan. Subsequently, the Local Frame Integration module performs continuous local refinement with a sliding window and constructs local-frame submaps. Then, the Rotation-Aware Submap Integration module aggregates local-frame submaps into rotation-aware submaps based on the trajectory. Finally, the Loop-Optimized Submap Integration module finds loops between rotation-aware submaps and performs global pose graph optimization. The performance of MLS-SLAM has been rigorously evaluated on public and self-collected datasets. Experiments show that MLS-SLAM achieves state-of-the-art precision and real-time performance at 10 Hz on datasets with diverse scales and environments.
While Multimodal Large Language Models (MLLMs) excel in cross-modal reasoning, they often struggle to perceive fine-grained details in complex high-resolution images. Recent training-free methods address this through image scaling and localized cropping. However, applying these manipulations indiscriminately introduces computational redundancy for simple queries and can degrade accuracy by truncating essential global context or introducing irrelevant background noise. To this end, we propose LazyMCoT, a dynamic and training-free framework that adaptively allocates visual grounding efforts based on sample difficulty. The framework features an Adaptive Routing mechanism that evaluates predictive uncertainty using first-token statistics from a single forward pass. This efficiently bypasses confident cases while ensuring the recall of difficult samples via conformal calibration. For these challenging cases, a Collaborative Grounding module integrates the inherent cross-modal attention of the model with an external visual expert through a two-stage refinement process. This refinement process generates a precise localized display to recover small or occluded targets. Extensive experiments across diverse benchmarks demonstrate that LazyMCoT rivals training-based approaches by simultaneously improving reasoning accuracy and reducing average inference latency. Our code is availble at https://github.com/TencentBAC/LazyMCoT.
Low-light image enhancement (LLIE) is essential for robust visual perception in tasks such as autonomous driving and surveillance. However, frame-based approaches often suffer from noise amplification, motion blur, and loss of fine details under extreme illumination. Event cameras, which asynchronously record pixel-wise brightness changes, provide high dynamic range sensing and excel in challenging lighting conditions. To exploit their complementary strengths, we propose Event-RGB Fusion Transformer (ERFormer), which integrates event streams with RGB frames through event-guided feature enhancement and adaptive brightness control based on Retinex decomposition. ERFormer introduces two key innovations: (1) an SNR-guided fusion module that adaptively combines spatial RGB features with event-derived structural cues to preserve multi-scale context, specifically enhancing skip-layer features in the U-Net architecture to preserve multiscale contextual information, and (2) an adaptive brightness control module that dynamically adjusts illumination via learnable orthogonal embeddings. Extensive experiments on the SDE and SDSD datasets show that ERFormer surpasses state-of-the-art methods by 0.26 dB and 0.61 dB in PSNR, respectively, while reducing FLOPs by 63.72%, delivering superior edge preservation and fewer artifacts under extreme lighting.
Rapid advances in audio-video (AV) generation have enabled high-fidelity synthesis with synchronized sound, particularly for human-related scenarios involving speech and interactions. Yet evaluation for AV generation remains at an early stage, with only a few coarse-grained benchmarks for human-related scenarios and relying on limited preset evaluations with generic multimodal LLMs, leading to inaccurate assessments of model capabilities. To address these issues, we introduce AVBench, a fully automated benchmark tailored for human-centric AV generation. AVBench is built on two key designs for comprehensive and accurate evaluation: (i) Human-centric and fine-grained metrics. AVBench integrates ten evaluation dimensions designed for human-centered real-world scenarios, covering visual quality, audio quality, and multi-level consistency across modalities. These practical metrics capture human-related details that existing benchmarks often overlook. (ii) Specialized evaluators via preference learning. To address the lack of specialized training data, we construct large-scale supervision by transforming real-world videos into diverse training pairs with controlled perturbations. After fine-tuning on this high-quality dataset, the evaluators learn to reliably detect subtle cross-modal inconsistencies. Crucially, instead of producing discrete textual judgment, AVBench derives continuous evaluation scores from the model's prediction confidence on binary decisions. This probabilistic scoring mechanism enables a more reliable assessment than traditional VQA-style evaluation and aligns closely with human judgment. Taken together, AVBench offers automated evaluation for AV generation, demonstrates strong potential for data filtering, and serves as a differentiable reward signal for Reinforcement Learning from Human Feedback (RLHF).
The creation of cinematic-quality animal effects necessitates the precise modeling of muscle and fur dynamics, a process that remains both labor-intensive and computationally expensive within traditional production workflows. While generative diffusion models have shown promise in diverse artistic workflows, their capacity for high-fidelity animal simulation remains largely unexploited. We present MoZoo, a generative dynamics solver that bypasses conventional refinement to synthesize high-fidelity animal videos from coarse meshes under multimodal guidance. We propose Role-Aware RoPE (RAR-RoPE) which employs role-based index remapping to synchronize motion alignment while decoupling reference information via fixed temporal offsets. Complementing this, Asymmetric Decoupled Attention partitions the latent sequence to enforce a unidirectional information flow, effectively preventing feature interference and improving computational efficiency. To address the scarcity of high-quality training data, we introduce MoZoo-Data, a synthetic-to-real pipeline that leverages a rendering engine and an inverse mapping approach to construct a large-scale dataset of paired sequences. Furthermore, we establish MoZooBench, a comprehensive benchmark with 120 mesh-video pairs. Experimental results demonstrate that MoZoo achieves high-fidelity fur simulation across diverse animal skeletons and layouts, preserving superior temporal and structural consistency.
Recently, two-stage 3D human pose estimation using monocular cameras has gained significant attention. However, the inherent uncertainty in the upscaling process from 2D to 3D often compromises the accuracy of deterministic methods. To address this, we propose a novel diffusion-based refinement framework (DRPose) which models the uncertainty during the upscaling process by introducing stochastic noise to the initially predicted 3D poses. This approach facilitates the generation of more realistic predictions through iterative refinement with multiple noise samples, ultimately producing multi-hypothesis predictions that better align with ground truth. Our framework incorporates two key components: a Graph Convolution Transformer module (SGCT), which integrates scaling and displacement adjustments based on conditional information with a joint temporal-spatial feature separation mechanism, and a Pose Refinement Module (PRM), which balances the initial and refined poses. This design allows DRPose to effectively refine pose estimation for both individual frames and sequential data. Furthermore, our framework establishes new benchmarks for performance in both frame2frame and seq2frame scenarios. Extensive experiments demonstrate that our method achieves state-of-the-art performance on the Human3.6M and MPI-INF-3DHP datasets. Notably, when applied to the current state-of-the-art single-frame 3D pose extractor, our multi-hypothesis optimization achieves an 18.8% reduction in Mean Per Joint Position Error (MPJPE) and a 16.9% reduction in Procrustes MPJPE (P-MPJPE). Code is available at https://github.com/KHB1698/DRPose.
Emerging multi-modal world models attempt to jointly generate videos across diverse modalities (e.g., RGB, depth, and mask), yet they fail to fully exploit the rich priors of existing foundation models. We propose M^2-REPA, the first representation alignment method tailored for multi-modal video generation. Our key insight is that foundation models trained on different modality spaces naturally capture distinct domain-specific priors, acting as complementary "experts." Specifically, we first decouple modality-specific features from the diffusion model's intermediate representations, then align each with its corresponding expert foundation model. To this end, we design two synergistic objectives: a multi-modal representation alignment loss that enforces feature-to-expert matching, and a modality-specific decoupling regularization that encourages complementarity across different modalities. This design enables joint optimization, fully exploiting priors from multiple foundation models. Extensive experiments demonstrate that our method significantly outperforms baselines in visual quality and long-term consistency.