
Face clustering aims to group unlabeled face images into identity-consistent clusters and is essential for large-scale face retrieval and dataset construction. However, existing graph-based methods often suffer from false positive connections and limited recall in complex feature spaces. This paper introduces a robust face clustering framework that addresses these challenges through two key components. First, the proposed local identity relation estimation network (LIRENet) verifies edge reliability by modeling both local context and identity-level interactions. Second, a neighborhood edge extension module adaptively expands the graph using shared neighbor constraints and similarity thresholds. The combination of these components enhances connectivity prediction while preserving cluster quality. Experimental results on MS-Celeb-1M, IJB-B, and DeepFashion demonstrate consistent improvements over previous state-of-the-art methods on both closed- and open-set benchmarks.
In the field of computer vision, accurate estimation of 3D human poses and shapes plays a vital role in applications such as motion analysis, virtual try-on, and human-computer interaction. While most existing methods rely on monocular RGB images and deep learning regression networks, they often suffer from ambiguities caused by occlusions, self-similasrity, and missing visual cues. Leveraging publicly available multi-view image datasets offers an effective way to mitigate these limitations and improve reconstruction accuracy.In this work, we propose a novel framework for 3D human reconstruction that jointly lever-ages multi-view consistency and iterative refinement of SMPL parameters. Specifically, our method takes images from multiple viewpoints as input, extracts SMPL pose and shape parameters from each view using a CNN-based regressor, and integrates them through a multi-view parameter fusion strategy. To further refine the fused parameters, we introduce an Online Evolutive Learning(OEL) mechanism, which applies a SMPLify-inspired optimization process to iteratively minimize a reprojection-based evaluation loss without requiring ground-truth 3D annotations. This iterative optimization loop progressively enhances structural plausibility and alignment consistency across views. The im-proved parameters are then used to optimize the CNN regressor in an online evolutive manner, forming a feedback loop that continuously improves the model generalization ability.
Image inpainting, the task of synthesizing visually plausible content in missing or occluded regions of an image, has seen substantial progress in the context of 2D real-world imagery. However, existing methods often struggle to generalize to 3D virtual environments due to fundamental differences in geometric structure, rendering pipelines, and the diversity of visual styles and abstraction levels inherent in virtual scenes. To address these limitations, we propose a novel inpainting framework specifically adapted for 3D virtual environments. Our approach builds upon conventional inpainting architectures but introduces domain-specific enhancements to better capture the structural and semantic characteristics of virtual scenes. We evaluate our method across multiple widely-used 3D virtual environment datasets and additional in-the-wild virtual samples to assess its generalizability. Experimental results demonstrate that our model not only achieves competitive performance in filling missing regions but also enables practical applications such as scene editing and content creation, thereby validating the effectiveness and utility of the proposed framework.
Low-light photography is one of the most common use cases in the current smartphone industry. Today, multiple smartphones offer some form of night-mode photography that allows images to be captured in low-light conditions without the use of flash. However, most solutions that exist today depend on increasing exposure times and capturing multiple frames, which leads to long waiting times along with other issues such as motion blur. While other deep learning solutions exist, they are too heavy to be integrated in a smartphone. In this paper, we propose a lightweight low-light enhancement method that can work with a single frame, a custom loss function tailored for low-light enhancement and a multi-stage training process to enhance our model’s quality. Our testing demonstrates that our strategy can yield outputs that are comparable to SOTA techniques while requiring a significantly lesser latency and memory footprint, making them ideal for smartphones.
Video object detection (VOD) is vital in edge intelligence applications such as surveillance, autonomous systems, and wearable devices. Although high-performance still-image detectors such as YOLO are commonly employed in VOD tasks, applying detection to every video frame has redundant computation and a higher energy cost, making such a solution less attractive for resource-constrained platforms. This paper proposes MA-YOLO (Motion-Assisted YOLO), a lightweight VOD framework tailored for edge environments. Instead of inferring on every frame, MA-YOLO executes complete YOLO detection only on sparse keyframes and propagates the results to intermediate frames using motion information derived from precomputed H. 264 motion vectors and geometric offsets from reference detections. We introduce a lightweight, XGBoost-based decision module tailored to each geometric offset regression, realizing efficient detection propagation from the keyframes to non-keyframes. Experiments on the ImageNet VID dataset demonstrate that MA-YOLO reduces inference cost while maintaining competitive accuracy, offering a practical and efficient solution for edge-based video analysis.
More and more multimedia data, such as images and videos, is transmitted over digital networks and stored or shared in the Cloud. For reasons of confidentiality or secret information, it is increasingly necessary to protect multimedia content directly. Although many image obscuration methods have been developed to protect the semantic content of images, few of them are both reversible and non-visible, and therefore detectable visually or by trained classifiers. In this paper, we propose a new image obscuration method based on variational autoencoders and a secret key, that transforms images from their source class into a target class, in a non-visible and reversible way, allowing the original image content to be recovered. In the experimental results we present whether the obscured images generated by our method are detectable.
Forests are critical ecosystems that sustain biodiversity and regulate global environmental processes. Effective monitoring of these ecosystems is essential for conservation and management. Traditional techniques, including infrared and thermal imaging, often fail to capture the complexity of wildfire dynamics. To address these limitations, we propose a novel unsupervised segmentation framework, i.e. CAE-CME framework that integrates a Convolutional Autoencoder (CAE) with a novel Cluster Mask Enhancer (CME). A key innovation of our CAE model is the incorporation of a Color-Aware Regularization (CAR) loss, which acts as a contextual prior, by prioritizing red-intense pixels indicative of fire regions, thereby integrating both visual and contextual cues to enhance segmentation accuracy. In addition, a channel attention mechanism is also incorporated within the CAE to emphasize the most relevant feature representations, improving the model’s focus on fire-related patterns. Further, the proposed approach also integrates clustering techniques and a novel Cluster-based Image Masking technique within the Cluster Mask Enhancer (CME) module. Using clustering algorithm, it facilitates an efficient means of unsupervised image segmentation from its latent representations. The results show the superior performance (89.98% IoU score) of our proposed CAE-CME model compared with the state-of-the-art models.
While recent advances in artificial intelligence have accelerated the development of pathology image analysis, most existing models operate in a black-box manner, relying solely on visual features of images and lacking interpretability and explainability. To address this limitation, we propose a novel approach that leverages Vision-Language Models (VLMs) to enhance both interpretability and predictive performance in pathology image classification. Specifically, we construct a task-specific word pool composed of expert-defined pathology terms and generate a concept-aware image representation based on the semantic similarity between images and textual descriptions. This enables the integration of pathological knowledge into visual representations. Experimental results show that our method outperforms not only conventional CNN and Transformer-based models but also existing pathology-specific VLMs in classification performance. Furthermore, it allows model predictions to be interpreted using meaningful pathology terms and provides a quantitative means to analyze the model’s decision-making process based on similarity distributions between images and pathology concepts.
Advancements in remote sensing technology and deep learning techniques have paved the way for accurate aerial object detection in urban environments. However, object detection in these settings remains challenging due to dense scenes, small and occluded objects, and high variability across geographic domains. To tackle these challenges, we propose a tri-axial scaling framework for aerial object detection that systematically improves performance along three dimensions: model size, dataset size and quality, and inference strategy. First, we explore the use of larger backbone architectures to enhance feature representation. Second, we apply diffusion-based data augmentation and balanced class sampling to improve training data diversity and address class imbalance. Third, we incorporate test-time augmentation and ensemble models to increase robustness during inference. Our solution ranks first on the leaderboard in the IEEE ICIP 2025 - CADOT challenge. The source code and pretrained models are available at https://github.com/yjwong1999/Double_J_CADOT_Challenge.
In this work, we propose a single-shot transient imaging framework that combines sparsity-based signal modeling with compressive frequency-domain sampling. The temporal response function of a scene encodes rich information about depth and light transport, and its accurate reconstruction is critical for various imaging applications. To enable transient recovery from a limited number of measurements, we exploit learned sparse representations in an optimized dictionary basis. We compare multiple sampling strategies in the frequency domain and show that both transient profiles and depth maps can be reconstructed under highly compressed acquisition. Notably, we achieve full transient reconstruction using only 16 modulation frequencies, based on real correlation functions, enabling practical singleshot acquisition through spatial frequency multiplexing.
Waste management and maintenance of urban infrastructure are central challenge for modern smart cities. Image-based re-identification systems can help to automate condition monitoring and to build intelligent solutions. However, most existing image-based re-identification research targets persons and vehicles, leaving other objects such as urbane elements largely unexplored. To close this gap, we adapt a state-of-the-art re-identification model to the Urban Elements ReID Challenge 2025, which encompasses infrastructure elements like containers, crosswalks, and rubbish bins. The introduced model (i) fuses intermediate backbone features to enrich global embeddings with fine-grained, low-level cues, and (ii) incorporates view-specific embeddings that exploit the direction of travel during image recording to enhance robustness to viewpoint changes. Combined with a post-hoc re-ranking method, our system achieves a mAP of 28.1% and ranks second on the public leaderboard.
Modern image sensors frequently suffer from various types of defective pixels that critically degrade imaging quality, including isolated point defects, spatially clustered defects, and linear row/column defects. Existing detection and correction algorithms typically focus on single defect types and require manual parameter tuning, limiting their practical applicability across diverse imaging scenarios. This paper presents TRIMED-DPDC, a novel unified algorithm based on triple-stage median filtering for comprehensive detection and correction of all defect types without manual parameter optimization. The algorithm employs an innovative 3 x 5 neighborhood window configuration that provides enhanced spatial coverage while maintaining computational efficiency. Through a three-stage hierarchical filtering process, horizontal filtering, regional filtering, and diagonal filtering, the algorithm generates a comprehensive reference set for adaptive threshold determination. Experimental validation on the NWPU VHR-10 dataset demonstrates superior performance across all defect categories: the algorithm achieves detection sensitivity values of 99.8%, 99.7%, 98.0%, and 99.9% for point defects, cluster defects, column defects, and row defects, respectively. The peak signal-to-noise ratio improvements reach 5.18 dB, 3.23 dB, 4.68 dB, and 2.68 dB for the corresponding defect types, with consistent structural similarity index enhancements indicating excellent preservation of image details.
With the increasing availability of AI-based image generators, it has become remarkably easy to create realistic synthesized images. While this technological progress enables many creative and practical applications, it also raises serious concerns over misuse, such as the spread of misinformation and fraudulent activities. As a result, reliable detection of synthesized images has become critical. Existing detection methods often rely on supervised learning using labeled datasets that include both authentic and synthesized images, which limits their generalization to unseen AI generators. In this paper, we propose a novel one-class detection method that leverages only real images during training. Our approach is based on exploiting the Bayer pattern, a characteristic signature of digital camera sensors. We incorporate a Variational Autoencoder (VAE) to effectively extract statistical differences in pixel variances introduced by the demosaicing process. Furthermore, we design our method to be invariant to different Bayer pattern variants. Experimental evaluations showed that our approach outperforms other state-of-the-art methods in detecting authentic and synthesized images.
Recent advancements in text-to-video and video-to-audio generative AI (GenAI) have spurred excitement around their ability to potentiate human creativity. However, the simulated imagery and audio produced by state-of-the-art GenAI models may contain artifacts and anomalies that diminish their realism. Our multimodal experiment investigates 1) the relative weighting of auditory versus visual information humans utilize to distinguish real user-generated content (UGC) from AI-generated content (AIGC) and 2) which visual artifacts produced by GenAI are most conspicuous to human observers. Observers completed a 2-interval forced-choice (2IFC) task with eye tracking, choosing which of two sequentially presented video clips was more realistic. Observers were more accurate at discerning real from fake imagery than real from fake audio. Eye movement analysis revealed that top-down motion-related artifacts, such as violations of biological motion, were more conspicuous than bottom-up artifacts, such as the intermingling of real-world textures with data compression artifacts.
The growing prevalence of machine-driven visual data consumption underscores the need to meet the unique requirements of Video Coding for Machines (VCM). In this paper, we propose to enhance the coding efficiency of Versatile Video Coding (VVC) for machine consumption by adaptively adjusting the dynamic range of the input luma channel prior to encoding. The visual input is characterized using the introduced input analyzer that predicts the optimal dynamic range and provides the corresponding 1) luma down-scaling factors applied before encoding and 2) luma up-scaling factors used after decoding to restore the dynamic range. Our input analyzer is implemented as a lightweight neural network. For the network training, we introduce a training framework incorporating a codec proxy module that enables end-to-end optimization by simulating a conventional non-differentiable video codec. The proposed method has been evaluated as part of the conventional VVC pipeline, where VVC test model (VTM) is used for encoding and decoding. Our experimental results show that integrating the proposed solution into the pipeline improves coding efficiency by up to 28.0% on image datasets and up to 45.4% on video dataset for object detection tasks.
The Urban Elements ReID Challenge was created as a contribution to smart city development and infrastructure monitoring. It evaluates methods for the re-identification of urban elements, such as waste containers, rubbish bins, and crosswalks. The dataset includes annotated instances recorded over time along repeated trajectories of the same routes, in both forward and inverse directions, capturing variations in viewpoint, weather, and season. Evaluation results are split between a public leaderboard (47% of the test set) and a private leaderboard (53% of test set), revealed postcompetition for final ranking, promoting generalization and reducing overfitting to the public set. The baseline reaches 22% and 15% mean average precision on the public and private leaderboard. In contrast, top participants achieved up to 31% mean average precision public and 25% private, surpassing the baseline and highlighting the complexity of the task, which still poses significant challenges. A total of 12 teams participated in the challenge with 253 valid submissions.
Quantization is widely used to reduce the computational and memory demands of modern deep neural networks, enabling their deployment on resource-constrained edge devices. Most quantization methods assume a simple and direct relationship between local quantization error (e.g., layer-wise $L_{2}$ error) and global task performance degradation. However, we empirically demonstrate that this assumption does not hold in practice, as the relationship between local quantization error and task-level loss is nonlinear and varies significantly across layers and architectures. Our comprehensive empirical analysis covers both convolutional and transformer-based models, as these architectures also underlie many multi-modal generative frameworks. Moreover, local quantization errors exhibit a monotonic relationship with task performance increase in most cases, we show that accurate mixed-precision bit allocation requires more than monotonicity-it demands precise estimation of quantization-induced loss. We propose a practical empirical method to directly measure and estimate these task-level losses efficiently.
Facial image restoration is an essential task in human-centric visual enhancement, as facial regions convey rich identity cues and are highly sensitive to perceptual inconsistencies. These characteristics make face restoration fundamentally different from general image restoration. In this work, we present a robust blind face restoration framework built upon the generative power of diffusion models. Our method incorporates personalized priors from user photo galleries and extracts 3D facial geometry along with facial surface to recover the subject’s physical identity from degraded inputs. This integration enables our system to handle a wide range of facial degradations, particularly excelling in challenging scenarios where existing blind restoration techniques often fail.
Fluorescence Lifetime Imaging Microscopy (FLIM) is a powerful tool for investigating biological and physiological processes down to the subcellular level, by leveraging the temporal properties of weak fluorescence signals. Among photon detection techniques, Time-Correlated Single-Photon Counting (TCSPC) is the gold standard for achieving high temporal resolution and sensitivity. However, traditional TCSPC systems historically suffer from slow acquisition speed, especially for large-field imaging. Wide-field acquisition using parallel detector arrays increases hardware complexity, while single-pixel detectors offer a simpler alternative but rely on spatial scanning, which prolongs acquisition times. In this context, Single-Pixel Cameras (SPCs) combined with Compressive Sensing (CS) have emerged as a cost-effective solution, significantly reducing the number of required measurements. However, acquisition speed in SPC-FLIM is fundamentally limited by pile-up effects inherent to conventional TCSPC systems, which restrict photon count rates to below 5% of the laser repetition frequency. In this work, we demonstrate that integrating a custom ultra-low dead time detection module as SPC into a wide-field FLIM system, along with a novel pile-up correction algorithm, enables video-rate imaging at 20 frames per second achieving photon count rates exceeding 200% of the excitation frequency, with minimal fluorescence lifetime estimation error.
Video Anomaly Detection (VAD) is vital for enhancing surveillance safety but faces challenges from scarce anomalies, complex temporal patterns, and real-time requirements. We propose SAVES, an efficient VAD framework that integrates multi-scale depthwise separable convolutions, compact temporal pooling, and lightweight attention mechanisms to achieve robust detection with minimal computational over-head. Using the X3D backbone, SAVES captures short-term motion and long-range dependencies effectively.On the UCF-Crime dataset, it achieves ROC-AUC scores of 92.64%(1.4M parameters) at 45 FPS and 91.7% (32K parameters) at 78 FPS, enabling real-time applications.Ablation studies and robustness evaluations confirm SAVES’s balance of accuracy, efficiency, and practical deployment suitability for scalable surveillance systems.