The growing computational demand for deep neural networks (DNNs) has raised concerns about their energy consumption and carbon footprint, particularly as the size and complexity of the models continue to increase. To address these challenges, energy-efficient hardware and custom accelerators have become essential. Additionally, adaptable DNNs are being developed to dynamically balance performance and efficiency. The use of these strategies became more common to enable sustainable AI deployment. However, these efficiency-focused designs may also introduce vulnerabilities, as attackers can potentially exploit them to increase latency and energy usage by triggering their worst-case-performance scenarios. This new type of attack, called energy-latency attacks, has recently gained significant research attention, focusing on the vulnerability of DNNs to this emerging attack paradigm, which can trigger denial-of-service (DoS) attacks. This article provides a comprehensive overview of current research on energy-latency attacks, categorizing them using the established taxonomy for traditional adversarial attacks. We explore different metrics used to measure the success of these attacks and provide an analysis and comparison of existing attack strategies. We also analyze existing defense mechanisms and highlight current challenges and potential areas for future research in this developing field. The GitHub page for this work can be accessed at https://github.com/hbrachemi/Survey_energy_attacks/.
Today, video content constitutes a significant portion of internet traffic. This video can be viewed by a wide range of devices with varying characteristics and under different network conditions. Video transcoding is a crucial mechanism for adapting video content to this diverse array of devices and bandwidth requirements while ensuring the best possible user experience. However, video transcoding is a computationally intensive process, requiring scalable infrastructure like cloud computing to efficiently handle the complexity and volume of tasks. In this paper, we propose a novel method to predict transcoding time across different types of platforms (CPU and GPU) and codecs (H.264/AVC, H.265/HEVC). Unlike existing approaches that focus mainly on CPU-based transcoding, the proposed model explicitly considers hardware-accelerated (GPU) transcoding, where accelerators significantly influence video transcoding performance in cloud computing. The predicted transcoding time can be utilized to optimize the scheduling of transcoding tasks in cloud computing, helping to ensure optimal load balancing and minimize total transcoding time while maintaining the highest video quality. The proposed solution consists of two essential phases: (i) dataset construction and (ii) model construction. The first phase involves video selection, segmentation, and video transcoding. The second phase focuses on analyzing the most important features that influence the prediction of transcoding time and developing a machine learning-based model for accurate video transcoding time prediction. Experimental results demonstrate that the XGBoost model achieves superior prediction accuracy across both software and hardware codecs, achieving a global coefficient of determination of R^2=0.993 when evaluated on the complete dataset, which includes video segments transcoded using H.264/AVC and H.265/HEVC codecs on CPU and GPU platforms. This performance represents an improvement of approximately 7.45
This paper presents the results of the Grand Challenge on Video Quality Assessment for Asymmetric Encoded Videos, held at QoMEX 2026 in Cardiff, UK. The challenge addresses the growing need for video quality metrics (VQM) capable of accurately predicting the perceptual quality of asymmetrically encoded videos, where saliency-driven or semantic-based encoding allocates different quality levels to different spatial regions. Participants were provided with the Sport-ROI dataset containing subjective quality scores and were invited to develop both full-reference (FR) and no-reference (NR) VQM models. We describe the challenge design, the dataset, the evaluation methodology, and summarize the submitted approaches and their performance.
Deepfake detection models often fail to generalize across unseen manipulations and real-world video degradations. Despite numerous proposed methods, prior work lacks a systematic assessment of how data augmentation strategies affect forensic robustness. In this study, we perform a comprehensive evaluation of 14 data augmentation techniques for deepfake detection. Using the Xception backbone across FaceForensics++, DFDC-P, and Celeb-DF, we show that data augmentation improves both in-domain accuracy and cross-dataset generalization. In particular, frequency-aware strategies, such as FourierMix, substantially improve accuracy by up to +2–3
Recent years have witnessed remarkable progress in developing Vision-Language Models (VLMs) capable of processing both textual and visual inputs. These models have demonstrated impressive performance, leading to their widespread adoption in various applications. However, this widespread raises serious concerns regarding user privacy, particularly when models inadvertently process or expose private visual information. In this work, we frame the preservation of privacy in VLMs as an adversarial attack problem. We propose a novel attack strategy that selectively conceals information within designated Region Of Interests (ROIs) in an image, effectively preventing VLMs from accessing sensitive content while preserving the semantic integrity of the remaining image. Unlike conventional adversarial attacks that often disrupt the entire image, our method maintains high coherence in unmasked areas. Experimental results across three state-of-the-art VLMs namely LLaVA, Instruct-BLIP, and BLIP2-T5 demonstrate up to 98
The rise of deep learning (DL) has increased computing complexity and energy use, prompting the adoption of application specific integrated circuits (ASICs) for energy-efficient edge and mobile deployment. However, recent studies have demonstrated the vulnerability of these accelerators to energy attacks. Despite the development of various inference time energy attacks in prior research, backdoor energy attacks remain unexplored. In this paper, we design an innovative energy backdoor attack against deep neural networks (DNNs) operating on sparsity-based accelerators. Our attack is carried out in two distinct phases: backdoor injection and backdoor stealthiness. Experimental results using ResNet-18 and MobileNet-V2 models trained on CIFAR-10 and Tiny ImageNet datasets show the effectiveness of our proposed attack in increasing energy consumption on trigger samples while preserving the model's performance for clean/regular inputs. This demonstrates the vulnerability of DNNs to energy backdoor attacks. The source code of our attack is available at: https://github.com/hbrachemi/energy_backdoor.
Cardiovascular diseases are a global health concern and developing automatic detection tools ensures fast diagnosis and helps improving the patient’s conditions. In this paper, we address the problem of PCG classification for the purpose of cardiovascular diseases detection, where we employ deep neural networks for classification. We pay particular attention to the audio segmentation modules. Specifically, we investigate two audio segmentation approaches: windowing and heat cycles. We design a hybrid CNN-LSTM neural network which operates on Mel-frequency cepstral coefficients of the PCG segments. We evaluate our contributions on the publicly available heart sounds dataset PhysioNet 2016. The obtained results are promising, provides useful insights on the PCG segmentation approach.
HTTP adaptive streaming (HAS) has emerged as a prevalent approach for over-the-top (OTT) video streaming services due to its ability to deliver a seamless user experience. A fundamental component of HAS is the bitrate ladder, which comprises a set of encoding parameters (e.g., bitrate-resolution pairs) used to encode the source video into multiple representations. This adaptive bitrate ladder enables the client’s video player to dynamically adjust the quality of the video stream in real-time based on fluctuations in network conditions, ensuring uninterrupted playback by selecting the most suitable representation for the available bandwidth. The most straightforward approach involves using a fixed bitrate ladder for all videos, consisting of pre-determined bitrate-resolution pairs known as one-size-fits-all . Conversely, the most reliable technique relies on intensively encoding all resolutions over a wide range of bitrates to build the convex hull , thereby optimizing the bitrate ladder by selecting the representations from the convex hull for each specific video. Several techniques have been proposed to predict content-based ladders without performing a costly, exhaustive search encoding. This article provides a comprehensive review of various convex hull prediction methods, including both conventional and learning-based approaches. Furthermore, we conduct a benchmark study of several handcrafted- and deep learning (DL)-based approaches for predicting content-optimized convex hulls across multiple codec settings. The considered methods are evaluated on our proposed large-scale dataset, which includes 300 UHD video shots encoded with software and hardware encoders using three state-of-the-art video standards, including AVC/H.264, HEVC/H.265, and VVC/H.266, at various bitrate points. Our analysis provides valuable insights and establishes baseline performance for future research in this field ( Dataset URL : https://nasext-vaader.insa-rennes.fr/ietr-vaader/datasets/br_ladder ).
Recently, with the growing popularity of mobile devices as well as video sharing platforms (e.g., YouTube, Facebook, TikTok, and Twitch), User-Generated Content (UGC) videos have become increasingly common and now account for a large portion of multimedia traffic on the internet. Unlike professionally generated videos produced by filmmakers and videographers, typically, UGC videos contain multiple authentic distortions, generally introduced during capture and processing by naive users. Quality prediction of UGC videos is of paramount importance to optimize and monitor their processing in hosting platforms, such as their coding, transcoding, and streaming. However, blind quality prediction of UGC is quite challenging, because the degradations of UGC videos are unknown and very diverse, in addition to the unavailability of pristine reference. Therefore, in this article, we propose an accurate and efficient Blind Video Quality Assessment (BVQA) model for UGC videos, which we name 2BiVQA for double Bi-LSTM Video Quality Assessment. 2BiVQA metric consists of three main blocks, including a pre-trained Convolutional Neural Network to extract discriminative features from image patches, which are then fed into two Recurrent Neural Networks for spatial and temporal pooling. Specifically, we use two Bi-directional Long Short-term Memory networks, the first is used to capture short-range dependencies between image patches, while the second allows capturing long-range dependencies between frames to account for the temporal memory effect. Experimental results on recent large-scale UGC VQA datasets show that 2BiVQA achieves high performance at lower computational cost than most state-of-the-art VQA models. The source code of our 2BiVQA metric is made publicly available at https://github.com/atelili/2BiVQA .
In today's digital landscape, video streaming holds an important role in internet traffic, driven by the pervasive use of mobile devices and the surge in streaming platform popularity. In light of this, the imperative to gauge energy consumption takes center stage, paving the way for eco-conscious and sustainable video streaming solutions with a minimal Carbon Dioxide (CO2) footprint. This paper meticulously examines the energy consumption and CO2 emissions of five popular open-source and fast video encoders: x264, x265, VVenC, libvpx-vp9, and SVT-AV1. These encoders are optimized software implementations of three video coding standards (AVC/H.264, HEVC/H.265, VVC/H.266) and two video formats (VP9andAV1).Toensureafaircomparison,wealsoassesscodingefficiencyacross these encoders at four distinct presets, applying three objective quality metrics. Additional factors like computing density and memory usage are considered. Our experiments employ the JVET-CTC video dataset, encompassing video sequences of diverse content and resolutions. Encoding is executed on an Intel x86 multi-core processor, while CO2 emissions are computed based on the energy mix data from a server situated in France, reflecting an average emission rate of 101 g/kWh. Our findings underscore that the x264 and SVT-AV1 encoders, especially at fast and faster presets, exhibit the lowest energy consumption and CO2 emissions. Notably, x264 boasts the most energy-efficient performance, yielding CO2 emissions of 0.28 g, 0.91 g, 2.07 g, and 9.74 g when encoding videos using faster, fast, medium, and slower presets, respectively. Furthermore, SVT-AV1 and VVenC encoders operating at a slower preset demonstrate superior coding efficiency, albeit at the cost of higher computational complexity and CO2 emissions of 60.5 g and 406 g, respectively. A salient observation from our study is that resolution and encoder presets serve as crucial parameters for curbing energy consumption and CO2 emissions, albeit with an inherent trade-off in video quality. Comprehensive results from this research are publicly accessible https://chachoutaieb.github.io/encoding energy co2.
Recent advances in generative models and the availability of large-scale benchmarks have made deepfake video generation and manipulation easier. Nowadays, the number of new hyper-realistic deepfake videos used for negative purposes is dramatically increasing, thus creating the need for effective deepfake detection methods. Although many existing deepfake detection approaches, particularly CNN-based methods, show promising results, they suffer from several drawbacks. In general, poor generalization results have been obtained under unseen/new deepfake generation methods. The crucial reason for the above defect is that CNN-based methods focus on the local spatial artifacts, which are unique for every manipulation method. Therefore, it is hard to learn the general forgery traces of different manipulation methods without considering the dependencies that extend beyond the local receptive field. To address this problem, this article proposes a framework that combines Convolutional Neural Network (CNN) with Vision Transformer (ViT) to improve detection accuracy and enhance generalizability. Our method, named HCiT, exploits the advantages of CNNs to extract meaningful local features, as well as the ViT's self-attention mechanism to learn discriminative global contextual dependencies in a frame-level image explicitly. In this hybrid architecture, the high-level feature maps extracted from the CNN are fed into the ViT model that determines whether a specific video is fake or real. Experiments were performed on Faceforensics++, DeepFake Detection Challenge preview, Celeb datasets, and the results show that the proposed method significantly outperforms the state-of-the-art methods. In addition, the HCiT method shows a great capacity for generalization on datasets covering various techniques of deepfake generation. The source code is available at: https://github.com/KADDAR-Bachir/HCiT
Abstract. Recent developments in advanced generative deep learning techniques have led to considerable progress in deepfake technology. CNN-based deepfake detection approaches have demonstrated superior performance. The ability to learn meaningful representations generated by convolutional multilayer nonlinear structures is the key to success. However, the black-box nature of such approaches has been a major concern for exploring hidden and complex characteristics as well as potential limitations of CNN-based models. To gain insights into the scope of the deepfake detection task, we investigate the effectiveness of handcrafted feature-based methods for deepfake video detection. First, we experiment with six top-performing handcrafted descriptors to extract the discriminating image features and then train SVMs on the extracted features to learn a suitable model. We also study the effect of selecting specific facial components on the detection performance. Specifically, we consider features extracted from the left eye, right eye, mouth, and entire face. Moreover, we propose a combination of these features and highlight the importance of this combination in terms of detection performance. Experimental results show that the SIFT feature descriptor achieves the best performance on deepfake videos generated by the neural texture technique, with a detection accuracy of 83.50%, which is better than deep learning-based methods. This is in contrast to the conventional understanding that deep learning methods systematically outperform handcrafted feature-based approaches. In addition, the obtained results on the FaceForensics++ dataset highlight the benefit of using some facial components to further boost the detection performance. Moreover, motivated by the effectiveness of the LBPTOP and SIFT in the deepfake detection task, we combined the LBPTOP and SIFT to best characterize the specific spatiotemporal inconsistencies commonly found in fake videos for boosting deepfake detection performance. Finally, we show the strengths and weaknesses of methods based on handcrafted features for deepfake detection and provide directions for future research.
This paper reports on the NTIRE 2023 Quality Assessment of Video Enhancement Challenge, which will be held in conjunction with the New Trends in Image Restoration and Enhancement Workshop (NTIRE) at CVPR 2023. This challenge is to address a major challenge in the field of video processing, namely, video quality assessment (VQA) for enhanced videos. The challenge uses the VQA Dataset for Perceptual Video Enhancement (VDPVE), which has a total of 1211 enhanced videos, including 600 videos with color, brightness, and contrast enhancements, 310 videos with deblurring, and 301 deshaked videos. The challenge has a total of 167 registered participants. 61 participating teams submitted their prediction results during the development phase, with a total of 3168 submissions. A total of 176 submissions were submitted by 37 participating teams during the final testing phase. Finally, 19 participating teams submitted their models and fact sheets, and detailed the methods they used. Some methods have achieved better results than baseline methods, and the winning methods have demonstrated superior prediction performance.
Recently, HTTP adaptive streaming (HAS) has become a standard approach for over-the-top (OTT)-based video streaming services due to its ability to provide smooth streaming. In HAS, stream representations are encoded to target a specific bitrate providing a wide range of operating bitrates known as the bitrate ladder. In the past, a fixed bitrate ladder approach for all videos has been widely used. However, such a method does not consider video content, which can vary considerably in motion, texture, and scene complexity. Moreover, building a per-title bitrate ladder based on an exhaustive encoding is quite expensive due to the large encoding parameter space. Thus, alternative solutions allowing accurate and efficient per-title bitrate ladder prediction are in great demand. On the other hand, self-attention-based architectures have achieved tremendous performance in large language models (LLMs) and particularly vision transformers (ViTs) in computer vision tasks. Therefore, this paper investigates ViT’s capabilities in building an efficient bitrate ladder without performing any encoding process. We provide the first in-depth analysis of the prediction accuracy and the complexity overhead induced by the ViTs model in predicting the bitrate ladder on a large and diverse video dataset. The source code of the proposed solution and the dataset will be made publicly available.
Recent studies have discovered that Deep Learning (DL) models are vulnerable to adversarial attacks in image classification tasks. While most studies have focused on DL models for image classification, only a few works have addressed this issue in the context of Image Quality Assessment (IQA). This paper investigates the robustness of different Convolutional Neural Network (CNN) models against adversarial attacks when used for an IQA task. We propose an adaptation of state-of-the-art image classification attacks in both targeted and untargeted modes for an IQA regression task. We also analyze the correlation between the perturbation’s visibility and the attack’s success. Our experimental results show that DL-based IQA methods are vulnerable to such attacks, with a significant decrease in correlation scores. Consequently, the development of countermeasures against such attacks is essential for improving the reliability and accuracy of DL-based IQA models. To support the principle of reproducible research and fair comparison, we make the codes publicly available on https://github.com/hbrachemi/IQA_AttacksSurvey.
The estimation of energy consumption has become vital in developing eco-friendly and sustainable video streaming solutions to monitor CO2 emissions. In this paper, we seek to evaluate and compare the energy consumption and CO2 emissions of the decoding process related to three popular video coding standards, namely AVC, HEVC, VVC, along with two video formats VP9, and AV1 through their real-time software decoders, including h264, hevc, VVdeC/OpenVVC, vp9, and libdav1d. The evaluation is conducted on two types of consumer hardware, desktop PC and laptop. To ensure a fair evaluation, we also assess the coding efficiency of software encoder implementations using three objective quality metrics. The experimental results revealed that the h264 decoder consumes the lowest energy and is associated with the lowest CO2 emissions compared to other decoders on both hardware platforms. On the other hand, the VVenC encoder enhances coding efficiency at the cost of increased decoding energy consumption and CO2 emissions, particularly noticeable in the case of the OpenVVC decoder. Meanwhile, x265/hevc achieves a compelling balance between coding efficiency and decoding energy consumption. The full results of this work are available at https://decodingenergy.github.io/decoding_energy_co2.html.
HTTP adaptive streaming (HAS) has emerged as a widely adopted approach for over-the-top (OTT) video streaming services, due to its ability to deliver a seamless streaming experience. A key component of HAS is the bitrate ladder, which provides the encoding parameters (e.g., bitrate-resolution pairs) to encode the source video. The representations in the bitrate ladder allow the client's player to dynamically adjust the quality of the video stream based on network conditions by selecting the most appropriate representation from the bitrate ladder. The most straightforward and lowest complexity approach involves using a fixed bitrate ladder for all videos, consisting of pre-determined bitrate-resolution pairs known as one-size-fits-all. Conversely, the most reliable technique relies on intensively encoding all resolutions over a wide range of bitrates to build the convex hull, thereby optimizing the bitrate ladder for each specific video. Several techniques have been proposed to predict content-based ladders without performing a costly exhaustive search encoding. This paper provides a comprehensive review of various methods, including both conventional and learning-based approaches. Furthermore, we conduct a benchmark study focusing exclusively on various learning-based approaches for predicting content-optimized bitrate ladders across multiple codec settings. The considered methods are evaluated on our proposed large-scale dataset, which includes 300 UHD video shots encoded with software and hardware encoders using three state-of-the-art encoders, including AVC/H.264, HEVC/H.265, and VVC/H.266, at various bitrate points. Our analysis provides baseline methods and insights, which will be valuable for future research in the field of bitrate ladder prediction. The source code of the proposed benchmark and the dataset will be made publicly available upon acceptance of the paper.
Over past decades, many image encryption algorithms have been proposed, among which we can cite the perceptual/selective encryption methods which have attracted wide attention. Such methods allow for adjusting the scrambling intensity, it is therefore essential to have a reliable visual security metric to adjust the scrambling intensity on the one hand and to evaluate the visual security of encrypted images on the other hand. Usually, these tasks are performed based on classical randomness-based measures or image quality assessment metrics. However, these methods have shown their inadequacy as a visual security metric, as they do not address content intelligibility, which represents an essential security requirement. Moreover, these methods are either dedicated to the prediction of visual security (VS) or visual quality (VQ), but not both. In this paper, we propose a no-reference (NR) visual security metric for perceptually encrypted images based on deep multi-task learning, which we dub the Multi-Task Visual Security (MTVS) metric. The proposed metric consists of one shared convolutional neural network (CNN) followed by two separate sub-networks of fully-connected (FC) layers, where one sub-network is responsible for predicting the VS score, while the other is for predicting the VQ score. Experiments were performed on two publicly perceptually encrypted image databases and the results show that the proposed metric yields superior performance on both VS and VQ prediction tasks. The source code and models are available at: https://github.com/Mamadou-Keita/MTVS.
Thanks to the remarkable advances in generative adversarial networks (GANs), it is becoming increasingly easy to generate/manipulate images. The existing works have mainly focused on deepfake in face images and videos. However, we are currently witnessing the emergence of fake satellite images, which can be misleading or even threatening to national security. Consequently, there is an urgent need to develop detection methods capable of distinguishing between real and fake satellite images. To advance the field, in this paper, we explore the suitability of several convolutional neural network (CNN) architectures for fake satellite image detection. Specifically, we benchmark four CNN models by conducting extensive experiments to evaluate their performance and robustness against various image distortions. This work allows the establishment of new baselines and may be useful for the development of CNN-based methods for fake satellite image detection.
Deep learning (DL) has shown great success in many human-related tasks, which has led to its adoption in many computer vision based applications, such as security surveillance systems, autonomous vehicles and healthcare. Such safety-critical applications have to draw their path to success deployment once they have the capability to overcome safety-critical challenges. Among these challenges are the defense against or/and the detection of the adversarial examples (AEs). Adversaries can carefully craft small, often imperceptible, noise called perturbations to be added to the clean image to generate the AE. The aim of AE is to fool the DL model which makes it a potential risk for DL applications. Many test-time evasion attacks and countermeasures, i.e., defense or detection methods, are proposed in the literature. Moreover, few reviews and surveys were published and theoretically showed the taxonomy of the threats and the countermeasure methods with little focus in AE detection methods. In this paper, we focus on image classification task and attempt to provide a survey for detection methods of test-time evasion attacks on neural network classifiers. A detailed discussion for such methods is provided with experimental results for eight state-of-the-art detectors under different scenarios on four datasets. We also provide potential challenges and future perspectives for this research direction.