3D Gaussian Splatting (3DGS) has recently emerged as a promising contender to Neural Radiance Fields (NeRF) in 3D scene reconstruction and real-time novel view synthesis. 3DGS outperforms NeRF in training and inference speed but has substantially higher storage requirements. To remedy this downside, we propose POTR, a post-training 3DGS codec built on two novel techniques. First, POTR introduces a novel pruning approach that uses a modified 3DGS rasterizer to efficiently calculate every splat's individual removal effect simultaneously. This technique results in 2-4x fewer splats than other post-training pruning techniques and as a result also significantly accelerates inference with experiments demonstrating 1.5-2x faster inference than other compressed models. Second, we propose a novel method to recompute lighting coefficients, significantly reducing their entropy without using any form of training. Our fast and highly parallel approach especially increases AC lighting coefficient sparsity, with experiments demonstrating increases from 70
Generative AI has made text-guided inpainting a powerful image editing tool, but at the same time a growing challenge for media forensics. Existing benchmarks, including our text-guided inpainting forgery (TGIF) dataset, show that image forgery localization (IFL) methods can localize manipulations in spliced images but struggle in fully regenerated (FR) images, while synthetic image detection (SID) methods can detect fully regenerated images but cannot perform localization. With new generative inpainting models emerging and the open problem of localization in FR images remaining, updated datasets and benchmarks are needed. We introduce TGIF2, an extended version of TGIF, that captures recent advances in text-guided inpainting and enables a deeper analysis of forensic robustness. TGIF2 augments the original dataset with edits generated by FLUX.1 models, as well as with random non-semantic masks. Using the TGIF2 dataset, we conduct a forensic evaluation spanning IFL and SID, including fine-tuning IFL methods on FR images and generative super-resolution attacks. Our experiments show that both IFL and SID methods degrade on FLUX.1 manipulations, highlighting limited generalization. Additionally, while fine-tuning improves localization on FR images, evaluation with random non-semantic masks reveals object bias. Furthermore, generative super-resolution significantly weakens forensic traces, demonstrating that common image enhancement operations can undermine current forensic pipelines. In summary, TGIF2 provides an updated dataset and benchmark, which enables new insights into the challenges posed by modern inpainting and AI-based image enhancements. TGIF2 is available at https://github.com/IDLabMedia/tgif-dataset.
We present Generative Anchored Fields (GAF), a generative model that learns independent endpoint predictors J (noise) and K (data) rather than a trajectory predictor. The velocity field v=K-J emerges from their time-conditioned disagreement. This factorization enables Transport Algebra: algebraic operation on learned {(J_n,K_n)}_n=1^N heads for compositional control. With class-specific K_n heads, GAF supports a rich family of directed transport maps between a shared base distribution and multiple modalities, enabling controllable interpolation, hybrid generation, and semantic morphing through vector arithmetic. We achieve strong sample quality (FID 7.5 on CelebA-HQ 64× 64) while uniquely providing compositional generation as an architectural primitive. We further demonstrate, GAF has lossless cyclic transport between its initial and final state with LPIPS=0.0. Code available at https://github.com/IDLabMedia/GAF
Background: Patients (pts) with endocrine therapy (ET)-resistant, hormone receptor+ (HR+), HER2– metastatic breast cancer (mBC) have a poor prognosis and may derive less benefit when treated with ET + a cyclin-dependent kinase 4/6 inhibitor (CDK4/6i) in the first-line (1L) setting. PIK3CA mutations occur in 35–40% of pts with HR+, HER2– mBC and are associated with poor prognosis. In the INAVO120 (NCT03006172) study, the addition of inavolisib, a PI3Kα inhibitor and mutant degrader, to fulvestrant plus palbociclib (a CDK4/6i) conferred a substantially longer progression-free survival (PFS) improvement in pts with PIK3CA-mutated (PIK3CAmut), HR+, HER2– mBC in the 1L, ET-resistant setting. This analysis aimed to evaluate baseline characteristics, PIK3CA mutation prevalence, and treatment outcomes in pts with HR+, HER2– mBC treated in the 1L, real-world (RW) setting, including those pts who met further INAVO120 criteria. Methods: This is a retrospective, observational study of de-identified, electronic health record-derived data from the United States Flatiron Health Network (thus is not human subjects research, which would have required institutional review board assessment/approval). Pts selected for inclusion in the study had recurrent HR+, HER2– mBC diagnosed between 2015 and 2023, and started 1L treatment within 90 days of mBC diagnosis. A cohort of pts with recurrent disease, prior ET, and any 1L treatment for mBC was defined (‘broad cohort’), as well as ‘INAVO120-like’ populations (subdivided based on receipt of fulvestrant + palbociclib, fulvestrant + CDK4/6i, or ET + CDK4/6i) that included further criteria such as ET resistance, Eastern Cooperative Oncology Group Performance Status 0–2, no bone-only disease, and no elevated HbA1c or glucose levels at baseline. Results: The broad cohort included 7096 pts, of whom, 922 met the further INAVO120 criteria (overall INAVO120-like population). All pts in the overall INAVO120-like population had received 1L CDK4/6i + ET; 514/922 pts received CDK4/6i + fulvestrant, of whom 339/514 received palbociclib as the CDK4/6i partner. 39–41% of the INAVO120-like populations (212/922, 120/514, and 82/339) tested positive for PIK3CA mutations. Fast progressors (relapse within ≤6 months [mo] of 1L treatment start) were more frequent in the INAVO120-like populations (35–38%) vs the broad cohort (25%). Race/ethnicity distribution, body mass index at mBC diagnosis, duration of adjuvant treatment and use of different CDK4/6is were similar between the ‘INAVO120-like’ populations and the broad cohort. Compared with the broad cohort, INAVO120-like pts were more likely to have visceral metastases (75–78% vs 48%), stage 3 disease at initial diagnosis (32–37% vs 29%) and a shorter time to mBC diagnosis (3.9–4.8 years vs 5.7 years). The INAVO120-like populations also had shorter RW median PFS (8.2–8.5 vs 13.0 mo), shorter RW median time to chemotherapy (16.2–21.7 vs 33.0 mo), and shorter RW median overall survival (27.0–29.7 vs 36.8 mo) compared with the broad cohort. PIK3CA mutation testing at any point was more frequent in pts in the INAVO120-like populations compared with the broad cohort (59–61% vs 45%); however, the proportion of pts tested for mutations prior to 1L treatment was numerically lower, but generally similar across cohorts: 9.5–11.8% (INAVO120-like) vs 12.4% (broad cohort). The PIK3CA mutation rates were similar across cohorts (39–41% vs 41%). Conclusions: These RW data from a large US, community-based oncology network suggest that pts meeting the INAVO120 criteria have poor outcomes with 1L ET+ CDK4/6i treatment, regardless of the type of ET and CDK4/6i partner. Pts with PIK3CAmut, HR+, HER2– mBC may derive more benefit from therapies targeting the PI3K pathway (since PIK3CA mutations are a predictive biomarker of response), as demonstrated in the INAVO120 study. Citation Format: Peter Lambert, Eirini Thanopoulou, Preet Dhillon. Real-world observational study of patients with endocrine-resistant, hormone receptor+, HER2- metastatic breast cancer in the first-line setting: Patient characteristics, PIK3CA mutation prevalence, treatments, and clinical outcomes [abstract]. In: Proceedings of the San Antonio Breast Cancer Symposium 2024; 2024 Dec 10-13; San Antonio, TX. Philadelphia (PA): AACR; Clin Cancer Res 2025;31(12 Suppl):Abstract nr P4-07-27.
Images and videos allow us to explore places and connect to people all around the world, in the present or past. What if we could break through the glass screen in front of us and step into those camera captures. Although challenging, in recent years, light field technology has developed some promising techniques, such as 3D Gaussian Splatting, that is able to render high-quality views of a scene reconstructed using only camera captures. However, creating the scene model takes a significant amount of time and compute power, which makes it unviable for the multimedia industry which outputs terabytes of new content daily. In this paper, we present a method of speeding up the modeling process, not by optimizing the training, but by initializing the pipeline with an already semi-finished reconstruction. This is done by estimating the depth maps of the camera images, fusing them and converting this to a dense set of Gaussian splats which already closely resembles the scene. Afterwards, the default training process is applied to fine-tune and quickly synthesize new high-quality views. We show that our method on average, after 1000 iterations, improves PSNR by +1.27 dB, SSIM by +0.065 and LPIPS by -0.10, compared to the default initialization.
Real-time video encoding requires efficient and accurate quality metrics to optimize performance under strict computational and latency constraints. Traditional low-complexity metrics such as PSNR and SSIM often fall short in perceptual alignment, while accurate metrics such as VMAF are too computationally intensive for real-time deployment.We present x265-pVMAF, a low-complexity perceptual quality metric integrated into the x265 encoding loop. By leveraging machine learning and efficiently extracted encoder features, x265-pVMAF bridges the gap between computational efficiency and perceptual accuracy. It replicates VMAF predictions with a correlation of 0.99, while delivering a 37× speed-up. These results establish x265-pVMAF as a practical solution for real-time video quality assessment in next-generation encoding workflows.
Deepfakes have raised significant concerns due to their potential to spread false information and compromise the integrity of digital media. Current deepfake detection models often struggle to generalize across a diverse range of deepfake generation techniques and video content. In this work, we propose a Generative Convolutional Vision Transformer (GenConViT) for deepfake video detection. Our model combines ConvNeXt and Swin Transformer models for feature extraction, and it utilizes an Autoencoder and Variational Autoencoder to learn from latent data distributions. By learning from the visual artifacts and latent data distribution, GenConViT achieves an improved performance in detecting a wide range of deepfake videos. The model is trained and evaluated on DFDC, FF++, TM, DeepfakeTIMIT, and Celeb-DF (v2) datasets. The proposed GenConViT model demonstrates strong performance in deepfake video detection, achieving high accuracy across the tested datasets. While our model shows promising results in deepfake video detection by leveraging visual and latent features, we demonstrate that further work is needed to improve its generalizability when encountering out-of-distribution data. Our model provides an effective solution for identifying a wide range of fake videos while preserving the integrity of media.
The COM-PRESS dashboard enables image manipulation analysis for fact checkers.
Several frameworks have been proposed for delivering interactive, panoramic, camera-captured, six-degrees-of-freedom video content. However, it remains unclear which framework will meet all requirements the best. In this work, we focus on a Steered Mixture of Experts (SMoE) for 4D planar light fields, which is a kernel-based representation. For SMoE to be viable in interactive light-field experiences, real-time view synthesis is crucial yet unsolved. This paper presents two key contributions: a mathematical derivation of a view-specific, intrinsically 2D model from the original 4D light field model and a GPU graphics pipeline that synthesizes these viewpoints in real time. Configuring the proposed GPU implementation for high accuracy, a frequency of 180 to 290 Hz at a resolution of 2048×2048 pixels on an NVIDIA RTX 2080Ti is achieved. Compared to NVIDIA’s instant-ngp Neural Radiance Fields (NeRFs) with the default configuration, our light field rendering technique is 42 to 597 times faster. Additionally, allowing near-imperceptible artifacts in the reconstruction process can further increase speed by 40%. A first-order Taylor approximation causes imperfect views with peak signal-to-noise ratio (PSNR) scores between 45 dB and 63 dB compared to the reference implementation. In conclusion, we present an efficient algorithm for synthesizing 2D views at arbitrary viewpoints from 4D planar light-field SMoE models, enabling real-time, interactive, and high-quality light-field rendering within the SMoE framework.
As the demand for high-quality video content continues to rise, accurately assessing the visual quality of digital videos has become more crucial than ever before. However, evaluating the perceptual quality of an impaired video in the absence of the original reference signal remains a significant challenge. To address this problem, we propose a novel No-Reference (NR) video quality metric called NR-VMAF. Our method is designed to replicate the popular Full-Reference (FR) metric VMAF in scenarios where the reference signal is unavailable or impractical to obtain. Like its FR counterpart, NR-VMAF is tailored specifically for measuring video quality in the presence of compression and scaling artifacts. The proposed model utilizes a deep convolutional neural network to extract quality-aware features from the pixel information of the distorted video, thereby eliminating the need for manual feature engineering. By adopting a patch-based approach, we are able to process high-resolution video data without any information loss. While the current model is trained solely on H.265/HEVC videos, its performance is verified on subjective datasets containing mainly H.264/AVC content. We demonstrate that NR-VMAF outperforms current state-of-the-art NR metrics while achieving a prediction accuracy that is comparable to VMAF and other FR metrics. Based on this strong performance, we believe that NR-VMAF is a viable approach to efficient and reliable No-Reference video quality assessment.
Digital image manipulation has become increasingly accessible and realistic with the advent of generative AI technologies. Recent developments allow for text-guided inpainting, making sophisticated image edits possible with minimal effort. This poses new challenges for digital media forensics. For example, diffusion model-based approaches could either splice the inpainted region into the original image, or regenerate the entire image. In the latter case, traditional image forgery localization (IFL) methods typically fail. This paper introduces the Text-Guided Inpainting Forgery (TGIF) dataset, a comprehensive collection of images designed to support the training and evaluation of image forgery localization and synthetic image detection (SID) methods. The TGIF dataset includes approximately 75k forged images, originating from popular open-source and commercial methods, namely SD2, SDXL, and Adobe Firefly. We benchmark several state-of-the-art IFL and SID methods on TGIF. Whereas traditional IFL methods can detect spliced images, they fail to detect regenerated inpainted images. Moreover, traditional SID may detect the regenerated inpainted images to be fake, but cannot localize the inpainted area. Finally, both IFL and SID methods fail when exposed to stronger compression, while they are less robust to modern compression algorithms, such as WEBP. In conclusion, this work demonstrates the inefficiency of state-of-the-art detectors on local manipulations performed by modern generative approaches, and aspires to help with the development of more capable IFL and SID methods. The dataset and code can be downloaded at https://github.com/IDLabMedia/tgif-dataset.
Image manipulation is easier than ever, often facilitated using accessible AI-based tools. This poses significant risks when used to disseminate disinformation, false evidence, or fraud, which highlights the need for image forgery detection and localization methods to combat this issue. While some recent detection methods demonstrate good performance, there is still a significant gap to be closed to consistently and accurately detect image manipulations in the wild. This paper aims to enhance forgery detection and localization by combining existing detection methods that complement each other. First, we analyze these methods' complementarity, with an objective measurement of complementariness, and calculation of a target performance value using a theoretical oracle fusion. Then, we propose a novel fusion method that combines the existing methods' outputs. The proposed fusion method is trained using a Generative Adversarial Network architecture. Our experiments demonstrate improved detection and localization performance on a variety of datasets. Although our fusion method is hindered by a lack of generalization, this is a common problem in supervised learning, and hence a motivation for future work. In conclusion, this work deepens our understanding of forgery detection methods' complementariness and how to harmonize them. As such, we contribute to better protection against image manipulations and the battle against disinformation.
Anterior cruciate ligament (ACL) injuries are common in sports such as soccer, often requiring extensive rehabilitation post-surgery. Return-to-sport (RTS) rehabilitation protocols typically involve clinical and kinematic evaluations to ensure safety. This study investigates the added value of non-immersive (Extended Reality XR) and immersive (Virtual Reality VR) on movement quality during RTS assessment of male soccer players post ACL reconstruction, compared to healthy control players. Our XR tests were performed using a projection on a big screen on 11 male soccer players with a soccer-related ACL tear. Additionally, our VR tests were performed using a Head Mounted Display (HMD) on 31 male soccer players, including 14 with ACL reconstruction and 17 healthy control players. We performed clinical tests and kinematic analyses with statistical comparisons between testing conditions. The utilization of XR seems to elicit the mechanism associated with ACL rupture more significantly, suggesting its potential value in incorporating XR into RTS screening and rehabilitation protocols. Analysis of the VR-results revealed greater knee flexion and valgus range compared to non-VR, demonstrating VR's potential for enhanced kinematic sensitivity. As such, VR-based kinematic analysis may improve sensitivity in detecting movement deviations during RTS evaluations post-ACLR. Further research is needed across diverse athletic populations. In conclusion, this study highlights the potential for both XR and VR applications in identifying ACLR-related deficits, and motivates their integration with conventional clinical RTS screening batteries.
A crucial competence for mentor teachers is the ability to analyse classroom practices as they are expected to model effective teaching practices and to provide feedback to student teachers. This ability is referred to in the literature as professional vision. The present study assesses mentor teachers' (n = 137) professional vision regarding teacher-student interactions and differentiated instruction, using a validated video-based comparative judgement measurement instrument. The results indicate that mentor teachers have a high professional vision. It can thus be assumed that mentor teachers can support student teachers. Additionally, their professional vision is compared with that of classroom teachers (n = 996) and student teachers (n = 2168), expecting it to be significantly higher than that of classroom teachers and student teachers. The results show no significant difference between mentor teachers and classroom teachers but a significant difference between mentor teachers and student teachers. Hence, mentor teachers and classroom teachers are equally able to identify and interpret crucial aspects of effective teaching behaviour and both groups are better able than student teachers in this regard. This study contributes to the current state of the art on mentor teachers from a theoretical, empirical and methodological point of view.
The rapid evolution of digital image circulation has necessitated robust techniques for image identification and comparison, particularly for sensitive applications such as detecting Child Sexual Abuse Material (CSAM) and preventing the spread of harmful content online. Traditional perceptual hashing methods, while useful, fall short when exposed to some common image transformations, or when images are doctored to avoid detection, rendering them ineffective for nuanced comparisons. Addressing this challenge, this paper introduces a novel pretrained vision transformer artificial intelligence (AI) model approach that enhances the robustness and accuracy of perceptual hashing. Leveraging a pretrained Vision Transformer (ViT-L/14), our approach integrates visual and textual data processing to generate feature arrays that represent perceptual image hashes. Through a comprehensive evaluation using a dataset of 50,000 images, we demonstrate that our method offers significant improvements in detecting similarities for certain complex image transformations, aligning more closely with human visual perception than conventional methods. While our method presents certain initial drawbacks such as larger hash sizes and high computational complexity, its ability to better handle perceptual nuances presents a forward step in the realm of image forensics. The potential applications of this research extend to law enforcement, digital media management, and the broader domain of content verification, setting the stage for more secure and efficient digital content analysis.
Deepfakes are hyper-realistic videos in which the faces are replaced, swapped, or forged using deep-learning models. This potent media manipulation techniques hold promise for applications across various domains. Yet, they also present a significant risk when employed for malicious intents like identity fraud, phishing, spreading false information, and executing scams. In this work, we propose a novel and improved Deepfake video detector that uses a Convolutional Vision Transformer (CViT2), which builds on the concepts of our previous work (CViT). The CViT architecture consists of two components: a Convolutional Neural Network that extracts learnable features, and a Vision Transformer that categorizes these learned features using an attention mechanism. We trained and evaluted our model on 5 datasets, namely Deepfake Detection Challenge Dataset (DFDC), FaceForensics++ (FF++), Celeb-DF v2, Deep-fakeTIMIT, and TrustedMedia. On the test sets unseen during training, we achieved an accuracy of 95%, 94.8%, 98.3% and 76.7% on the DFDC, FF++, Celeb-DF v2, and TIMIT datasets, respectively. In conclusion, our proposed Deepfake detector can be used in the battle against misinformation and other forensic use cases.
Sports rehabilitation exercises and return-to-sports screening are typically conducted in confined spaces that lack the highly motivating environments found in stadiums filled with people. Virtual Reality (VR) can be employed to virtually recreate those settings. A challenge in VR sports is that running exercises with six degrees of freedom can pose a risk due to potential collisions with obstacles. Redirected walking addresses this issue by utilizing techniques that adjust the user's virtual motion in relation to their physical motion. This study evaluates the established positive perceptual bounds of redirected walking and extends its application to a reaction-based exercise. In this proposed use case, the user must chase the ball from one point to another in the room. Our findings indicate that rotational manipulations impact user experience at +44% rotation, which is stricter than the established +49% rotation bound. For translational manipulations, our research shows that the established +26% translational bound can be extended to +41% without a significant difference in self-reported subjective experience. This study demonstrates that redirected walking manipulations do not lead to significant changes in participants' total reaction times. In conclusion, the findings of this study illustrate that redirected walking techniques, with careful consideration of the bounds on rotational and translational manipulations, can be effectively applied in a Virtual Reality exergame designed for return-to-sports screening. This offers a more engaging alternative to traditional rehabilitation environments.
Imperceptible image watermarking methods strive to balance visibility with robustness requirements in order to protect copyright images from copyright infringement without impacting users’ viewing experience. Additionally, image generators such as Stable Diffusion and Google’s Imagen utilize watermarking to enable the detection of synthetically generated images. To guarantee a sufficient level of robustness, watermarking methods are evaluated against traditional attacks. We recognize diffusion models as a potential disruptive technology in the watermark robustness field and therefore investigate this approach. This paper presents Diffusion Denoising Watermark Removal Models (DDWRM) as a watermark attack. Instead of generating images from noise, our approach strategically diffuses watermarked images, followed by a denoising process to effectively remove embedded watermarks while preserving the original content. The methodology is designed to be generic, showcasing its potential as a robust and versatile watermark removal technique. Experimental results highlight the proposed DDWRM’s success in removing watermarking information, outperforming the traditional JPEG compression attack method for specific watermarking techniques (DWT-DCT-SVD and Riva-GAN). However, variations in resilience, particularly with DWT-DCT watermarking method, prompt further exploration needs. In conclusion, our proposed attack emerges as an innovative and effective watermark removal method.