
The Corporate Sustainability Due Diligence Directive (CSDDD) by the European Union (EU) mandates companies to identify and address environmental impacts and human rights issues within their supply chains. Procurement organizations play a pivotal role in both the effective implementation of and ongoing compliance with the CSDDD. This paper investigates which challenges procurement organizations must overcome along the stages of implementation through a Systematic Literature Review (SLR). The results indicate that the largest number of challenges arise during initial implementation, mainly related to communication and collaboration with suppliers as well as their willingness and ability to share data. Moreover, it highlights that substantial efforts will be required to ensure ongoing compliance. The results contribute to a systematic understanding of challenges imposed on procurement organizations and lay the groundwork for future research on potential solutions.
Recent advances in Large Language Models (LLMs) have enabled new opportunities in Computer-Aided Design (CAD) through natural language-driven code generation. We present Prompt2CAD, a lightweight and replicable framework that combines prompt-based macro generation with optional visual, human-in-the-loop refinement. Users provide a natural language description of a 3D part, and a general-purpose LLM (GPT-4o in this study) interprets and generates executable CAD scripts. If the output is incorrect, users can submit screenshots of the rendered CAD output and clarifying feedback for iterative correction—without requiring fine-tuning, scoring mechanisms, or additional Artificial Intelligence (AI) models. We evaluate Prompt2CAD on 63 modeling tasks of varying complexity. The framework achieves a 90.48% overall success rate, with 77.78% success on the first attempt and 57.14% resolved through iterative visual feedback. These results highlight the potential of general-purpose LLMs as collaborative design partners in CAD—bridging the gap between code automation and human-guided refinement.
In line with the Operator 5.0 vision of a resilient, human-centric workforce, this study presents a Virtual Reality (VR) training framework for reskilling and upskilling industrial operators in manual assembly tasks. The system is designed as a modular and immersive learning environment where users engage in hands-on procedural training using intuitive hand tracking and step-by-step guidance. The framework aims to reduce training time, improve task readiness, and enable safe, repeatable learning experiences without disrupting physical workflows. To explore the feasibility of the approach, the VR environment was validated using a standardized furniture assembly scenario based on the Human-Robot Collaboration (HRC) Model Set as a reproducible testbed for training transfer. Preliminary results from pilot sessions suggest that users trained in VR were able to assemble physical furniture components more confidently and accurately. Although full experimental trials are pending, this exploratory work demonstrates the viability and educational potential of VR-based training for operator skill development. The proposed solution contributes a flexible, cost-effective strategy for immersive workforce training aligned with the principles of Industry 5.0.
Multimedia AI now characterizes multiple concurrent streams at high resolutions, spanning video, audio, language, and multimodal fusion. Energy, bandwidth, and even water for cooling have become first-order constraints for real deployments. Training a single large NLP model has been estimated to emit up to 284 tons of CO2, while infrastructure and model choices can shift emissions by orders of magnitude; at the same time, data centers consume huge volumes of water for cooling (hundreds of thousands). In this position paper, we argue that we cannot continue to choose models based on accuracy and throughput alone. We propose a Sustainability Card for Multimedia AI-Correct Outputs per kWh, Joules per Sample, Bytes per Output, and optional Water per 1,000 Outputs-and outline a simple measurement protocol.
We present "Dialogue-Pseudo", a dialogue-aware pseudonymization framework that (i) maps each real speaker to a distinct pseudo-speaker via a dialogue-aware PartitionAssignment-Selection procedure to preserve inter-speaker separability, (ii) replaces personally identifiable information with syllable- and accent-matched surrogates to maintain phonorhythmic timing, and (iii) applies bounded prosody control that normalizes global pitch/energy while preserving relative contours and turn-taking cues. Experiments on the RWCP conversational corpus with a pseudo-speaker pool derived from Common Voice show improved privacy under an ECAPA-TDNN speaker verification model: Equal Error Rate (EER) increases from 37.38% (original) to 42.22% (pseudonymized). Prosody-oriented statistics (F0 mean/variance, RMS energy, pause ratio, turn length) remain close across conditions, and UMAP visualizations indicate that pseudo-speakers are well separated from both the sources and from one another. These results support Dialogue-Pseudo as a practical approach to privacy-preserving dialogue analytics that extends pseudonymization beyond timbre conversion to dialogue-relevant linguistic and prosodic cues, thereby balancing unlinkability with the retention of task-critical information.
Large language models (LLMs) can be used to solve various tasks based on text inputs, e.g., video quality estimation. We explore the usage of LLMs for video quality prediction based on metadata (video codec, bitrate, resolution), which has not been addressed before. For the evaluation we use test #1 from our AVT-VQDB-UHD-1 dataset. We generated text prompts based on the metadata and collected answers from 17 different LLMs. The evaluation indicates that especially larger LLMs could be used to simulate human raters. However, a pure model prediction with one model has lower performance than SoA video quality models. Thus, we further investigated combinations of LLMs, which resulted in comparable performance to state-of-the-art models. Our work is a proof-of-concept, considering that LLMs are slower for the prediction than the traditional metadata-based models.
This study focuses on improving the safety of Vulnerable Road Users (VRUs) in traffic by leveraging multispectral imaging and deep learning. We introduce ScaleFuse, a novel RGB-thermal fusion architecture specifically designed to address the limitations of existing multispectral detection methods. Unlike conventional approaches, ScaleFuse performs multiscale feature fusion at intermediate layers of the backbone, adaptively learning spatial and channel-wise importance from both modalities. To ensure robust fusion, we employ the SuperGlue network for precise image alignment, mitigating the common issue of misregistration between RGB and thermal inputs. ScaleFuse is implemented within the YOLOv10 framework, enabling efficient and accurate detection in challenging conditions such as low-light environments, adverse weather, and occlusions. Experimental evaluations on LLVIP, FLIR, and our newly introduced VeNIT-MSD dataset demonstrate that ScaleFuse consistently outperforms single-modality baselines and previous fusion strategies in terms of accuracy, precision, and recall. The proposed system achieves real-time inference while remaining fully compatible with deployment frameworks such as TensorRT, making it suitable for intelligent transportation applications.
Malaria remains a significant public health challenge, particularly in developing countries. Traditional machine learning and deep learning classification approaches typically rely on large, labeled datasets that include both normal (malariafree) and abnormal (infected) samples. However, in the medical domain, the acquisition of annotated infected samples is difficult due to privacy restrictions and limited availability, which restrict data access and sharing. Anomaly detection techniques offer a compelling alternative by reducing the dependency on labeled abnormal data. In this paper, we propose a one-class autoencoder model trained solely on microscopic images of healthy (normal) blood cells. The model learns the underlying distribution and structural patterns of normal cells, enabling it to detect deviations indicative of malaria infection. Structural similarity (SSIM) is employed to measure reconstruction error, which is then used to classify cells as normal or infected. Experimental results demonstrate that the proposed method achieves high accuracy in detecting malaria-infected cells in blood smear images, highlighting its potential as a scalable and effective solution for automated malaria diagnosis.
dThis paper investigates the potential of Multi-access Edge Computing (MEC)-enabled architectures to expand Content Delivery Networks (CDNs) features by exploiting real-time telemetry across the radio access, core, and edge layers. We present experimental results from three representative use cases each posing distinct challenges in terms of mobility, traffic surges, and content pre-positioning. The proposed system leverages telemetry-informed decisions to orchestrate CDN resources dynamically, including a Proxy for Local Cache and Segment Prefetching to reduce latency and improve delivery predictability, a Multi-CDN Steering Proxy that selects optimal content paths based on live Quality of Service (QoS) metrics, and an Orchestrator for Proxy Scaling and Load Balancing that ensures elastic resource adaptation in response to demand fluctuations. These capabilities collectively enable low-latency, resilient media delivery even under network congestion conditions. The experimentation also conducts the identification of key operational challenges for MEC-based CDN deployments, including virtualization gaps, lack of telemetry standards, proxy discovery limitations, and cross-domain orchestration needs critical for scalable and interoperable 5G streaming solutions. This paper provides a practical evaluation of telemetry-driven MEC-based CDNs, offering insights to guide future designs and standardization for reliable, immersive streaming in next-generation cloud-native networks.
Anomalous sound detection is vital for predictive maintenance, enabling early fault detection to prevent costly failures. However, existing systems often struggle under realworld conditions due to background noise and low signal-to-noise ratios (SNR). To address this, we propose a denoising autoencoder (DAE) model that leverages hybrid noise perturbation, including white, pink, and brown noise types. Based on our evaluations on seven diverse machine-sound datasets across multiple SNR levels, the proposed approach consistently improves source-domain reconstruction AUC source by 0.9% to 1.6% over the baseline autoencoder. While improvements in AUC target and pAUC are less consistent, brown noise performs best in high-noise settings, and hybrid noise perturbation outperforms under low-noise conditions. These results underscore the need for noise-aware perturbation strategies tailored to acoustic conditions and evaluation goals.
Parking-lot monitoring is a facility-management task in which a system infers the occupancy and accessibility of predefined spaces from fixed-camera imagery. The scene typically contains many visually similar objects arranged in a stable spatial layout, which remains challenging for current multimodal large language models (MLLMs) and vision-language models (VLMs) that often misbind objects and locations. In prior work, we proposed a binding-aware, grid-based spatial prompting scheme that embeds the parking layout directly into the prompt and anchors vehicles to discrete spatial “slots”, improving slot-occupancy and obstruction detection without retraining or additional sensors. This paper extends that architecture with a layout-aware self-correcting loop around the same general-purpose MLLM. A judgment module evaluates grid-based predictions against basic layout constraints and, when violations occur, augments the prompt with explicit textual feedback and requests a single revised prediction. Experiments on real parking scenes comparing vanilla prompting, the previous grid-based method, and the extended method show that layout-aware self-correcting prompts reduce both occupancy errors and logically inconsistent configurations, while preserving low deployment cost. This provides a simple way to couple foundation models with domain-specific spatial rules in safety-relevant facility monitoring.
In this paper, we propose a voxel-based rendering framework specifically designed to improve the quality of compressed point clouds. Compressed point clouds often suffer from sparsity, uneven density, and geometric degradation, which lead to visual artifacts such as holes, noise, and the loss of fine structural details. To address these problems, our method converts geometry-based point cloud compression (G-PCC) data into a structured 3D voxel grid that preserves local geometric information while providing a consistent spatial representation for convolutional processing. This voxelization alleviates sparsity and density imbalance, enabling the network to effectively learn local and global geometric features without introducing holes. The framework integrates voxel-level structural reasoning with point-level radiance restoration through a hybrid pipeline that combines voxel networks and multilayer perceptrons. Structural voxel features are projected into 2D space and refined by a U-Net-based image network to restore high-frequency details and reduce compression artifacts. The training process employs a combined L2 and perceptual loss to balance structural fidelity and perceptual quality. Experiments conducted on the NeRFSynthetic dataset, converted into the G-PCC format, demonstrate that the proposed method achieves stable and accurate rendering even under sparse conditions. Compared with existing baselines such as BPCR, NeRF, and FreqPCR, our framework consistently achieves higher PSNR and SSIM values while yielding the lowest LPIPS, confirming its effectiveness for high-quality rendering of compressed point clouds. These results indicate that the proposed voxel-based approach provides a robust and efficient solution for immersive media and 3D reconstruction applications.
For decades, RGB data has been central in human-pose estimation (HPE). The growing legislative pressure now limits their use during deployment, as RGB frames reveal sensitive personal information (SPI). Recent work often designs and trains new anonymization models, at a high computational cost. We ask instead whether adaptive filtering is sufficient. We introduce an adaptive obfuscation pipeline that converts RGB datasets into privacy-aware counterparts without training new networks. For each image, we detect the largest face, estimate a re-identification risk from its relative area, and raise blur or pixelation until an InsightFace recognizer fails to match the original embedding; that intensity is then applied to the whole image. Applied to MS-COCO, the pipeline generates two new sets, COCO-BLUR and COCO-PIX. On COCO-BLUR, three off-the-shelf HPE models retain more than 87% of their baseline AP. We release our code at this link (GitHub repository) offering a sustainable way to anonymize RGB data.
Despite progress in visual reconstruction from fMRI signals, evaluating reconstruction quality remains challenging due to noisy, low-resolution data and semantic ambiguity. Conventional metrics often overlook perceptual and structural alignment with the original stimuli. To address this, we propose Graph-based Semantic and Structural Similarity (GSS), a novel evaluation approach that represents both stimuli and reconstructions as patch-wise graphs using CLIP-derived features. By applying graph matching, GSS captures spatial and semantic relationships beyond pixel-level similarities. Our approach aligns with neuroscientific models of visual processing and demonstrates robust, interpretable results that complement existing metrics.
Recent advances in Gaussian splatting have enabled real-time and photorealistic rendering of volumetric scenes, showing potential to become a mainstream 3D representation. Through efficient learning, Gaussian splatting transforms multiview images into a 3D representation, which can be interpreted as a point cloud with additional attributes. One of the challenges of Gaussian splatting is the large file size needed for storing its raw representation, which may be in the order of 1GB to 10GB for objects and scenes. Fast and mass market adoption of this technology requires a standard that offers interoperability and allows to efficiently reduce the size of the Gaussian splat representation to manageable bitrates at high visual quality. In this paper, we present how the ISO/IEC 23090-5 visual volumetric video-based coding (V3C) and video-based point cloud compression (V-PCC) standard can be utilized for coding scenes and objects represented with Gaussian splats. We show that a new profile of V-PCC provides better compression gains compared to other existing methods. Moreover, as V-PCC leverages video coding at its core, the video decoding hardware acceleration ecosystem paves the way towards streaming, decoding and realtime rendering of static Gaussian Splat scenes on mobile devices.
Video analysis is a key tool in modern sports, but its costs and complexity pose challenges, especially for clubs in indoor disciplines like futsal or handball. The R&D project ISVP.AI (Automated System for Indoor Sport Video Production) was initiated to democratize access to these technologies by automating the video production and analysis process using AI. The system utilizes computer vision algorithms (including YOLO) for object detection and tracking in 4K recordings, along with a dedicated expert system for recognizing match events, enabling the generation of professional video materials (broadcasts with a virtual cameraman, highlights, statistics) without specialized equipment. A key outcome of the ISVP.AI project was the implementation of the developed R&D results into business operations, laying the foundation for the commercial platform Stellis One - All your sports organization's tools integrated into one platform. Stellis One integrates the analytical module (based on ISVP.AI) with a wide range of functions, including Ticketing, Shop, Marketing Automation, Mobile Application, Competition modules with a statistics and analysis system, and the crucial One FanHub - the heart of the system where fan information is collected. The Stellis One platform is currently successfully used by over 50 sports organizations, confirming the effectiveness of technology transfer from the R&D project to the market and providing real support for the digital transformation of sport. The presentation will outline the path from ISVP.AI research to the integration of this technology within Stellis One, showcasing the practical application of AI in sports.
Virtual Reality (VR) transforms daily activities into immersive experiences, bridging the gap between physical and virtual environments. The growing use of VR applications to handle sensitive personal data necessitates the development of authentication methods that are both secure and practical. Traditional PIN-based systems remain susceptible to theft, spoofing, and credential sharing, while previous behavior-based biometrics often restrict trajectories to fixed lengths and overlook device heterogeneity. This study presents the first VR authentication framework that integrates knowledge-based PIN entry with behavioral biometrics derived from cross-device input modalities, incorporating both controller and hand-tracking trajectories to enhance the security. We evaluate a Siamese network designed to process variable-length motion data using a pilot dataset of 26 participants performing four distinct PIN-entry tasks. The proposed approach achieves authentication accuracies up to 96%, with Equal Error Rates (EERs) as low as 2.95%. By preserving natural temporal variability and generating match scores, this method enables authentication to generalize across devices and users, thereby establishing a foundation for secure and deployable VR access control systems.
The use of Smart TVs and companion devices is gaining momentum, enabling personalised and enhanced media experiences. A new wave of enriched services, and even new business models, is now possible. This is particularly relevant to TV operators and other stakeholders, since the (linear) broadcast TV content can be augmented by on-demand media content delivered via broadband technologies. Nevertheless, despite the efforts from some initiatives to drive hybrid broadcast/broadband TV media services, the current solutions still do not fully exploit the potential and opportunities that hybrid TV media services can offer. Existing companion-screen (CS) apps typically remain confined to 2D interfaces. However, AR/VR head-mounted displays have matured sufficiently to enable a transition from the use of physical 2D interfaces to a more immersive scenario, allowing many virtual companion screens around the smart TV. This paper presents the preliminary results of a recently conducted study on how users perceive the enhancement of linear TV services with AR/VR in hybrid broadcast/broadband TV scenarios. A prototype of a typical ARTV service has been implemented and tested by 41 participants, who declared ARTV services as the future of TV, considering them useful, informative, desirable, and fun.
Japanese SNS sentiment analysis has broad applications in marketing, public opinion surveys, and mental health support. However, because posts are often short, colloquial, and elliptical, they tend to be misclassified. To address this issue, we propose three methods that incorporate syntactic information into a Japanese BERT classifier. Using the GiNZA dependency parser, we extract four types of syntactic features—POS tags, dependency labels, parent position, and tree depth—and encode them via: (1) an embedding-based method that adds syntactic embeddings, (2) a syntactic positional encoding method that applies sinusoidal encodings to positional information, and (3) a hybrid method that combines these two methods. Experiments on the WRIME ver. 2 dataset show that the embedding variant utilizing POS tags and parent position achieves the best performance in 5-class polarity classification, while the hybrid method combining all syntactic features yields the best results in 8-class emotion classification. To the best of our knowledge, this is the first systematic study demonstrating the effectiveness of syntactic integration for Japanese SNS sentiment analysis and identifying optimal combinations of syntactic features.
In this paper, we present an analysis of the dual role of Large Language Models (LLMs) in the context of multimedia security, examining how these models can both strengthen defensive capabilities and introduce new vulnerabilities within modern content ecosystems. As LLMs increasingly interface with multimedia workflows-ranging from text-image generation pipelines to cross-modal retrieval and content moderation-their susceptibility to attacks such as data leakage, prompt injection, and jailbreaks raises critical concerns for the integrity and trustworthiness of multimedia platforms. We focus our analysis on privacy-oriented threats, including model inversion and membership inference, and discuss their implications for systems that process or generate multimedia content. Building on this risk landscape, we outline principles for improving the robustness and resilience of LLM-driven multimedia applications, highlighting strategies suited to in-the-wild threat scenarios. Indeed, we explore the constructive application of LLMs within cybersecurity frameworks, such as the Cyber Kill Chain, demonstrating how these models can be leveraged for threat detection, risk assessment, and automated defensive operations.