
In collaborative and distributed intelligent systems, compressed intermediate features are routinely transmitted and reused, making semantic quality assessment (SQA) crucial for reliable deployment. Recent compressed feature quality assessment (CFQA) benchmarks, however, show that conventional similarity measures often correlate poorly with downstream semantic utility and lack robustness across diverse feature codecs. In this paper, we propose a token-wise, attention-guided method for assessing the semantic quality of compressed features. First, motivated by the observation that many downstream heads normalize and process tokens largely independently, we assess quality at the token level. This token-wise formulation exploits the intrinsic correspondence between the original and reconstructed tokens while reducing cross-token interference. Second, since tokens contribute unequally to downstream task performance, we adopt an attention-guided aggregation scheme: we derive task-adaptive importance weights from DINOv2 self-attention and use them to pool token-wise quality predictions into a global semantic quality score. Third, to accommodate heterogeneous supervision across tasks, we cast CFQA as a regression problem and rescale classification-based rank targets to mitigate label imbalance. Experiments on the CFQA benchmark demonstrate that our method consistently improves PLCC and SROCC across three tasks and four codecs, yielding a practical, codec-agnostic quality interface for next-generation intelligent systems.
Current Image Quality Assessment (IQA) models are often constrained by the inconsistency of Mean Opinion Score (MOS) scales across different datasets. While traditional regression-based approaches (MSE) require complex MOS alignment, traditional ranking losses often ignore the magnitude of quality differences between stimuli. In this paper, we propose a Weighted Pairwise Hinge Loss (WPHL) that enables seamless training across multiple datasets without manual score calibration. By incorporating a dynamic margin proportional to the ground-truth quality gap and an importance-weighting mechanism, our loss function captures the nuanced topology of image quality perception. We demonstrate that WPHL allows lightweight architectures to rival significantly larger ones. Our approach offers a scalable path toward a “foundation model” for IQA by unlocking the use of diverse, multi-source training data. In the future, WPHL might be used to increase the performance of existing state-of-the-art models or help in the introduction of new ones.
Evaluating 8K ultra-high-definition (UHD) video quality is challenging, as viewer expertise strongly influences visual behavior. This study investigates the strategic differences between experts and non-experts during single-stimulus 8K video quality assessment (VQA). First, analysis using Coefficient K, a metric to distinguish ambient/focal attention, shows that experts adopt a focal strategy to scrutinize localized artifacts, whereas non-experts maintain an ambient pattern. Second, spatial analysis through bivariate contour ellipse area (BCEA), an indicator of spatial gaze concentration, and gaze density heatmaps show that experts' gaze is more localized and targets artifact-prone textures from the initial phase, while non-experts tend to be influenced by center bias. Finally, we demonstrate that gaze velocity, integrated with a gaze-based local quality metric, significantly predicts subjective scores. While this combination improves overall performance, variations across content types highlight the potential of temporal gaze dynamics to further refine quality models. Our findings advocate for a dual-axis framework in 8K VQA that incorporates conventional spatial gaze distribution (“Where”) with expertise-driven temporal gaze dynamics (“How”).
This paper investigates the impact of interactioninduced body movement on cybersickness in Virtual Reality (VR). Therefore we designed a controlled VR environment with three interaction conditions of increasing movement complexity. Data from 24 participants, combining subjective measures and motionderived features, show a consistent increase in cybersickness with higher movement demands, particularly in conditions involving vertical displacement and full-body interaction. Moreover, correlation analysis reveals a significant negative relationship between the coefficient of variation of head velocity and cybersickness. This suggests that more variable and adaptive motion may mitigate the discomfort. These findings highlight the importance of movement characteristics, beyond movement intensity alone, for designing full-body interactive VR experiences.
Visual artefacts in 3D meshes can act as local distractors and perturb visual attention, which may affect how users inspect and interpret interactive 3D content in applications such as technology-enhanced learning and training. An important open question is how localized geometric distortions translate into measurable perceptual disruption during interactive viewing, as understanding this relationship is necessary for developing perceptually grounded evaluation methods for 3D assets. In this work, we investigate whether gaze can serve as an implicit, spatially localized indicator of such effects. To this end, we introduce a surface-based gaze analysis approach that maps eye-tracking samples onto the visible 3D mesh geometry under user-driven interaction, enabling view-consistent analysis of localized attention. This enables defect-level, interaction-aware attention analysis that is not supported by conventional image-plane or fixed-viewpoint methods. In a study with 22 participants viewing 27 textureless meshes with four defect types, we observe that self-intersection defects produce significantly higher fixation counts, while time-to-first fixation (TFF) does not show significant differences compared to corresponding regions in unmodified meshes. In contrast, smoothing defects shows weaker trends, while low-polygon and semantic defects do not produce reliable localized effects under the present conditions. These results suggest that visual attention reflects the perceptual severity of localized mesh distortions, positioning gaze as a complementary signal for perceptually grounded quality-of-experience assessment in 3D content. The findings are conditioned by the controlled viewing setup and defect placement, and should be interpreted accordingly.
Despite the numerous approaches available in the state-of-the-art, the design of a full-reference image quality assessment metric aligned to human perception remains an open challenge. One possible reason is that most existing metrics fail to faithfully reproduce human judgments, as they disregard the global semantic structure of a scene. To fill this gap, we introduce a new method that evaluates image quality by explicitly modeling the preservation of semantic content. Segmentation masks are used to estimate semantic information loss, which is then integrated into the LPIPS framework, thus enabling the perceptual distance to account for semantic consistency. Experimental results show consistent improvements over the LPIPS baseline and competitive performance with respect to state-of-the-art methods, suggesting that semantic information loss is a useful cue for perceptual quality evaluation.
Crowdsourced speech quality evaluation following ITU-T P. 808 has become standard for codec assessment. Prior work has studied the number of votes per condition needed for reliable per-stimulus scores, but guidance on the number of stimuli per codec for detecting codec-level MOS differences, and cross-session test-retest reliability across independent listener pools, remain open questions. Drawing on 89,852 ratings from one of the largest crowdsourced codec evaluations reported to date (three independent rounds, 922 unique listeners), we make two contributions. First, we establish test-retest reliability by comparing Opus codec quality scores across three evaluations spanning six months with independent listener pools (144, 498, and 280 listeners), finding mean differences $\leq 0.10$ MOS across all three pairwise comparisons. Second, we derive practical power guidelines: detecting a 0.2 MOS difference at 90% power requires ∼100 stimuli per codec with 10 listeners each; detecting 0.1 MOS requires ∼400.
Determining the bitrate at which compression artifacts become just visible is crucial for achieving visually lossless compression of ultra-high-definition (UHD) videos. This perceptual boundary, defined as the just-noticeable difference (JND), provides a principled target for perceptually lossless compression control. Despite its practical importance, limited work has systematically compared JND behavior across fundamentally different codecs. In this paper, we conduct a controlled psychophysical study to measure JND thresholds for two distinct compression approaches: the AV1 video codec and the JPEG 2000 image codec. By comparing a modern hybrid video codec with a wavelet-based image codec, we examine how coding principles influence artifact visibility at the perceptual limit. To reliably capture near-threshold distortions, we use a flicker-based experimental procedure that amplifies subtle differences between reference and compressed videos. We combine this with a collaborative QUEST+ framework to estimate shared psychometric functions across observers, allowing us to quantify the thresholds for different codecs and content types. Finally, we evaluate the ability of widely used objective quality metrics to predict measured JNDs across codecs and content types. Among the evaluated metrics, FSIMc, ColorVideoVDP, and ColorVideoVDP-ML-Saliency demonstrate the strongest agreement with perceptual data.
The rapid growth of Augmented Reality (AR) in a variety of applications such as mobile and retail sectors has intensified the need for perceptually accurate Quality of Experience (QoE) assessment. However, existing image quality assessment (IQA) models often fail to account for the unique degradations found in AR, particularly the interaction between virtual objects and real-world backgrounds. Existing works are largely limited by a lack of comprehensive datasets and an inability to jointly model object-level fidelity and photometric harmony. In this paper, we propose the AR composition quality assessment (AR-COMPQ) model, a novel dual-path deep learning framework designed to evaluate AR image composition quality by explicitly separating and then integrating object-level structural fidelity and appearance-level harmony. To support this research, we introduce the ARC-IQA dataset, a curated benchmark featuring 193 AR scenes with controlled geometric, texture, and illumination distortions, accompanied by Mean Opinion Scores (MOS) derived from a systematic subjective study. Experimental results demonstrate that AR-COMPQ achieves superior alignment with human perception, yielding higher correlation and lower prediction error with respect to MOS. The proposed method significantly outperforms traditional metrics and recent perceptually driven deep models, thereby facilitating robust automated quality control in AR rendering pipelines.
Video streaming accounts for the majority of today's Internet traffic and contributes substantially to global carbon emissions and, consequently, to climate change. Increasing sustainability in video streaming requires reducing these emissions while maintaining, or even improving, service quality. A promising approach is finding tradeoffs between Quality of Experience (QoE) and sustainability by lowering video resolution, which decreases bandwidth demands and thus carbon emissions, especially in the transmission network. Additionally, users' ecoconsciousness may influence their perception of service quality, suggesting that environmentally aware users could experience comparable QoE even at lower, more sustainable resolutions. This work addresses these topics by measuring and modeling potential emission savings associated with video streaming and investigates users' willingness to compromise on streaming quality for environmental sustainability. We conduct a crowdsourced experiment with 150 participants across two platforms, capturing not only self-reported willingness but also participants' actual resolution choices and QoE ratings after being informed about the emissions associated with video streaming. Our findings show that while most users prefer the highest resolution, providing carbon emission information partially mitigates the QoE decline for a subset of users, demonstrating that sustainability can influence resolution choices.
Prior work has shown that QoE-aware resource sharing for real-time interactive video can support up to three times more simultaneous sessions at acceptable quality compared to rate-fair allocation. However, the required capabilities (QoE-targeted encoding, runtime spatial complexity estimation, and rich application-network APIs) are not yet available in commercial deployments. In this paper, we take an evolutionary approach: we design a system that delivers QoE-aware resource allocation using only capabilities that can be assembled in a lab today. We extend the utility-based allocation framework to the radio resource domain by introducing composite spatial complexity, which combines a session's video spatial complexity with its time-variant spectral efficiency into a single resource demand function. To operate with commercial real-time video streaming applications that use rate-based congestion control and lack capability to measure QoE, we use external tooling for QoE measurements. We develop an incremental reallocation algorithm with per-interval limits that encode both the congestion control algorithm's speed constraint and that spatial complexity estimates are reliable only near the current rate. The resulting prototype combines external QoE measurements with congestion-signal-based rate steering and does not require modification to commercial applications. We chart an evolution path from this prototype toward full QoE-aware resource sharing, mapping emerging standards (IETF SCONE, CAMARA, Media over QUIC) to the progressive capabilities they enable.
Subjective quality assessment (QA) remains the de facto standard for evaluating image and video coding performance, in which perceptual quality is determined by human viewers, often in comparison to a reference. Rating-based protocols naturally provide this reference anchoring, but often exhibit limited sensitivity to subtle perceptual differences. Pairwise comparison (PC) methods offer superior discriminability, yet require substantially more comparisons and produce inherently relative quality scales that lack direct semantic interpretation with respect to a reference, which is an essential aspect in fidelity-oriented coding scenarios. In this paper, we propose a hybrid subjective QA framework that combines the complementary strengths of rating-based and PC methodologies. A reference-anchored rating stage first structures the quality space, after which PC is selectively applied only to perceptually ambiguous stimulus pairs. The framework is evaluated through analytical reconstruction using independently conducted rating-based and PC experiments on light field content. Results indicate that the proposed approach preserves much of the discriminative capability of full PC while reducing the number of required comparisons and providing reference-anchored quality scores. Although motivated by light field coding assessment, the proposed design principles are applicable to other QA scenarios.
The increasing adoption of volumetric media in immersive applications has created a need for dynamic point-cloud and mesh datasets that are scalable, diverse, and reproducible. Most publicly available datasets rely on captured data acquired through LiDAR or multi-camera pipelines, which inherently contain sensor noise, reconstruction artifacts, and uncertain surface geometry. Such characteristics complicate or even prevent the development of accurate full-reference objective quality metrics, as the true underlying surface geometry is uncertain and cannot serve as a reliable distortion reference. To address these gaps, we introduce Prompt2Point 11Dataset Available: https://github.com/IN2GM-Lab/Prompt2Point, a prompt-driven, LiDAR-free dataset for semantically controlled dynamic point-cloud and mesh video generation. Prompt2Point leverages prompt-based 3D content generation to synthesize animated meshes without specialized capture hardware, which are then processed through an automated pipeline to produce synchronized perframe mesh and dynamic point-cloud sequences. The dataset includes five distinct avatars, each paired with five animations, and additionally provides composite multi-object scenes with realistic occlusions and inter-object interactions. By providing noise-free, temporally consistent mesh ground truths alongside point clouds, Prompt2Point enables precise full-reference calculation of geometric and textural distortions introduced by compression algorithms. The dataset supports rigorous objective quality assessment, perceptual modeling, and distortion analysis, serving as a foundation for next-generation Quality of Experience (QoE) research in dynamic volumetric media.
In short-form video streaming, overlays, i.e., user interface (UI) elements like creator information or captions, are superimposed onto the video during playback. By potentially obscuring content or redirecting users' attention, they become an integral part of the overall streaming experience. However, the impact of such overlays on subjectively perceived Quality of Experience (QoE) and specifically video quality has not been explored yet. In this work, we thus investigate how overlays affect subjectively perceived video quality through two crowdsourcing studies. To ensure comparability to existing research, we first replicate the experimental setup and reuse short-form video content annotated by experts from a prior study. In a first study, we aim to partially reproduce the results of the prior study to validate our experimental setup and establish a baseline. The results indicate a strong agreement between experts and crowdworkers in relative quality ranking, although crowdworkers generally rated more conservative, thereby introducing a systematic bias between the studies. The second study then explores the impact of overlays and finds no significant differences in perceived video quality between the baseline study and the second study. Additionally, 62% of the participants reported perceiving overlays as annoying, which, however, did not affect their rating behavior. These findings suggest that overlays have no impact on perceived video quality, but may affect other QoE dimensions.
While AI-generated content has advanced rapidly, assessing the quality of AI-generated images (AGIs) remains challenging. Existing methods often separately evaluate perceptual quality, semantic alignment, and authenticity, overlooking their intrinsic correlations. Moreover, most CLIP-based approaches rely on global similarity between independently encoded features, limiting fine-grained vision-language interactions. To address this, we propose a cross-modal interaction framework for multidimensional AI-generated image quality assessment (CMIQA). Specifically, a bidirectional cross-modal interaction and fusion module improves fine-grained alignment by enabling bidirectional interactions between visual regions and textual semantics, while a consistency-aware loss weighting module adaptively adjusts the importance of different quality objectives based on the reliability of perceptual quality prediction. Experimental results demonstrate that our method achieves competitive performance on two public datasets.
This paper examines how interaction modality affects User Experience (UX) in a signal-based multi-user Extended Reality (XR) smart environment, where an Internet of Things (IoT) device acts as the sole communication medium. We developed a Cross-Reality (CR) scenario where a Mixed Reality (MR) user interacts with a physical laboratory and a Virtual Reality (VR) user occupies its Digital Twin (DT), exchanging information solely through a smart lamp. In a Mastermind-inspired task, 25 MR participants used three lamp-control modalities: hand gestures (HG), a virtual remote controller (RC), and a virtual user interface (UI). Two games per modality were completed and interaction quality, usability, and overall Quality of Experience (QoE) were rated. Differences were clearest in interaction quality and usability: UI was the most responsive and precise, RC was rated lowest, while HG achieved comparable QoE to UI.
Video conferencing is ubiquitous, yet whether richer display formats meaningfully alter how people engage during mediated interaction remains poorly understood. We present a controlled within-subjects study $(\mathbf{N} \boldsymbol{=} \mathbf{3 3})$ comparing three video conferencing formats, a standard monitor, a life-size television, and a novel quasi-holographic display in which the remote interlocutor is back-projected onto a transparent plexiglass screen (ARCADE), against an unmediated face-to-face baseline. Participants engaged in a structured social interaction task while wearing a head-mounted eye tracker. Subjective social presence, rapport, and engagement did not differ significantly across the three mediated conditions, though a consistent directional trend with small-to-medium effect sizes suggests the study was underpowered rather than the effects absent. Gaze behavior told a different story: display format did not influence facedirected gaze rate, but did influence the character of gaze. The quasi-holographic display elicited significantly more face-directed fixations than the television and sustained fixation sequences more than twice as long as the monitor, a large, Bonferronicorrected effect robust to outlier removal. Strikingly, sustained fixation sequences in the quasi-holographic condition numerically exceeded those in face-to-face interaction, a pattern inconsistent with equivalence and tentatively attributed to novelty-driven evaluative attention rather than enhanced social engagement. These results suggest that behavioral gaze metrics capture displaydriven differences in visual engagement that standard self-report instruments miss, and advocate for their inclusion in telepresence quality evaluation.
This short positioning paper asks whether the Quality of Experience (QoE) modeling toolbox developed for video services can be transferred to large language model (LLM) systems. We argue that the analogy is useful, but only if it is reformulated around interaction instead of content, compute and memory constraints instead of mainly network constraints, and multidimensional user outcomes instead of a single quality axis. Using the video-oriented theoretical model of Koniuch et al. as a point of departure, together with local LLM performancequality evidence, we identify what can be directly reused from multimedia QoE, what breaks in LLM settings, and which methodological questions should define a future LLM QoE agenda.
While spatial feature extraction has advanced rapidly in Video Quality Assessment (VQA), the temporal aggregation that processes frame-level scores into a global quality judgment remains an overlooked “black box.” Dominant strategies rely on heuristic pooling or generic sequence models, failing to account for the non-linear, retrospective nature of human Quality-of-Experience (QoE). In this paper, we first present a comprehensive Temporal Quality Aggregation Modeling (TQAM) Benchmark, evaluating 13 aggregation architectures across nine typical IQA backbones on four diverse VQA datasets. Our findings reveal that standard Transformers, despite their global reach, lack the necessary inductive bias to capture perceptual anchor points. To bridge this gap, we propose the Peak-End sensitive TransfER (PETER), a psychology-grounded framework that formalizes the foundational Peak-End Rule by prioritizing perceptually salient peak and end frames while preserving global temporal context for TQAM. Extensive experiments demonstrate consistent gains across diverse IQA backbones. Furthermore, as a plug-and-play module, PETER significantly enhances existing VQA frameworks, validating the robustness and generalizability of psychology-guided temporal modeling.
Vision-language models (VLMs) have recently been adapted to image quality assessment (IQA), using language-defined quality levels to produce interpretable scores. This design introduces a new attack surface: readable text inside an image may be interpreted as evidence about quality rather than as ordinary scene content. However, the standard typographic attack paradigm, which pastes deceptive text directly onto an image, is poorly suited to IQA. A pasted overlay injects a positive semantic cue but simultaneously introduces a visible artifact that degrades image fidelity, creating a self-contradictory attack. We propose Naturalistic Typographic Attack (NTA), a black-box method that uses a text-guided image editor to embed quality-related words into scene-consistent carriers such as signs, posters, and labels. NTA preserves the semantic influence of typographic cues while avoiding the visual degradation of naive overlays. Experiments on AVA images scored by Q-Align show that NTA nearly doubles the mean score gain of direct overlay while achieving higher success rates and greater stability across images. These results indicate that scene-consistent text insertion exposes a more potent semantic vulnerability in VLM-based IQA than conventional pasted text.