
The rapid advancement of text-to-3D (T23D) generation technologies has created an urgent demand for reliable objective quality assessment approaches. However, existing methods typically provide only a coarse-grained overall score, failing to capture fine-grained flaws across different perspectives, such as geometric distortion, texture blur, and text-misalignment. Furthermore, jointly learning these diverse dimensions often leads to severe feature conflicts, where strong semantic signals dominate weak geometric features. To address these challenges, we propose GateScore, a novel condition-gated multi-scale fusion framework for fine-grained multi-dimensional T23D quality assessment. Specifically, GateScore projects 3D assets into 2D multi-view images and extracts multi-modal features using CLIP encoders for cross-modal fusion. To effectively decouple conflicting tasks, we introduce a condition-aware detail gating mechanism, which conditionally filters and fuses scale-specific features based on the targeted evaluation dimension, enabling task-specific feature routing before final quality prediction. We conduct extensive experiments on three benchmark datasets with varying numbers of evaluation dimensions. Experimental results demonstrate that GateScore achieves state-of-the-art performance on multi-dimensional 3D mesh benchmarks and maintains highly competitive accuracy on video-based assets. Ablation studies further validate the individual contributions of its core components.
Rotary motion deblurring is an unavoidable process for high-level vision applications when the camera and scenes relatively rotate during the exposure time. The model-based methods construct an interpretative degradation model along each circle, but two-dimensional (2D) image information is neglected. The learning-based methods directly estimate the deblurred image from the 2D RMB image but do not exploit the one-dimensional (1D) degradation model. Although the progressive framework can simultaneously utilize the 1D degradation model and 2D information, it is not an end-to-end network structure. To combine the benefits of the 1D degradation model and 2D information in an end-to-end way, we present the frequency transformation function (FTF) and inverse FTF (IFTF), which implement mutual transformation between space and frequency domains for a 2D matrix. Furthermore, we develop an end-to-end network named FTF-Net, which is mainly composed of inversion and denoising modules. In the inversion module, the 1D degradation model is solved in the frequency domain. As for the denoising module, a convolutional neural network (CNN) is designed to suppress noise. Additionally, we present a CNN-based filler module to fill the ‘missing pixels’. Extensive experiments well demonstrate the superior performance of FTF-Net over other state-of-the-art (SOTA) methods. The code and models are available at https://github.com/Jinhui-Qin/FTF-Net/.
Initial trust in public-facing medical consultation AI is formed during early encounters, yet the safety judgments underlying this process remain insufficiently understood. Using a vignette-based survey of prospective users from the general public (N = 525) and partial least squares structural equation modeling , this study tested a dual-safety model in which perceived clinical safety (PCS) and perceived data security (PDS) were specified as two proximal safety judgments associated with initial trust (IT) and advice-following intention (INT). PCS and PDS were both positively associated with IT, and IT was positively associated with INT. Perceived explainability and key-evidence visibility were positively associated with PCS, whereas perceived system performance and stability was strongly associated with PDS and privacy invasion concern was negatively associated with PDS. Structural assurance showed an unexpected negative association with PCS, and scenario risk perception was not significantly associated with PCS or IT despite successful risk manipulation. The core PCS/PDS–IT pathways remained stable in supplementary robustness analyses. These findings clarify dual safety judgments as proximal correlates of initial trust and provide human-factors-oriented evidence for future work on safety communication, data-use explanation, and calibrated trust in public-facing medical AI.
As the use of head-mounted displays (HMDs) becomes increasingly common, concerns about visual fatigue in virtual reality (VR) environments have grown, as HMDs are known to induce greater fatigue than traditional 2D displays. To enable timely mitigation of this issue, it is essential to predict subjective visual fatigue in real time from continuously monitored objective measures. To address this need, we designeda passive HMD-based video-viewing experiment that manipulated motion (high/low) and brightness (high/low) to induce varying levels of visual fatigue and aimed to predict dynamic visual fatigue using eye-tracking data. Fifty-two participants took part in the study, providing continuous and post-task fatigue ratings while their eye movements were recorded as they watched video content in VR. Based on the recorded data, we developed a temporal attention based deep learning model that predicts fatigue 10 s ahead using a 30-second window of eye-related metrics.The proposed BiGRU-Attn model achieved competitive predictive performance—a mean absolute error (1.156), root mean square error (1.253), and two-point ordinal accuracy of 91% under leave-one-subject-out cross-validation–indicating that the predictive signal resides primarily in the eye-tracking features. Moreover, explainable feature importance analysis converged on the number of blinks (NB) as the single most informative ocular signal across zero-order correlation, Integrated Gradients, and feature ablation, with fixation duration (FD) emerging as a supporting contributor through temporal multivariate interactions. These findings demonstrate that eye-tracking data alone can predict dynamic visual fatigue duringpassive HMD-based video viewing and highlight the potential for a real-time fatigue prediction system that can enable timely application of fatigue reduction techniques.
Efficient search-and-rescue (SAR) operations rely heavily on timely and reliable navigational support in complex environments. Smart glasses have emerged as a promising wearable platform for delivering hands-free navigation guidance. However, the human–computer interaction implications of different navigation modalities for SAR-related field tasks remain insufficiently understood. As an early-stage feasibility study, this work focuses on simplified, navigation-related subtasks derived from SAR workflows and evaluates three smart-glasses navigation modalities —visual-only (V), auditory-only (A), and combined audio–visual (AV)—in outdoor navigation tasks inspired by SAR scenarios. Two task types, including a sequential search task and a return-to-base transfer task, were conducted in an unstructured environment without public roads. Six performance and workload metrics were collected: success rate of target finding, task completion time, actual-to-shortest path ratio, number of corrective reminders, heart rate, and subjective workload (NASA-TLX). Results show that audio–visual navigation significantly improves task efficiency, evidenced by shorter completion times and reduced path deviations, while maintaining physiological and subjective workload levels comparable to unimodal conditions. The AV modality also supported more stable navigation performance and fewer corrective interventions during task execution. Based on these findings, the study discusses design implications for smart-glasses navigation systems intended for SAR-inspired tasks and related outdoor search scenarios. Overall, the results provide an initial empirical basis for understanding how multimodal feedback can support wearable navigation interfaces in outdoor search scenarios.
Although object detection has been widely adopted in critical applications such as autonomous driving and security monitoring, its performance often deteriorates significantly under adverse weather conditions, including fog, rain, and snow. To address the challenges of object detection in foggy environments, this paper proposes the AFS-YOLO detection model. The proposed model improves detection accuracy and robustness without relying on additional dehazing preprocessing steps or external supervisory signals. Specifically, it incorporates an Adaptive Image Processing Module (AIPM) mechanism to better preserve image details, utilizes a Fog-Aware Adaptive Modulation (FAAM) module to dynamically optimize feature fusion, and introduces a Dynamic Shape-Context Intersection over Union (DSCIoU) loss function to enhance localization precision. Extensive experiments were conducted on two benchmark datasets—RTTS and Foggy-Cityscapes. The results demonstrate that AFS-YOLO achieves an overall mAP50 of 80.8% on RTTS and 56.4% on Foggy-Cityscapes, it demonstrates superior performance compared to the majority of existing methods for object detection in foggy conditions. Ablation studies further validate the effectiveness of each designed component in improving model performance. These findings comprehensively confirm the accuracy and robustness of AFS-YOLO in complex foggy scenarios.
Underwater image enhancement plays a crucial role in improving the perceptual capabilities of underwater vision systems. This paper proposes a dual-color-space, two-stage decoupled network for underwater image enhancement, where the task is decomposed into two independent stages: structural recovery and color correction. The dual-branch network is designed to efficiently mitigate color distortion while dehazing. The structural recovery stage is carried out within the YCbCr color space, focusing on the dehazing and contrast enhancement of the luminance (Y) channel. The color correction stage transitions to the lαβ opponent color space, where global non-linear correction is applied exclusively to the chromaticity channels α and β, for the sake of preservation of prior structural information. Furthermore, a cross-space hybrid loss function is introduced, which maps the predicted results back to the RGB space to enforce global constraints and improve color fidelity. Experimental results show that the proposed method performs well on metrics like PSNR, SSIM, and UCIQE. This process is well-balanced between dehazing and color restoration and therefore improves generalization performance robustly.
Image restoration aims to remove degradations while recovering both high- and low-frequency details of the image. Consequently, in addition to operations in the spatial and channel domains, optimization in the frequency domain is essential for restoring details and textures. Although some work has employed frequency-domain processing for image restoration, they have overlooked the importance of interactions among features across different frequency bands. To address these challenges, we designed a frequency learning module based on wavelet transform, which directly decouples the spectrum of the degraded image to extract the feature content of specific frequencies in the feature map, and then performs interactive enhancement to achieve the recovery of both global structure and fine details. Furthermore, we propose a Divide-and-Conquer U-Net model framework, which selectively applies different domain processing strategies at various depths of the network, thereby improving restoration performance. The proposed architecture, Divide-and-Conquer Net (DCNet), demonstrates superior performance over state-of-the-art methods across eleven synthetic and real-world datasets, excelling in tasks such as derain, desnow, dehazing, and raindrop removal. Implementation code and pre-trained weights will be available at https://github.com/Youqiang-Gui/DCNet.
We propose a novel mixture-of-experts (MoE) based omnidirectional image quality assessment (OIQA) method, named MoE-OIQA. MoE-OIQA first uses a designed adaptive feature representation module (AFRM) to integrate MoE into multiple stages of a backbone and leverage distortion information extracted from viewports as the prior information for adaptively representing multi-level features of viewports. Then, a quality prediction module (QPM) is designed to obtain the overall quality by weighting the quality scores of individual viewports. Extensive experiments demonstrate that MoE-OIQA effectively improves the model representation capability in diverse distortions and achieves superior performance on both uniform and non-uniform distortions.
A head-mounted display (HMD) delivers a virtual reality (VR) environment that allows an immersive experience. An understanding of visual performance with an HMD can improve visualization quality in VR applications. In this study, we designed visual performance tasks and recruited 36 participants to examine their dynamic visual acuity (DVA) when observing moving objects with an HMD. The results indicated that poorer visual acuity (higher DVA scores) was observed as the optotype speed increased from 20°/s to 80°/s. Particularly, participants had better DVA when they observed objects moving at a low speed (20°/s) than when they observed static objects. The corrected vision and sex difference made little difference (p > 0.05) in DVA scores. Furthermore, we found a significant speed-suppressing effect on identifying a moving object when it was smaller than a specific size. We also found an anisotropy effect in capturing objects when moving in the horizontal or vertical direction. Questionnaire results indicate that the experience of identifying objects in the IVE was highly consistent with that in the real world. These findings provide useful guidelines for graphic design in HMD-based VR applications.
Underwater object detection serves as both an engine for developing the “blue economy” and a cornerstone for safeguarding maritime sovereignty and ecological balance. However, underwater images often suffer from color deviation and blur, which lead to pronounced confusion between backgrounds and targets, posing significant challenges for detection. What is more, the aggregation of small underwater objects is also a tough problem to solve. To address these issues, this paper proposes a Dynamic Cross-Fusion Network for Underwater Object Detection (DCF-Net), which mitigates target-background confusion and challenges in detecting small objects. First, it achieves scene-adaptive enhancement through a dynamic fusion strategy. The enhancement branch exchanges feature information with the main branch through bidirectional dynamic cross-interaction. Second, an adaptive multi-granularity attention fusion mechanism and a hierarchical residual window attention mechanism are designed, which extends hierarchical residual window attention units to windowed statistical computation and design a dual-attention collaborative mechanism. This facilitates hierarchical fusion from coarse-grained screening to fine-grained optimization. Finally, we employ WIoU v3 to effectively reduce harmful gradients from low-quality samples and mitigate their influence. Experimental results demonstrate that DCF-Net achieves competitive detection performance, with mAP@0.5:0.95 scores of 70.2% and 52.1% on the DUO and UTDAC2020 datasets, respectively. Compared with existing underwater object detection methods, DCF-Net demonstrates improved detection capability under the evaluated underwater scenarios.
The reliability of underwater optical imaging is often compromised by image degradation caused by light scattering and absorption, which significantly affects visual quality and subsequently challenges downstream analysis and measurement-related tasks. While Underwater Image Quality Assessment (UIQA) and Underwater Image Enhancement (UIE) provide foundational tools, they are often developed in isolation, lacking effective perceptual guidance to bridge the gap between them. This work proposes DualNet, a dual-branch network that functions as a perceptual quality assessor by explicitly quantifying global chroma and local sharpness distortions-key perceptual dimensions of degradation. Leveraging frequency domain analysis and spatial domain processing, DualNet effectively captures color shifts, contrast attenuation, and detail loss. We further introduce cross-branch interaction mechanisms and feature correlation constraints to strengthen the representational capacity of these quality-aware features, ensuring robust integration of perceptual cues. To leverage these representations for enhancement, the learned perceptual features from DualNet are embedded into an enhancement model, termed DualNet-UIE, enabling a quality-guided enhancement process that adaptively improves images based on perceptual degradation. Extensive experiments on benchmark datasets demonstrate that our method achieves competitive performance in perceptual quality assessment, while also delivering strong enhancement results. The source code is available at: https://github.com/akihoti/DualNet-UIE.
Graph continual learning (GCL) is essential for deploying graph neural networks (GNNs) in dynamic environments, where models must mitigate catastrophic forgetting while adapting to evolving graph data. Existing replay-based methods face a fundamental trade-off between memory efficiency and the retention of useful structural context: storing isolated nodes discards relational information, whereas preserving explicit edges incurs substantial memory overhead. To alleviate this trade-off, we propose Spectral Prototype Generation Network (SPGNet), a replay-based framework that constructs compact, edge-free memory prototypes for graph continual learning. Specifically, SPGNet includes a prototype memory construction stage and a replay-based continual training stage. In the memory construction stage, SPGNet first allocates class memory budgets adaptively according to class size, then augments the original node features with neighborhood information via K-step low-pass graph propagation, and finally condenses the resulting structure-aware features into prototypes via clustering. This design yields more informative memory representations under a bounded replay budget, while avoiding the need to store explicit replayed subgraph edges. In the continual training stage, we further introduce a class-frequency reweighting objective to alleviate the imbalance between abundant current-task samples and limited replay prototypes. Extensive experiments under both class-incremental and task-incremental settings show that SPGNet achieves competitive and often superior performance compared with representative replay-based and learnable-memory baselines.
The widespread adoption of head-up display technology is constrained by the bulky free-form mirror assemblies in its optical system, which are difficult to miniaturize and thus hinder integration into modern space-constrained automotive cockpits. In this work, we present, for the first time, a metasurface-based windshield head-up display with two-dimensional exit pupil expansion, where the integration of a metasurface enables the windshield itself to function as a waveguide. A metasurface composed of a silicon nanopillar array approximately 570 nm thick enables an expanded exit pupil with a calculated of approximately 178 mm × 78 mm. By replacing the bulky free-form mirror assembly with the proposed metasurface-based architecture, the estimated system volume is reduced from approximately 10 L to approximately 1 L, corresponding to a reduction by a factor of approximately 10. Experiments on a fabricated metasurface sample demonstrate a diffraction angle of approximately 49°, which aligns with the optimized simulation results. Furthermore, angular tolerance studies confirm that effective light guidance can be maintained within a ± 10° range of incident angles. This design offers a promising pathway toward compact and superior-performance optical systems for next-generation head-up display applications
The operational instability of blue fluorescent organic light-emitting diodes (OLEDs), manifested as luminance decay and driving-voltage rise, remains a critical reliability issue for display applications. Here, a physics-informed prediction framework based on technology computer-aided design (TCAD) simulations using a drift–diffusion model was developed to predict the luminance–time (L–t) and voltage–time (V–t) characteristics of blue fluorescent OLEDs under different temperatures and current densities. For L–t, a saturation-type defect-generation model was used to describe effective polaron-induced defect buildup in the emissive layer. By linking the defect-generation term to TCAD-extracted polaron-related quantities, the model reproduced the measured L–t curves with fitted R2 of 0.996–0.998 at 25 °C and 0.994–0.999 at 35 °C for 7, 12, and 20 mA cm−2. It further achieved cross-current L–t prediction at 25 °C and 35 °C, with R2 values of 0.991–0.994 and 0.941–0.986, respectively. For V–t, a physics-based voltage-rise model was established by considering polaron-induced defects, triplet–polaron annihilation-generated defects, and impurity-related defects. A time-dependent TPA-defect buildup term enabled the model to capture the rapid initial voltage increase and subsequent slowdown, with fitted R2 values above 0.99 under all tested conditions. Using only the 7 and 12 mA cm−2 datasets, the model predicted the 20 mA cm−2 V–t behavior with R2 of 0.952 and 0.971 at 25 °C and 35 °C, respectively. These results demonstrate that TCAD-assisted physics-informed modeling can provide a practical route for evaluating blue OLED lifetime, predicting degradation under unmeasured current conditions, and supporting reliability-oriented optimization of OLED display devices.