Nowadays, high-quality images are pursued by both humans for better viewing experience and by machines for more accurate visual analysis. However, images are usually compressed before being consumed, decreasing their quality. It is meaningful to predict the perceptual quality of compressed images for both humans and machines, which guides the optimization for compression. This issue remains underexplored. In this paper, we propose a unified approach to fill this gap. Specifically, we create a deep learning-based model to predict Satisfied User Ratio (SUR) and Satisfied Machine Ratio (SMR) of compressed images simultaneously. We first pre-train a feature extractor network on a large-scale SMR-annotated dataset with human perception-related quality labels generated by diverse image quality models, which simulates the acquisition of SUR labels. Then, we leverage and fuse the extracted multi-layer features to predict SUR and SMR with a more capable network architecture. We propose a Difference Feature Residual Learning (DFRL) module to learn more discriminative difference features. We also design a Multi-Head Attention Aggregation and Pooling (MHAAP) layer to aggregate difference features and reduce their redundancy. We further introduce an MLP-Mixer module to integrate global spatial and channel information for subsequent SUR and SMR regression. Experimental results indicate that the proposed model significantly outperforms state-of-the-art SUR and SMR prediction methods by up to 48% on several challenging datasets. Moreover, our joint learning scheme of human and machine perceptual quality prediction tasks is effective at improving the performance of both.
The popularity of template-generated videos has recently experienced a significant increase on social media platforms. In general, videos from the same template share similar temporal characteristics, which are unfortunately ignored in the current compression schemes. In view of this, we aim to examine how such temporal priors from templates can be effectively utilized during the compression process for template-generated videos. First, a comprehensive statistical analysis is conducted, revealing that the coding decisions, including the merge, non-affine, and motion information, across template-generated videos are strongly correlated. Subsequently, leveraging such correlations as prior knowledge, a simple yet effective prior-driven compression scheme for template-generated videos is proposed. In particular, a mode decision pruning algorithm is devised to dynamically skip unnecessarily advanced motion vector prediction (AMVP) or affine AMVP decisions. Moreover, an improved AMVP motion estimation algorithm is applied to further accelerate reference frame selection and the motion estimation process. Experimental results on the versatile video coding (VVC) platform VTM-23.0 demonstrate that the proposed scheme achieves moderate time reductions of 14.31% and 14.99% under the Low-Delay P (LDP) and Low-Delay B (LDB) configurations, respectively, while maintaining negligible increases in Bj & oslash;ntegaard Delta Rate (BD-Rate) of 0.15% and 0.18%, respectively.
Efficient and large-scale evaluation of antibody-antigen neutralization is critical for accelerating antibody drug development. To address this need, we propose SPAAN, a deep-learning framework that predicts neutralization directly from antibody and antigen sequences. Rather than relying on experimentally determined structures, SPAAN learns from structural knowledge and biologically relevant molecular properties during training, enabling accurate predictions using sequence information alone. On the SARS-CoV-2 neutralization dataset, SPAAN consistently outperforms existing state-of-the-art methods. The model also shows strong interpretability by capturing key interaction patterns underlying antibody-antigen recognition. Furthermore, on the HIV neutralization dataset, SPAAN achieves state-of-the-art performance in multiple challenging scenarios involving previously unseen antibodies or antigens, demonstrating robust generalization ability. Overall, SPAAN provides an accurate, interpretable, and broadly applicable framework for antibody-antigen neutralization prediction, offering a practical tool to support large-scale antibody engineering and therapeutic discovery.
The growing popularity of virtual reality (VR) provides new opportunities for cultural heritage protection and online education. Many studies have investigated how to guide users to move to a specific view/position. However, the coguided interaction of position and gaze is still under researched. In this article, a novel interaction model named GazeTance Guidance (GTG) is proposed. This model utilizes head rotation and viewing distance to improve the effect of room-scale user guidance. To explore the efficiency of GTG, we conducted two within-subject studies with 25 and 31 participants navigating through virtual museums: first, we evaluated user acceptance of audio commentary and virtual guidance markers in terms of presence and user experience, and assessed the feasibility of memory tasks. second, we separately investigated the effects of gaze and distance guidance on motion sickness, cognitive load, memory efficiency, and user behavior. The experimental results prove that the guidance markers do not affect user experience and sense of presence. Compared with independent effects, the combination of gaze and distance guidance can reduce the user's motion sickness and significantly improve users' memory effect. These findings provide new inspiration for the interaction design of complex VR tours.
Multiview video streaming requires much higher bandwidth than conventional 2D video streaming. Current multi-view video streaming systems mainly leverage view interpolation to reduce redundancy between adjacent 2D views, and do not consider 3D spatial redundancy among all views. In this paper, we propose Argus, a multiview video streaming system that reduces 3D spatial redundancy by transmitting only a sparse subset of views and synthesizing the remaining views through 3D Gaussian reconstruction and splatting. Specifically, we propose: (i) a 3D Gaussian splatting-assisted multiview video coding method featuring real-time 3D Gaussian reconstruction and view synthesis, and (ii) a content-adaptive view selection method that dynamically selects a subset of views to optimize the visual quality of the synthesized views. We develop a prototype system of Argus, supporting up to 50-view autostereoscopic 3D display. Comprehensive experiments show that Argus achieves bitrate reduction by 41.33% and 44.12% compared with two baselines on two datasets, respectively, and supports real-time multiview video streaming at over 30 FPS.
In this article, we propose a multimodal robotic agent that addresses the limitations of passive perception and fixed-model deployment through a flexible large-model architecture. This architecture enables dynamic selection of multimodal large language models based on computational constraints. Building on this foundation, we introduce a logic-guided active perception strategy that decides which skills (e.g., knock and weigh) to employ based on intermediate reasoning, rather than exhaustively executing all possible actions. Our work focuses on cohesive skill integration within a unified control loop, optimizing both perception and action. This allows the agent to strategically probe objects’ visual, auditory, tactile, and weight attributes for accurate material inference and robust task completion. Extensive evaluations in the simulation environment highlight the efficiency and adaptability of our method, especially in active perception and latent information inference for robotic systems.
High-fidelity reconstruction of dynamic urban environments is a cornerstone of autonomous driving simulation and large-scale world modeling. While 3D Gaussian Splatting (3DGS) has established a new standard for real-time rendering, its reliance on expensive per-scene optimization limits scalability. Conversely, recent feedforward methods that infer Gaussian parameters offer faster speed but face fundamental bottlenecks: they are memory-prohibitive at high resolutions and struggle to fuse dense multi-view observations consistently. This paper presents L2D2-GS, a unified framework that reformulates generalizable reconstruction not as a one-shot regression, but as a robust iterative process of optimization and densification. To resolve the ambiguity of supervision in primitive generation, we propose a self-supervised densification policy that derives explicit reward signals from global reconstruction gains to guide local densification. Furthermore, we mitigate irreversible early-stage artifacts through a geometric regularization mechanism, utilizing reparameterization to constrain the optimization manifold and prevent convergence to poor local optima. Extensive experiments on the PandaSet and Waymo datasets demonstrate that our method achieves state-of-the-art reconstruction fidelity and strong zero-shot generalization, while using fewer primitives than competing baselines.
Free-Viewpoint Video (FVV) has emerged as a cornerstone of next-generation immersive media systems and attracted widespread attention. Previous methods primarily focus on short video sequences and suffer from significant performance degradation when processing long-horizon free-viewpoint video (LFVV). Motivated by bit allocation theory, we analyze dynamic-anchor-based volumetric video representation within a rate-distortion optimization framework and propose SoLAR, which is the first error-resilient streamable FVV framework that maintains stable reconstruction quality on long sequences without requiring group-of-pictures partitioning. We propose the Anchor Activation Dynamics (AAD), which enables dynamic anchors to model non-rigid transformations by dynamically activating informative anchors and suppressing redundant ones. Furthermore, we introduce Latent Discrepancy Aware Recalibration (LaDAR), which is a mechanism to identify discrepancies between latent representations and recalibrate the correspondences encoded in the network, effectively mitigating error propagation in LFVV without compromising real-time performance or storage compactness. Extensive experiments demonstrate that SoLAR achieves state-of-the-art reconstruction performance while maintaining minimum storage overhead, which provides a new direction for LFVV reconstruction and advances the practical deployment of immersive systems. Demo free-viewpoint videos are provided in the supplementary material.
This study addresses multi-robot collaborative anomaly handling challenges in self-driving scientific laboratories, where anomalies cause distorted results, equipment damage and safety risks. For Artificial Intelligence (AI) contributions, we design a hierarchical large language model (LLM)-driven intelligent task planning framework for dynamic action decomposition and adjustment, and propose an anomaly-triggered cross-behavior tree expansion algorithm with a spatiotemporal propagation suppression strategy to realize real-time anomaly localization and management. For engineering applications, we build a high-fidelity simulation engine for silicone preparation and a physical multi-robot experimental system, designing three representative anomaly tasks (spatial, temporal, hybrid) for validation. Experimental results show that at 50% anomaly probability, the framework achieves a 78.67% task success rate (74.45% higher than the baseline) and shortens the task duration by 150 s (1476.67→1326.67 s). Both simulated and real-world tests verify that the proposed AI methods effectively enhance the system’s fault tolerance and flexibility. This work provides a novel AI-driven engineering solution for the full automation transformation of scientific laboratories.
Text-video retrieval plays a pivotal role in cross-modal tasks, aiming to match textual descriptions with corresponding video content accurately. Existing methods often employ fine-grained feature matching to improve retrieval accuracy, but such approaches consume extensive computational resources. Conversely, coarse-grained feature matching between entire sentences and videos offers computational efficiency but may overlook the heterogeneous semantic concepts embedded within the data. To overcome these challenges, we develop the Disentangled Concept Matching (DCM) framework, designed as an imitation of human semantic perception processes. The framework utilizes disentangled representation learning to divide coarse-grained features into distinct semantic concepts represented as latent factors, effectively generating finer-grained features while reducing computational demands. To improve the accuracy of retrieval, we first propose the Composed Spatial-temporal Module (CSTM) to optimize the quality of multimodal feature extraction. Utilizing a branch-structured temporal modeling approach, CSTM effectively enhances the DCM model's comprehension of video content and temporal information, leading to the extraction of refined video features. Secondly, building on the optimized features, we propose the Adaptive Pooling Module (APM) to measure the confidence level of each latent factor matching during the process of decoupling concepts. APM enhances the fidelity of text and video concepts, thereby further ensuring the accuracy of matching after decoupling. With CSTM and APM, DCM accurately matches latent factors in lower dimensions, achieving significant improvements in computing efficiency and retrieval performance. Our experimental evaluations across standard datasets, namely MSR-VTT, LSMDC, MSVD, ActivityNet, and DiDeMo, demonstrate that the DCM framework achieves state-of-the-art performance, with Recall@1 scores of 48.7%, 25.6%, 48.4%, 45.0%, and 48.6% respectively. Compared to our previous model, the DCM framework shows improvements of 2.54%, 0.08%, 2.11%, 6.89%, and 6.35% respectively.
AI alignment aims to make AI systems behave in line with human intentions and values. As AI systems grow more capable, so do risks from misalignment. To provide a comprehensive and up-to-date overview of the alignment field, in this survey, we delve into the core concepts, methodology, and practice of alignment. First, we identify four principles as the key objectives of AI alignment: Robustness, Interpretability, Controllability, and Ethicality (RICE). Guided by these four principles, we outline the landscape of current alignment research and decompose them into two key components: forward alignment and backward alignment. The former aims to make AI systems aligned via alignment training, while the latter aims to gain evidence about the systems' alignment and govern them appropriately to avoid exacerbating misalignment risks. On forward alignment, we discuss techniques for learning from feedback and learning under the distribution shift. Specifically, we survey traditional preference modeling methods and reinforcement learning from human feedback and further discuss potential frameworks to reach scalable oversight for tasks where effective human oversight is hard to obtain. Within learning under distribution shift, we also cover data distribution interventions such as adversarial training that helps expand the distribution of training data and algorithmic interventions to combat goal misgeneralization. On backward alignment, we discuss assurance techniques and governance practices. Specifically, we survey assurance methods of AI systems throughout their lifecycle, covering safety evaluation, interpretability, and human value compliance. We discuss current and prospective governance practices adopted by governments, industry actors, and other third parties, aimed at managing existing and future AI risks. This survey aims to provide a comprehensive yet beginner-friendly review of alignment research topics. Based on this, we also release and continually update the website www.alignmentsurvey.com which features tutorials, collections of papers, blog posts, and other resources.
The recent development of feedforward 3D Gaussian Splatting (3DGS) presents a new paradigm to reconstruct 3D scenes. Using neural networks trained on large-scale multi-view datasets, it can directly infer 3DGS representations from sparse input views. Although the feedforward approach achieves high reconstruction speed, it still suffers from the substantial storage cost of 3D Gaussians. Existing 3DGS compression methods relying on scene-wise optimization are not applicable due to architectural incompatibilities. To overcome this limitation, we propose TinySplat, a complete feedforward approach for generating compact 3D scene representations. Built upon standard feedforward 3DGS methods, TinySplat integrates a training-free compression framework that systematically eliminates key sources of redundancy. Specifically, we introduce View-Projection Transformation (VPT) to reduce geometric redundancy by projecting geometric parameters into a more compact space. We further present Visibility-Aware Basis Reduction (VABR), which mitigates perceptual redundancy by aligning feature energy along dominant viewing directions via basis transformation. Lastly, spatial redundancy is addressed through an off-the-shelf video codec. Comprehensive experimental results on multiple benchmark datasets demonstrate that TinySplat achieves over $100 imes $ compression for 3D Gaussian data generated by feedforward methods. Compared to the state-of-the-art compression approach, we achieve comparable quality with only 6% of the storage size. Meanwhile, our compression framework requires only 25% of the encoding time and 1% of the decoding time.
Lossless compression of volumetric medical images is of paramount importance for clinical and research applications where data fidelity is essential. Traditional compression methods are often limited in efficiency due to rigid, handcrafted models. Conversely, deep neural network (DNN)-based compression methods, while effective, demand substantial computational resources, hindering deployment in resource-constrained settings. To address these challenges, we propose a novel tri-plane context tree (TCT)-based method for lossless volumetric medical image compression that delivers high performance without relying on DNNs or external training data. To exploit intra-slice and interslice redundancies, we introduce a compact tri-plane context representation that decomposes complex 3D context modeling into efficient 2D modeling on three orthogonal planes. By integrating this representation with a context tree framework, we develop an input-specific TCT model employing an adaptive binary tree structure. At each tree node, the model dynamically selects from a suite of tri-plane based predictors and contextual feature extractors, enabling data-adaptive context modeling tailored to local structural characteristics. Instead of offline training, we sample a subset of the input volume to learn the TCT model by optimizing the minimum description length (MDL) through iterative construction and pruning. With the learned TCT model, each pixel retrieves its corresponding context, computes the prediction residual using the predictor dictated by the context, and performs entropy encoding based on the associated histograms. Experimental results demonstrate that the proposed method achieves compression performance on par with recent DNN-based methods on multiple datasets, while maintaining low computational cost and fast coding speeds, making it highly applicable in practice.
While Large Vision-Language Models (LVLMs), represented by LLaVA and GPT-4V, have demonstrated remarkable capabilities, their visual inputs remain vulnerable to adversarial attacks, posing significant security risks. Existing defense methods predominantly target single-task scenarios (e.g., zero-shot classification) and consequently lack generalizability across various multimodal tasks. To address this limitation, we propose a dual adversarial fine-tuning framework that jointly optimizes visual and semantic supervision signals from two modalities, enhancing model robustness while generalizing across multiple downstream tasks. The proposed framework comprises two core components, i.e., Visual supervision branch and Semantic supervision branch. The former branch leverages features from clean images, extracted via a frozen original vision encoder, to guide adversarial robustness while the latter incorporates caption-image alignment as a contextual signal to preserve semantic coherence under attack. Moreover, our method achieves cross-task robustness by simply replacing the CLIP vision encoder in the original model, with no need of separate task-specific retraining or architecture modifications.Extensive experiments demonstrate that our approach outperforms the state-of-the-art method in adversarial robustness evaluation across zero-shot classification, image captioning, and visual question answering (VQA) tasks.
Learned image compression (LIC) methods have shown promising results and achieved superior performance compared to traditional image compression methods. Due to the neglect of the utilization of cross-component correlations, there is still a potential for further performance improvement. In this paper, we first explore the inter-channel correlations of different color spaces and transform the image compression problem in RGB color space into that in YUV color space, which has cross-component prior information. We propose a novel image compression method that leverages local-to-global cross-component prior modeling, utilizing a cross-component attention mechanism to improve coding performance. First, we design the cross-component prior gate (CPG) to model the cross-component prior information based on attention mechanism. Inspired by common knowledge in data compression, luma component (Y) contains more details and textural/structural information compared to chroma components (UV). The proposed method can make full use of the cross-component guidance information from luma to chroma components to achieve effective image compression. Experimental results demonstrate that the proposed method can achieve superior performance compared to existing learned image compression methods. The proposed method can achieve 9.20% rate savings compared to the image compression standard Versatile Video Coding (VVC) Test Model (VTM-11.0) on Kodak dataset.
Accurate and consistent evaluation is crucial for decision-making across numerous fields, yet it remains challenging due to inherent subjectivity, variability, and scale. Large language models (LLMs) have achieved remarkable success, leading to "LLM-as-a-judge," where LLMs serve as evaluators for complex tasks. With their ability to process diverse data types and provide scalable assessments, LLMs present a compelling alternative to traditional expert-driven evaluations. However, ensuring the reliability of LLM-as-a-judge systems remains a significant challenge requiring careful design and standardization. This paper provides a comprehensive survey of LLM-as-a-judge, offering a formal definition and detailed classification while addressing the core question of how to build reliable LLM-as-a-judge systems. We explore strategies to enhance reliability, including improving consistency, mitigating biases, and adapting to diverse scenarios. We propose methodologies for evaluating reliability, supported by a novel benchmark. To advance development and deployment, we discuss practical applications, challenges, and future directions. Our contributions span multiple levels: we establish conceptual boundaries, reorganize fragmented literature into a unified framework, and propose a reliability-oriented benchmark. We articulate a forward-looking research agenda, offering theoretical foundations and practical guidance for constructing reliable and trustworthy LLM-as-a-judge systems.
Learned video compression (LVC) designed for energy efficiency on resource-constrained platforms is widely recognized as a critical challenge hindering its practical deployment. Entropy coding, the primary bottleneck for efficient inference, is essential for the transformation between symbols and bitstreams. Among various entropy algorithms, range-Asymmetric Numeral Systems (rANS) is preferred for its low computational complexity and high compression efficiency. However, hardware acceleration for rANS remains limited, leading to computational bottlenecks on edge CPUs and impairing symbol-bitstream conversion. This paper presents HANS, an FPGA-based hardware implementation of rANS for LVC, which addresses these challenges through two key innovations. First, we introduce the “Integer Division Transfer Theorem,” which replaces resource-intensive division operations with multiplication and bit-shifting, significantly boosting encoder throughput. Second, we propose a Block-based Retrieval (BBR) strategy for the inverse CDF lookup in the decoder, balancing both storage efficiency and access latency. Experimental results demonstrate that HANS achieves up to 81.51% higher throughput and 66.59% lower energy consumption per symbol compared to CPU-based edge solutions. When integrated into various cutting-edge LVC algorithms, our framework supports real-time 4K@80FPS entropy encoding and over 4K@30FPS entropy decoding, ensuring high-throughput practical deployment. This work represents the first robust, large-scale hardware implementation of a rANS encoder/decoder, enabling efficient ultra-high-definition LVC on edge devices and providing valuable insights for future edge-oriented data compression systems.
Traditional image compression prioritizes pixel fidelity but often preserves details irrelevant to downstream vision tasks. Compressing task-specific representations instead better aligns with task semantics, yet redundant information persists across correlated tasks. Existing multi-task compression methods typically rely on static dependency structures, leading to redundant bit allocation across correlated tasks and suboptimal rate-distortion performance. We present Adaptive Task Dependency Compression (ATDC), a framework that models per-image task relationships and encodes representations following an adaptive directed acyclic graph (DAG). ATDC infers pairwise task predictability via a learned correlation matrix, constructs a dynamic DAG to determine the optimal compression order, and encodes each task conditionally on its predecessors, achieving predictive redundancy removal and asymmetric information sharing across tasks. Experiments on the Taskonomy dataset demonstrate consistent gains in rate–distortion efficiency and task accuracy over both human-oriented codecs and state-of-the-art multi-task compression methods.The learned DAGs reveal interpretable, content-dependent task hierarchies, establishing adaptive dependency modeling as a principled paradigm for multi-task representation compression.
The emergence of Large Vision-Language Models (LVLMs) marks significant strides towards achieving general artificial intelligence. However, these advancements are accompanied by concerns about biased outputs, a challenge that has yet to be thoroughly explored. Existing benchmarks are not sufficiently comprehensive in evaluating biases due to their limited data scale, single questioning format and narrow sources of bias. To address this problem, we introduce VLBiasBench, a comprehensive benchmark designed to evaluate biases in LVLMs. VLBiasBench features a dataset that covers nine distinct categories of social biases, including age, disability status, gender, nationality, physical appearance, race, religion, profession, social economic status, as well as two intersectional bias categories: race × gender and race × social economic status. To build a large-scale dataset, we use Stable Diffusion XL model to generate 46,848 high-quality images, which are combined with various questions to create 128,342 samples. These questions are divided into open-ended and close-ended types, ensuring thorough consideration of bias sources and a comprehensive evaluation of LVLM biases from multiple perspectives. We conduct extensive evaluations on 15 open-source models as well as two advanced closed-source models, yielding new insights into the biases present in these models.
BackgroundDeep learning faces a significant bottleneck in medical image analysis due to its reliance on large-scale, expert-annotated datasets. This challenge is acute in ophthalmology, particularly for detecting early-stage diseases like mild Diabetic Retinopathy (DR1), where subtle lesions and a scarcity of annotations limit supervised learning approaches.MethodsWe propose a generalizable eye disease detection framework based on Zero-shot Learning (ZSL) that mimics clinical reasoning. Using the LCFP-14M dataset, a large-scale fundus image resource we present in this work, our method first identifies disease correlations via a Siamese network. It then transfers knowledge by segmenting DR1-specific lesions from a highly correlated source disease and employs a ResNet-Agglomerative clustering pipeline to enable unsupervised detection of DR1 without using any labeled DR1 cases.ResultsHere we show that the proposed framework enables effective DR1 detection without annotated DR1 data. The model achieves an accuracy of 0.8337, precision of 0.8700, recall of 0.7456, F1 score of 0.8030, and ROC-AUC of 0.9226, outperforming most supervised baselines on external test datasets.ConclusionsOur findings demonstrate that ZSL can simulate clinical diagnostic logic and generalize to unseen eye diseases, offering a promising approach for automated screening where labeled data are scarce.