ICASSP 2026 - 2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)(2026)
Department of Computer Science and Communication Engineering
被引用0|浏览2
摘要
Video Coding for Machines (VCM) is an emerging topic aiming to bridge compression for human and machine vision tasks. While existing VCM methods excel at spatial tasks (e.g., object detection), they have largely neglected temporal-dependent tasks (e.g., action recognition). Furthermore, the common cascaded pipeline connecting compression and analysis networks induces information bottlenecks and wasteful computation. This paper proposes a VCM framework that addresses these issues by exploiting information entirely within the compressed domain. We introduce a Spatial-Temporal Decoupled Latent Composition module (STDLC) that intercepts and processes features directly from the compression pipeline. The spatial path is composed of Swin-transformer blocks, and in the temporal path, a novel Temporal-Swin (T-Swin) block, finalized with a strong latent composition. Our framework demonstrates superior post-compression action recognition accuracy compared to traditional and prior VCM approaches on 2 standard datasets, with reduced computational costs while maintaining perceptual quality.
更多
查看译文
关键词
Video Coding for Machines,Action Recognition,Learned Video Compression,Deep Learning