2025 19TH INTERNATIONAL CONFERENCE ON MACHINE VISION AND APPLICATIONS, MVA(2025)
Univ Tsukuba
被引用0|浏览7
摘要
We propose an efficient and generalizable method for skeleton-based action recognition using a novel representation called the Superposed Shape Subspace (SSS). Our approach encodes both structural and temporal dynamics of skeletal sequences by modeling multiple frames as a unified subspace in a high-dimensional space. Unlike deep learning methods that require large datasets and GPU resources, our method enables fast, CPU-based training and inference by comparing canonical angles between subspaces.To further improve efficiency, we introduce an optimal frame selection strategy that identifies the most informative frames, significantly reducing computational cost without compromising accuracy. Experiments on the First-Person Hand Action Benchmark show that our method achieves competitive performance (89% accuracy across 45 classes), processes over 100 frames per second on CPU, and supports one-shot learning with minimal data. These characteristics make it highly suitable for real-time and edge computing applications.