2025 IEEE International Workshop on Multimedia Signal Processing (MMSP)(2025)
Tsinghua University
被引用0|浏览0
摘要
Document layout analysis, a critical process in automated document processing, traditionally relies on object detection techniques, primarily focusing on the structural segmentation of documents. However, these approaches often fall short in comprehensively understanding the semantic content within the text, leading to a disjointed analysis of document structure and content. To address this, we propose a novel methodology that combines text clustering with multi-modal graph convolution networks, aiming to integrate structural detection with semantic understanding. Our approach starts with text detection, followed by encoding using a large language model. Subsequently, we integrate visual and positional data using Graph Neural Networks to perform clustering, creating a synergy between the textual and structural aspects of documents. Extensive experiments on mainstream datasets demonstrate that our method significantly outperforms existing approaches, especially in understanding text-centric document layouts. This paper contributes to the field by offering a novel, semantically-enriched approach to document layout analysis, enhancing the capabilities of automated document processing systems in handling diverse and complex document formats.