2025 3rd International Conference on Foundation and Large Language Models (FLLM)(2025)
Contents Convergence Software Research Institute
被引用0|浏览3
摘要
This paper presents a novel big-data generation framework, COME-HVLM (Cognitive Object–Motion–Environment Hybrid Vision–Language Model), which integrates multiple deep learning–based vision–language architectures into a unified hybrid model. Built upon the foundations of established and advanced VLMs—such as YOLO, BLIP, CLIP, and VILA—COME-HVLM is designed to automatically detect and interpret active contextual elements including objects, motions, and environmental features within CCTV surveillance video streams. Operating primarily in edge-computing environments of smart city surveillance systems, the framework systematically extracts and encodes these contextual elements—termed CCTV-Surveillance Active Contexts—into a structured representation format known as COME-Code (Contextual Object–Motion–Environment Code), expressed in JSON. Furthermore, COME-HVLM supports the augmentation of these COME-Code datasets with additional semantic attributes and descriptive properties, thereby enabling enriched and machine-processable contextual representations for large-scale CCTV video analytics and intelligent surveillance applications.
更多
查看译文
关键词
Bigdata Lake Platforms,CCTV-Surveillance,Video-Contents,Active Contexts,Cognitive Model,Hybrid Vision Language Models,CCTV-Video Contextualization