Toward Abstraction-Level Event Retrieval in Large Video Collections: Leveraging Human Knowledge and LLM-Based Reasoning in the Ho Chi Minh City AI Challenge 2025 | AMiner
Toward Abstraction-Level Event Retrieval in Large Video Collections: Leveraging Human Knowledge and LLM-Based Reasoning in the Ho Chi Minh City AI Challenge 2025
Large-scale video collections require retrieval systems that understand both semantic content and temporal structure. While vision–language models have improved cross-modal retrieval, searching long-form videos remains challenging, especially for abstract queries and event sequences. This paper presents an overview of the Ho Chi Minh City AI Challenge 2025, a large-scale evaluation campaign inspired by the Video Browser Showdown and the Lifelog Search Challenge. The challenge includes multiple query types: Textual and Visual Known-Item Search and Question Answering. It also introduces a new task, Temporal Retrieval and Alignment of Key Events (TRAKE), which involves retrieving a relevant video and aligning multiple key events in time. Progressive hints are provided to simulate realistic query refinement. We describe the dataset, task design, evaluation protocol, and analyze team performance and retrieval behavior. The results highlight both recent progress and remaining challenges in temporally-aware video retrieval.
更多
查看译文
关键词
Video Retrieval,Multimodal Search,Temporal Event Localization,Vision–Language Models,Interactive Multimedia Retrieval