This paper presents the approach of the KsLab NUT team in the TRECVID 2020[1] VTT Task. We propose a method that focuses on reducing the processing time. By extracting only important frames from videos and using them for processing, we were able to drastically reduce the number of frames to be processed while achieving certain levels of accuracy. Furthermore, we also applied the methods used for text summarization to examine their performance.
更多
查看译文
关键词
Video Summarization,Image Captioning,Key Frame Extraction,Scene Graph Generation,Video Description