The growth of online educational content, particularly slide-based video lectures, has created a need for tools that enhance navigation, comprehension, and accessibility. Many existing systems for video analysis are closed-source, hindering reproducibility and extension. To address this, we present the Lecture Video Analysis Toolkit, an open-source, proof-of-concept application designed for the multimodal analysis of slide video lectures. The toolkit integrates a processing pipeline that includes scene detection, visual entity extraction, transcription, optical character recognition (OCR), and semantic linking between spoken and visual content using embeddings. A key contribution is its interactive interface, motivated by early user feedback, that allows for customization of the viewing experience to suit individual preferences. The entire system is openly available and serves as a research prototype for validating the potential of multimodal analysis in creating more inclusive and improved learning experiences. A live demo is accessible at https://travis-seng.fr/svla, and the source code is openly available at https://github.com/travisseng/svla-toolkit.
The increasing amount of slide presentations in various sectors has amplified the need for effective slide layout and semantic analysis. However, we found that current slide datasets contain inconsistencies, mislabels, and incomplete annotations. Using them as a basis for developing deep learning-based slide analysis models could lead to models that are not robust and suboptimal. Addressing these challenges, we introduce SlideCraft, a tool for creating synthetic slide datasets that imitate real-world presentations. This tool overcomes the drawbacks of existing datasets by allowing users to create balanced, diverse, and accurately annotated slide data. We demonstrate SlideCraft's efficacy in enhancing slide layout analysis algorithms, focusing on its capability to improve dataset quality and object detection performance.
The old film restoration process involves many operations, one of which is the ability to identify defects that altered the film. This operation can be formulated as a binary segmentation problem and solved using state-of-the-art segmentation networks such as DeepLab v3+ or NAS-FPN. While being very powerful at describing the spatial characteristics of defects, these methods fail to take into account the fact that defects are also temporal anomalies. We therefore propose an architecture that builds on the correlation layer introduced in FlowNet to compensate for motion and eliminate potential false positives, features that look like defects but can be tracked over multiple images and are actually part of the scene. We also introduce a self-supervised pre-training process of the network, which precedes a fine-tuning phase to specifically adapt the detector to each film. Results show that our architecture, while being more compact and less resource-consuming than state-of-the-art methods, achieves higher precision and recall.
Educational content is increasingly available online. The most common media types for such content are audio and video. Many of these educational video are unprocessed -- they are simply captured with a camera and then uploaded onto video servers for viewers to watch. We believe that automatic video analysis that recovers the structure of educational videos would allow learners and creators to fully exploit the semantics of the captured content. In this paper, we outline our research plan towards better semantic understanding of educational videos. We highlight our work done so far: (i) we extended the FitVid Dataset to exploit the semantics of any type of learning video; and (ii) we improved upon existing state-of-the-art slide classification techniques.
We propose to detect defects in old movies, as the first step of a larger framework of old movies restoration by inpainting techniques. The specificity of our work is to learn a film restorer's expertise from a pair of sequences, composed of a movie with defects, and the same movie which was semi-automatically restored with the help of a specialized software. In order to detect those defects with minimal human interaction and further reduce the time spent for a restoration, we feed a U-Net with consecutive defective frames as input to detect the unexpected variations of pixel intensity over space and time. Since the output of the network is a mask of defect location, we first have to create the dataset of mask frames on the basis of restored frames from the software used by the film restorer, instead of classical synthetic ground truth, which is not available. These masks are estimated by computing the absolute difference between restored frames and defectuous frames, combined with thresholding and morphological closing. Our network succeeds in automatically detecting real defects with more precision than the manual selection with an all-encompassing shape, including some the expert restorer could have missed for lack of time.
Jean-Denis Durou合作论文数Universite Paul Sabatier4