Hallucinations in video-capable vision-language models (Video-VLMs) remain frequent and high-confidence, while existing uncertainty metrics often fail to align with correctness. We introduce VideoHEDGE, a modular framework for hallucination detection in video question answering that extends entropy-based reliability estimation from images to temporally structured inputs. Given a video-question pair, VideoHEDGE draws a baseline answer and multiple high-temperature generations from both clean clips and photometrically and spatiotemporally perturbed variants, then clusters the resulting textual outputs into semantic hypotheses using either Natural Language Inference (NLI)-based or embedding-based methods. Cluster-level probability masses yield three reliability scores: Semantic Entropy (SE), RadFlag, and Vision-Amplified Semantic Entropy (VASE). We evaluate VideoHEDGE on the SoccerChat benchmark using an LLM-as-a-judge to obtain binary hallucination labels. Across three 7B Video-VLMs (Qwen2-VL, Qwen2.5-VL, and a SoccerChat-finetuned model), VASE consistently achieves the highest ROC-AUC, especially at larger distortion budgets, while SE and RadFlag often operate near chance. We further show that embedding-based clustering matches NLI-based clustering in detection performance at substantially lower computational cost, and that domain fine-tuning reduces hallucination frequency but yields only modest improvements in calibration. The hedge-bench PyPI library enables reproducible and extensible benchmarking, with full code and experimental resources available at https://github.com/Simula/HEDGE#videohedge .
Shot Boundary Detection (SBD), which identifies scene (or "shot") changes, is a core step in video analysis pipelines such as summarization and highlight generation. Yet, it remains challenging in sports broadcasts because rapid camera motion, frequent camera switches, sport-specific transitions and graphic overlays often cause false detections and poor cross-domain generalization. In this paper, we address shot boundary detection in professional sports broadcasts, focusing on ice hockey and soccer. We propose a sports-oriented model based on a fine-tuned R(2+1)D 3D CNN, trained to detect hard cuts, gradual transitions, and logo-based replay effects. The model is evaluated on both goal-centered clips and full-match broadcast footage, and is benchmarked against two state-of-the-art baselines: TransNetV2 for accuracy and PySceneDetect for efficiency. Our approach consistently achieves higher precision, recall, and F1-score across all evaluation settings, while demonstrating strong cross-league generalization within ice hockey and cross-sport generalization to soccer. We release our pretrained sports-specific SBD model as an open-source Python package, enabling straightforward integration into existing video analysis pipelines.
Injury occurrence in football poses significant challenges for athletes and teams, carrying personal, competitive, and financial consequences. While machine learning has been applied to injury prediction before, existing approaches often rely on static pre-season data and binary outcomes, limiting their real-world utility. This study investigates the feasibility of using a DeepHit neural network to forecast time-to-injury from longitudinal athlete monitoring data, while providing interpretable predictions. The analysis utilised the publicly available SoccerMon dataset, containing two seasons of training, match, and wellness records from elite female footballers. Data was pre-processed through cleaning, feature engineering, and the application of three imputation strategies. Baseline models (Random Forest, XGBoost, Logistic Regression) were optimised via grid search for benchmarking, while the DeepHit model, implemented with a multilayer perceptron backbone, was evaluated using chronological and leave-one-player-out (LOPO) validation. DeepHit achieved a concordance index of 0.762, outperforming baseline models and delivering individualised, time-varying risk estimates. Shapley Additive Explanations (SHAP) identified clinically relevant predictors consistent with established risk factors, enhancing interpretability. Overall, this study provides a novel proof of concept: survival modelling with DeepHit shows strong potential to advance injury forecasting in football, offering accurate, explainable, and actionable insights for injury prevention across competitive levels.
Understanding player orientation is a critical component of sports analytics, offering insights into gameplay strategies and player behavior. This paper presents HockeyOrient, a novel dataset for classifying the orientation of ice hockey players based on their poses. The dataset comprises 9,700 manually annotated frames, selected randomly and non-sequentially, taken from Swedish Hockey League (SHL) games during the 2023 and 2024 seasons. Each player image is cropped from game footage and categorized into one of eight orientation classes: top, top-right, right, bottom-right, bottom, bottom-left, left, and top-left. The dataset includes diverse scenarios, such as different teams, jersey colors, referees, and goaltenders with unique protective gear. Alongside the dataset, we provide an open-source classification model trained on the dataset using the SqueezeNet architecture, achieving an F1 score of 75% across all the classes. This work addresses a significant gap in ice hockey analytics, enabling advanced player tracking and gameplay analysis. The dataset and the trained model can be accessed publicly on Hugging Face under an open-source license (https://huggingface.co/datasets/SimulaMet-HOST/HockeyOrient).
VoiceVision is an AI-powered system designed for intelligent speaker-focused video cropping and speaker-aware content indexing. Built on top of the TalkNet audio-visual speaker diarization backbone, Voice Vision detects active speakers in multi-speaker videos and dynamically crops and reframes the video to center on the current speaker, creating smooth visual transitions. In addition to smart cropping, the system integrates automatic speech recognition (ASR) using Whisper to generate accurate transcriptions, which are further processed through a transcript attribution module to associate spoken segments with specific speakers. A dedicated speech search module enables efficient retrieval and indexing of content based on keywords or speaker identity. Voice Vision supports automatic aspect ratio adaptation (9:16, 1:1, 4:5) to generate social-media- optimized outputs. By combining speaker-aware video cropping with searchable, speaker-attributed transcriptions, Voice Vision simplifies content creation, indexing, and sharing for interview- style or conversational videos on platforms such as TikTok, Instagram, and YouTube Shorts. A demonstration of the system's capabilities is presented at https://youtu.be/SBSqOyMpe60
Accurate localization of players and objects in ice hockey is essential for advanced analytics, tactical analysis, and automated content generation. Traditional homography estimation approaches rely heavily on explicit camera calibration or heuristic methods, which are computationally intensive and sensitive to camera variations. This paper introduces Hockey2D, a robust framework leveraging YOLO-based pose estimation to identify rink keypoints, combined with RANSAC-based homography optimization for precise spatial mapping. To handle broadcast variations and occlusions effectively, a shot-type classifier filters input frames, selectively applying homography estimation to suitable scenes. By integrating keypoint and object detection, our approach achieves accurate 2D localization of players, referees, and goalkeepers. Evaluations conducted on ice hockey-specific datasets confirm both computational efficiency and localization accuracy. The proposed method highlights the practical viability of AI-driven localization for interactive storytelling, game summarization, and tactical coaching. A video demonstration is available at https://youtu.be/JCnX4N4fi8I.
The integration of artificial intelligence in sports analytics has transformed soccer video understanding, enabling real-time, automated insights into complex game dynamics. Traditional approaches rely on isolated data streams, limiting their effectiveness in capturing the full context of a match. To address this, we introduce SoccerChat, a multimodal conversational AI framework that integrates visual and textual data for enhanced soccer video comprehension. Leveraging the extensive SoccerNet dataset, enriched with jersey color annotations and automatic speech recognition (ASR) transcripts, SoccerChat is fine-tuned on a structured video instruction dataset to facilitate accurate game understanding, event classification, and referee decision making. We benchmark SoccerChat on action classification and referee decision-making tasks, demonstrating its performance in general soccer event comprehension while maintaining competitive accuracy in referee decision making. Our findings highlight the importance of multimodal integration in advancing soccer analytics, paving the way for more interactive and explainable AI-driven sports analysis. https://github.com/simula/SoccerChat
In this paper, we present SoccerNet-Echoes, an extension of the SoccerNet dataset which has been curated by augmenting the 550 games in the original dataset with multilingual audio commentary transcriptions, with a pipeline utilizing OpenAI's Whisper models for transcription and Google Translate for translation to English. We demonstrate the potential of SoccerNet-Echoes through several applications. Our experiments reveal that incorporating ASR-generated transcripts as a third modality alongside audio and video can improve the performance of multimodal event detection, with our audio-video-text model achieving a top F1-score of 0.7175. We also introduce a novel framework that leverages Large Language Models (LLMs) to extract both predefined, official events, as well as unscripted, unofficial events directly from the commentary. Our evaluation shows that the Gemini-1.5-Pro model effectively identifies official events from text alone, and that LLM-generated game summaries are more descriptive and accurate when using SoccerNet-Echoes compared to using only structured event data. Surprisingly, our experiments also show that feeding powerful LLMs like Gemini-1.5-Pro with visual data may not improve results compared to their text-only counterpart, but rather degrade performance, for which we analyze the potential reasons. By releasing SoccerNet-Echoes, we provide a resource for the scientific community and offer benchmarks that highlight the current capabilities and limitations of ASR and LLM technologies in the domain of multimodal sports analysis.
Quantifying sponsor visibility in sports broadcasts is a critical marketing task traditionally hindered by manual, subjective, and unscalable analysis methods. While automated systems offer an alternative, their reliance on axis-aligned Horizontal Bounding Box (HBB) leads to inaccurate exposure metrics when logos appear rotated or skewed due to dynamic camera angles and perspective distortions. This paper introduces ExposureEngine, an end-to-end system designed for accurate, rotation-aware sponsor visibility analytics in sports broadcasts, demonstrated in a soccer case study. Our approach predicts Oriented Bounding Box (OBB) to provide a geometrically precise fit to each logo regardless of the orientation on-screen. To train and evaluate our detector, we developed a new dataset comprising 1,103 frames from Swedish elite soccer, featuring 670 unique sponsor logos annotated with OBBs. Our model achieves a mean Average Precision (mAP@0.5) of 0.859, with a precision of 0.96 and recall of 0.87, demonstrating robust performance in localizing logos under diverse broadcast conditions. The system integrates these detections into an analytical pipeline that calculates precise visibility metrics, such as exposure duration and on-screen coverage. Furthermore, we incorporate a language-driven agentic layer, enabling users to generate reports, summaries, and media content through natural language queries. The complete system, including the dataset and the analytics dashboard, provides a comprehensive solution for auditable and interpretable sponsor measurement in sports media. An overview of the ExposureEngine is available online.
The fast paced nature of ice hockey presents unique challenges for object detection, particularly in tracking the puck-a small, fast moving object that is critical to gameplay analysis. This paper introduces HockeyAI, a novel open source dataset specifically designed for multi-class object detection in ice hockey. The dataset includes 2,101 high resolution frames extracted from professional games in the Swedish Hockey League (SHL), annotated in the You Look Only Once (YOLO) format. Annotations span 7 classes, covering dynamic objects such as the players and the puck, as well as static rink elements such as goalposts and face-off circles. The dataset is derived from diverse SHL games across multiple seasons and teams, ensuring a rich variety of scenarios reflective of real world gameplay. A fine tuned YOLOv8 medium model is also provided, demonstrating high performance across all classes. Key comparisons highlight the dataset's advancements over existing resources, addressing their limitations such as incomplete class coverage, low resolution, and in-consistent annotations. The dataset, model, and an interactive demo are publicly available on Hugging Face under an open source license, fostering further collaboration in sports related computer vision applications (https://huggingface.co/SimulaMet-HOST/HockeyAI).
One of the key metrics for player and team performance in football is speed. Traditionally, measuring player speed either requires extensive human effort or expensive equipment. In this work, we investigate methods for manual and automatic player speed extraction directly from broadcast football videos. We implement a pipeline combining player detection, tracking, field mapping, and position measurement. We experiment with different configurations of the pipeline, concluding that a setting with YOLOv11 for detection, StrongSORT for tracking, PnL-Calib for field mapping and keypoint detection, and position measurement through homography yields the best results for our current scope and test datasets. This work serves as a proof-of-concept analytics solution for player speed extraction, which can be adopted by teams of all sizes and means, and built upon by researchers. Our pipeline implementation and two newly curated player speed datasets (Begnadalen and TACDEC++) are openly available for the sports and scientific community.
Precise mapping of ice hockey rinks is critical for applications such as player tracking, game strategy analysis, and broadcast enhancements. Traditional methods often rely on manual annotations or simplistic models that fail to account for the rink's complex geometry and dynamic in-game conditions. To address these limitations, we present HockeyRink, a novel dataset comprising 56 meticulously annotated keypoints corresponding to significant landmarks on a standard hockey rink, including face-off dots, goalposts, and blue lines. Leveraging the YOLOv8-Large pose estimation architecture, we adapted the model to treat the rink as a single 'pose' object, enabling accurate keypoint predictions tailored to the nuances of hockey rink imagery. Our dataset, derived from diverse hockey game footage from the Swedish Hockey League (SHL), facilitates applications in homography estimation, 2D/3D scene mapping, and tactical overlays. By detecting and mapping rink keypoints, one can compute precise transformations between the image plane and the rink's physical dimensions, enabling the overlay of player trajectories and tactical insights onto broadcast footage. This work addresses the unique challenges of ice hockey, such as occlusions, rapid camera movements, and varying lighting conditions, and offers a robust foundation for future research and innovation in sports analytics. The HockeyRink dataset and trained model are openly available at: https://huggingface.co/SimulaMet-HOST/HockeyRink.
Data analysis for athletic performance optimization and injury prevention is of tremendous interest to sports teams and the scientific community. However, sports data are often sparse and hard to obtain due to legal restrictions, unwillingness to share, and lack of personnel resources to be assigned to the tedious process of data curation. These constraints make it difficult to develop automated systems for analysis, which require large datasets for learning. We therefore present SoccerMon, the largest soccer athlete dataset available today containing both subjective and objective metrics, collected from two different elite women’s soccer teams over two years. Our dataset contains 33,849 subjective reports and 10,075 objective reports, the latter including over six billion GPS position measurements. SoccerMon can not only play a valuable role in developing better analysis and prediction systems for soccer, but also inspire similar data collection activities in other domains which can benefit from subjective athlete reports, GPS position information, and/or time-series data in general.
Social media plays a significant role for sports organizations with millions of active fans, but publishing highlights is often a tedious manual operation. With the development of AI, new tools are available for content generation and personalization to engage audiences. We propose an AI-based multimedia production framework for the automatic publishing of soccer and ice hockey highlights on social media, and disseminate our experiences with developing dedicated pipelines for event detection and classification, player detection and tracking, highlight clipping, cropping, thumbnail generation, game summarization, caption generation, and social media sharing.
We introduce Kvasir-VQA, an extended dataset derived from the HyperKvasir and Kvasir-Instrument datasets, augmented with question- and-answer annotations to facilitate advanced machine learning tasks in Gastrointestinal (GI) diagnostics. This dataset comprises 6,500 annotated images spanning various GI tract conditions and surgical instruments, and it supports multiple question types including yes/no, choice, location, and numerical count. The dataset is intended for applications such as image captioning, Visual Question Answering (VQA), text-based generation of synthetic medical images, object detection, and classification. Our experiments demonstrate the dataset's effectiveness in training models for three selected tasks, showcasing significant applications in medical image analysis and diagnostics. We also present evaluation metrics for each task, highlighting the usability and versatility of our dataset. The dataset and supporting artifacts are available at https://datasets.simula.no/kvasir-vqa.
Extracting meaningful insights from large and complex datasets poses significant challenges, particularly in ensuring the accuracy and relevance of retrieved information. Traditional data retrieval methods such as sequential search and index-based retrieval often fail when handling intricate and interconnected data structures, resulting in incomplete or misleading outputs. To overcome these limitations, we introduce Structured-GraphRAG, a versatile framework designed to enhance information retrieval across structured datasets in natural language queries. Structured-GraphRAG utilizes multiple knowledge graphs, which represent data in a structured format and capture complex relationships between entities, enabling a more nuanced and comprehensive retrieval of information. This graph-based approach reduces the risk of errors in language model outputs by grounding responses in a structured format, thereby enhancing the reliability of results. We demonstrate the effectiveness of Structured-GraphRAG by comparing its performance with that of a recently published method using traditional retrieval-augmented generation. Our findings show that Structured-GraphRAG significantly improves query processing efficiency and reduces response times. While our case study focuses on soccer data, the framework's design is broadly applicable, offering a powerful tool for data analysis and enhancing language model applications across various structured domains.
We present SoccerGuard, a novel framework for predicting injuries in women's soccer using Machine Learning (ML). This framework can ingest data from multiple sources, including subjective wellness and training load reports from players, objective GPS sensor measurements, third-party player statistics, and injury reports verified by medical personnel. We experiment with a number of different settings related to synthetic data generation, input and output window sizes, and ML models for prediction. Our results show that, given the right configurations and feature combinations, injury event prediction can be undertaken with considerable accuracy. The optimal results are achieved when input windows are reduced and larger combined output windows are defined, in combination with an ideally balanced data set. The framework also includes a dashboard with a user-friendly Graphical User Interface (GUI) to support interactive analysis and visualization.
In the realm of soccer analytics, the need for efficient and accurate information retrieval is crucial. In this paper, we introduce SoccerGraphRAG, a framework designed to facilitate the retrieval of soccerrelated information through natural language queries. This system leverages knowledge graphs, created from the recently released SoccerNetEchoes dataset which includes transcriptions of soccer game audio commentaries. Soccer-GraphRAG aims to streamline the retrieval, access, and analysis of soccer data, providing insights with precision and contextual relevance. This framework is ideally suited for analyzing player performance, as well as for engaging in question answering (Q&A) and summarizing tasks.
Deep learning has achieved immense success in computer vision and has the potential to help physicians analyze visual content for disease and other abnormalities. However, the current state of deep learning is very much a black box, making medical professionals skeptical about integrating these methods into clinical practice. Several methods have been proposed to shed some light on these black boxes, but there is no consensus on the opinion of medical doctors that will consume these explanations. This paper presents a study asking medical professionals about their opinion of current state-of-the-art explainable artificial intelligence methods when applied to a gastrointestinal disease detection use case. We compare two different categories of explanation methods, intrinsic and extrinsic, and gauge their opinion of the current value of these explanations. The results indicate that intrinsic explanations are preferred and that physicians see value in the explanations. Based on the feedback collected in our study, future explanations of medical deep neural networks can be tailored to the needs and expectations of doctors. Hopefully, this will contribute to solving the issue of black box medical systems and lead to successful implementation of this powerful technology in the clinic.