Dry eye disease is common and heterogeneous, but whether retrospective blood metabolite profiles capture dry-eye signal remains unclear. We performed a retrospective machine-learning analysis of untargeted serum metabolomics in TwinsUK, using approximately 4,200 serum samples from approximately 1,400 participants and two questionnaire-derived outcomes: a broad Women’s Health Study dry-eye definition and a more specific highly symptomatic phenotype. Serum samples were collected from 1996 to 2014, whereas dry-eye outcomes were defined from 2013 questionnaires; the median sample-to-outcome gap in the main repeated-row analysis was 7.0 years. In family-aware analyses, metabolomics added statistically detectable but limited discrimination beyond covariates, increasing receiver operating characteristic area under the curve from 0.43 to 0.57 for the broader outcome and from 0.42 to 0.60 for the highly symptomatic phenotype. Preprocessing strongly influenced discrimination. After false-discovery correction, 13 metabolites were associated with highly symptomatic dry eye, and none with the broader outcome; five remained significant in low-value imputation sensitivity analyses. These findings suggest modest dry-eye-related serum metabolomic signal, shaped by preprocessing and most evident for the highly symptomatic phenotype.
Large Language Models (LLMs) rely on various decoding strategies to generate text, and these choices can significantly affect output quality. In healthcare, where accuracy is critical, the impact of decoding strategies remains underexplored. We investigate this effect in five open-ended medical tasks, including translation, summarization, question answering, dialogue, and image captioning, evaluating 11 decoding strategies with medically specialized and general-purpose LLMs of different sizes. Our results show that deterministic strategies generally outperform stochastic ones: beam search achieves the highest scores, while η and top-k sampling perform worst. Slower decoding methods tend to yield better quality. Larger models achieve higher scores overall but have longer inference times and are no more robust to decoding. Surprisingly, while medical LLMs outperform general ones in two of the five tasks, statistical analysis shows no overall performance advantage and reveals greater sensitivity to decoding choice. We further compare multiple evaluation metrics and find that correlations vary by task, with MAUVE showing weak agreement with BERTScore and ROUGE, as well as greater sensitivity to the decoding strategy. These results highlight the need for careful selection of decoding methods in medical applications, as their influence can sometimes exceed that of model choice.
Video object segmentation is vital for applications like medical diagnostics, but acquiring dense pixel-level annotations, especially for specialized domains like polyp segmentation, remains a major bottleneck. Foundational models offer zero-shot segmentation but typically require manual prompting, which is impractical for long videos. We propose Map2VidSeg, a novel pipeline that automatically generates prompts from imagelevel classification labels. It leverages localization cues (attention maps/CAMs) from a trained image classifier (ViT/CNN) to create bounding box prompts. These guide an efficient model (YOLOE) with tracking (BOT-SORT) and bidirectional propagation for initial segmentation. Optionally, a high-fidelity model (SAM-2) refines these masks using temporal memory and fusion. Demonstrated on the challenging SUN-SEG benchmark, finetuned DINOv2 (ViT) prompts significantly outperform DenseNet-121 (CNN). Our best configuration (DINOv2+YOLOE+SAM-2 Bidirectional) achieves Dice/mIoU 0.76/0.70 (Easy Unseen) and 0.66/0.60 (Hard Unseen), showcasing the viability of robust video segmentation without segmentation training data.
Medical Visual Question Answering (MedVQA) is a promising field for developing clinical decision support systems, yet progress is often limited by the available datasets, which can lack clinical complexity and visual diversity. To address these gaps, we introduce Kvasir-VQA-x1, a large-scale dataset for gastrointestinal (GI) endoscopy. Our work significantly expands upon the original Kvasir-VQA by incorporating 159,549 new question-answer pairs that are designed to test deeper clinical reasoning. We developed a systematic method using large language models to generate these questions, which are stratified by complexity to better assess a model’s inference capabilities. To ensure our dataset prepares models for real-world clinical scenarios, we have also introduced a variety of visual augmentations that mimic common imaging artifacts. The dataset is structured to support two main evaluation tracks: one for standard VQA performance and another to test model robustness against these visual perturbations. By providing a more challenging and clinically relevant benchmark, Kvasir-VQA-x1 aims to accelerate the development of more reliable and effective multimodal AI systems for use in clinical settings. The dataset follows FAIR data principles and is fully accessible, with accompanying code and documentation available at: https://github.com/simula/Kvasir-VQA-x1 .
The Medico 2025 challenge addresses Visual Question Answering (VQA) for Gastrointestinal (GI) imaging, organized as part of the MediaEval task series. The challenge focuses on developing Explainable Artificial Intelligence (XAI) models that answer clinically relevant questions based on GI endoscopy images while providing interpretable justifications aligned with medical reasoning. It introduces two subtasks: (1) answering diverse types of visual questions using the Kvasir-VQA-x1 dataset, and (2) generating multimodal explanations to support clinical decision-making. The Kvasir-VQA-x1 dataset, created from 6,500 images and 159,549 complex question-answer (QA) pairs, serves as the benchmark for the challenge. By combining quantitative performance metrics and expert-reviewed explainability assessments, this task aims to advance trustworthy Artificial Intelligence (AI) in medical image analysis. Instructions, data access, and an updated guide for participation are available in the official competition repository: https://github.com/simula/MediaEval-Medico-2025
Lifestyle diseases significantly contribute to the global health burden, with lifestyle factors playing a crucial role in the development of depression. The COVID-19 pandemic has intensified many determinants of depression. This study aimed to identify lifestyle and demographic factors associated with depression symptoms among Indians during the pandemic, focusing on a sample from Kolkata, India. An online public survey was conducted, gathering data from 1,834 participants (with 1,767 retained post-cleaning) over three months via social media and email. The survey consisted of 44 questions and was distributed anonymously to ensure privacy. Data were analyzed using statistical methods and machine learning, with principal component analysis (PCA) and analysis of variance (ANOVA) employed for feature selection. K-means clustering divided the pre-processed dataset into five clusters, and a support vector machine (SVM) with a linear kernel achieved 96% accuracy in a multi-class classification problem. The Local Interpretable Model-agnostic Explanations (LIME) algorithm provided local explanations for the SVM model predictions. Additionally, an OWL (web ontology language) ontology facilitated the semantic representation and reasoning of the survey data. The study highlighted a pipeline for collecting, analyzing, and representing data from online public surveys during the pandemic. The identified factors were correlated with depressive symptoms, illustrating the significant influence of lifestyle and demographic variables on mental health. The online survey method proved advantageous for data collection, visualization, and cost-effectiveness while maintaining anonymity and reducing bias. Challenges included reaching the target population, addressing language barriers, ensuring digital literacy, and mitigating dishonest responses and sampling errors. In conclusion, lifestyle and demographic factors significantly impact depression during the COVID-19 pandemic. The study’s methodology offers valuable insights into addressing mental health challenges through scalable online surveys, aiding in the understanding and mitigation of depression risk factors.
Assisted reproductive technologies (ART) are fundamental for cattle breeding and sustainable food production. Together with genomic selection, these technologies contribute to reducing the generation interval and accelerating genetic progress. In this paper, we discuss advancements in technologies used in the fertility evaluation of breeding animals, and the collection, processing, and preservation of the gametes. It is of utmost importance for the breeding industry to select dams and sires of the next generation as young as possible, as is the efficient and timely collection of gametes. There is a need for reliable and easily applicable methods to evaluate sexual maturity and fertility. Although gametes processing and preservation have been improved in recent decades, challenges are still encountered. The targeted use of sexed semen and beef semen has obliterated the production of surplus replacement heifers and bull calves from dairy breeds, markedly improving animal welfare and ethical considerations in production practices. Parallel with new technologies, many well-established technologies remain relevant, although with evolving applications. In vitro production (IVP) has become the predominant method of embryo production. Although fundamental improvements in IVP procedures have been established, the quality of IVP embryos remains inferior to their in vivo counterparts. Improvements to facilitate oocyte maturation and development of new culture systems, e.g. microfluidics, are presented in this paper. New non-invasive and objective tools are needed to select embryos for transfer. Cryopreservation of semen and embryos plays a pivotal role in the distribution of genetics, and we discuss the challenges and opportunities in this field. Finally, machine learning (ML) is gaining ground in agriculture and ART. This paper delves into the utilization of emerging technologies in ART, along with the current status, key challenges, and future prospects of ML in both research and practical applications within ART.
BACKGROUND AND AIMS:The American Society for Gastrointestinal Endoscopy (ASGE) AI Task Force along with experts in endoscopy, technology space, regulatory authorities, and other medical subspecialties initiated a consensus process that analyzed the current literature, highlighted potential areas, and outlined the necessary research in artificial intelligence (AI) to allow a clearer understanding of AI as it pertains to endoscopy currently. METHODS:A modified Delphi process was used to develop these consensus statements. RESULTS:Statement 1: Current advances in AI allow for the development of AI-based algorithms that can be applied to endoscopy to augment endoscopist performance in detection and characterization of endoscopic lesions. Statement 2: Computer vision-based algorithms provide opportunities to redefine quality metrics in endoscopy using AI, which can be standardized and can reduce subjectivity in reporting quality metrics. Natural language processing-based algorithms can help with the data abstraction needed for reporting current quality metrics in GI endoscopy effortlessly. Statement 3: AI technologies can support smart endoscopy suites, which may help optimize workflows in the endoscopy suite, including automated documentation. Statement 4: Using AI and machine learning helps in predictive modeling, diagnosis, and prognostication. High-quality data with multidimensionality are needed for risk prediction, prognostication of specific clinical conditions, and their outcomes when using machine learning methods. Statement 5: Big data and cloud-based tools can help advance clinical research in gastroenterology. Multimodal data are key to understanding the maximal extent of the disease state and unlocking treatment options. Statement 6: Understanding how to evaluate AI algorithms in the gastroenterology literature and clinical trials is important for gastroenterologists, trainees, and researchers, and hence education efforts by GI societies are needed. Statement 7: Several challenges regarding integrating AI solutions into the clinical practice of endoscopy exist, including understanding the role of human-AI interaction. Transparency, interpretability, and explainability of AI algorithms play a key role in their clinical adoption in GI endoscopy. Developing appropriate AI governance, data procurement, and tools needed for the AI lifecycle are critical for the successful implementation of AI into clinical practice. Statement 8: For payment of AI in endoscopy, a thorough evaluation of the potential value proposition for AI systems may help guide purchasing decisions in endoscopy. Reliable cost-effectiveness studies to guide reimbursement are needed. Statement 9: Relevant clinical outcomes and performance metrics for AI in gastroenterology are currently not well defined. To improve the quality and interpretability of research in the field, steps need to be taken to define these evidence standards. Statement 10: A balanced view of AI technologies and active collaboration between the medical technology industry, computer scientists, gastroenterologists, and researchers are critical for the meaningful advancement of AI in gastroenterology. CONCLUSIONS:The consensus process led by the ASGE AI Task Force and experts from various disciplines has shed light on the potential of AI in endoscopy and gastroenterology. AI-based algorithms have shown promise in augmenting endoscopist performance, redefining quality metrics, optimizing workflows, and aiding in predictive modeling and diagnosis. However, challenges remain in evaluating AI algorithms, ensuring transparency and interpretability, addressing governance and data procurement, determining payment models, defining relevant clinical outcomes, and fostering collaboration between stakeholders. Addressing these challenges while maintaining a balanced perspective is crucial for the meaningful advancement of AI in gastroenterology.
In this demonstration paper, we present "e2evideo" a versatile Python package composed of domain-independent modules. These modules can be seamlessly customised to suit specialised tasks by modifying specific attributes, allowing users to tailor functionality to meet the requirements of a targeted task. The package offers a variety of functionalities, such as interpolating missing video frames, background subtraction, image resizing, and extracting features utilising state-of-the-art machine learning techniques. With its comprehensive set of features, "e2evideo" stands as a facilitating tool for developers in the creation of image and video processing applications, serving diverse needs across various fields of computer vision.
To maintain and improve an amateur athlete's fitness throughout training and to achieve peak performance in sports events, good nutrition and physical activity (general and training specifically) must be considered as important factors. In our context, the terminology "amateur athletes" represents those who want to practice sports to protect their health from sickness and diseases and improve their ability to join amateur athlete events (e.g., marathons). Unlike professional athletes with personal trainer support, amateur athletes mostly rely on their experience and feeling. Hence, amateur athletes need another way to be supported in monitoring and recommending more efficient execution of their activities. One of the solutions to (self-)coaching amateur athletes is collecting lifelog data (i.e., daily data captured from different sources around a person) to understand how daily nutrition and physical activities can impact their exercise outcomes. Unfortunately, not all factors of the lifelog data can contribute to understanding the mutual impact of nutrition, physical activities, and exercise frequency on improving endurance, stamina, and weight loss. Hence, there is no guarantee that analyzing all data collected from people can produce good insights towards having a good model to predict what the outcome will be. Besides, analyzing a rich and complicated dataset can consume vast resources (e.g., computational complexity, hardware, bandwidth), and this therefore does not suit deployment on IoT or personal devices. To meet this challenge, we propose a new method to (i) discover the optimal lifelog data that significantly reflect the relation between nutrition and physical activities and training performance and (ii) construct an adaptive model that can predict the performance for both large-scale and individual groups. Our suggested method produces positive results with low MAE and MSE metrics when tested on large-scale and individual datasets and also discovers exciting patterns and correlations among data factors.
This study investigated the potential of recognising arousal in motor activity collected by wrist-worn accelerometers. We hypothesise that emotional arousal emerges from the generalised central nervous system which embeds affective states within motor activity. We formulate arousal detection as a statistical problem of separating two sets - motor activity under emotional arousal and motor activity without arousal. We propose a novel test regime based on machine learning assuming that the two sets can be distinguished if a machine learning classifier can separate the sets better than random guessing. To increase the statistical power of the testing regime, the performance of the classifiers is evaluated in a cross-validation framework, and to test if the classifiers perform better than random guessing, a repeated cross-validation corrected t-test is used. The classifiers were evaluated on the basis of accuracy and Matthew’s correlation coefficient. The suggested procedures were further compared against a traditional multivariate paired Hotelling’s T-squared test. The classifiers achieved an accuracy of about 60%, and according to the proposed t-test were significantly better than random guessing. The suggested test regime demonstrated higher statistical power than Hotelling’s T-squared test, and we conclude that we can distinguish between motor activity under emotional arousal and without it.
Abstract Study question Can deep learning be used to detect and track spermatozoa and the different parts of an ICSI procedure? Summary answer Deep learning can be used as a tool to assist and organize the contents of an ICSI procedure. What is known already Sperm tracking has been a topic of research and practice for many years, especially in the context of computer-aided sperm analysis (CASA). Recent studies have proposed using deep learning algorithms to track spermatozoa for spermatozoon selection in human and animal samples. One critical part of performing ICSI involves the selection of the “best” spermatozoon for injection, but other parts of the procedure may also be of importance. However, as far as we know, tracking using deep learning has not been applied to the ICSI procedure, where detecting instruments and the oocyte could also be helpful in post-analysis and training. Study design, size, duration The study was performed using three anonymized videos of the ICSI procedure. The frames of the videos were manually annotated by data scientists and verified by an embryologist. The annotations were bounding boxes around specific parts of the ICSI procedure, including sperm, pipettes, and the oocyte. We trained a YOLOv5 model on the collected data, where two videos were used for training and one video for validation. Participants/materials, setting, methods The videos of the ICSI procedure were captured at 200x magnification with a DeltaPix camera at Fertilitetssenteret in Oslo, Norway. ICSI was performed using a Nikon ECLIPSE TE2000-S microscope connected with Eppendorf TransferMan 4m micromanipulators. The spermatozoa were immobilised in 5 µl Polyvinylpyrrolidone (PVP; CooperSurgical). The videos had a resolution of 1920x1080 and were resized to 640x640 before being processed by the YOLOv5 model. The data will be made public in a later study. Main results and the role of chance Mean average precision (mAP) with the threshold of 0.5 (mAP@.5) is the main quantitative parameter measured in the YOLOv5 model. All the experiments were performed using three-fold cross-validation, where we present the average metrics calculated over the three folds. Overall, the method showed an average mAP@.5 of 0.50 across all predicted classes, which means that the method can track the different components with good accuracy. Looking closer at the individual classes, we see that instruments like the holding pipette and ICSI pipette are detected with high accuracy with a mAP@.5 of 0.87 and 0.94, respectively. The oocyte is also easily tracked with a mAP@.5 of 0.92. The first polar body is well detected with a mAP@.5 of 0.65. The model has issues detecting and tracking individual sperm (both outside and within the pipette), where the method achieved a mAP@.5 of 0.46 for tracking sperm outside the pipette and 0.03 for the sperm inside the pipette. The low score of detecting the sperm in the pipette can be explained by the often unclear visibility of the sperm through the pipette and the low number of training samples. Limitations, reasons for caution The limited sample size makes the generalizability of the method difficult to determine. A more extensive evaluation is necessary. Moreover, as the currency study focuses on tracking, patient information and clinical outcome were not included in the analysis. Wider implications of the findings Deep learning has the potential to aid embryologists to perform successful ICSI through tracking and detection of spermatozoa, pipettes, and the oocyte. This could potentially lead to better internal quality control and teaching possibilities, and hopefully better results. Trial registration number not applicable
Methods based on convolutional neural networks have improved the performance of biomedical image segmentation. However, most of these methods cannot efficiently segment objects of variable sizes and train on small and biased datasets, which are common for biomedical use cases. While methods exist that incorporate multi-scale fusion approaches to address the challenges arising with variable sizes, they usually use complex models that are more suitable for general semantic segmentation problems. In this paper, we propose a novel architecture called Multi-Scale Residual Fusion Network (MSRF-Net), which is specially designed for medical image segmentation. The proposed MSRF-Net is able to exchange multi-scale features of varying receptive fields using a Dual-Scale Dense Fusion (DSDF) block. Our DSDF block can exchange information rigorously across two different resolution scales, and our MSRF sub-network uses multiple DSDF blocks in sequence to perform multi-scale fusion. This allows the preservation of resolution, improved information flow and propagation of both high- and low-level features to obtain accurate segmentation maps. The proposed MSRF-Net allows to capture object variabilities and provides improved results on different biomedical datasets. Extensive experiments on MSRF-Net demonstrate that the proposed method outperforms the cutting-edge medical image segmentation methods on four publicly available datasets. We achieve the Dice Coefficient (DSC) of 0.9217, 0.9420, and 0.9224, 0.8824 on Kvasir-SEG, CVC-ClinicDB, 2018 Data Science Bowl dataset, and ISIC-2018 skin lesion segmentation challenge dataset respectively. We further conducted generalizability tests and achieved DSC of 0.7921 and 0.7575 on CVC-ClinicDB and Kvasir-SEG, respectively.
In this work, we argue that the search for Artificial General Intelligence should start from a much lower level than human-level intelligence. The circumstances of intelligent behavior in nature resulted from an organism interacting with its surrounding environment, which could change over time and exert pressure on the organism to allow for learning of new behaviors or environment models. Our hypothesis is that learning occurs through interpreting sensory feedback when an agent acts in an environment. For that to happen, a body and a reactive environment are needed. We evaluate a method to evolve a biologically-inspired artificial neural network that learns from environment reactions named Neuroevolution of Artificial General Intelligence, a framework for low-level artificial general intelligence. This method allows the evolutionary complexification of a randomly-initialized spiking neural network with adaptive synapses, which controls agents instantiated in mutable environments. Such a configuration allows us to benchmark the adaptivity and generality of the controllers. The chosen tasks in the mutable environments are food foraging, emulation of logic gates, and cart-pole balancing. The three tasks are successfully solved with rather small network topologies and therefore it opens up the possibility of experimenting with more complex tasks and scenarios where curriculum learning is beneficial.
AbstractContact tracing applications generally rely on Bluetooth data. This type of data works well to determine whether a contact occurred (smartphones were close to each other) but cannot offer the contextual information GPS data can offer. Did the contact happen on a bus? In a building? And of which type? Are some places recurrent contact locations? By answering such questions, GPS data can help develop more accurate and better-informed contact tracing applications. This chapter describes the ideas and approaches implemented for GPS data within the Smittestopp contact tracing application.We will present the pipeline used and the contribution of GPS data for contextual information, using inferred transport modes and surrounding POIs, showcasing the opportunities in the use of GPS information. Finally,we discuss ethical and privacy considerations, as well as some lessons learned.