PURPOSE:Social deprivation is a predictor of delivered care and clinical outcomes in many areas of medicine. It is not known whether social deprivation affects quality of care received by patients admitted as medical emergencies to hospitals in the United Kingdom (UK). METHODS:The Society for Acute Medicine's Benchmarking Audit (SAMBA) measures performance of Acute Medical Units in the UK during a 24-hour period. SAMBA23 took place on Thursday, 22 June 2023. Social deprivation was measured by matching the first half of patient postcodes with publicly available data about the Index of Multiple Deprivation for that geographic area. RESULTS:SAMBA23 included 161 UK hospitals and 9612 patients. Data on social deprivation were available for 7855 patients. Patients from areas with the highest social deprivation were more frail and more likely to be admitted via the emergency department, as opposed to being referred by a General Practitioner. Patients from areas with the highest and lowest levels of deprivation received a similar standard of care as measured by three quality indicators; there were no statistically significant differences in the proportion of patients who received timely measurements of the National Early Warning Score, were seen within 4 hours by a clinician, and were reviewed in a timely manner by a consultant. CONCLUSIONS:In this single 24-hour snapshot we were unable to identify differences in the delivery of the Acute Medical take in relation to the residential area index of multiple deprivation. Further research is required using patient level data. Key messages What is already known on this topic: Social deprivation is a co-factor for poor outcomes in most areas of medicine. It is unclear whether social deprivation influences the process of care and outcomes of patients admitted as medical emergencies. What this study adds: We found no evidence for an effect of the residential area index of multiple social deprivation on quality indicators for medical emergencies in data a single 24-hour snapshot as part of a national audit in Acute Medicine. How this study might affect research, practice or policy: This is the first published data from acute medical care in the UK looking at social deprivation, but more data is needed at individual patient level.
Quality improvement activities in healthcare are limited by the substantial time burden associated with manual clinical text review. To address this limitation within an established hospital discharge summary improvement project, we aimed to automate quality monitoring using large language models. Models were trained to identify ‘perfect’ content using clinician-graded data from 1,876 discharge summaries. Performance was evaluated on a held out validation subset and then applied to 107,000 summaries covering the full project period. The models showed strong agreement with clinician-graded data, achieving F1 scores of 87 to 95 percent across targeted text fields. Automated processing enabled near real time evaluation of the entire dataset and revealed trends that were not detectable through traditional sampling methods. These findings demonstrate the feasibility of using large language models to increase the efficiency, coverage, and analytical depth of quality improvement and audit activities that rely on free-text review.
BACKGROUND:Medical research depends on access to high quality data that protects patient privacy. Free text in health records contains valuable clinical detail, yet it often includes sensitive personal information that must be removed before use. Current approaches rely on manually created training data and focus mainly on narrow domains. They are difficult to scale to new medical fields and languages. This study aims to address these limitations by developing a framework that supports privacy-preserving use of medical text across diverse settings. METHODS:The study introduces an annotation-free framework for training and adapting LLM-based anonymization models across diverse medical domains. Our reproducible framework includes the development of a generative medical anonymization model, leveraging synthetic data and instruction tuning of generative LLMs. Performance is evaluated on both synthetic test sets and on patient requests from a digital triage service. Accuracy, recall, precision, and the ability to maintain the original meaning of non-sensitive text are assessed. RESULTS:Here we show that generative models trained with the synthetic framework reach performance that exceeds strong baseline systems across several medical domains. The models preserve non-sensitive text with high fidelity and anonymize sensitive information with high accuracy. They perform well even when trained on small datasets, generalize to unseen clinical fields, and support anonymization in multiple languages without requiring additional training data in those languages. CONCLUSIONS:The study presents a reproducible, annotation-free approach that enables the development of effective anonymization models for medical text. The framework reduces reliance on real patient data, lowers the cost of adaptation to new settings, and supports wider use of unstructured clinical information for research and service improvement.
Text-to-image (T2I) customization empowers users to adapt the T2I diffusion model to new concepts absent in the pre-training dataset. On this basis, capturing multiple new concepts from a single image has emerged as a new task, allowing the model to learn multiple concepts simultaneously or discard unwanted concepts. However, multiple-concept disentanglement remains a key challenge. Existing disentanglement models often exhibit two main issues: feature fusion and asynchronous learning across different concepts. To address these issues, we propose AttenCraft, an attention-based method for multiple-concept disentanglement. Our method uses attention maps to generate accurate masks for each concept in a single initialization step, aiding in concept disentanglement without requiring mask preparation from humans or specialized models. Moreover, we introduce an adaptive algorithm based on attention scores to estimate sampling ratios for different concepts, promoting balanced feature acquisition and synchronized learning. AttenCraft also introduces a feature-retaining training framework that employs various loss functions to enhance feature recognition and prevent fusion. Extensive experiments show that our model effectively mitigates these two issues, achieving state-of-the-art image fidelity and comparable prompt fidelity to baseline models.
Computer Aided Diagnosis/Detection (CAD) systems for skin lesion analysis are challenged by limited and imbalanced dermoscopic datasets, necessitating the use of advanced data augmentation techniques. In this paper, we introduce DiDGen, an innovative method employing text-to-image Diffusion models for high-quality Dermoscopic image Generation to enhance CAD performance. Specifically, we propose a dynamic prompting framework, DermPrompt, that leverages large language models to produce attribute-rich text prompts, improving image generation quality. We further refine generation control by incorporating a novel region-aware fine-tuning approach to build visual-textual alignments and a training-free pipeline for synthesizing lesion-mask pairs. Extensive experiments reveal that our proposed method outperforms existing generative methods in image fidelity and diversity, with downstream classifiers and segmentation models showing average improvements of 2.32% in F1 score and 3.16% in IoU score-all achieved with a single finetuning process. This approach offers an efficient solution for augmenting dermoscopic datasets and advancing skin lesion diagnosis.
Radiology Report Generation (RRG) has advanced considerably with the development of multimodal generative models. Despite the progress, the field still faces significant challenges in evaluation, as existing metrics lack robustness and fairness. We reveal that, RRG with high performance on existing lexical-based metrics (e.g. BLEU) might be more of a mirage - a model can get a high BLEU only by learning the template of reports. This has become a pressing issue for RRG due to the highly patternized nature of these reports. In addition, standard radiology reports are often highly technical. Helping patients understand these reports is crucial from a patient's perspective, yet this has been largely overlooked in previous work. In this work, we un-intuitively approach these problems by proposing the Layman's RRG framework that can systematically improve RRG with day-to-day language. Specifically, our framework first contributes a translated Layman's terms dataset. Building upon the dataset, we then propose a semantics-based evaluation method, which is effective in mitigating the inflated numbers of BLEU and provides more robust evaluation. We show that training on the layman's terms dataset encourages models to focus on the semantics of the reports, as opposed to overfitting to learning the report templates. Last, we reveal a promising scaling law between the number of training examples and semantics gain provided by our dataset, compared to the inverse pattern brought by the original formats.
Healthy Life Expectancy (HLE) considers the years an individual has lived free of disease. Minimising time a population spends in ill-health will directly impact national healthcare budgets, alongside individuals’ personal quality of life. We utilise Personal Health Records (PHR), consisting of Electronic Health Records with lifestyle and wellness data, to investigate loss of healthy life. We leveraged Survival Analysis (SA) to directly estimate HLE without the Sullivan Method. Additionally, a multiple imputation ensemble Machine Learning (ML) model was trained to successfully predict the loss of healthy life within one year with an AUPRC of almost double that of random of the unbalanced data. ML explainability techniques enabled investigation of the model’s learned relationships, providing insights about the effects of lifestyle on maintaining healthy life. Finally, a novel method of combining the ML model’s prediction with a SA’s estimated hazards was proposed, enabling the formulation of a tailored conditioned Survival Function.
Historically, veterinary studies screening for breed, age and sex predisposition to disease have relied on collating small-scale studies of clinical datasets. The availability of larger datasets through groups such as the Small Animal Veterinary Surveillance Network (SAVSNET) promise access to information regarding a wide range of clinical presentations at scale, however, methodological limitations surrounding the extraction of specific disease information or screening for disease predispositions result in a substantial reduction in the number of animals studied. These studies often address very focused hypotheses - only leveraging a small fraction of the intrinsic value of the data at any one time. Here, we implemented an unsupervised machine learning methodology, creating a representation of a large volume of clinical notes collected by SAVSNET from veterinary practices across the UK. We utilise BERTopic, a topic-modelling tool based on Bidirectional Encoder Representations using Transformers (BERT) architecture, and show it is able to surface known phenotypes, such as breed predispositions to hypoadrenocorticism, diabetes mellitus and mitral valve disease, as well as potential novel patterns of disease phenotypes. This scalable and granular modelling technique facilitates the rapid interrogation of large clinical datasets, enabling the identification of a broad range of phenotypes within the population and the early detection of temporal changes indicative of emerging infectious or environmental diseases.
Large Language Models (LLMs) have advanced zero-shot Information Extraction (IE), particularly in Sentence-level Relation Extraction (SentRE), through in-context learning and instruction tuning. However, the current evaluation of LLMs’ zero-shot ability on IE tasks remains fragile and unreliable. In this work, we provide a systematic examination of the fragility underlying current evaluation practices across three interrelated levels. At the data level, we demonstrate that the commonly adopted random sampling strategy introduces significant biases in class-imbalanced datasets, whereas balanced sampling provides more stable and faithful assessments of LLMs performance. At the task level, we reveal that three domain prompt frameworks on SentRE transfer inconsistently to Document-level Relation Extraction (DocRE) and Named Entity Recognition (NER), showing partial effectiveness on NER but notable limitations on DocRE due to long contexts and complex entity structures. At the method level, through extensive experiments on three IE tasks and seven datasets, we conduct the first comprehensive comparison of five general prompt frameworks, including Chain-of-Thought, Self-Improvement, and Self-Debate, showing that prompt effectiveness is highly task-dependent, with no single strategy dominating across tasks. For each task, the CoT prompt framework achieves the best performance on SentRE, the Vanilla prompt framework performs best on DocRE, and the Self-Consistency prompt framework excels on NER. These insights challenge current landscape of information extraction, providing guidelines for robust evaluation and prompt designs.
Objective To explore administrators’ and clinicians’ views on the factors that influence their use and adoption of a machine learning clinical decision support system (ML-CDSS) to predict patients’ risk of hepatic and renal deterioration during chemotherapy.Methods and analysis This was a qualitative study that used purposive sampling. 18 participants with administration and clinical backgrounds working in cancer care in England were recruited. Qualitative data were collected by conducting semi-structured interviews and a focus group. Data were analysed thematically using the framework method to identify key themes.Results Participants acknowledged that monitoring blood chemistry is a core component of chemotherapy as it helps clinicians assess patient fitness and treatment response. The ML-CDSS was perceived as a potentially valuable tool for identifying patients at increased risk of hepatic and renal deterioration, supporting clinical decision-making and enhancing care efficiency. However, several concerns were raised regarding its potential implementation in practice. Participants questioned clinicians’ willingness and capacity to integrate the tool into their existing workflows. Participants also believed it was important to demonstrate the ML-CDSS’s sensitivity, specificity and validity in accurately predicting patients’ risk to build clinicians’ trust in the tool, demonstrating evidence of its efficacy and effectiveness in practice.Conclusion Administrators and clinicians recognised the potential benefits of the ML-CDSS to enhance the delivery of chemotherapy by identifying patients at risk for hepatic and renal deterioration. Successful adoption in practice depends on building trust with the tool by being transparent in its development, its effectiveness and impact. Future work should demonstrate the ML-CDSS being used in practice to generate real-world evidence.
Background Skin cancer is one of the most prevalent cancers globally, with early detection critical to ensure reduced mortality risk. To aid early detection, machine learning (ML) skin cancer detection models have been proposed, currently with a focus on dermatoscopic imaging only. However, freetext may provide extra diagnostic information that is not present in images alone. Methods We constructed a multimodal dataset comprising 5481 dermatoscopic images from 4538 patients, including patient metadata and clinical notes, with binary labels (benign vs. malignant, 7% malignant). To assess and mitigate bias from leading language, we developed a clinical text preprocessing pipeline combining regular expressions and large language models, enabling multiple levels of filtering. We train multimodal ML models on this dataset to explore the effect of freetext on model performance. Results Our results show that incorporating unfiltered text significantly improves classification performance (0.970 AUROC) compared to visual data alone (0.909 AUROC); even with leading language removed, performance gains persist (0.948 AUROC). Conclusions This work benchmarks clinical freetext inclusion in skin lesion classification, demonstrating that clinical text contributes predictive value beyond that available in images alone. The model's high performance on unfiltered clinical text highlights the high levels of bias, and possible shortcutting, present in this text which may make it unsuitable for inclusion in some ML models. By systematically filtering clinical notes via our proposed technique, we show that multimodal models retain improved accuracy while reducing bias. These results provide practical guidance for integrating clinical text into real-world skin cancer detection systems and establish a foundation for future multimodal research in dermatology.
Vector quantization approaches (VQ-VAE, VQ-GAN) learn discrete neural representations of images, but these representations are inherently position-dependent: codes are spatially arranged and contextually entangled, requiring autoregressive or diffusion-based priors to model their dependencies at sample time. In this work, we ask whether positional information is necessary for discrete representations of spatially aligned data. We propose the permutation-invariant vector-quantized autoencoder (PI-VQ), in which latent codes are constrained to carry no positional information. We find that this constraint encourages codes to capture global, semantic features, and enables direct interpolation between images without a learned prior. To address the reduced information capacity of permutation-invariant representations, we introduce matching quantization, a vector quantization algorithm based on optimal bipartite matching that increases effective bottleneck capacity by 3.5× relative to naive nearest-neighbour quantization. The compositional structure of the learned codes further enables interpolation-based sampling, allowing synthesis of novel images in a single forward pass. We evaluate PI-VQ on CelebA, CelebA-HQ and FFHQ, obtaining competitive precision, density and coverage metrics for images synthesised with our approach. We discuss the trade-offs inherent to position-free representations, including separability and interpretability of the latent codes, pointing to numerous directions for future work.
Multimodal learning, which involves integrating information from various modalities such as text, images, audio, and video, is pivotal for numerous complex tasks like visual question answering, cross-modal retrieval, and caption generation. Traditional approaches rely on modality-specific encoders and late fusion techniques, which can hinder scalability and flexibility when adapting to new tasks or modalities. To address these limitations, we introduce a novel framework that extends the concept of task reformulation beyond natural language processing (NLP) to multimodal learning. We propose to reformulate diverse multimodal tasks into a unified next-frame prediction problem, allowing a single model to handle different modalities without modality-specific components. This method treats all inputs and outputs as sequential frames in a video, enabling seamless integration of modalities and effective knowledge transfer across tasks. Our approach is evaluated on a range of tasks, including text-to-text, image-to-text, video-to-video, video-to-text, and audio-to-text, demonstrating the model's ability to generalize across modalities with minimal adaptation. We show that task reformulation can significantly simplify multimodal model design across various tasks, laying the groundwork for more generalized multimodal foundation models.
Sparse Autoencoders (SAEs) are a popular method for decomposing Large Language Model (LLM) activations into interpretable latents, however they have a substantial training cost and SAEs learned on different models are not directly comparable. Motivated by relative representation similarity measures, we introduce Inference-Time Decomposition of Activation models (ITDAs). ITDAs are constructed by greedily sampling activations into a dictionary based on an error threshold on their matching pursuit reconstruction. ITDAs can be trained in 1% of the time of SAEs, allowing us to cheaply train them on Llama-3.1 70B and 405B. ITDA dictionaries also enable cross-model comparisons, and outperform existing methods like CKA, SVCCA, and a relative representation method on a benchmark of representation similarity. Code available at https://github.com/pleask/itda.
This paper presents the setup and results of the third edition of the BioLaySumm shared task on Lay Summarization of Biomedical Research Articles and Radiology Reports, hosted at the BioNLP Workshop at ACL 2025. In this task edition, we aim to build on the first two editions' successes by further increasing research interest in this important task and encouraging participants to explore novel approaches that will help advance the state-of-the-art. Specifically, we introduce the new task of Radiology Report Generation with Layman's terms, which is parallel to the task of lay summarization of biomedical articles in the first two editions. Overall, our results show that a broad range of innovative approaches were adopted by task participants, including inspiring explorations of latest RL techniques adopted in the training of general-domain large reasoning models.
We introduce PetEVAL, the first benchmark dataset derived from real -world, free -text veterinary electronic health records (EHRs). PetEVAL comprises 17,600 professionally annotated EHRs from first -opinion veterinary practices across the UK, partitioned into training (11,000). evaluation (1,600), and test (5,000) sets with distinct clinic distributions to assess model generalisability. Each record is annotated with International Classification of Disease 11 (ICD-11) syndromic chapter labels (20,408 labels), disease Named Entity Recognition (NER) tags (429 labels), and anonymisation NER tags (8,244 labels). PetEVAL enables evaluating Natural Language Processing (NLP) tools across applications, including syndrome surveillance and disease outbreak detection. We implement a multistage anonymisation protocol, replacing identifiable information with clinically relevant pseudonyms while establishing the first definition of identifiers in veterinary free text. PetEVAL introduces three core tasks: syndromic classification, disease entity recognition, and anonymisation. We provide baseline results using BERT-base, PetBERT, and LLaMA 3.1 8B generative models. Our experiments demonstrate the unique challenges of veterinary text, showcasing the importance of domain -specific approaches. By fostering advancements in veterinary informatics and epidemiology, we envision PetEVAL catalysing innovations in veterinary care, animal health, and comparative biomedical research through access to real -world, annotated veterinary clinical data.
Image representations are often evaluated through disjointed, task-specific protocols, leading to a fragmented understanding of model capabilities. For instance, it is unclear whether an image embedding model adept at clustering images is equally good at retrieving relevant images given a piece of text. We introduce the Massive Image Embedding Benchmark (MIEB) to evaluate the performance of image and image-text embedding models across the broadest spectrum to date. MIEB spans 38 languages across 130 individual tasks, which we group into 8 high-level categories. We benchmark 50 models across our benchmark, finding that no single method dominates across all task categories. We reveal hidden capabilities in advanced vision models such as their accurate visual representation of texts, and their yet limited capabilities in interleaved encodings and matching images and texts in the presence of confounders. We also show that the performance of vision encoders on MIEB correlates highly with their performance when used in multimodal large language models. Our code, dataset, and leaderboard are publicly available at https://github.com/embeddings-benchmark/mteb.
Inspired by the performance and scalability of autoregressive large language models (LLMs), transformer-based models have seen recent success in the visual domain. This study investigates a transformer adaptation for video prediction with a simple end-to-end approach, comparing various spatiotemporal self-attention layouts. Focusing on causal modeling of physical simulations over time; a common shortcoming of existing video-generative approaches, we attempt to isolate spatiotemporal reasoning via physical object tracking metrics and unsupervised training on physical simulation datasets. We introduce a simple yet effective pure transformer model for autoregressive video prediction, utilizing continuous pixel-space representations for video prediction. Without the need for complex training strategies or latent feature-learning components, our approach significantly extends the time horizon for physically accurate predictions by up to 50% when compared with existing latent-space approaches, while maintaining comparable performance on common video quality metrics. In addition, we conduct interpretability experiments to identify network regions that encode information useful to perform accurate estimations of PDE simulation parameters via probing models, and find that this generalizes to the estimation of out-of-distribution simulation parameters. This work serves as a platform for further attention-based spatiotemporal modeling of videos via a simple, parameter efficient, and interpretable approach.
Although large language models excel across many tasks, they can memorise training data and thereby expose private or copyrighted text. Most defences target the pre-training stage, leaving memorisation during fine-tuning, especially for domain adaptation and instruction tuning, poorly understood. We fine-tune Pythia, Llama3, and Mistral models spanning 1.4B-70B parameters on common evaluation datasets and track verbatim memorisation throughout training. We find that memorisation increases dramatically in the first few epochs, often significantly before either validation perplexity or evaluation performance is optimised. We use a simple but effective n-gram memorisation score which reliably precedes verbatim memorisation; using it as an early-stopping criterion mitigates memorisation with minimal performance loss. Further, we introduce an n-gram-aware loss regulariser and show that it reduces memorisation across all model families tested by up to 40
Facial recognition is one of the most academically studied and industrially developed areas within computer vision where we readily find associated applications deployed globally. This widespread adoption has uncovered significant performance variation across subjects of different racial profiles leading to focused research attention on racial bias within face recognition spanning both current causation and future potential solutions. In support, this study provides an extensive taxonomic review of research on racial bias within face recognition exploring every aspect and stage of the face recognition processing pipeline. Firstly, we discuss the problem definition of racial bias, starting with race definition, grouping strategies, and the societal implications of using race or race-related groupings. Secondly, we divide the common face recognition processing pipeline into four stages: image acquisition, face localisation, face representation, face verification and identification, and review the relevant corresponding literature associated with each stage. The overall aim is to provide comprehensive coverage of the racial bias problem with respect to each and every stage of the face recognition processing pipeline whilst also highlighting the potential pitfalls and limitations of contemporary mitigation strategies that need to be considered within future research endeavours or commercial applications alike.