Background/Objectives: Accurate data subcategorising is vital for reliability and traceability in the training and validation of all artificial intelligence (AI) models. Methods: In this paper we show the complexity of clinical and technical features likely to affect the appearance and interpretation of mammography images and in turn affect the output of AI software used to aid clinical decisions. Results: Using mammography as a case study, the equitability covers screened population characteristics (e.g., women’s age and ethnicity) and image acquisition key factors (e.g., brand of system, exposure factors, image processing). We examine some studies and available datasets of mammography images, summarising the metadata available. Conclusions: We recommend that, where possible, AI models are trained and evaluated using data that includes subcategories based on these features, ensuring increased equitability in the data and coverage of image heterogeneities; or, where not possible, that the subcategories for which the AI model is valid are clearly defined. Such practices can easily be implemented in a wide range of AI applications but are illustrated here with mammography. Clinical Relevance: Reliable AI holds invaluable potential for both clinical efficiency and accuracy in diagnosis. With appropriately categorised training data, a reduction in subjective assessment can be achieved, leading to trustworthy and rapid assessment.
We introduce ControlAugment (Ctrl-A), an automated data augmentation algorithm for image-vision tasks, which incorporates principles from control theory for online adjustment of augmentation strength distributions during model training. Ctrl-A eliminates the need for initialization of individual augmentation strengths. Instead, augmentation strength distributions are dynamically, and individually, adapted during training based on a control-loop architecture and what we define as relative operation response curves. Using an operation-dependent update procedure provides Ctrl-A with the potential to suppress augmentation styles that negatively impact model performance, alleviating the need for manually engineering augmentation policies for new image-vision tasks. Experiments on the CIFAR-10, CIFAR-100, and SVHN-core benchmark datasets using the common WideResNet-28-10 architecture demonstrate that Ctrl-A is highly competitive with existing state-of-the-art data augmentation strategies.
Standardised digital records can improve access and quality of health services and enable meta-analysis of data which may reveal unknown patterns, which improve our understanding of the data. In this study, we report on meta-analysis of data curated from advanced radiotherapy dosimetry audits conducted at hospitals across the UK using test objects developed at the National Physical Laboratory. This meta-analysis highlights hospitals which are performing within expectation, or outside of expected intervals and would benefit from measurement support. Anonymised hospitals with low precision or accuracy treatment plans are identified, enabling support and improvements where appropriate. The results presented may be used to provide insight to hospitals and inform areas of focus for improved predictions of radiation dose that are tailored to a given hospital. Moreover, this analysis can enable auditors and regulators to provide additional services or recommendations, and potentially identify previously unknown patterns or dependencies in the data.
Data reduction and data mining are common practices for handling large-scale data from wide-ranging sources, but high-dimensional omics and imaging data sets present difficult challenges for feature extraction and data mining due to the large number of features that cannot be simultaneously examined. The sample numbers and variables in these methods are constantly growing as new technologies are developed, and computational analysis needs to evolve to keep up with growing demand. In recent years, there has been a rapid uptake of nonlinear dimensionality reduction via methods such as t-distributed stochastic neighbor embedding and uniform manifold approximation and projection. These approaches have revolutionized our ability to visualize and interpret high-dimensional data and have rapidly become preferred methods for analysis of data sets containing an extremely high number of variables. Further to this is the emerging interest in combining information from multiple omics sources to gain a more holistic view of systems biology. Current state-of-the-art algorithms can perform data mining, visualization, and classification on routine data sets but struggle when data sets grow above a certain size. We present a new approach to large and multiomic data integration to extract, mine, and integrate large multiomics data sets that were previously considered prohibitively large. Here, we demonstrate the use of deep learning on subsampled nonlinear dimensionality reduction using t-SNE and UMAP to extract features from large complex data sets including mass spectrometry imaging and chromosome conformation capture. We then go on to demonstrate how this method can be used to learn embeddings from the fusion of different omics data, allowing metabolomics data to be projected into a reduced transcriptomics representation.
The COVID-19 pandemic and national lockdowns have had profound impacts on population mental health in Scotland. In this study, we examine the impact on the relationship between the number of patients receiving in-patient care for mental health related diagnoses and the number of patients on waiting lists for treatment. The relationship between waiting lists and in-patient treatment changed during the COVID-19 pandemic, and this may be contributing to additional pressures on hospitals and mental health services in Scotland.
Worldwide, pancreatic cancer has a poor prognosis, with less than 13 % of those diagnosed surviving beyond five years. Early diagnosis can enable curative intervention thus improving survival rates, but many people are diagnosed too late. To improve this situation, researchers are trying to utilise routinely collected healthcare data. Over 50 risk prediction algorithms for the early diagnosis of pancreatic cancer have been developed to date. However, few have undergone external validation – a crucial step before implementation in clinical practice. The Enriching New-Onset Diabetes for Pancreatic Cancer (ENDPAC) model, developed using data from the United States of America, is one of the few that has been externally validated. ENDPAC shows promise for use in primary care, both due to its predictive power and because it uses changes in routinely collected patient data – age, blood glucose and body weight – to calculate a risk score for having pancreatic cancer. In theory, this simplicity makes ENDPAC readily applicable in clinical settings. In reality, there are significant barriers to using even this simple model. There are numerous primary care computer systems, and these all differ in how data can be accessed and analysed in order to calculate the ENDPAC risk score. Furthermore, the blood glucose measurement units used by the model differ from those used in countries such as the United Kingdom (UK) as reporting is not globally standardised, meaning that values must be converted. However, this results in uncertainties being introduced. Potential metrological implications include: multiple scores per patient, artificial score inflation and incalculable scores, resulting in unclear follow-up actions and a lack of confidence in the model’s outputs. We have developed a method to enable reproducible data extraction from one of the main systems used by UK primary care practices. Thus far, five out of the intended 20 primary care practices have extracted data for analysis. Ninety (24 %) of the 371 people extracted had sufficient blood glucose and weight data for an ENDPAC score to be calculated. However, 87 % of these (78/90) could not have an ENDPAC score calculated due to problems introduced by the conversion: blood glucose results were artificially inflated by the conversion in 86 % (67/78), and 14 % (11/78) straddled two or more risk score boundaries due to blood glucose result uncertainty. We will demonstrate how we are overcoming these issues by quantifying the impact of blood glucose result conversion on ENDPAC score calculation and developing a method to calculate confidence intervals for ENDPAC scores. Our aim is to increase patients’ and clinicians’ confidence in the outputs provided by the ENDPAC model, which in turn will aid in the detection, subsequent diagnosis and treatment of this devastating disease.
Pancreatic cancer is the twelfth most common cancer worldwide, but high mortality rates make it the sixth leading cause of cancer deaths. Diagnosis is frequently too late for curative intervention. Risk assessment tools incorporating diagnostic prediction models may assist early pancreatic cancer detection by primary care clinicians. This mixed methods systematic review aims to identify risk assessment tools which can be used for the detection of pancreatic cancer and have been investigated in primary care. It also seeks to synthesise the qualitative and quantitative evidence relating to the patient and clinician perspectives and experiences with these tools. Ten studies were included with five risk assessment tools identified: ‘QCancer’, ‘eRATs’ (electronic risk assessment tools), ‘CaDet’, ‘Future Health Today’ and ‘C the Signs’. No tools were found for pancreatic cancer alone. Thematic synthesis of stakeholder perspectives resulted in three themes: impact on clinical decision-making, impact on patient consultations and implementation barriers and facilitators. Overall, experiences and impacts were positive, especially if used by less experienced clinicians. There is little evidence for the inclusion of many developed pancreatic cancer diagnostic prediction models in risk assessment tools in primary care, and limited research into stakeholder perceptions, especially patient perceptions. This review can inform future tool development, and further research should be undertaken assessing these tools’ clinical validity to encourage uptake in clinical practice. This study was registered prospectively as PROSPERO CRD42024488160.
AI-based MRI reconstruction techniques improve efficiency by reducing acquisition times whilst maintaining or improving image quality. Recent recommendations from professional bodies suggest centres should perform quality assessments on AI tools. However, monitoring long-term performance presents challenges, due to model drift or system updates. Radiologist-based assessments are resource-intensive and may be subjective, highlighting the need for efficient quality control (QC) measures. This study explores using image quality metrics (IQMs) to assess AI-based reconstructions. 58 patients undergoing standard-of-care rectal MRI were imaged using AI-based and conventional T2-weighted sequences. Paired and unpaired IQMs were calculated. Sensitivity of IQMs to detect retrospective perturbations in AI-based reconstructions was assessed using control charts, and statistical comparisons between the four MR systems in the evaluation were performed. Two radiologists evaluated the image quality of the perturbed images, giving an indication of their clinical relevance. Paired IQMs demonstrated sensitivity to changes in AI-reconstruction settings, identifying deviations outside ± 2 standard deviations of the reference dataset. Unpaired metrics showed less sensitivity. Paired IQMs showed no difference in performance between 1.5 T and 3 T systems (p > 0.99), whilst minor but significant (p < 0.0379) differences were noted for unpaired IQMs. IQMs are effective for QC of AI-based MR reconstructions, offering resource-efficient alternatives to repeated radiologist evaluations. Future work should expand this to other imaging applications and assess additional measures.
Capsule networks are a relatively unexplored type of neural network architecture that preserve spatial information of the input by replacing the pooling layers with convolutional strides and dynamic routing, which allow part-whole relationships of the data to be retained. One disadvantage is the computational complexity of dynamic routing, where each capsule must route to all capsules in a layer. It is common practice to use many capsules with a smaller feature space and there has been little attention in the exploration of using fewer, wider capsules. This reduces the number of routes the network must make, making the network train faster while still accounting for the same learnable space. This paper presents an ablation study on a 3-layer capsule network architecture by changing the primary capsule dimensions to assess the impact on performance and training time. Experiments were performed on capsule networks with capsule sizes: 32× 8 , 8× 32 , 16× 8 and 8× 16 (number × width), on 11 benchmark datasets: MNIST, CIFAR-10, PCAM, fashionMNIST, BreastMNIST, BloodMNIST, OrganMNIST, PathMNIST and OCTMNIST and SVHN. For all of our datasets we observe capsule network structures that obtain accuracy that exceeds that of the 32× 8 structures and are least 40
Background. Preimplantation biopsy combines measurements of injury into a composite index to inform organ acceptance. The uncertainty in these measurements remains poorly characterized, raising concerns variability may contribute to inappropriate clinical decisions. Methods. We adopted a metrological approach to evaluate biopsy score reliability. Variability was assessed by performing repeat biopsies (n = 293) on discarded allografts (n = 16) using 3 methods (core, punch, and wedge). Uncertainty was quantified using a bootstrapping analysis. Observer effects were controlled by semi-blinded scoring, and the findings were validated by comparison with standard glass evaluation. Results. The surgical method strongly determined the size (core biopsy area 9.04 mm2, wedge 37.9 mm2) and, therefore, yield (glomerular yield r = 0.94, arterial r = 0.62) of each biopsy. Core biopsies yielded inadequate slides most frequently. Repeat biopsy of the same kidney led to marked variation in biopsy scores. In 10 of 16 cases, scores were contradictory, crossing at least 1 decision boundary (ie, to transplant or to discard). Bootstrapping demonstrated significant uncertainty associated with single-slide assessment; however, scores were similar for paired kidneys from the same donor. Conclusions. Our investigation highlights the risks of relying on single-slide assessment to quantify organ injury. Biopsy evaluation is subject to uncertainty, meaning each slide is better conceptualized as providing an estimate of the kidney’s condition rather than a definitive result. Pooling multiple assessments could improve the reliability of biopsy analysis, enhancing confidence. Where histological quantification is necessary, clinicians should seek to develop new protocols using more tissue and consider automated methods to assist pathologists in delivering analysis within clinical time frames.
Clustering algorithms are used extensively in data analysis for data exploration and discovery. Technological advancements lead to continually growth of data in terms of volume, dimensionality and complexity. This provides great opportunities in data analytics as the data can be interrogated for many different purposes. This however leads challenges, such as identification of relevant features for a given task. In supervised tasks, one can utilise a number of methods to optimise the input features for the task objective (e.g. classification accuracy). In unsupervised problems, such tools are not readily available, in part due to an inability to quantify feature relevance in unlabeled tasks. In this paper, we investigate the sensitivity of clustering performance noisy uncorrelated variables iteratively added to baseline datasets with well defined clusters. We show how different types of irrelevant variables can impact the outcome of a clustering result from $k$-means in different ways. We observe a resilience to very high proportions of irrelevant features for adjusted rand index (ARI) and normalised mutual information (NMI) when the irrelevant features are Gaussian distributed. For Uniformly distributed irrelevant features, we notice the resilience of ARI and NMI is dependent on the dimensionality of the data and exhibits tipping points between high scores and near zero. Our results show that the Silhouette Coefficient and the Davies-Bouldin score are the most sensitive to irrelevant added features exhibiting large changes in score for comparably low proportions of irrelevant features regardless of underlying distribution or data scaling. As such the Silhouette Coefficient and the Davies-Bouldin score are good candidates for optimising feature selection in unsupervised clustering tasks.
Performing a mitosis count (MC) is the diagnostic task of histologically grading canine Soft Tissue Sarcoma (cSTS). However, mitosis count is subject to inter- and intra-observer variability. Deep learning models can offer a standardisation in the process of MC used to histologically grade canine Soft Tissue Sarcomas. Subsequently, the focus of this study was mitosis detection in canine Perivascular Wall Tumours (cPWTs). Generating mitosis annotations is a long and arduous process open to inter-observer variability. Therefore, by keeping pathologists in the loop, a two-step annotation process was performed where a pre-trained Faster R-CNN model was trained on initial annotations provided by veterinary pathologists. The pathologists reviewed the output false positive mitosis candidates and determined whether these were overlooked candidates, thus updating the dataset. Faster R-CNN was then trained on this updated dataset. An optimal decision threshold was applied to maximise the F1-score predetermined using the validation set and produced our best F1-score of 0.75, which is competitive with the state of the art in the canine mitosis domain.
Introduction Worldwide, pancreatic cancer has a poor prognosis. Early diagnosis may improve survival by enabling curative treatment. Statistical and machine learning diagnostic prediction models using risk factors such as patient demographics and blood tests are being developed for clinical use to improve early diagnosis. One example is the Enriching New-onset Diabetes for Pancreatic Cancer (ENDPAC) model, which employs patients’ age, blood glucose and weight changes to provide pancreatic cancer risk scores. These values are routinely collected in primary care in the UK. Primary care’s central role in cancer diagnosis makes it an ideal setting to implement ENDPAC but it has yet to be used in clinical settings. This study aims to determine the feasibility of applying ENDPAC to data held by UK primary care practices.Methods and analysis This will be a multicentre observational study with a cohort design, determining the feasibility of applying ENDPAC in UK primary care. We will develop software to search, extract and process anonymised data from 20 primary care providers’ electronic patient record management systems on participants aged 50+ years, with a glycated haemoglobin (HbA1c) test result of ≥48 mmol/mol (6.5%) and no previous abnormal HbA1c results. Software to calculate ENDPAC scores will be developed, and descriptive statistics used to summarise the cohort’s demographics and assess data quality. Findings will inform the development of a future UK clinical trial to test ENDPAC’s effectiveness for the early detection of pancreatic cancer.Ethics and dissemination This project has been reviewed by the University of Surrey University Ethics Committee and received a favourable ethical opinion (FHMS 22-23151 EGA). Study findings will be presented at scientific meetings and published in international peer-reviewed journals. Participating primary care practices, clinical leads and policy makers will be provided with summaries of the findings.
Dosimetry audits are carried out to determine how well radiotherapy is delivered to the patient. It is also used to understand the uncertainty introduced into the measurement result when using different computational models. As measurement procedures are becoming increasingly complex with technological advancements, it is harder to establish sources of variability in measurements and understand if they stem from true differences in measurands or in the measurement pipelines themselves. The gamma index calculation is a widely accepted metric used for the comparison of measured and predicted doses in radiotherapy. However, various steps in the measurement pipeline can introduce variation in the measurement result. In this paper, we perform a sensitivity and correlation analysis to investigate the influence of various input factors (i.e. setting) in gamma index calculations on the uncertainty introduced in dosimetry audits. We identify a number of factors where standardization will improve measurements by reducing variability in outputs. Furthermore, we also compare gamma index metrics and similarities across audit sites.
Self-supervised learning (SSL) has become a popular method for generating invariant representations without the need for human annotations. Nonetheless, the desired invariant representation is achieved by utilizing prior online transformation functions on the input data. As a result, each SSL framework is customized for a particular data type, for example, visual data, and further modifications are required if it is used for other dataset types. On the other hand, autoencoder (AE), which is a generic and widely applicable framework, mainly focuses on dimension reduction and is not suited for learning invariant representation. This article proposes a generic SSL framework based on a constrained self-labeling assignment process that prevents degenerate solutions. Specifically, the prior transformation functions are replaced with a self-transformation mechanism, derived through an unsupervised training process of adversarial training, for imposing invariant representations. Via the self-transformation mechanism, pairs of augmented instances can be generated from the same input data. Finally, a training objective based on contrastive learning is designed by leveraging both the self-labeling assignment and the self-transformation mechanism. Despite the fact that the self-transformation process is very generic, the proposed training strategy outperforms a majority of state-of-the-art representation learning methods based on AE structures. To validate the performance of our method, we conduct experiments on four types of data, namely visual, audio, text, and mass spectrometry data and compare them in terms of four quantitative metrics. Our comparison results demonstrate that the proposed method is effective and robust in identifying patterns within the tested datasets.
IntroductionRenal transplant biopsies provide insights into graft health and support decision making. The current evidence on links between biopsy scores and transplant outcomes suggests there may be numerous factors affecting biopsy scores. Here we adopt measurement science approach to investigate the sources of uncertainty in biopsy assessment and suggest techniques to improve its robustness.MethodsHistological assessments, Remuzzi scores, biopsy processing and clinical variables are obtained from 144 repeat biopsies originating from 16 deceased-donor kidneys. We conducted sensitivity analysis to find the morphometric features with highest discriminating power and studied the dependencies of these features on biopsy and stain type. The analysis results formed a basis for recommendations on reducing the assessment variability.ResultsMost morphometric variables are influenced by the biopsy and stain types. The variables with the highest discriminatory power are sclerotic glomeruli counts, healthy glomeruli counts per unit area, percentages of interstitial fibrosis and tubular atrophy as well as diameter and lumen of the worst artery. A revised glomeruli adequacy score is proposed to improve the robustness of the glomeruli statistics, whereby a minimum of 104 µm2 of cortex tissue is recommended to keep type 1 and type 2 error probabilities below 0.15 and 0.2.DiscussionThe findings are transferable to several biopsy scoring systems. We hope that this work will help practitioners to understand the sources of statistical uncertainty and improve the utility of renal biopsy.
This study explores the efficacy of diffusion probabilistic models for generating synthetic histopathological images, specifically canine Perivascular Wall Tumours (cPWT), to supplement limited datasets for deep learning applications in digital pathology. This research evaluates an open-source medical domain-focused diffusion model called Medfusion, where the model was trained on a small (1,000 patches) and a large dataset (17,000 patches) of cPWT images to compare performance on the different sized datasets. A Receiver Operating Characteristic (ROC) study was implemented to investigate the ability of six veterinary medical professionals and pathologists to discern between generated and real cPWT patch images. The participants engaged in two separate rounds, where each round corresponded to models that had been trained on the two different sized datasets. The ROC study revealed mean average Area Under the Curve (AUC) values close to 0.5 for both rounds. The results from this study suggests that diffusion models can create histopathological patch images that are convincingly realistic where our participants often struggled to reliably differentiate between generated and real images. This underscores the potential of these models as a valuable tool for augmenting digital pathology datasets.
Gamma Indices are a widely used metric for the comparison of measured and predicted doses in dosimetry audits for advanced radiotherapy. Various steps in the measurement pipeline and analysis approach, as well as different computational models to calculate the gamma index, can introduce variability in the results. In this paper the authors investigate the sensitivity of gamma index measurements to changes in the factors in the dosimetry audit pipeline using the in house developed Versatile Independent Gamma Analysis sOftware (VIGO). Our results indicate that the choice of calibrated size and shape for the region of interest introduce the most variability in the measurement. Standardising these factors can help reduce overall variability in the measurement results, and therefore the radiation received by a patient, improving the accuracy and efficiency of radiotherapy.
In addressing the challenges in real-time Fluorescence Lifetime Imaging (FLIm)-Optical Endomicroscopy (OEM), particularly motion artefacts, this study introduces a comprehensive framework designed to enhance FLIm processing in in vivo studies. The framework focuses on improving image quality by selectively discarding uninformative frames and employing a novel registration technique. This technique integrates Normalised Cross Correlation (NCC) and Channel and Spatial Reliability Tracker (CSRT) to consistently track the dominant correlation peak across temporal sequences of images, thus enhancing the reliability and precision of subsequent analyses. This approach has shown a significant improvement upon existing registration methods in handling temporal FLIm motion artefacts. Our method overcomes the optimisation issues inherent in similarity-based registration and demonstrates a 17% enhancement in Quality of Alignment (QA) metric and a 25% increase in Structural Similarity Index Measure (SSIM) across various datasets.Clinical relevance— Our study introduces a significant advancement in FLIm imaging, with a novel method that increases the precision and reliability of the registration. This enhancement is crucial for the translational clinical research sphere, where precise, real-time imaging underpins the development of more effective diagnostics and treatments in pulmonary medicine.
Molecular imaging is a key tool in the diagnosis and treatment of prostate cancer (PCa). Magnetic Resonance (MR) plays a major role in this respect with nuclear medicine imaging, particularly, Prostate-Specific Membrane Antigen-based, (PSMA-based) positron emission tomography with computed tomography (PET/CT) also playing a major role of rapidly increasing importance. Another key technology finding growing application across medicine and specifically in molecular imaging is the use of machine learning (ML) and artificial intelligence (AI). Several authoritative reviews are available of the role of MR-based molecular imaging with a sparsity of reviews of the role of PET/CT. This review will focus on the use of AI for molecular imaging for PCa. It will aim to achieve two goals: firstly, to give the reader an introduction to the AI technologies available, and secondly, to provide an overview of AI applied to PET/CT in PCa. The clinical applications include diagnosis, staging, target volume definition for treatment planning, outcome prediction and outcome monitoring. ML and AL techniques discussed include radiomics, convolutional neural networks (CNN), generative adversarial networks (GAN) and training methods: supervised, unsupervised and semi-supervised learning.