Sparse autoencoders (SAEs) have been applied to large language models and protein language models, but not systematically to electronic health record (EHR) foundation models. We train TopK SAEs on FlatASCEND, a 14.5-million-parameter autoregressive clinical sequence model, at all 10 residual stream extraction points on INSPECT (outpatient) and MIMIC-IV (ICU). SAE decomposition reveals progressive abstraction across transformer depth: layer-0 features are near-perfect token detectors (45.7
Artificial Intelligence (AI) is increasingly influencing chemical risk assessment, enabling faster, more comprehensive, and potentially more ethical assessments. The application of AI in chemical risk assessment refers to both generative and predictive algorithms encompassing machine learning, to analyse complex chemical, biological, and environmental data and provide insights into adverse effect potential for humans and ecosystems. AI systems support the prediction of chemical hazards, exposure levels, and adverse effects by learning from experimental results, mechanistic models, and regulatory datasets, thereby enhancing the efficiency of safety evaluations.In October 2024, ECETOC held an international workshop, with experts from academia, industry, and regulatory bodies, to reflect upon the historical challenges in integrating multidimensional omics technologies into chemical regulation and explore the current capabilities and future potential of AI in toxicology and regulatory science. Discussions emphasised that implementation of Findable, Accessible, Interoperable, and Reusable (FAIR) data principles is not just a best practice but rather a prerequisite for building transparent, reliable, and unbiased AI systems. The reliability of AI in producing scientifically valid and socially responsible outcomes depends fundamentally on the availability of FAIR data. However, ensuring trustworthiness also requires robust governance frameworks that go beyond data and human oversight. Critical enablers of responsible AI in chemical risk assessment are rigorous governance, explainability, fit-for-purpose applications, and human oversight. ECETOC supports the development of flexible and iterative frameworks advancing development, validation, transparency, accountability, and trust in AI applications in chemicals regulation.
Autoregressive models can predict clinical events, but generating patient-conditioned multi-step trajectories that respond to intervention tokens and testing whether those responses preserve known pharmacological associations has received limited attention. We present FlatASCEND, a 14.5M-parameter autoregressive clinical sequence model using flat composite tokens and a zero-inflated log-normal time head. Standard distributional metrics (Jaccard 0.889-0.954) do not distinguish FlatASCEND from trivial baselines; the model's value lies in conditional generation from patient-specific prefixes. A prompt-shuffle ablation shows patient-specific conditioning amplifies mechanistic pharmacological effects (2.0-2.2x for steroid to glucose, diuretic to potassium) while leaving confounding-driven associations unchanged (0.9x for insulin to glucose). An incident-user framework assesses directional consistency against prior pharmacological knowledge on MIMIC-IV (N=500 per comparison): 4/10 recover correct mechanistic directions, 2 reproduce treatment-context associations, 4 are incorrect (9/10 significant, Wilcoxon p<0.05). This pattern - partial recovery under residual confounding - is consistent with learned observational associations without causal distinction. Direct preference optimisation with surrogate reward destroys all correct associations (3/3 to 0/3), illustrating reward exploitation when reward and evaluation share an outcome domain. Generative evidence is strongest for short-horizon ICU data; outpatient temporal fidelity is weaker (median 10 vs 154 days on INSPECT), and zero-shot cross-site transfer degrades without adaptation.
We present ASCENDgpt, a transformer-based model specifically designed for cardiovascular risk prediction from longitudinal electronic health records (EHRs). Our approach introduces a novel phenotype-aware tokenization scheme that maps 47,155 raw ICD codes to 176 clinically meaningful phenotype tokens, achieving 99.6% consolidation of diagnosis codes while preserving semantic information. This phenotype mapping contributes to a total vocabulary of 10,442 tokens - a 77.9% reduction when compared with using raw ICD codes directly. We pretrain ASCENDgpt on sequences derived from 19402 unique individuals using a masked language modeling objective, then fine-tune for time-to-event prediction of five cardiovascular outcomes: myocardial infarction (MI), stroke, major adverse cardiovascular events (MACE), cardiovascular death, and all-cause mortality. Our model achieves excellent discrimination on the held-out test set with an average C-index of 0.816, demonstrating strong performance across all outcomes (MI: 0.792, stroke: 0.824, MACE: 0.800, cardiovascular death: 0.842, all-cause mortality: 0.824). The phenotype-based approach enables clinically interpretable predictions while maintaining computational efficiency. Our work demonstrates the effectiveness of domain-specific tokenization and pretraining for EHR-based risk prediction tasks.
We reevaluate the pairwise learning to rank approach based on neural nets, called RankNet, and present a theoretical analysis of its architecture. We show mathematically that the model can, under certain conditions, learn reflexive, antisymmetric, and transitive relations, enabling simplified training and improved performance. Experimental results on the LETOR MSLR-WEB10K, MQ2007 and MQ2008 datasets show that the model outperforms numerous state-of-the-art methods (including a listwise approach), while being inherently simpler in structure and using a pairwise approach only.
Consumer-grade wearable technology has the potential to support clinical research and patient management. Here, we report results from the RATE-AF trial wearables study, which was designed to compare heart rate in older, multimorbid patients with permanent atrial fibrillation and heart failure who were randomized to treatment with either digoxin or beta-blockers. Heart rate (n = 143,379,796) and physical activity (n = 23,704,307) intervals were obtained from 53 participants (mean age 75.6 years (s.d. 8.4), 40% women) using a wrist-worn wearable linked to a smartphone for 20 weeks. Heart rates in participants treated with digoxin versus beta-blockers were not significantly different (regression coefficient 1.22 (95% confidence interval (CI) -2.82 to 5.27; P = 0.55); adjusted 0.66 (95% CI -3.45 to 4.77; P = 0.75)). No difference in heart rate was observed between the two groups of patients after accounting for physical activity (P = 0.74) or patients with high activity levels (>= 30,000 steps per week; P = 0.97). Using a convolutional neural network designed to account for missing data, we found that wearable device data could predict New York Heart Association functional class 5 months after baseline assessment similarly to standard clinical measures of electrocardiographic heart rate and 6-minute walk test (F1 score 0.56 (95% CI 0.41 to 0.70) versus 0.55 (95% CI 0.41 to 0.68); P = 0.88 for comparison). The results of this study indicate that digoxin and beta-blockers have equivalent effects on heart rate in atrial fibrillation at rest and on exertion, and suggest that dynamic monitoring of individuals with arrhythmia using wearable technology could be an alternative to in-person assessment. ClinicalTrials.gov identifier: NCT02391337. In a substudy of the RATE-AF trial, which compared heart rate control therapy using digoxin or the beta-blocker bisoprolol, heart rate and physical activity data collected using a wearable device showed equivalent heart rate control by the two drugs and could be used to predict future heart failure functional class as well as standard clinical measurements.
IntroductionThe echocardiographic measurement of left ventricular ejection fraction (LVEF) is fundamental to the diagnosis and classification of patients with heart failure (HF).MethodsThis paper aimed to quantify LVEF automatically and accurately with the proposed pipeline method based on deep neural networks and ensemble learning. Within the pipeline, an Atrous Convolutional Neural Network (ACNN) was first trained to segment the left ventricle (LV), before employing the area-length formulation based on the ellipsoid single-plane model to calculate LVEF values. This formulation required inputs of LV area, derived from segmentation using an improved Jeffrey’s method, as well as LV length, derived from a novel ensemble learning model. To further improve the pipeline’s accuracy, an automated peak detection algorithm was used to identify end-diastolic and end-systolic frames, avoiding issues with human error. Subsequently, single-beat LVEF values were averaged across all cardiac cycles to obtain the final LVEF.ResultsThis method was developed and internally validated in an open-source dataset containing 10,030 echocardiograms. The Pearson’s correlation coefficient was 0.83 for LVEF prediction compared to expert human analysis (p < 0.001), with a subsequent area under the receiver operator curve (AUROC) of 0.98 (95% confidence interval 0.97 to 0.99) for categorisation of HF with reduced ejection (HFrEF; LVEF<40%). In an external dataset with 200 echocardiograms, this method achieved an AUC of 0.90 (95% confidence interval 0.88 to 0.91) for HFrEF assessment.ConclusionThe automated neural network-based calculation of LVEF is comparable to expert clinicians performing time-consuming, frame-by-frame manual evaluations of cardiac systolic function.
BackgroundThe limitations of the traditional TNM system have spurred interest in multivariable models for personalized prognostication in laryngeal and hypopharyngeal cancers (LSCC/HPSCC). However, the performance of these models depends on the quality of data and modelling methodology, affecting their potential for clinical adoption. This systematic review and meta-analysis (SR-MA) evaluated clinical predictive models (CPMs) for recurrence and survival in treated LSCC/HPSCC. We assessed models’ characteristics and methodologies, as well as performance, risk of bias (RoB), and applicability.MethodsLiterature searches were conducted in MEDLINE (OVID), Embase (OVID) and IEEE databases from January 2005 to November 2023. The search algorithm used comprehensive text word and index term combinations without language or publication type restrictions. Independent reviewers screened titles and abstracts using a predefined Population, Index, Comparator, Outcomes, Timing and Setting (PICOTS) framework. We included externally validated (EV) multivariable models, with at least one clinical predictor, that provided recurrence or survival predictions. The SR-MA followed PRISMA reporting guidelines, and PROBAST framework for RoB assessment. Model discrimination was assessed using C-index/AUC, and was presented for all models using forest plots. MA was only performed for models that were externally validated in two or more cohorts, using random-effects model. The main outcomes were model discrimination and calibration measures for survival (OS) and/or local recurrence (LR) prediction. All measures and assessments were preplanned prior to data collection.ResultsThe SR-MA identified 11 models, reported in 16 studies. Seven models for OS showed good discrimination on development, with only one excelling (C-index >0.9), and three had weak or poor discrimination. Inclusion of a radiomics score as a model parameter achieved relatively better performance. Most models had poor generalisability, demonstrated by worse discrimination performance on EV, but they still outperformed the TNM system. Only two models met the criteria for MA, with pooled EV AUCs 0.73 (95% CI 0.71-0.76) and 0.67 (95% CI 0.6-0.74). RoB was high for all models, particularly in the analysis domain.ConclusionsThis review highlighted the shortcomings of currently available models, while emphasizing the need for rigorous independent evaluations. Despite the proliferation of models, most exhibited methodological limitations and bias. Currently, no models can confidently be recommended for routine clinical use.Systematic review registrationhttps://www.crd.york.ac.uk/prospero/display_record.php?ID=CRD42021248762, identifier CRD42021248762.
Objective:To enable reproducible research at scale by creating a platform that enables health data users to find, access, curate, and re-use electronic health record phenotyping algorithms.Materials and Methods:We undertook a structured approach to identifying requirements for a phenotype algorithm platform by engaging with key stakeholders. User experience analysis was used to inform the design, which we implemented as a web application featuring a novel metadata standard for defining phenotyping algorithms, access via Application Programming Interface (API), support for computable data flows, and version control. The application has creation and editing functionality, enabling researchers to submit phenotypes directly.Results:We created and launched the Phenotype Library in October 2021. The platform currently hosts 1049 phenotype definitions defined against 40 health data sources and >200K terms across 16 medical ontologies. We present several case studies demonstrating its utility for supporting and enabling research: the library hosts curated phenotype collections for the BREATHE respiratory health research hub and the Adolescent Mental Health Data Platform, and it is supporting the development of an informatics tool to generate clinical evidence for clinical guideline development groups.Discussion:This platform makes an impact by being open to all health data users and accepting all appropriate content, as well as implementing key features that have not been widely available, including managing structured metadata, access via an API, and support for computable phenotypes.Conclusions:We have created the first openly available, programmatically accessible resource enabling the global health research community to store and manage phenotyping algorithms. Removing barriers to describing, sharing, and computing phenotypes will help unleash the potential benefit of health data for patients and the public.
Annotation of biomedical entities with ontology classes provides for formal semantic analysis and mobilisation of background knowledge in determining their relationships. To date, enrichment analysis has been routinely employed to identify classes that are over-represented in annotations across sets of groups, such as biosample gene expression profiles or patient phenotypes, and is useful for a range of tasks including differential diagnosis and causative variant prioritisation. These approaches, however, usually consider only univariate relationships, make limited use of the semantic features of ontologies, and provide limited information and evaluation of the explanatory power of both singular and grouped candidate classes. Moreover, they are not designed to solve the problem of deriving cohesive, characteristic, and discriminatory sets of classes for entity groups. We have developed a new tool, called Klarigi, which introduces multiple scoring heuristics for identification of classes that are both compositional and discriminatory for groups of entities annotated with ontology classes. The tool includes a novel algorithm for derivation of multivariable semantic explanations for entity groups, makes use of semantic inference through live use of an ontology reasoner, and includes a classification method for identifying the discriminatory power of candidate sets, in addition to significance testing apposite to traditional enrichment approaches. We describe the design and implementation of Klarigi, including its scoring and explanation determination methods, and evaluate its use in application to two test cases with clinical significance, comparing and contrasting methods and results with literature-based and enrichment analysis methods. We demonstrate that Klarigi produces characteristic and discriminatory explanations for groups of biomedical entities in two settings. We also show that these explanations recapitulate and extend the knowledge held in existing biomedical databases and literature for several diseases. We conclude that Klarigi provides a distinct and valuable perspective on biomedical datasets when compared with traditional enrichment methods, and therefore constitutes a new method by which biomedical datasets can be explored, contributing to improved insight into semantic data.
Introduction and aim: Artificial Intelligence (AI) is already being successfully employed to aid the interpretation of multiple facets of burns care. In the light of the growing influence of AI, this systematic review and diagnostic test accuracy meta-analyses aim to appraise and summarise the current direction of research in this field.Method: A systematic literature review was conducted of relevant studies published between 1990 and 2021, yielding 35 studies. Twelve studies were suitable for a Diagnostic Test MetaAnalyses. Results: The studies generally focussed on burn depth (Accuracy 68.9%-95.4%, Sensitivity 90.8% and Specificity 84.4%), burn segmentation (Accuracy 76.0%-99.4%, Sensitivity 97.9% and specificity 97.6%) and burn related mortality (Accuracy > 90%-97.5% Sensitivity 92.9% and specificity 93.4%). Neural networks were the most common machine learning (ML) algorithm utilised in 69% of the studies. The QUADAS-2 tool identified significant heterogeneity between studies. Discussion: The potential application of AI in the management of burns patients is promising, especially given its propitious results across a spectrum of dimensions, including burn depth, size, mortality, related sepsis and acute kidney injuries. The accuracy of the results analysed within this study is comparable to current practices in burns care. Conclusion: The application of AI in the treatment and management of burns patients, as a series of point of care diagnostic adjuncts, is promising. Whilst AI is a potentially valuable tool, a full evaluation of its current utility and potential is limited by significant variations in Crown Copyright (c) 2022 Published by Elsevier Ltd on behalf of British Association of Plastic, Reconstructive and Aesthetic Surgeons. All rights reserved.
Background: Glycated haemoglobin (HbA1c) measurement is used for diagnosis, management and remission of type 2 diabetes (T2DM), with measurements comparable worldwide and the World Health Organization listing medical conditions that affect its accuracy. Admission glucose is in the ‘diabetes’ range in 5% of emergency hospital admissions without prior diagnosis, with literature searches indicating inconsistent practice on using HbA1c to confirm diagnosis. As oral glucose tolerance tests (OGTT) were not possible during the COVID-19 pandemic, guidance was issued by the Royal College of Obstetrics and Gynaecology on using HbA1c for gestational diabetes mellitus. Aims: This study explores use of HbA1c at Queen Elizabeth Hospital Birmingham, a large university hospital serving a multi- ethnic adult population. Methods: Information is presented on comparability, clinical audits, research studies and current practice, and is illustrated by case reports. Results: Data from the National Glycohemoglobin Standardization Program show comparability of laboratoryHbA1c and point-of-care testing methods from 1993 to 2023. Although HbA1c was used to diagnose gestational diabetes during the COVID-19 pandemic, hospitals have reverted to OGTT post pandemic. In contrast, HbA1c is now being used to assess T2DM remission. Case reports illustrate these scenarios and highlight the complexity of decision-making when the accuracy of the HbA1c reading is affected by multiple co- morbidities. Conclusions: This wider use of HbA1c includes remission of T2DM but the diagnosis of gestational diabetes has reverted to OGTT post pandemic. A pictorial representation of HbA1c range is presented to aid understanding of this test. It is suitable for diagnosis of diabetes in most people except those with some variant haemoglobins or abnormal red blood cell turnover.
Artificial intelligence (AI) is increasingly being utilized in healthcare. This article provides clinicians and researchers with a step-wise foundation for high-value AI that can be applied to a variety of different data modalities. The aim is to improve the transparency and application of AI methods, with the potential to benefit patients in routine cardiovascular care. Following a clear research hypothesis, an AI-based workflow begins with data selection and pre-processing prior to analysis, with the type of data (structured, semi-structured, or unstructured) determining what type of pre-processing steps and machine-learning algorithms are required. Algorithmic and data validation should be performed to ensure the robustness of the chosen methodology, followed by an objective evaluation of performance. Seven case studies are provided to highlight the wide variety of data modalities and clinical questions that can benefit from modern AI techniques, with a focus on applying them to cardiovascular disease management. Despite the growing use of AI, further education for healthcare workers, researchers, and the public are needed to aid understanding of how AI works and to close the existing gap in knowledge. In addition, issues regarding data access, sharing, and security must be addressed to ensure full engagement by patients and the public. The application of AI within healthcare provides an opportunity for clinicians to deliver a more personalized approach to medical care by accounting for confounders, interactions, and the rising prevalence of multi-morbidity.
In the extensive search for new physics, the precise measurement of the Higgs boson continues to play an important role. To this end, machine learning techniques have been recently applied to processes like the Higgs production via vector-boson fusion. In this paper, we propose to use algorithms for learning to rank, i.e., to rank events into a sorting order, first signal, then background, instead of algorithms for the classification into two classes, for this task. The fact that training is then performed on pairwise comparisons of signal and background events can effectively increase the amount of training data due to the quadratic number of possible combinations. This makes it robust to unbalanced data set scenarios and can improve the overall performance compared to pointwise models like the state-of-the-art boosted decision tree approach. In this work we compare our pairwise neural network algorithm, which is a combination of a convolutional neural network and the DirectRanker, with convolutional neural networks, multilayer perceptrons or boosted decision trees, which are commonly used algorithms in multiple Higgs production channels. Furthermore, we use so-called transfer learning techniques to improve overall performance on different data types.
Respiratory irritation is an important human health endpoint in chemical risk assessment. There are two established modes of action of respiratory irritation, 1) sensory irritation mediated by the interaction with sensory neurons, potentially stimulating trigeminal nerve, and 2) direct tissue irritation. The aim of our research was to, develop a QSAR method to predict human respiratory irritants, and to potentially reduce the reliance on animal testing for the identification of respiratory irritants. Compounds are classified as irritating based on combined evidence from different types of toxicological data, including inhalation studies with acute and repeated exposure. The curated project database comprised 1997 organic substances, 1553 being classified as irritating and 444 as non-irritating. A comparison of machine learning approaches, including Logistic Regression (LR), Random Forests (RFs), and Gradient Boosted Decision Trees (GBTs), showed, the best classification was obtained by GBTs. The LR model resulted in an area under the curve (AUC) of 0.65, while the optimal performance for both RFs and GBTs gives an AUC of 0.71. In addition to the classification and the information on the applicability domain, the web-based tool provides a list of structurally similar analogues together with their experimental data to facilitate expert review for read-across purposes.
Much of the knowledge and information needed for enabling high-quality clinical research is stored in free-text format. Natural language processing (NLP) has been used to extract information from these sources at scale for several decades. This paper aims to present a comprehensive review of clinical NLP for the past 15 years in the UK to identify the community, depict its evolution, analyse methodologies and applications, and identify the main barriers. We collect a dataset of clinical NLP projects ( n = 94; £ = 41.97 m) funded by UK funders or the European Union’s funding programmes. Additionally, we extract details on 9 funders, 137 organisations, 139 persons and 431 research papers. Networks are created from timestamped data interlinking all entities, and network analysis is subsequently applied to generate insights. 431 publications are identified as part of a literature review, of which 107 are eligible for final analysis. Results show, not surprisingly, clinical NLP in the UK has increased substantially in the last 15 years: the total budget in the period of 2019–2022 was 80 times that of 2007–2010. However, the effort is required to deepen areas such as disease (sub-)phenotyping and broaden application domains. There is also a need to improve links between academia and industry and enable deployments in real-world settings for the realisation of clinical NLP’s great potential in care delivery. The major barriers include research and development access to hospital data, lack of capable computational resources in the right places, the scarcity of labelled data and barriers to sharing of pretrained models.
Amanda Clare合作论文数Department of Computer Science,
Aberystwyth University13