In supervised learning model development, domain experts are often used to provide the class labels (annotations). Annotation inconsistencies commonly occur when even highly experienced clinical experts annotate the same phenomenon (e.g., medical image, diagnostics, or prognostic status), due to inherent expert bias, judgments, and slips, among other factors. While their existence is relatively well-known, the implications of such inconsistencies are largely understudied in real-world settings, when supervised learning is applied on such 'noisy' labelled data. To shed light on these issues, we conducted extensive experiments and analyses on three real-world Intensive Care Unit (ICU) datasets. Specifically, individual models were built from a common dataset, annotated independently by 11 Glasgow Queen Elizabeth University Hospital ICU consultants, and model performance estimates were compared through internal validation (Fleiss' κ = 0.383 i.e., fair agreement). Further, broad external validation (on both static and time series datasets) of these 11 classifiers was carried out on a HiRID external dataset, where the models' classifications were found to have low pairwise agreements (average Cohen's κ = 0.255 i.e., minimal agreement). Moreover, they tend to disagree more on making discharge decisions (Fleiss' κ = 0.174) than predicting mortality (Fleiss' κ = 0.267). Given these inconsistencies, further analyses were conducted to evaluate the current best practices in obtaining gold-standard models and determining consensus. The results suggest that: (a) there may not always be a "super expert" in acute clinical settings (using internal and external validation model performances as a proxy); and (b) standard consensus seeking (such as majority vote) consistently leads to suboptimal models. Further analysis, however, suggests that assessing annotation learnability and using only 'learnable' annotated datasets for determining consensus achieves optimal models in most cases.
Machine Learning systems rely heavily on annotated instances. Such annotations are frequently done by human experts, or by tools developed by experts, and so the central message of this book, Noise: A Flaw in Human Judgment (Kahneman, Sibony, and Sunstein 2021) is of considerable importance to AI/Machine Learning community. The core message is that if a number of experts are asked to annotate tasks that involve judgments, these responses will frequently differ. This observation poses a problem for how analysts choose a particular annotated dataset (from the group), or process the set of responses to give a “balanced” response, or whether to reject all the annotated datasets. A further important aspect of this book is the case studies which demonstrate that differences in judgments between fellow experts have been reported in a significant number of disciplines including, business, the law, government, and medicine. Kahneman, Sibony and Sunstein (2021), referred to as KSS subsequently, discuss how Expert Biases can be reduced, but the main focus of this book is a discussion of Noise, that is, differences that often occur between fellow experts, and how Noise can often be reduced. To address the last point KSS have formulated a set of six decision hygiene principles which include the recommendation that complex tasks should be subdivided, and then each subtask should be solved separately. A further principle is that each task should be solved by individual experts before the various judgments are discussed with fellow experts. Effectively, the book being reviewed covers three main topics: First, it reports several motivating studies that show how judgments of fellow experts varied significantly in the pricing of insurance premiums, and in setting the lengths of custodial sentences. These motivating studies very effectively illustrate the central concepts of Judgment, Noise, and Bias; that section also provides definitions of these core concepts and discusses how Noise is often amplified in group meetings. Secondly, the authors provide detailed discussion of further studies, in a variety of domains, which report the levels of disagreement between experts. Thirdly, KSS discusses how to reduce the levels of Noise between experts, as noted above, the authors refer to these as Principles of Noise Hygiene. These three parts are interwoven in a complex way throughout the book; in our view, the best overview of the book is given in the section Review and Conclusions: Taking Noise Seriously (KSS, p. 361).
Inconsistency in real-world judgments can cause random unfairness, injustice and misallocation of resources. In their recent monograph Kahneman, Sibony, and Sunstein (2021) analyse judgment inconsistency or “Noise,” examine its sources and propose remedies. In this commentary on Kahneman et al., we reflect on the major concepts (such as “judgment,” “noise,” “error,” and “bias”) used in analysing inconsistency. We place this work in the broader context of applied cognitive psychology, relating it to error typologies and to dual-systems views of thinking. We also compare Kahneman et al.'s heuristics based approach to the linear combination of attributes based approach of Social Judgment Theory (SJT), with particular reference to judgment noise. We conclude that the main contributions of Kahneman et al.'s book are (a) to raise awareness of the pervasiveness of judgment noise across a range of important real-world areas, (b) to provide a taxonomy of types of noise in terms of system noise versus occasion noise, and level noise versus pattern noise, and (c) to outline useful ways of reducing noise, and thus overall levels of error.
Argumentation theory is particularly well suited to support clinical decision-making due to its ability to reason with uncertain knowledge and derive defeasible and understandable conclusions. Subsequently, models of argumentation are being increasingly deployed in clinical decisionmaking systems to facilitate reasoning. However, challenges remain, including the development of more human-like argumentation which has the potential to improve the effectiveness of systems. As a step towards addressing this challenge, we have examined real-life clinical discourse during which clinical conflict was resolved. The dialogues captured from these interviews have been analysed to determine the methods of information selection and argument generation used. Further, we describe several argument schemes and associated critical questions which have been based on this real-life clinical discourse.
Knowledge intensive clinical systems, as well as machine learning algorithms, have become more widely used over the last decade or so. These systems often need access to sizable labelled datasets which could be more useful if their instances are accurately labelled / annotated. A variety of approaches, including statistical ones, have been used to label instances. In this paper, we discuss the use of domain experts, in this case clinicians, to perform this task. Here we recognize that even highly rated domain experts can have differences of opinion on certain instances; we discuss a system inspired by the Delphi approaches which helps experts resolve their differences of opinion on classification tasks. The focus of this paper is the IS-DELPHI tool which we have implemented to address the labelling issue; we report its use in a medical domain in a study involving 12 Intensive Care Unit clinicians. The several pairs of experts initially disagreed on the classification of 11 instances but as a result of using IS-DELPHI all those disagreements were resolved. From participant feedback (questionnaires), we have concluded that the medical experts understood the task and were comfortable with the functionality provided by IS-DELPHI. We plan to further enhance the system’s capabilities and usability, and then use IS-DELPHI, which is a domain independent tool, in a number of further medical domains.
BACKGROUND Endometriosis is a condition with relatively non-specific symptoms, and in some cases a long time elapses from first-symptom presentation to diagnosis. AIM To develop and test new composite pointers to a diagnosis of endometriosis in primary care electronic records. DESIGN AND SETTING This is a nested case-control study of 366 cases using the Practice Team Information database of anonymised primary care electronic health records from Scotland. Data were analysed from 366 cases of endometriosis between 1994 and 2010, and two sets of age and GP practice matched controls: (a) 1453 randomly selected females and (b) 610 females whose records contained codes indicating consultation for gynaecological symptoms. METHOD Composite pointers comprised patterns of symptoms, prescribing, or investigations, in combination or over time. Conditional logistic regression was used to examine the presence of both new and established pointers during the 3 years before diagnosis of endometriosis and to identify time of appearance. RESULTS A number of composite pointers that were strongly predictive of endometriosis were observed. These included pain and menstrual symptoms occurring within the same year (odds ratio [OR] 6.5, 95% confidence interval [CI] = 3.9 to 10.6), and lower gastrointestinal symptoms occurring within 90 days of gynaecological pain (OR 6.1, 95% CI = 3.6 to 10.6). Although the association of infertility with endometriosis was only detectable in the year before diagnosis, several pain-related features were associated with endometriosis several years earlier. CONCLUSION Useful composite pointers to a diagnosis of endometriosis in GP records were identified. Some of these were present several years before the diagnosis and may be valuable targets for diagnostic support systems.
We propose that access to data and knowledge be controlled through fine-grained, user-specified explicitly represented policies. Fine-grained policies allow stakeholders to have a more precise level of control over who, when, and how their data is accessed. We propose a representation for policies and a mechanism to control data access within a fully distributed system, creating a secure environment for data sharing. Our proposal provides guarantees against standard attacks, and ensures data security across the network. We present and justify the goals, requirements, and a reference architecture for our proposal. We illustrate through an intuitive example how our proposal supports a typical data-sharing transaction. We also perform an analysis of the various potential attacks against this system, and how they are countered. Additionally, we provide details of a proof-of-concept prototype which we used to refine our mechanism.
Treatment protocols for patients who have suffered traumatic brain injury (TBI) specify that hypoxia should be avoided and specifically that brain oxygen tension (PbtO2) should be maintained above a particular level (20mmHg). Results from several specialized Neuro ICUs world-wide suggest that such guidelines are not achieved in at least 24% of patients (Shafi et al. 2014). Many physiological and therapeutic factors can influence PbtO2. Furthermore, clinical staff in ICUs have many calls on their time and may not be able to direct sufficient attention to the management of a single parameter in one of several patients. We believe that automated analysis of complex data sets could be a useful step to developing software capable of assisting in the better maintenance of PbtO2 in these complex patients. A "manual" analysis of 5 patients' complete temporal records showed that certain distinct rules including a number of important descriptors (CPP, PbtO2, PaO2 & the correlation coefficient between CPP and PbtO2) with their values segmented into discrete ranges, cover 98.2% (SD 1.6) of the available dataset once the rules have been "fine-tuned". This study was then validated in a second set of five patients, in which 92.5% (SD 12.5) of the data set was covered without additional "tuning" of the rules. Moreover, it was noted that these patterns, once some account was taken of noise, occur in "blocks". As a result of these observations we developed a correlation module for the existing Temporal Discovery workbench to replicate the "manual" analysis. Using the dataset for the 10 patients, we have now obtained effectively perfect agreement between the manual analysis and that produced by the workbench. Subsequently, the expert provided clinical actions which correspond to each of the (48) patterns potentially created in this study, and so the correlation module is now able to produce for each time-point (and each time "block") a pattern, and the correct clinical action. Further work includes enhancing how the correlation module deals with noise, evaluating the approach across a much larger dataset, and evaluating the effectiveness of the module's recommendations in clinical settings.
Black-box classifiers are able to classify unseen instances, once they have been trained on an appropriate (domain) dataset. Such classifiers have the advantage of being generally very efficient but the disadvantage of not being able to explain their processes to a user. For these reasons, over the last decade or so, a number of rule extraction algorithms have been developed which are able to extract a rule-set from classifiers. The focus of this project has been to re-implement a state-of-the-art rule extraction system, OSRE [1], and then to show that when the extracted rules are refined by the Knowledge Refinement system, FIXIT, that the refinement process, in virtually all cases, improves the fidelity of the refined rule-set when compared with the rule-set extracted by OSRE. A statistically significant difference between these two approaches has been demonstrated. We investigated 4 classifiers (2 blackbox (Neural Networks & SVM), 1 Bayesian classifier & 1 (Decision-Tree-based) whitebox) and 4 domains, so a total of 16 Classifier-Dataset combinations were considered. In only 1 case (6.25%) was the result slightly worse; 5 cases (31.25%) were the same (these could not be improved), and the remaining 10 cases (62.5%) show significant improvements. In the future, we intend using similar approaches to improve the accuracy of the classification; this study focuses on fidelity.
Learning Objectives: Current scoring systems for use in critical illness do not take into account levels of simultaneous physiological or pharmacological support. Yet, clinicians at the bedside subconsciously interpret cardiovascular parameters in the context of this support. In the acquisition of subconscious expertise it is established that it is better to observe an expert solve a problem in real time than to ask them to describe what they do in the abstract. Hypothesis: Using knowledge capture techniques it is possible to design a scale of the overall state of a critically ill patient underpinned by a sophisticated physiological rule base. Methods: Data sets were prepared from the Electronic Records of 10 patients with 2761 time points of routinely collected physiological and pharmacological parameters. A clinician scored each time point as stable (A) through to unstable (E) whilst simultaneously describing a rule set of ranges (A-E) of derangement for each parameter. The same time points were annotated automatically using the rule set described in the abstract and inconsistencies between the two sets of annotations compared in a confusion matrix. Each disagreement was analyzed and changes made to the rule base where appropriate to better capture clinical expertise. The process was repeated with two other clinicians, their clinical annotations being tested against the previous clinician’s rule set, resulting in further refinements. Results: Agreement between clinician 1’s final rule set and annotations post refinement was 96.7%. Agreement between clinician 1’s final rule set clinician 2’s initial annotations and was 10.7% (97.6% after refinement). Initial agreement between clinician 2’s final rule set and clinician 3’s initial annotations was 90.6% (98.1% after refinement - the final A to E rule set for the new score).Conclusions: There was a higher agreement between the final rule set of the first clinician and the initial annotations of each subsequent clinician as the refinement proceeded. It is possible to design a sophisticated rule base underpinning a physiological score using knowledge capture techniques.
This research explores relations between software artefacts and explicitly represented (domain) knowledge. More specifically, we investigate ways in which domain knowledge (represented as ontologies) can support software engineering activities and, conversely, how software artefacts (e.g., programs, methods, and UML diagrams) can support the creation of ontologies. In our approach, class names, and class properties are the principal entities which are extracted from both sources. We implemented a tool, called Facilitator, to support programmers and knowledge engineers when they develop ontologies or programs. This tool provides a list of connections between the ontology and Java project, and provides reasons why these connections have been identified. These connections are created by matching names, types, and superclass-subclass relationships. Facilitator provides a range of semantic web enabled functionalities.
An earlier study asked two experts to discuss conditions under which Myocardial Damage (MD) can occur in ICU patients; these experts were shown temporal records where some sequences resulted in MD and others which did not. The resulting model was quite complex as it contained temporal constraints as well as the usual conjunctive and disjunctive terms. This was a classical KC Study. We have since implemented a Temporal Discovery Workbench (TDWB) to process the same temporal datasets to see if TDWB can discover simpler patterns to explain the same datasets. Subsequently, we have shown that the sets of patterns produced by TDWB generally have better “coverage”, than those produced by the original model. We then investigated whether some of the TDWB-created patterns might not be clinically acceptable. Recently we ran a pilot study in which we asked a single clinician to evaluate the patterns produced by TDWB, and to say whether they were acceptable, and why. This further information has now been implemented in TDWB; the resulting set of filtered patterns still has better coverage than the initial set of “manual” patterns.
Reuse has long been a major goal of the knowledge engineering community. We present a case study of the reuse of constraint knowledge acquired for one problem solver, by two further problem solvers. For our analysis, we chose a well-known benchmark knowledge base (KB) system written in CLIPS, which was based on the propose and revise problem-solving method and which had a lift/elevator KB. The KB contained four components, including constraints and data tables, expressed in an ontology that reflects the propose and revise task structure. Sufficient trial data was extracted manually to demonstrate the approach on two alternative problem solvers: a spreadsheet (Excel) and a constraint logic solver (ECLiPSe). The next phase was to implement ExtrAKTor, which automated the process for the whole KB. Each KB that is processed results in a working system that is able to solve the corresponding configuration task (and not only for elevators). This is in contrast to earlier work, which produced abstract formulations of the problem-solving methods but which were unable to perform reuse of actual KBs. We subsequently used the ECLiPSe solver on some more demanding vertical transport configuration tasks. We found that we had to use a little-known propagation technique described by Le Provost and Wallace (1991). Further, our techniques did not use any heuristic "fix"' information, yet we successfully dealt with a "thrashing" problem that had been a key issue in the original vertical transit work. Consequently, we believe we have developed a widely usable approach for solving this class of parametric design problem, by applying novel constraint-based problem solvers to data and formulae stored in existing KBs.
Temporal datasets are now often collected and curated in industry, scientific labs, and healthcare. A considerable amount of work has been done to analyze these for trends, inconsistencies and to make use of them for prediction. In an earlier study we discussed datasets for 15 or so Intensive Care Unit patients with clinicians, and asked them if they could detect when a particular harmful event (a myocardial infarction) had occurred. After 3 lengthy knowledge capture sessions we formulated a complex model to identify myocardial damage which we then implemented and tested against a test dataset; a detection rate of approximately 80% was achieved. (Specifically the model suggests that several events generally occur in a temporal sequence before the event-to-be-predicted occurs.) This work reports the design of a Temporal Discovery Workbench (TDWB) to address this class of tasks and has reproduced the results of the initial model acquired from the experts. Further we have now run TDWB's pattern discovery module with a range of settings to see if further clinically useful patterns are reported. Initial results are encouraging.
To retrieve information from documents, there are many Information Retrieval (IR) techniques. Current IR techniques are not so advanced that they can be able to exploit semantic knowledge within documents and give precise results. IR technology is major factor responsible for handling annotations in Semantic Web (SW) languages. With the rate of growth of web and huge amount of information available on the web which may be in unstructured, semi structured or structured form, it has become increasingly difficult to identify the relevant pieces of information on the internet.In this paper, implementation of new proposed model, “Mining in Ontology with Multi Agent Systems” has been discussed and analyzed the model for comparative study in the search of RDBMS system and Ontology based system. In this model, the Semantic Web addresses the first part of this challenge by trying to make the data also machine understandable in the form of Ontology, while MultiAgent addresses the second part by semi-automatically extracting the useful knowledge hidden in these data, and making it available.
Objective: While EIRA has proved to be successful in the detection of anomalous patient responses to treatments in the Intensive Care Unit, it could not describe to clinicians the rationales behind the anomalous detections. The aim of this paper is to address this problem. Methods: Few attempts have been made in the past to build knowledge-based medical systems that possess both argumentation and explanation capabilities. Here we propose an approach based on Dung's seminal calculus of opposition. Results: We have developed a new tool, arguEIRA, which is an extension of the existing EIRA system. In this paper we extend EIRA by providing it with an argumentation-based justification system that formalizes and communicates to the clinicians the reasons why a patient response is anomalous. Conclusion: Our comparative evaluation of the EIRA system against the newly developed tool highlights the multiple benefits that the use of argumentation-logic can bring to the field of medical decision support and explanation.
We present a Bayesian analysis of ordinal annotations made by clinicians of patients in intensive care. In particular, we investigate the different ways in which clinicians can disagree and how their disagreement is reduced once they take part in a recently proposed procedure (INSIGHT) that aims at improving consistency. The model combines a nonparametric function (loosely interpretable as the health of the patient) with clinician-specific generative procedures for producing the observed ordinal values. Our analysis provides valuable details of the rating behavior of the individual clinicians and shows that the INSIGHT procedure is particularly effective at removing (some) clinician-specific inconsistencies and biases.
To develop knowledge bases for use in knowledge-based (or expert) systems, domain experts are often interviewed or are asked to perform knowledge acquisition tasks. Although domain experts are highly regarded, they can still display cognitive biases (i.e. errors in judgement, knowledge, and reasoning) which can affect their performance and lead to differences in both the way several domain experts perform the task, and in some cases the conclusions drawn. Consequently, developing accurate knowledge bases is still a challenge for knowledge engineers. In this paper we illustrate this challenge with a case study from the Intensive Care Unit (ICU) domain. In this task the ICU clinicians were asked to identify possible anomalies from a set of patient datasets. In total, the clinicians identified 83 anomalies, of which there were only 9 instances where an anomaly was identified by more than one clinician. A further investigation explores whether individual problem solving strategies or biases are possibly responsible for the differences.
Introduction: We have aimed to describe haemodynamic changes when haemodialysis is instituted in the critically ill. 3 hypotheses are tested: 1)The initial session is associated with cardiovascular instability, 2)The initial session is associated with more cardiovascular instability compared to subsequent sessions, and 3)Looking at unstable sessions alone, there will be a greater proportion of potentially harmful changes in the initial sessions compared to subsequent ones. Methods: Data was collected for 209 patients, identifying 1605 dialysis sessions. Analysis was performed on hourly records, classifying sessions as stable/unstable by a cutoff of >+/-20% change in baseline physiology (HR/MAP). Data from 3 hours prior, and 4 hours after dialysis was included, and average and minimum values derived. 3 time comparisons were made (pre-HD:during, during HD:post, pre-HD:post). Initial sessions were analysed separately from subsequent sessions to derive 2 groups. If a session was identified as being unstable, then the nature of instability was examined by recording whether changes crossed defined physiological ranges. The changes seen in unstable sessions could be described as to their effects: being harmful/potentially harmful, or beneficial/potentially beneficial. Results: Discarding incomplete data, 181 initial and 1382 subsequent sessions were analysed. A session was deemed to be stable if there was no significant change (>+/-20%) in the time-averaged or minimum MAP/HR across time comparisons. By this definition 85/181 initial sessions were unstable (47%, 95% CI SEM 39.8-54.2). Therefore Hypothesis 1 is accepted. This compares to 44% of subsequent sessions (95% CI 41.1-46.3). Comparing these proportions and their respective CI gives a 95% CI for the standard error of the difference of -4% to 10%. Therefore Hypothesis 2 is rejected. In initial sessions there were 92/1020 harmful changes. This gives a proportion of 9.0% (95% CI SEM 7.4-10.9). In the subsequent sessions there were 712/7248 harmful changes. This gives a proportion of 9.8% (95% CI SEM 9.1-10.5). Comparing the two unpaired proportions gives a difference of -0.08% with a 95% CI of the SE of the difference of -2.5 to +1.2. Hypothesis 3 is rejected. Fisher’s exact test gives a result of p=0.68, reinforcing the lack of significant variance. Conclusions: Our results reject the claims that using haemodialysis is an inherently unstable choice of therapy. Although proportionally more of the initial sessions are classed as unstable, the majority of MAP and HR changes are beneficial in nature.
Wamberto Weber Vasconcelos合作论文数Department of Computing Science,University of Aberdeen12
Eugenio Alberdi合作论文数Centre for Software Reliability8
Kieron O'Hara合作论文数University of Southampton3
Fabio Ciravegna合作论文数Aeqora Ltd;Department of Computer Science, The University of Sheffield3
P. M. D. Gray合作论文数University of Aberdeen;Department of Computing Science3
Craig Mckenzie合作论文数University of Aberdeen, Computing Science, Aberdeen, UK3