The long-standing goal of creating a comprehensive, multi-purpose knowledge resource, reminiscent of the 1984 Cyc project, still persists in AI. Despite the success of knowledge resources like WordNet, ConceptNet, Wolfram|Alpha and other commercial knowledge graphs, verifiable, general-purpose widely available sources of knowledge remain a critical deficiency in AI infrastructure. Large language models struggle due to knowledge gaps; robotic planning lacks necessary world knowledge; and the detection of factually false information relies heavily on human expertise. What kind of knowledge resource is most needed in AI today? How can modern technology shape its development and evaluation? A recent AAAI workshop gathered over 50 researchers to explore these questions. This paper synthesizes our findings and outlines a community-driven vision for a new knowledge infrastructure. In addition to leveraging contemporary advances in knowledge representation and reasoning, one promising idea is to build an open engineering framework to exploit knowledge modules effectively within the context of practical applications. Such a framework should include sets of conventions and social structures that are adopted by contributors.
The paper presents a spatio-temporal ontology guided by a particular methodology, in which the semantics is constructed within a spatio-temporal interpretation structure that is built up in three stages. The first, stage stipulates a standard classical model of time and space. This structure forms the grounding for the interpretation. The next stage is the specification of domains of entities, which are either elements of the grounding structure (time points and regions) or constructions from these elements (mappings from time to space associated with individuals existing within the spatio-temporal structure). The final stage is the definition of conceptual vocabulary in terms of the grounding structure and the specified domains. This definitional stage can be further subdivided into three types of specification: direct grounding of primitives onto the underlying structure, indirect grounding by defining additional vocabulary in terms of grounded primitives, partial grounding by specifying semantics types and axioms to constrain the meaning of vocabulary that is not explicitly defined. The main goal of the paper is to advocate a methodology rather than a specific ontology. We suggest that building up in this way, results in robust ontologies, whose assumptions can be clearly seen, since they are encapsulated within the grounding stage and domain specifications. Although the definitional stage may incorporate a diverse and expressive vocabulary, its terms are essentially just labels for properties and relations that were already implicit within the grounding structure.
Classical semantics assumes that one can model reference, predication and quantification with respect to a fixed domain of precise referent objects. Non-logical terms and quantification are then interpreted directly in terms of elements and subsets of this domain. We explore ways to generalise this classical picture of precise predicates and objects to account for variability of meaning due to factors such as vagueness, context and diversity of definitions or opinions. Both names and predicative expressions can be given either multiple semantic referents or be associated with semantic referents that incorporate some model of variability. We present a semantic framework, Variable Reference Semantics, that can accommodate several modes of variability in relation to both predicates and objects.
The objective of this study was to assess the accuracy of two methods for predicting EQ-5D-5L utilities from EQ-5D-3L data.
To explore the effects of age and gender on mapped EQ-5D-5L utilities derived using algorithms developed by the National Institute for Health and Care Excellence Decision Support Unit (DSU) and EuroQol Group.
The Winograd Schema Challenge is a general test for Artificial Intelligence, based on problems of pronoun reference resolution. I investigate the semantics and interpretation of Winograd Schemas, concentrating on the original and most famous example. This study suggests that a rich ontology, detailed commonsense knowledge as well as special purpose inference mechanisms are all required to resolve just this one example. The analysis supports the view that a key factor in the interpretation and disambiguation of natural language is the preference for coherence. This preference guides the resolution of co-reference in relation to both explicitly mentioned entities and also implicit entities that are required to form an interpretation of what is being described. I suggest that assumed identity of implicit entities arises from the expectation of coherence and provides a key mechanism that underpins natural language understanding. I also argue that conceptual ontologies can play a decisive role not only in directly determining pronoun references but also in identifying implicit entities and implied relationships that bind together components of a sentence.
The Winograd Schema Challenge (WSC) is a commonsense reasoning task introduced as an alternative to the Turing Test. While machine learning approaches using language models show high performance on the original WSC data set, their performance degrades when tested on larger data sets. Moreover, they do not provide an interpretable explanation for their answers. To address these limitations, we present KARaML , a novel asymmetric method for integrating knowledge-based and machine learning approaches to tackle the WSC. A central idea in our work is that semantic roles are key for the high-level commonsense reasoning involved in the WSC. We extract semantic roles using a knowledge-based reasoning system. For this, we use relational representations of natural language sentences and define high-level patterns encoded in Answer Set Programming to identify relationships between entities based on their semantic roles. We then use the BERT language model to find the semantic role that best matches the pronoun. BERT performs better at this task than on the general WSC. We apply our ensemble method to a restricted domain of the large WSC data set, WinoGrande, and demonstrate that it achieves better performance than a state of the art pure machine learning approach.
Classical semantics assumes that one can model reference, predication and quantification with respect to a fixed domain of possible referent objects. Non-logical terms and quantification are then interpreted in relation to this domain: constant names denote unique elements of the domain, predicates are associated with subsets of the domain and quantifiers ranging over all elements of the domain. The current paper explores the wide variety of different ways in which this classical picture of precisely referring terms can be generalised to account for variability of meaning due to factors such as vagueness, context and diversity of definitions or opinions. Both predicative expressions and names can be given either multiple semantic referents or be associated with semantic referents that have some structure associated with variability. A semantic framework Variable Reference Semantics (VRS) will be presented that can accommodate several different modes of variability that may occur either separately or in combination. Following this general analysis of semantic variability, the phenomenon of co-predication will be considered. It will be found that this phenomenon is still problematic, even within the very flexible VRS framework.
The EQ-5D is widely used to inform economic evaluations of health technologies. However, the extent of its use for clinical outcome assessment (COA) is unclear. This review identified the prevalence with which EQ-5D data are evaluated by health authorities in clinical benefit assessments. Drug technology assessments (TAs) published by HTA agencies in England, France, Germany and the US during the last 2 years were identified. Product labelling for drugs approved by the European Medicines Agency (EMA) and US Food and Drug Administration (FDA) over the last 5 years were also reviewed. Only documents reporting EQ-5D as a COA measure were included. EQ-5D data were reported for COA in 139 TAs with the majority reported for Germany (n=78) and the remainder for England (n=46), France (n=12) and the US (n=3). Visual analogue scale (VAS) scores were presented most frequently (n=111) followed by utility index scores (n=48) and dimension levels (n=1). The VAS accounted for 99% of EQ-5D reports in Germany. Minimally important differences (MIDs) were discussed in 51 TAs: 34% and 24% of VAS and index score reports, respectively. Three-hundred twenty drugs were approved by the EMA and 735 by the FDA, and among these 15 and 35, respectively, presented EQ-5D data for COA. All EQ-5D data submitted to the FDA were reported in supporting documentation. Index scores, VAS scores and dimension levels were cited for 32, 26 and 5 drugs, respectively. Discussion of MIDs was more frequent in EMA documents (35%) than FDA documents (11%). The EQ-5D has been used for COA in HTA submissions; most frequently the VAS in German TAs. EQ-5D was also used to support labelling claims in a minority of EMA and FDA decisions. No EQ-5D data were reported in FDA product labelling, which suggests that the data were not considered material.
MPM is a rare and usually fatal malignancy. Severe, life-limiting symptoms include dyspnoea, pleural effusion, chest wall pain, and fatigue. The aim of this study was to explore the impact of MPM on the HRQoL of patients and their caregivers and to identify important treatment attributes. A mixed-methods design was employed, featuring a thematic analysis and the development of individual conceptual models for patients and caregivers. Individual, semi-structured interviews were conducted with people living with MPM and caregivers (N=45) in the UK and Australia. Participants also completed HRQoL questionnaires, including a newly developed EQ-5D-5L respiratory bolt-on. The study design and interim findings were reviewed by a panel of HEOR and mesothelioma experts. People living with MPM described themes including the experience of and debilitating burden of breathlessness, fatigue, and pain on daily life, as well as the distress that diagnosis can cause for themselves and their family, and their hopes for pursuing effective treatment despite uncertain outcomes and likely exacerbation of fatigue. Caregivers described themes including a range of caregiving duties and the impact on reducing or stopping work, as well as limitations to social and leisure activities. Caregivers described an emotional burden of providing care, the desire for additional support, and challenges when accommodating a patient’s needs while faced with their own difficulties, as well as positive relational impacts of caregiving. Questionnaire responses supported the patient and caregiver themes. Conceptual models will be presented using schematic diagrams. The themes and questionnaire responses demonstrate a substantial burden of MPM on patients and caregivers. Highlighted are increased impacts on both groups for more severely ill patients. This study offers perspectives on the most impactful aspects of living with MPM or caregiving for someone living with MPM, as well as the relative importance of treatment attributes for patients.
Achieving “commonsense reasoning” capabilities has been one of the goals of AI since its inception. However, as Marcus and Davis have recently argued, “Common sense is not just the hardest problem for AI; in the long run, it’s also the most important problem”. Moreover, it is generally accepted that space (and time) underlie much of what we regard as commonsense reasoning. Despite many successes in dealing with particular restricted types of spatial information, the development of a system capable of carrying out automated spatial reasoning of similar diversity to what one finds in ordinary natural language descriptions, seems to be a long way off. The chapter gives a general (though not comprehensive) overview of the goal of automating commonsense spatial reasoning by means of symbolic representations and reasoning. Existing work is surveyed, the nature of the goal clarified, and the problem analysed into seven interacting sub-problems.
In this paper, we present a formalism for handling polysemy in spatial expressions based on supervaluation semantics called standpoint semantics for polysemy (SSP). The goal of this formalism is, given a prepositional phrase, to define its possible spatial interpretations. For this, we propose to characterize spatial prepositions by means of a triplet $\langle $image schema, semantic feature, spatial axis$\rangle $. The core of SSP is predicate grounding theories, which are formulas of a first-order language that define a spatial preposition through the semantic features of its trajector and landmark. Precisifications are also established, which are a set of formulae of a qualitative spatial reasoning formalism that aims to provide the spatial characterization of the trajector with respect to the landmark. In addition to the theoretical model, we also present results of a computational implementation of SSP for the preposition ‘in’.
The Winograd Schema Challenge (WSC) is a common-sense reasoning task that requires background knowledge. In this paper, we contribute to tackling WSC in four ways. Firstly, we suggest a keyword method to define a restricted domain where distinctive high-level semantic patterns can be found. A thanking domain was defined by key-words, and the data set in this domain is used in our experiments. Secondly, we develop a high-level knowledge-based reasoning method using semantic roles which is based on the method of Sharma [2019]. Thirdly, we propose an ensemble method to combine knowledge-based reasoning and machine learning which shows the best performance in our experiments. As a machine learning method, we used Bidirectional Encoder Representations from Transformers (BERT) [Kocijan et al., 2019]. Lastly, in terms of evaluation, we suggest a "robust" accuracy measurement by modifying that of Trichelair et al. [2018]. As with their switching method, we evaluate a model by considering its performance on trivial variants of each sentence in the test set.
GBM patients experience debilitating neurological and physical symptoms often requiring caregiver support. This study explored burden among caregivers of GBM patients. Real-world data were drawn from the GBM Disease-Specific ProgrammeTM – a cross-sectional study administered to physicians and caregivers between May-July 2016 (EU5) and March-October 2019 (US). Caregivers completed a caregiver self-completion form (CSC) that captured demographics, care provided and caregiver burden (Zarit Burden Interview [ZBI]). ZBI scores range between 0-88; higher scores indicating greater burden. The corresponding burden for a score <20 is 'low', 20-40 is 'mild-moderate', and >40 is 'high'. CSCs were matched with corresponding patient record forms. Summary statistics were reported and descriptively analysed. 304 CSCs were completed. Mean (SD) age of caregivers was 55.7 (12.08), 31% were male, 70% were a partner/spouse and 82% resided with the patient. GBM patients with caregivers were mean (SD) age of 62.4 (12.40) years and 64% were male. Mean (SD) time since GBM diagnosis was 8.1 (6.6) months. Most patients had an ECOG score of 2 (37%) at time of data collection. 51% were on first line therapy, 55% had received both surgery and radiotherapy, and 45% had comorbidities. Mean (SD) overall ZBI score was 32.2 (15.76), 23% of caregivers had a ZBI score of <20 (low burden), 45% of 20–40 (mild-moderate burden) and 32% of >40 (high burden). Patients whose caregivers had high burden had a shorter mean time since diagnosis (230.1 days) compared with patients whose caregivers experienced mild-moderate or low burden (234.6 days and 305.9 days respectively). Among patients whose caregivers experienced high burden; 53% had comorbidities, 65% had an ECOG score ≥2 and 45% had received both surgery and radiotherapy. Most caregivers experience substantial burden. Consideration should be given to support GBM patients and their caregivers, particularly when patients have comorbidities.
Previous work on pooled data from five phase 3 nivolumab trials highlighted that use of the cancer-specific Quality of Life Utility Measure-Core 10 Dimensions (EORTC QLU-C10D) UK index yielded smaller quality-adjusted life year (QALY) differences between treatments than the generic 3-level EQ-5D (EQ-5D-3L) UK index on average. Since a cancer-specific utility measure might be expected to be more sensitive, this study aimed to explore the drivers of these differences. Associations between levels of each dimension were assessed through polychoric correlations, then a regularized partial correlation network was estimated to reduce spurious associations occurring due to background correlation between dimensions of the same instrument. These were supplemented by contingency tables, produced overall and by trial. Data from 2,625 patients at baseline were available. Polychoric correlations were highest between the following respective QLU-C10D and EQ-5D-3L dimensions: Pain and Pain/Discomfort (r=0.865), Role Functioning and Usual Activities (r=0.807), Physical Functioning and Mobility (r=0.784) and Emotional Functioning and Anxiety/Depression (r=0.755). These associations were consistently significant across trials as shown by contingency tables (p<0.001). The regularized network revealed an independent association between the QLU-C10D Social Functioning and EQ-5D-3L Usual Activities dimensions. QLU-C10D symptom-related dimensions (e.g., Sleep, Bowel Problems, Appetite Loss, Nausea and Fatigue) had limited to no independent association with EQ-5D-3L dimensions. This analysis facilitates a deeper understanding of potential uses of the cancer-specific QLU-C10D compared to the generic EQ-5D-3L in the evaluation of QALYs. Treatments with an impact on sleep, bowel function, appetite, nausea or fatigue will only influence QALYs as derived by the QLU-C10D; however, EQ-5D-3L utility decrements are often larger than the QLU-C10D on closely related dimensions. Therefore, increased sensitivity achieved through capturing additional dimensions in the cancer-specific measure can be counterbalanced by decreased sensitivity on the dimensions similar to the generic measure.