Many social and ecological problems require us to consider objectively verifiable phenomena as well as subjective states of knowledge and associated value systems. When approximating the facts of reality, the wisdom of crowds phenomenon demonstrates that many pooled estimates can be more accurate than individual or expert estimates. For complex and social systems, wisdom of crowd approaches are improved by aggregating knowledge over subpopulations. In this paper we consider subpopulations defined by different sets of shared values. We first discuss two approaches to qualitatively understanding differences in value sets held by individuals and groups, which in turn motivate our discussion of three unsupervised methods for identifying subpopulations based upon value-laden statements in narrative data from hyperlocal maternal and child health (MCH) contexts in Gombe State, Nigeria. We employ data science techniques and compare methods to assess the stability of inferences. We find the hypothesized groups to be method dependent and discuss implications for wisdom-of-crowd estimates in sustainable development contexts.
ObjectiveLanguage used by providers in medical documentation may reveal evidence of race-related implicit bias. We aimed to use natural language processing (NLP) to examine if prevalence of stigmatizing language in emergency medicine (EM) encounter notes differs across patient race/ethnicity.MethodsIn a retrospective cohort of EM encounters, NLP techniques identified stigmatizing and positive themes. Logistic regression models analyzed the association of race/ethnicity and themes within notes. Outcomes were the presence (or absence) of 7 different themes: 5 stigmatizing (difficult, non-compliant, skepticism, substance abuse/seeking, and financial difficulty) and 2 positive (compliment and compliant).ResultsThe sample included notes from 26,363 unique patients. NH Black patient notes were less likely to contain difficult (odds ratio (OR) 0.80, 95% confidence interval (CI), 0.73-0.88), skepticism (OR 0.87, 95% CI, 0.79-0.96), and substance abuse/seeking (OR 0.62, 95% CI, 0.56-0.70) compared to NH White patient notes but more likely to contain non-compliant (OR 1.26, 95% CI, 1.17-1.36) and financial difficulty (OR 1.14, 95% CI, 1.04-1.25). Hispanic patient notes were less likely to contain difficult (OR 0.68, 95% CI, 0.58-0.80) and substance abuse/seeking (OR 0.78, 95% CI, 0.66-0.93). NH NA/AI patient notes had twice the odds as NH White patient notes to contain a stigmatizing theme (OR 2.02, 95% CI, 1.64-2.49).ConclusionsUsing an NLP model to analyze themes in EM notes across racial groups, we identified several inequities in the usage of positive and stigmatizing language. Interventions to minimize race-related implicit bias should be undertaken.
Large language models (LLMs) generate diverse, situated, persuasive texts from a plurality of potential perspectives, influenced heavily by their prompts and training data. As part of LLM adoption, we seek to characterize - and ideally, manage - the socio-cultural values that they express, for reasons of safety, accuracy, inclusion, and cultural fidelity. We present a validated approach to automatically (1) extracting heterogeneous latent value propositions from texts, (2) assessing resonance and conflict of values with texts, and (3) combining these operations to characterize the pluralistic value alignment of human-sourced and LLM-sourced textual data.
With the growing prevalence of large language models, it is increasingly common to annotate datasets for machine learning using pools of crowd raters. However, these raters often work in isolation as individual crowdworkers. In this work, we regard annotation not merely as inexpensive, scalable labor, but rather as a nuanced interpretative effort to discern the meaning of what is being said in a text. We describe a novel, collaborative, and iterative annotator-in-the-loop methodology for annotation, resulting in a 'Bridging Benchmark Dataset' of comments relevant to bridging divides, annotated from 11,973 textual posts in the Civil Comments dataset. The methodology differs from popular anonymous crowd-rating annotation processes due to its use of an in-depth, iterative engagement with seven US-based raters to (1) collaboratively refine the definitions of the to-be-annotated concepts and then (2) iteratively annotate complex social concepts, with check-in meetings and discussions. This approach addresses some shortcomings of current anonymous crowd-based annotation work, and we present empirical evidence of the performance of our annotation process in the form of inter-rater reliability. Our findings indicate that collaborative engagement with annotators can enhance annotation methods, as opposed to relying solely on isolated work conducted remotely. We provide an overview of the input texts, attributes, and annotation process, along with the empirical results and the resulting benchmark dataset, categorized according to the following attributes: Alienation, Compassion, Reasoning, Curiosity, Moral Outrage, and Respect.
Understanding and modeling collective intelligence is essential for addressing complex social systems. Directed graphs called fuzzy cognitive maps (FCMs) offer a powerful tool for encoding causal mental models, but extracting high-integrity FCMs from text is challenging. This study presents an approach using large language models (LLMs) to automate FCM extraction. We introduce novel graph-based similarity measures and evaluate them by correlating their outputs with human judgments through the Elo rating system. Results show positive correlations with human evaluations, but even the best-performing measure exhibits limitations in capturing FCM nuances. Fine-tuning LLMs improves performance, but existing measures still fall short. This study highlights the need for soft similarity measures tailored to FCM extraction, advancing collective intelligence modeling with NLP.
The fields of AI current lacks methods to quantitatively assess and potentially alter the moral values inherent in the output of large language models (LLMs). However, decades of social science research has developed and refined widely-accepted moral value surveys, such as the World Values Survey (WVS), eliciting value judgments from direct questions in various geographies. We have turned those questions into value statements and use NLP to compute to how well popular LLMs are aligned with moral values for various demographics and cultures. While the WVS is accepted as an explicit assessment of values, we lack methods for assessing implicit moral and cultural values in media, e.g., encountered in social media, political rhetoric, narratives, and generated by AI systems such as LLMs that are increasingly present in our daily lives. As we consume online content and utilize LLM outputs, we might ask, which moral values are being implicitly promoted or undercut, or -- in the case of LLMs -- if they are intending to represent a cultural identity, are they doing so consistently? In this paper we utilize a Recognizing Value Resonance (RVR) NLP model to identify WVS values that resonate and conflict with a given passage of output text. We apply RVR to the text generated by LLMs to characterize implicit moral values, allowing us to quantify the moral/cultural distance between LLMs and various demographics that have been surveyed using the WVS. In line with other work we find that LLMs exhibit several Western-centric value biases; they overestimate how conservative people in non-Western countries are, they are less accurate in representing gender for non-Western countries, and portray older populations as having more traditional values. Our results highlight value misalignment and age groups, and a need for social science informed technological solutions addressing value plurality in LLMs.
Joan Zheng, Scott Friedman, Sonja Schmer-galunder, Ian Magnusson, Ruta Wheelock, Jeremy Gottlieb, Diana Gomez, Christopher Miller. Proceedings of the Sixth Workshop on Online Abuse and Harms (WOAH). 2022.
Qualitative causal relationships compactly express the direction, dependency, temporal constraints, and monotonicity constraints of discrete or continuous interactions in the world. In everyday or academic language, we may express interactions between quantities (e.g., sleep decreases stress), between discrete events or entities (e.g., a protein inhibits another protein's transcription), or between intentional or functional factors (e.g., hospital patients pray to relieve their pain). Extracting and representing these diverse causal relations are critical for cognitive systems that operate in domains spanning from scientific discovery to social science. This paper presents a transformer-based NLP architecture that jointly extracts knowledge graphs including (1) variables or factors described in language, (2) qualitative causal relationships over these variables, (3) qualifiers and magnitudes that constrain these causal relationships, and (4) word senses to localize each extracted node within a large ontology. We do not claim that our transformer-based architecture is itself a cognitive system; however, we provide evidence of its accurate knowledge graph extraction in real-world domains and the practicality of its resulting knowledge graphs for cognitive systems that perform graph-based reasoning. We demonstrate this approach and include promising results in two use cases, processing textual inputs from academic publications, news articles, and social media.
Understanding the cognitive structure of explanations – and the cognitive processes that assemble them – is a milestone for understanding how people learn and communicate. Recent research on explanatory coexistence and representational plurality suggests that people's causal beliefs are less globally coherent than previously thought: people use seemingly competing supernatural and biological causes to explain different aspects of the same phenomenon, or they assemble supernatural and biological causes into single, coherent explanations. This coexistence – and unexpected coherence – of diverse causal mechanisms poses interesting questions about the role of coherence and fragmentation in people's mental models and explanations. This chapter presents a computational model of explanatory coherence in the well-characterized domain of disease transmission, extending a previous cognitive model of explanation-based conceptual change. Our computational model (1) retrieves diverse causal model fragments based on the phenomenon to explain, (2) assembles coherent causal models using relevance-directed abductive reasoning and (3) selects explanatory paths that support within-explanation and within-scenario coherence. Our model simulates the three different types of explanatory coexistence detailed in the literature.
Although implicit cultural values are reflected in human narrative texts, few robust computational solutions exist to recognize values that resonate within these texts. In other words, given a statement text and a value text, the task is to predict the label that resonates, conflicts or is neutral with respect to the value. In this paper, we present a novel, annotated dataset and transformer-based model for Recognizing Value Resonance (RVR). We created the World Values Corpus (WVC): a labeled collection of [statement, value] pairs of text based on the World Values Survey (WVS), which is a well-validated, comprehensive survey for assessing values across cultures. Each pair expresses whether the value resonates with, conflicts with, or is neutral to the statement. The 384 values in the WVC are derived from the WVS to assure the WVC’s cross-cultural relevance. The statement pairs for each value were generated by a pool of six annotators across genders and cultural backgrounds. We demonstrate that off-the-shelf Recognizing Textual Entailment (RTE) models perform unfavorably on the RVR task. However, RTE models trained on the WVC achieve substantially higher accuracy on RVR, serving as a strong, replicable baseline for future RVR work, advancing the study of cultural values using computational NLP approaches. We also present results of applying our baseline model on the “World of Tales” corpus, an online repository of international folktales. The results suggest that such a model can provide useful anthropological insights, which in turn is an important step towards facilitating automated ethnographic modeling.
Author(s): Friedman, Scott E; Magnusson, Ian; Schmer-Galunder, Sonja; Wheelock, Ruta; Gottlieb, Jeremy; patel, pooja; Miller, Christopher | Abstract: Moral disengagement is a mechanism whereby people distance or disconnect their actions from their moral evaluation. This work presents a novel knowledge graph schema, dataset, and transformer-based NLP model to identify and represent indicators of moral disengagement in text. Our graph schema is informed by Albert Bandura’s psychosocial mechanisms of moral disengagement, including dehumanization, victimization, moral condemnation and justification, and attribution (or displacement) of responsibility. Our preliminary dataset is comprised of online posts from five different communities. We present initial evidence that (1) our theory-based schema can represent moral disengagement indicators across these communities and (2) our transformer-based NLP model can identify indicators of moral disengagement in text. As it matures, this thread of computational social science research can help us understand the spread of morally-disengaged language and its effect on online communities.
The replicability of research is crucial for building trust in the peer review process and transitioning knowledge to realworld applications. While manual peer review excels in some regards, the variability of reviewer expertise, publication requirements, and research domains brings about uncertainty in the process. Replicability, in particular, is not necessarily a priority; this is evidenced by repeated failures in replication attempts such as the Psychology Reproducibility Project, where 61 of 100 replications fail. Improving human comprehension of decisive factors is crucial for integrating automated systems for replicability prediction into the review process. We develop a robust, automated method for semantic parsing, information extraction, and replication prediction that operates directly on PDFs. We introduce features that have not been explored in prior work, construct argument structures to guide understanding, and provide preliminary results for replication prediction.
Building and evaluating explainable Artificial Intelligence (AI) systems that accommodate human cognition remains a challenge for Human-Computer Interaction (HCI) and the need for practical solutions increases with our reliability on machines to extract, classify, and process information. Recent work has proposed triggers and metrics for explainable AI based on human mental models and psychological explanation quality. We complement this previous work by (1) extending and supporting these triggers and metrics with existing directives for information integrity, transparency, and rigor, (2) outlining a provenance-based framework for recording human-machine collaboration, and (3) demonstrating that a provenance-based approach address many of these explainable AI triggers and metrics. We show that provenance-based analyses help address questions of foundations, alternatives, necessity vs. sufficiency, sensitivity (e.g., what-if analyses), impact, and rationale, and we provide concrete evidence using an implemented human-machine analytic workspace. We outline ways to empirically measure the ability of these additional interpretation strategies to improve human understanding.
Qualitative causal relationships compactly express the direction, dependency, temporal constraints, and monotonicity constraints of discrete or continuous interactions in the world. In everyday or academic language, we may express interactions between quantities (e.g., sleep decreases stress), between discrete events or entities (e.g., a protein inhibits another protein's transcription), or between intentional or functional factors (e.g., hospital patients pray to relieve their pain). This paper presents a transformer-based NLP architecture that jointly identifies and extracts (1) variables or factors described in language, (2) qualitative causal relationships over these variables, and (3) qualifiers and magnitudes that constrain these causal relationships. We demonstrate this approach and include promising results from in two use cases, processing textual inputs from academic publications, news articles, and social media.
Recent transformer-based approaches demonstrate promising results on relational scientific information extraction. Existing datasets focus on high-level description of how research is carried out. Instead we focus on the subtleties of how experimental associations are presented by building SciClaim, a dataset of scientific claims drawn from Social and Behavior Science (SBS), PubMed, and CORD-19 papers. Our novel graph annotation schema incorporates not only coarse-grained entity spans as nodes and relations as edges between them, but also fine-grained attributes that modify entities and their relations, for a total of 12,738 labels in the corpus. By including more label types and more than twice the label density of previous datasets, SciClaim captures causal, comparative, predictive, statistical, and proportional associations over experimental variables along with their qualifications, subtypes, and evidence. We extend work in transformer-based joint entity and relation extraction to effectively infer our schema, showing the promise of fine-grained knowledge graphs in scientific claims and beyond.
Ghost fishing in derelict blue crab traps is ubiquitous and causes incidental mortality which can be reduced by trap removal programs. In an effort to scale the benefits of such removal programs, in the context of restoring the Gulf of Mexico after the Deepwater Horizon oil spill, this paper calculates the ecological benefits of trap removal by estimating the extent of derelict blue crab traps across Gulf of Mexico waterbodies and combining these estimates with Gulf-specific crab and finfish mortality rates due to ghost fishing. The highest numbers and densities of traps are found in Louisiana, with estimates ranging up to 203,000 derelict traps across the state and up to 41 traps per square kilometer in areas such as Terrebonne Bay. Mortality rates are estimated at 26 crabs per trap per year and 8 fish per trap per year. The results of this analysis indicate a Gulf-wide removal program targeting 10% of derelict traps over the course of 5 years would lead to a combined benefit of more than 691,000 kg of crabs and fish prevented from mortality in ghost fishing traps. These results emphasize the importance of ongoing derelict trap removal programs. Future work could assess additional benefits of trap removal programs, such as fewer entanglements of marine organisms, improved esthetics, and increases in harvestable catch. Lastly, this model could be utilized by fishery managers to calculate the benefits of other management options designed to decrease the extent and impact of derelict fishing gear.
Analytic software tools and workflows are increasing in capability, complexity, number, and scale, and the integrity of our workflows is as important as ever. Specifically, we must be able to inspect the process of analytic workflows to assess (1) confidence of the conclusions, (2) risks and biases of the operations involved, (3) sensitivity of the conclusions to sources and agents, (4) impact and pertinence of various sources and agents, and (5) diversity of the sources that support the conclusions. We present an approach that tracks agents' provenance with PROV-O in conjunction with agents' appraisals and evidence links (expressed in our novel DIVE ontology). Together, PROV-O and DIVE enable dynamic propagation of confidence and counter-factual refutation to improve human-machine trust and analytic integrity. We demonstrate representative software developed for user interaction with that provenance, and discuss key needs for organizations adopting such approaches. We demonstrate all of these assessments in a multi-agent analysis scenario, using an interactive web-based information validation UI.
Robert P. Goldman合作论文数Computer Science Research8
Jason Fritts合作论文数Department of Computer Science, Saint Louis University2