In this paper, we survey measures aimed at quantifying the intrinsic quality of cases' contents, which we collectively refer to as case fidelity, and measures that capture their competence within the Case-Based Reasoning literature. We discuss how insights from the Truth Discovery and Item Response Theory literature can respectively inform advancements in estimating case fidelity and competence. Additionally, we highlight novel research directions that emerge from a deeper examination of case fidelity and competence.
There has been much recent interest in developing fair clustering algorithms that seek to do justice to the representation of groups defined along sensitive attributes such as race and sex. Within the centroid clustering paradigm, these algorithms are seen to generate clusterings where different groups are disadvantaged within different clusters with respect to their representativity, i.e., distance to centroid. In view of this deficiency, we propose a novel notion of cluster-level centroid fairness that targets the representativity unfairness borne by groups within each cluster, along with a metric to quantify the same. Towards operationalising this notion, we draw on ideas from political philosophy aligned with consideration for the worst-off group to develop Fair-Centroid; a new clustering method that focusses on enhancing the representativity of the worst-off group within each cluster. Our method uses an iterative optimisation paradigm wherein an initial cluster assignment is refined by reassigning objects to clusters such that the worst-off group in each cluster is benefitted. We compare our notion with a related fairness notion and show through extensive empirical evaluations on real-world datasets that our method significantly enhances cluster-level centroid fairness at low impact on cluster coherence.
There have been massive advances in AI technologies towards addressing the contemporary challenge of fake news identification. However, these technologies, as observed widely, have not had the same kind or depth in impact across global societies. In particular, the AI scholarship in fake news detection arguably has not been as beneficial or appropriate for Global South, bringing geo-political bias into the picture. While it is often natural to think of data bias as the potential reason for geo-political bias, other factors could be much more important in being more latent, and thus less visible. In this commentary, we investigate as to how the facet of affect, comprising emotions and sentiments, could be a potent vehicle for geo-political biases in AI. We highlight, through assembling and interpreting insights from literature, the overarching neglect of affect across methods for fake news detection AI, and how this could be a potentially important factor for geo-political bias within them. This exposition, we believe, also serves as a first effort in understanding how geo-political biases work within AI pipelines beyond the data collection stage.
This article critically examines the recent hype around AI safety. We first start with noting the nature of the AI safety hype as being dominated by governments and corporations, and contrast it with other avenues within AI research on advancing social good. We consider what 'AI safety' actually means, and outline the dominant concepts that the digital footprint of AI safety aligns with. We posit that AI safety has a nuanced and uneasy relationship with transparency and other allied notions associated with societal good, indicating that it is an insufficient notion if the goal is that of societal good in a broad sense. We note that the AI safety debate has already influenced some regulatory efforts in AI, perhaps in not so desirable directions. We also share our concerns on how AI safety may normalize AI that advances structural harm through providing exploitative and harmful AI with a veneer of safety.
From the latter half of the last decade, there has been a growing interest in developing algorithms for automatically solving mathematical word problems (MWP). It is a challenging and unique task that demands blending surface level text pattern recognition with mathematical reasoning. In spite of extensive research, we still have a lot to explore for building robust representations of elementary math word problems and effective solutions for the general task. In this paper, we critically examine the various models that have been developed for solving word problems, their pros and cons and the challenges ahead. In the last 2 years, a lot of deep learning models have recorded competing results on benchmark datasets, making a critical and conceptual analysis of literature highly useful at this juncture. We take a step back and analyze why, in spite of this abundance in scholarly interest, the predominantly used experiment and dataset designs continue to be a stumbling block. From the vantage point of having analyzed the literature closely, we also endeavor to provide a road-map for future math word problem research. This article is categorized under: Technologies > Machine Learning Technologies > Artificial Intelligence Fundamental Concepts of Data and Knowledge > Knowledge Representation
Web search engines arguably form the most popular data-driven systems in contemporary society. They wield a considerable power by functioning as gatekeepers of the Web. Since the late 1990s, search engines have been dominated by the paradigm of link-based web search. In this paper, we critically analyse the Political Economy of the paradigm of link-based web search, drawing upon insights and methodologies from Critical Political Economy. We illustrate how link-based web search has led to phenomena that favour capital through long-term structural changes on the Web, and how it has led to accentuating unpaid digital labour and ecologically unsustainable practices, among several others. We show how con-temporary observations on the degrading quality of link-based web search can be traced back to the internal contradictions with the paradigm, and how such socio-technical phenom-ena may lead to an eventual disutility of the link-based model. Our contribution is on enhanc-ing the understanding of the Political Economy of link-based web search, and laying bare the phenomena at work, towards catalysing the search for alternative models of content organi-sation and search on the Web.
Groundbreaking inventions and highly significant performance improvements in deep learning based Natural Language Processing are witnessed through the development of transformer based large Pre-trained Language Models (PLMs). The wide availability of unlabeled data within human generated data deluge along with self-supervised learning strategy helps to accelerate the success of large PLMs in language generation, language understanding, etc. But at the same time, latent historical bias/unfairness in human minds towards a particular gender, race, etc., encoded unintentionally/intentionally into the corpora harms and questions the utility and efficacy of large PLMs in many real-world applications, particularly for the protected groups. In this paper, we present an extensive investigation towards understanding the existence of “Affective Bias” in large PLMs to unveil any biased association of emotions such as anger, fear, joy, etc., towards a particular gender, race or religion with respect to the downstream task of textual emotion detection. We conduct our exploration of affective bias from the very initial stage of corpus level affective bias analysis by searching for imbalanced distribution of affective words within a domain, in large scale corpora that are used to pre-train and fine-tune PLMs. Later, to quantify affective bias in model predictions, we perform an extensive set of class-based and intensity-based evaluations using various bias evaluation corpora. Our results show the existence of statistically significant affective bias in the PLM based emotion detection systems, indicating biased association of certain emotions towards a particular gender, race, and religion.
Technological advancements in web platforms allow people to express and share emotions towards textual write-ups written and shared by others. This brings about different interesting domains for analysis; emotion expressed by the writer and emotion elicited from the readers. In this paper, we propose a novel approach for Readers' Emotion Detection from short-text documents using a deep learning model called REDAffectiveLM. Within state-of-the-art NLP tasks, it is well understood that utilizing context-specific representations from transformer-based pre-trained language models helps achieve improved performance. Within this affective computing task, we explore how incorporating affective information can further enhance performance. Towards this, we leverage context-specific and affect enriched representations by using a transformer-based pre-trained language model in tandem with affect enriched Bi-LSTM+Attention. For empirical evaluation, we procure a new dataset REN-20k, besides using RENh-4k and SemEval-2007. We evaluate the performance of our REDAffectiveLM rigorously across these datasets, against a vast set of state-of-the-art baselines, where our model consistently outperforms baselines and obtains statistically significant results. Our results establish that utilizing affect enriched representation along with context-specific representation within a neural architecture can considerably enhance readers' emotion detection. Since the impact of affect enrichment specifically in readers' emotion detection isn't well explored, we conduct a detailed analysis over affect enriched Bi-LSTM+Attention using qualitative and quantitative model behavior evaluation techniques. We observe that compared to conventional semantic embedding, affect enriched embedding increases ability of the network to effectively identify and assign weightage to key terms responsible for readers' emotion detection.
Since ChatGPT's debut, generative AI technologies have surged in popularity within the AI community. Recognized for their cutting-edge language processing capabilities, these excel in generating human-like conversations, enabling open-ended dialogues with end-users. We consider that the future adoption of generative AI for critical public domain applications transforms the accountability relationship. Previously characterized by the relationship between an actor and a forum, the introduction of generative systems complicates accountability dynamics as the initial interaction shifts from the actor to an advanced generative system. We conceptualise a dual-phase accountability relationship involving the actor, the forum, and the generative AI as a foundational approach to understanding public sector accountability in the context of these technologies. Focusing on integrating generative AI for assisting healthcare triaging, we identify potential challenges introduced for maintaining effective accountability relationships, highlighting concerns that these technologies relegate actors to a secondary phase of accountability and creates a disconnect between government actors and citizens. We suggest recommendations aimed at disentangling the complexities generative systems bring to the accountability relationship. As we speculate on the technologies disruptive impact on accountability, we urge public servants, policymakers, and system designers to deliberate on the potential accountability impact generative systems produce prior to their deployment.
Fact-checking of health-related claims has become necessary in this digital age, where any information posted online is easily available to everyone.The most effective way to verify such claims is by using evidences obtained from reliable sources of medical knowledge, such as PubMed.Recent advances in the field of NLP have helped automate such fact-checking tasks.In this work, we propose a domainspecific BERT-based model using a transfer learning approach for the task of predicting the veracity of claim-evidence pairs for the verification of health-related facts.We also improvise on a method to combine multiple evidences retrieved for a single claim, taking into consideration conflicting evidences as well.We also show how our model can be exploited when labelled data is available and how backtranslation can be used to augment data when there is data scarcity.
In this paper, we demonstrate that diverse CBR research contexts share a common thread, in that their origin can be traced to the problem of circularity . An example is where the knowledge of property A requires us to know property B , but B , in turn, is not known unless A is determined. We examine the root cause of such circularities and present fundamental impossibility results in this context. We show how a systematic study of circularity can motivate the quest for novel CBR paradigms and lead to novel approaches that address circularities in traditional CBR retrieval, adaptation, and maintenance tasks. Furthermore, such an analysis can help in extending the solution of one problem to solve an apparently unrelated problem, once we discover the commonality they share deep down in terms of the circularities they address.
It is well documented that there has been significant enthusiasm across the globe in respect of using AI for all forms of social activity. However, the electoral process – the time, place, and manner of elections within democratic nations – is one of few sectors in which there has been limited penetration of AI. Electoral management bodies in many countries have recently started exploring and deliberating over the use of AI in the electoral process. In this paper, we consider five avenues within the core electoral process which have potential for AI usage, and map the challenges involved in using AI within them. These five avenues are: voter list maintenance, determining polling booth locations, polling booth protection processes, voter authentication, and video monitoring of elections. Within each avenue, we lay down the context, illustrate current or potential usage of AI, and discuss extant or potential ramifications of AI usage, as well as potential directions for mitigating risks when considering AI usage. We believe that the scant current usage of AI within electoral processes provides a very rare opportunity to deliberate on the risks and mitigation possibilities prior to actual and widespread AI deployment. This paper is an attempt to map the horizons of risks and opportunities in using AI within electoral processes and to help shape the debate around the topic.
Solving kinematics word problems is a specialized task which is best addressed through bespoke logical reasoners. Reasoners, however, require structured input in the form of kinematics parameter values, and translating textual word problems to such structured inputs is a key step in enabling end-to-end automated word problem solving. Span detection for a kinematics parameter is the process of identifying the smallest span of text from a kinematics word problem that has the information to estimate the value of that parameter. A key aspect differentiating kinematics span detection from other span detection tasks is the presence of multiple inter-related parameters for which separate spans need to be identified. State-of-the-art span detection methods are not capable of leveraging the existence of a plurality of inter-dependent span identification tasks. We propose a novel neural architecture that is designed to exploit the inter-relatedness between the separate span detection tasks using a single joint model. This allows us to train the same network for span detection over multiple kinematics parameters, implicitly and automatically transferring knowledge across the kinematics parameters. We show that such a joint training delivers an improvement of accuracies over real-world datasets against state-of-the-art methods for span detection.
There has been a significant recent interest in algorithmic fairness within data-driven systems. In this paper, we consider group fairness within Case-based Reasoning. Group fairness targets to ensure parity of outcomes across pre-specified sensitive groups, defined on the basis of extant entrenched discrimination. Addressing the context of binary decision choice scenarios over binary sensitive attributes, we develop three separate fairness interventions that operate at different stages of the CBR process. These techniques, called Label Flipping (LF), Case Weighting (CW) and Weighted Adaptation (WA), use distinct strategies to enhance group fairness in CBR decision making. Through an extensive empirical evaluation over several popular datasets and against natural baseline methods, we show that our methods are able to achieve significant enhancements in fairness at low detriment to accuracy, thus illustrating effectiveness of our methods at advancing fairness.
This journal review provides an overview of the current state of cloud computing, including its definition, benefits, and challenges. It examines the various cloud computing models and architectures, and discusses the security and privacy issues associated with the cloud. It also looks at the potential of cloud computing to transform the IT industry and provide new opportunities for businesses. Finally, the review looks at the future of cloud computing, including potential use cases and applications Cloud computing is a paradigm that has revolutionized the way we store, process and access data and applications. At its core, it involves delivering computing resources such as servers, storage, and applications over the internet, allowing users to access them on-demand from anywhere and at any time. This technology has numerous benefits, including cost-effectiveness, scalability, flexibility, and ease of maintenance. It has become a crucial enabler for many businesses and organizations, offering them the ability to streamline their operations and remain competitive in a constantly changing market. In this abstract, we explore the fundamentals of cloud computing, its key features, and its potential impact on the future of technology.. KEYWORDS Cloud computing, technology, advantages, disadvantages, security, challenges, opportunities.
Online reviews have become critical in informing purchasing decisions, making the detection of fake reviews a crucial challenge to tackle.Many different Machine Learning based solutions have been proposed, using various data representations such as n-grams or document embeddings.In this paper, we first explore the effectiveness of different data representations, including emotion, document embedding, n-grams, and noun phrases in embedding format, for fake reviews detection.We evaluate these representations with various state-of-theart deep learning models, such as a BILSTM, LSTM, GRU, CNN, and MLP.Following this, we propose to incorporate different data representations and classification models using early and late data fusion techniques in order to improve the prediction performance.The experiments are conducted on four datasets: Hotel, Restaurant, Amazon, and Yelp.The results demonstrate that a combination of different data representations significantly outperforms any single data representation.
Groundbreaking inventions and highly significant performance improvements in deep learning based Natural Language Processing are witnessed through the development of transformer based large Pre-trained Language Models (PLMs). The wide availability of unlabeled data within human generated data deluge along with self-supervised learning strategy helps to accelerate the success of large PLMs in language generation, language understanding, etc. But at the same time, latent historical bias/unfairness in human minds towards a particular gender, race, etc., encoded unintentionally/intentionally into the corpora harms and questions the utility and efficacy of large PLMs in many real-world applications, particularly for the protected groups. In this paper, we present an extensive investigation towards understanding the existence of "Affective Bias" in large PLMs to unveil any biased association of emotions such as anger, fear, joy, etc., towards a particular gender, race or religion with respect to the downstream task of textual emotion detection. We conduct our exploration of affective bias from the very initial stage of corpus level affective bias analysis by searching for imbalanced distribution of affective words within a domain, in large scale corpora that are used to pre-train and fine-tune PLMs. Later, to quantify affective bias in model predictions, we perform an extensive set of class-based and intensity-based evaluations using various bias evaluation corpora. Our results show the existence of statistically significant affective bias in the PLM based emotion detection systems, indicating biased association of certain emotions towards a particular gender, race, and religion.
Research on fake reviews detection and review helpfulness prediction is prevalent, yet most studies tend to focus solely on either fake reviews detection or review helpfulness prediction, considering them separate research tasks.In contrast to this prevailing pattern, we address both challenges concurrently by employing a multi-task learning approach.We posit that undertaking these tasks simultaneously can enhance the performance of each task through shared information among features.We utilize pre-trained RoBERTa embeddings with a document-level data representation.This is coupled with an array of deep learning and neural network models, including Bi-LSTM, LSTM, GRU, and CNN.Additionally, we employ ensemble learning techniques to integrate these models, with the objective of enhancing overall prediction accuracy and mitigating the risk of overfitting.The findings of this study offer valuable insights to the fields of NLP and machine learning and present a novel perspective on leveraging multi-task learning for the twin challenges of fake reviews detection and review helpfulness prediction.
Decision-making algorithms are becoming intertwined with each aspect of society. As we automate tasks which result in outcomes that affect an individual’s life, the need for assessing and understanding the ethical consequences of these processes becomes vital. With bias often originating from the datasets imbalanced group distributions, we propose a novel approach to in-processing fairness techniques, by considering training at a group-level. Adapting the standard training process of the logistic regression, our approach considers aggregating coefficient derivatives at a group-level to produce fairer outcomes. We demonstrate on two real-world datasets that our approach provides groups with more equal weighting towards defining the model parameters and displays potential to reduce unfairness disparities in group imbalanced data. Our experimental results illustrate a stronger influence on improving fairness when considering binary sensitive attributes, which may prove beneficial in continuing to construct fair algorithms to reduce biases existing in decision-making practices. Whilst the results present our group-level approach achieving less fair results than current state-of-the-art directly optimized fairness techniques, we primarily observe improved fairness over fairness-agnostic models. Subsequently, we find our novel approach towards fair algorithms to be a small but crucial step towards developing new methods for fair decision-making algorithms.
Sahely Bhadra合作论文数Department of Computer Science and Automation of the Indian Institute of Science6
Raghu Krishnapuram合作论文数IBM India Research Lab5
Delip Rao合作论文数Johns Hopkins University3