Misinformation about climate science is a serious challenge for our society. This paper introduces CPIQA (Climate Paper Image Question-Answering), a new question-answer dataset featuring 4,551 full-text open-source academic papers in the area of climate science with 54,612 GPT-4o generated question-answer pairs. CPIQA contains four question types (numeric, figure-based, non-figure-based, reasoning), each generated using three user roles (expert, non-expert, climate sceptic). CPIQA is multimodal, incorporating information from figures and graphs with GPT-4o descriptive annotations. We describe Context-RAG, a novel method for RAG prompt decomposition and augmentation involving extracting distinct contexts for the question. Evaluation results for Context-RAG on the benchmark SPIQA dataset outperforms the previous best state of the art model in two out of three test cases. For our CPIQA dataset, Context-RAG outperforms our standard RAG baseline on all five base LLMs we tested, showing our novel contextual decomposition method can generalize to any LLM architecture. Expert evaluation of our best performing model (GPT-4o with Context-RAG) by climate science experts highlights strengths in precision and provenance tracking, particularly for figure-based and reasoning questions.
Digitizing historical tabular records is essential for preserving and analyzing valuable data across various fields, but it presents challenges due to complex layouts, mixed text types, and degraded document quality. This paper introduces a comprehensive framework to address these issues through three key contributions. First, it presents UoS_Data_Rescue, a novel dataset of 1,113 historical logbooks with over 594,000 annotated text cells, designed to handle the complexities of handwritten entries, aging artifacts, and intricate layouts. Second, it proposes a novel context-aware text extraction approach (TrOCR-ctx) to reduce cascading errors during table digitization. Third, it proposes an enhanced end-to-end OCR pipeline that integrates TrOCR-ctx with ByT5, combining OCR and post-OCR correction in a unified training framework. This framework enables the system to produce both the raw OCR output and a corrected version in a single pass, improving recognition accuracy, particularly for multilingual and degraded text, within complex table digitization tasks. The model achieves superior performance with a 0.049 word error rate and a 0.035 character error rate, outperforming existing methods by up to 41% in OCR tasks and 10.74% in table reconstruction tasks. This framework offers a robust solution for large-scale digitization of tabular documents, extending its applications beyond climate records to other domains requiring structured document preservation. The dataset and implementation are available as open-source resources.
Much of academic discussion of responsible innovation (RI) has focused on RI integration into research projects. In addition, significant attention has also been paid to RI structures and policies at the research policy and institutional level. This article reports experiences of RI implementation with a focus on the intermediate i.e. meso-level. The research described here included a series of interviews that aimed to clarify researchers' perspectives on RI as well as barriers to and benefits of RI implementation. Two cases of engagement with research projects, with the aim of promoting RI, were undertaken. The analysis of the data demonstrates the crucial contribution that the meso-level of a research programme can make in interpreting, implementing and perpetuating RI across related activities. The article provides strong evidence that the scholarly debate surrounding RI should pay more explicit attention to this meso-level, ultimately strengthening RI theory and practice.
This study uses a novel semi-supervised learning framework to explore Tabular Structure Recognition (TSR) for digitizing historical documents, specifically employing the CascadeTabNet model. TSR is crucial for transforming archival tabular data into digital formats, enhancing accessibility and analysis across various research fields. Challenges like physical degradation, inconsistent lighting, and non-standard handwriting hinder the generation of high-quality annotations of historical documents needed for effective model training. To address these issues, this research explores two research questions: (i) Can a semi-supervised training approach reduce the need for expensive data annotations? and (ii) Does semi-supervised training improve model robustness? We applied our methodology across three datasets: the GloSAT and ICDAR-2019 datasets based on historical documents, and the predominantly modern documents PubTabNet dataset. Our results indicate that semi-supervised learning substantially increases TSR accuracy and decreases dependency on extensive labelled datasets, providing a robust solution for large-scale digitization initiatives and contributing to the preservation and improved accessibility of historical data. All code from this paper is freely available on GitHub.
This paper discusses the ways in which complexity and degrees of autonomy in AI-based medical devices (AIaMD) may challenge the safety and performance of software for EU regulatory alignment and responsible AI regarding AI-induced harms. It examines the EU Commission proposals for an AI Liability Directive and a revised Product Liability Directive to identify two research challenges. These challenges relate to identifications of “defects” arising from algorithmic change and degrees of human oversight. Some suggestions will be made in how they can be addressed through causal modelling, counterfactuals, and responsibility reasoning.
Prompt-based models have gathered a lot of attention from researchers due to their remarkable advancements in the fields of zero-shot and few-shot learning. Developing an effective prompt template plays a critical role. However, prior studies have mainly focused on prompt vocabulary searching or embedding initialization within a predefined template with the prompt position fixed. In this empirical study, we conduct the most comprehensive analysis to date of prompt position for diverse Natural Language Processing (NLP) tasks. Our findings quantify the substantial impact prompt position has on model performance. We observe that the prompt positions used in prior studies are often sub-optimal, and this observation is consistent even in widely used instruction-tuned models. These findings suggest prompt position optimisation as a valuable research direction to augment prompt engineering methodologies and prompt position-aware instruction tuning as a potential way to build more robust models in the future.
In this paper, we demonstrate how social media technologies can co-produce data-related harms unless preventative measures are instituted. To this end, we draw on a passive ethnography of a public Facebook group in the UK practicing sharenting which occurs when parents and guardians post sensitive and identifying information about children in their care on social media. Theoretically, we draw on the ‘harm translation’ concept from digital criminology and the ‘seductions of crime’ perspective from cultural criminology. Further we analyse documents on the operations of Facebook's content filtering algorithms published by Meta (Facebook's parent company). With insights from these sources, we demonstrate how platform technologies go beyond facilitation to the inadvertent co-production of harm via embedded mediative properties that shape user perception and action. We show that, in the specific context of sharenting, the properties invite rather than simply facilitate the practice and can also invite subsequent misuses of child-centric data. Through our analysis of these dynamics, we set out an empirical basis for challenging reductive depictions of social media technologies as solely facilitative of human action including harmful conduct. We also outline our vision to integrate insights from the analysis into a new sociotechnical harm prevention framework informed by Natural Language Processing approaches.
Dataset and codes for the Multimodal USElecDeb60To16 dataset, released in the paper "Augmenting pre-trained language models with audio feature embedding for argumentation mining in political debates", published at the Findings of the 17th conference on European chapter of the Association for Computational Linguistics (EACL) in 2023. This is the version that contains the audio features and the artificial voices. If you want a (lighter) version without those large files, check out the first release v.1.0.0 with DOI 10.5281/zenodo.7628465.
The integration of multimodality in natural language processing (NLP) tasks seeks to exploit the complementary information contained in two or more modalities, such as text, audio and video. This paper investigates the integration of often under-researched audio features with text, using the task of argumentation mining (AM) as a case study. We take a previously reported dataset and present an audio-enhanced version (the Multimodal USElecDeb60To16 dataset). We report the performance of two text models based on BERT and GloVe embeddings, one audio model (based on CNN and Bi-LSTM) and multimodal combinations, on a dataset of 28,850 utterances. The results show that multimodal models do not outperform text-based models when using the full dataset. However, we show that audio features add value in fully supervised scenarios with limited data. We find that when data is scarce (e.g. with 10% of the original dataset) multimodal models yield improved performance, whereas text models based on BERT considerably decrease performance. Finally, we conduct a study with artificially generated voices and an ablation study to investigate the importance of different audio features in the audio models.
This work describes the classification system proposed for the Computational Linguistics and Clinical Psychology (CLPsych) Shared Task 2022. We propose the use of multitask learning approach with bidirectional long-short term memory (Bi-LSTM) model for predicting changes in user’s mood and their suicidal risk level. The two classification tasks have been solved independently or in an augmented way previously, where the output of one task is leveraged for learning another task, however this work proposes an ‘all-in-one’ framework that jointly learns the related mental health tasks. The experimental results suggest that the proposed multi-task framework outperforms the remaining single-task frameworks submitted to the challenge and evaluated via timeline based and coverage based performance metrics shared by the organisers. We also assess the potential of using various types of feature embedding schemes that could prove useful in initialising the Bi-LSTM model for better multitask learning in the mental health domain.
Digital platforms for mental health and wellbeing purposes have become increasingly common to help users exhibiting risk behaviours (e.g. self-harming, eating-related disorders) across all ages, opening new frontiers in supporting vulnerable users. This study stems from a larger project, which explores how responsible AI solutions can up-scale existing manual moderation approaches and better target interventions for young people who ask for help or engage in risk behaviours online. This research aims to better understand the challenges and needs of moderators and digital counsellors, i.e. the ‘behind the scenes’. Through this case study, the authors intend to contribute to the development of responsible AI tools that are fit for purpose and better understand the challenges. The key focus lies on Kooth.com, the UK’s leading free online confidential service offering counselling and emotional wellbeing support to young people in the UK through its online web-based and pseudo-anonymous digital platform.
intention to accept vulnerability based upon positive expectations of the intentions or behavior of another.” Trust is an attitude that an agent will behave as expected and can be relied upon to reach its goal. Trust breaks down after an error or a misunderstanding between the agent and the trusting individual. The psychological state of trust in AI is an emergent property of a complex system, usually involving many cycles of design, training, deployment, measurement of performance, regulation, redesign, and retraining. Trust matters, especially in critical sectors such as healthcare, defense, and security, where duty of care is foremost. Trustworthiness must be planned, rather than an afterthought. We can trust in AI, such as when a doctor uses algorithms to screen medical images.20 We can also trust with AI, such as when journalists reference a social network algorithm to analyze sources of a news story.37 Growing adoption of AI into institutional systems relies on citizens to trust in these systems and have confidence in the way these systems are designed and regulated. Regional approaches for managing trust in AI have recently emerged, leading to different regulatory regimes in the U.S., the European region, and China. We review these regulatory divergences. Within the European region, research programs are examining how trust impacts user acceptance of AI. Examples include the UKRI Trustworthy Autonomous Systems Hub,a the French Confiance. ai project,b and the German AI Breakthrough Hub.c Europe appears to be developing a “third way,” alongside the U.S. and China.19 Healthcare contains many examples of AI applications, including online harm risk identification,24 mental health behavior classification,29 and
Understanding and extracting tables from documents is a research problem that has been studied for decades. Table structure recognition is the labelling of components within a detected table, which can be detected automatically or manually provided. This paper presents the GloSAT historical measurement table dataset designed to train table structure recognition models for use in downstream historical data rescue applications. The dataset contains 500 scanned and manually annotated images of pages from meteorological measurement logbooks. We enhance standard full table and individual cell annotations by adding additional annotations for headings, headers, and table bodies. We also provide annotations for coarse segmentation cells consisting of multiple data cells logically grouped by ruling lines of ink or whitespace in the table, which often represent data cells that are semantically grouped. Our dataset annotations are provided in VOC2007 and ICDAR-2019 Competition on Table Detection and Recognition (cTDaR-19) XML formats, and our dataset can easily be aggregated with the cTDaR-19 dataset. We report results running a series of benchmark algorithms on our new dataset, concluding that post-processing is very important for performance, and that page style is not as significant a feature as table type on model performance.
In this paper, we present a new application-focused benchmark dataset and results from a set of baseline Natural Language Processing and Machine Learning models for prediction of match outcomes for games of football (soccer). By doing so we give a baseline for the prediction accuracy that can be achieved exploiting both statistical match data and contextual articles from human sports journalists. Our dataset is focuses on a representative time-period over 6 seasons of the English Premier League, and includes newspaper match previews from The Guardian. The models presented in this paper achieve an accuracy of 63.18% showing a 6.9% boost on the traditional statistical methods.
Siegfried Benkner合作论文数Head of Institute4