Drawing on insights from the transdisciplinary project 'SENSE: Sensory Explorations of Nature in School Environments', this paper articulates a novel approach that addresses current calls for meaningful participation of children in citizen science activities in schools. We scaffolded more typical data collection activities within diverse digital and natural haptic experiences aimed at developing observational skills through arts and science-based methods, such as clay modelling, digital haptic tree identification and textural mapping exercises of the school grounds. Data were collected in three primary schools in Scotland, through audio and video-recording, and observation notes in the field. Findings showed how incorporating touch focuses attention differently to vision, leading to different scientific questions and inquiries. In effect, touch experiences may serve to balance the aims of citizen science beyond the intentional identification and enumeration of species towards the more taxing, epistemic and ethical questions of 'who decides what matters in nature observation' and 'for whom' is the learning, as students are invited to participate and contribute on their own terms. Implications for this form of citizen science to open significantly new directions for children's participation, and make its way into existing teaching practices in schools, are discussed.
We are developing techniques to generate summary descriptions of sets of objects. In this paper, we present and evaluate a rule-based NLG technique for summarising sets of bibliographical references in academic papers. This extends our previous work on summarising sets of consumer products and shows how our model generalises across these two very different domains.
The iSpot citizen science platform*1 allows anyone anywhere to upload images of biodiversity and its community of users helps to identify observations. Key elements of any such system are the species dictionaries that tie together all observations of similar species, allow further information about the taxa to be shown to users and ensure that data collected can be passed on to other recording schemes and global databases. iSpot was launched in 2009 with the UK Species Inventory (UKSI) database as its UK dictionary and the Catalogue of Life (CoL) for its global dictionary. Later, the South African National Biodiversity Institute (SANBI) dictionary was added (Silvertown et al. 2015). Taxonomy has changed in many species groups since then and missing groups, particularly from the CoL database, are now included. Linking directly to a live version of CoL or UKSI as a web service or similar, had been considered multiple times over iSpot's development timeline but not implemented due to the complexity. Updating species dictionaries can be a very difficult task. First, these national or international dictionaries are comprised of data from a large number of organisations and individual taxonomists e.g., CoL has global contributions combining taxonomic databases in a variety of formats covering various parts of the taxonomic tree. Second, data comes in at various times and there may be differences over the ‘current’ name and possible synonyms. Third, are issues splitting and aggregating the taxa. Finally, are the number of levels in the taxonomy e.g., are subfamily, subphylum, suborder, used or not. This may differ in data from different organisations, different parts of the tree and may change between overall dictionary versions. In recent years, the use of DNA methods has revolutionised taxonomy in certain groups i.e., fungi; this has not only affected species-level identification but also parts of the higher levels in the tree. iSpot not only shows the species identification (ID), but also the full taxonomic tree with a built-in species browser. This helps with initial identification and education by showing similar taxa and how an observation fits in with the overall tree of life. In theory, all that is needed to update the dictionary is to match the taxon codes from the old to the new dictionary, and for the taxon codes that do not match, to attempt to match on taxon name. However, in practice other issues arose e.g., CoL changing all its taxon codes so additional layers of matching are required. Also not just the current name but all the synonyms and homonyms have to be dealt with. The trial included an initial large matchup, done via ChecklistBank, assisted by staff at the Global Biodiversity Information Facility (GBIF) and CoL. In addition to updating the tree itself, for the system to work, each taxon had to be allocated to an iSpot top level group: amphibians and reptiles, birds, fish, fungi and lichens, invertebrates, mammals, plants and other. For the CoL dictionary, this was a manual process for all five million taxa and synonyms. To facilitate this, the data were written with the full taxonomic hierarchy. This could then be sorted and taxonomic groups added relatively easily. However, there were still taxa that did not match. This is where the iSpot citizen science community provided help. iSpot volunteers attempted to manually work out matches, for example by checking the spelling. There were some spelling errors that led to lack of match, but other issues were more difficult. In some cases, there were pseudonyms that did not appear in the current dictionary or names that seemed to have vanished completely, but the observation could be redetermined. A common problem was where whole species groups were not present in the new dictionary, which was fed back so that the dictionary could be corrected. Examples of common groups where some or all species were missing included: ladybird beetles, gall wasps, cockroaches, dragonflies, shield bugs and some moths. The iSpot community involvement demonstrated citizen scientists helping to improve the overall global species dictionary, see example in Table 1. For each of the 991 instances that volunteer found, notes were provided with references either of the correct taxon to match to or an appropriate higher level group, if the taxon were missing. They also produced notes on, e.g., serious pest species present in the SANBI dictionary but missing from CoL. Some of the volunteers thought it was too difficult or too much responsibility and dropped out. The results were checked by the iSpot Curator, who also dealt with groups of organisms with no volunteers. Once the dictionary is updated, information needs to be propagated to all parts of the iSpot platform. During this process some of the manually entered names for taxa, can be automatically linked to the new dictionary.
Despite the impressive capability of large language models (LLMs), knowing when to trust their generations remains an open challenge. The recent literature on uncertainty quantification of natural language generation (NLG) utilizes a conventional natural language inference (NLI) classifier to measure the semantic dispersion of LLMs responses. These studies employ logits of NLI classifier for semantic clustering to estimate uncertainty. However, logits represent the probability of the predicted class and barely contain feature information for potential clustering. Alternatively, CLIP (Contrastive Language-Image Pre-training) performs impressively in extracting image-text pair features and measuring their similarity. To extend its usability, we propose Contrastive Semantic Similarity, the CLIP-based feature extraction module to obtain similarity features for measuring uncertainty for text pairs. We apply this method to selective NLG, which detects and rejects unreliable generations for better trustworthiness of LLMs. We conduct extensive experiments with three LLMs on several benchmark question-answering datasets with comprehensive evaluation metrics. Results show that our proposed method performs better in estimating reliable responses of LLMs than comparable baselines. The code are available at https://github.com/AoShuang92/css_uq_llms.
We investigate the potential of a new citizen science paradigm that facilitates collaborative learning between humans and artificial intelligence (AI). Recognising the potential of AI to support and empower rather than replace human participation, we explore the integration of image recognition as a ‘dialogic AI partner’ in citizen science (CS) projects, interacting with participants in real time. We study this in the context of a biodiversity monitoring project that relies on volunteers to identify biological species from images taken in the wild. Guided by the idea of Bakhtin’s dialogism and Bayesian inference principles, we developed a web interface that integrated an image recognition model, fine-tuned for classifying 22 UK bumblebee species, into an interactive interface based on visual feature keys to enable real-time dialogue between humans and AI. We report a significant improvement in identification accuracy for both humans and AI when they engage in such dialogue and retain the ability to reach independent conclusions rather than achieve consensus. Given the inherent need for convergence in decision-making within scientific processes such as species identification tasks, we augmented the dialogic process with a Bayesian model that unifies potentially divergent human and AI perspectives post collaboration to achieve a more accurate consensus decision than that achieved by either AI or citizens. Our work provides new understandings around the design of a dialogic space for CS practice that effectively builds on the complementary strengths of human and AI visual recognition approaches.
Over the last decades, the interdisciplinary field of the affective sciences has seen proliferation rather than integration of theoretical perspectives. This is due to differences in metaphysical and mechanistic assumptions about human affective phenomena (what they are and how they work) which, shaped by academic motivations and values, have determined the affective constructs and operationalizations. An assumption on the purpose of affective phenomenacan be used as a teleological principle to guide the construction of a common set of metaphysical and mechanistic assumptions—a framework for human affective research. In this capstone paper for the special issue “Towards an Integrated Understanding of the Human Affectome”, we gather the tiered purpose of human affective phenomena to synthesize assumptions that account for human affective phenomenacollectively. This teleologically-grounded framework offers a principled agenda and launchpad for both organizing existing perspectives and generating new ones. Ultimately, we hope Human Affectome brings us a step closer to not only an integrated understanding of human affective phenomena, but an integrated field for affective research.
Failure detection (FD) in AI systems is a crucial safeguard for the deployment for safety-critical tasks. The common evaluation method of FD performance is the Risk-coverage (RC) curve, which reveals the trade-off between the data coverage rate and the performance on accepted data. One common way to quantify the RC curve by calculating the area under the RC curve. However, this metric does not inform on how suited any method is for FD, or what the optimal coverage rate should be. As FD aims to achieve higher performance with fewer data discarded, evaluating with partial coverage excluding the most uncertain samples is more intuitive and meaningful than full coverage. In addition, there is an optimal point in the coverage where the model could achieve ideal performance theoretically. We propose the Excess Area Under the Optimal RC Curve (E-AUoptRC), with the area in coverage from the optimal point to the full coverage. Further, the model performance at this optimal point can represent both model learning ability and calibration. We propose it as the Trust Index (TI), a complementary evaluation metric to the overall model accuracy. We report extensive experiments on three benchmark image datasets with ten variants of transformer and CNN models. Our results show that our proposed methods can better reflect the model trustworthiness than existing evaluation metrics. We further observe that the model with high overall accuracy does not always yield the high TI, which indicates the necessity of the proposed Trust Index as a complementary metric to the model overall accuracy. The code are available at \url{https://github.com/AoShuang92/optimal_risk}.
Proper confidence calibration of deep neural networks is essential for reliable predictions in safety-critical tasks. Miscalibration can lead to model over-confidence and/or under-confidence; i.e., the model's confidence in its prediction can be greater or less than the model's accuracy. Recent studies have highlighted the over-confidence issue by introducing calibration techniques and demonstrated success on various tasks. However, miscalibration through under-confidence has not yet to receive much attention. In this paper, we address the necessity of paying attention to the under-confidence issue. We first introduce a novel metric, a miscalibration score, to identify the overall and class-wise calibration status, including being over or under-confident. Our proposed metric reveals the pitfalls of existing calibration techniques, where they often overly calibrate the model and worsen under-confident predictions. Then we utilize the class-wise miscalibration score as a proxy to design a calibration technique that can tackle both over and under-confidence. We report extensive experiments that show our proposed methods substantially outperforming existing calibration techniques. We also validate our proposed calibration technique on an automatic failure detection task with a risk-coverage curve, reporting that our methods improve failure detection as well as trustworthiness of the model. The code are available at \url{https://github.com/AoShuang92/miscalibration_TS}.
As the citizen science (CS) community flourishes, there is an opportunity to reflect on how practitioners can widen participation and work with participants as co-researchers to investigate and take action around global challenges. Through the lens of one CS case study, the X-Polli:Nation project, we report on how technologists, ecologists, and education specialists repurposed older projects by cross-pollinating ideas with children and teachers in the UK and in Italy to create Artificial Intelligence–enhanced tools appropriate for teaching sustainability in schools. Taking part in an actionable CS cycle, children learn about pollinating insects, record scientific data, create flowering habitats, and communicate their importance. Through this process, X-Polli:Nation demonstrates relevance across a number of Sustainable Development Goals (e.g., SDG 4, Quality Education; SDG 10, Reducing Inequality; and SDG 15, Life on Land), and applies the underlying SDG principle “leave no one behind.” We go on to investigate if, and how, young people would like to deepen their engagement with the SDGs, and we report that taking action and communicating the importance of the SDGs were of paramount interest. The challenge of building sustainability into an already crowded curriculum can be alleviated by understanding its value, considering the audience, and adapting to new contexts. The considerable benefits include raising awareness about global sustainability issues and giving children the confidence to become passionate environmental stewards, all the while extending the life of older projects and thus making CS methods sustainable too.
Despite the great success of state-of-the-art deep neural networks, several studies have reported models to be over-confident in predictions, indicating miscalibration. Label Smoothing has been proposed as a solution to the over-confidence problem and works by softening hard targets during training, typically by distributing part of the probability mass from a 'one-hot' label uniformly to all other labels. However, neither model nor human confidence in a label are likely to be uniformly distributed in this manner, with some labels more likely to be confused than others. In this paper we integrate notions of model confidence and human confidence with label smoothing, respectively Model Confidence LS and Human Confidence LS, to achieve better model calibration and generalization. To enhance model generalization, we show how our model and human confidence scores can be successfully applied to curriculum learning, a training strategy inspired by learning of 'easier to harder' tasks. A higher model or human confidence score indicates a more recognisable and therefore easier sample, and can therefore be used as a scoring function to rank samples in curriculum learning. We evaluate our proposed methods with four state-of-the-art architectures for image and text classification task, using datasets with multi-rater label annotations by humans. We report that integrating model or human confidence information in label smoothing and curriculum learning improves both model performance and model calibration. The code are available at https://github.com/AoShuang92/Confidence Calibration CL.
A number of initiatives invite members of the public to perform online classification tasks such as identifying objects in images. These tasks are crucial to numerous large-scale Citizen Science projects in different disciplines, with volunteers using their knowledge and online support tools to, for example, identify species of wildlife or classify galaxies by their shapes. However, for complex classification tasks, such as this case study on identifying species of bumblebee, reaching an agreement between volunteers - or even between experts~-~may require consensus-building processes. Collaboration and teamwork approaches to problem solving and decision-making have been widely documented to improve both task performance and user learning in the real world. Most of these processes and projects are mediated online through feedback delivered in an asynchronous manner, and this article thus addresses a central research question: How do participants involved in species identification tasks respond to different forms of feedback provided in online collaboration, designed to support peer-learning and improve task performance? We tested four different approaches to feedback within a collaboration task, where participants reviewed their previously annotated data based on information curated from their peers on a long running online citizen science initiative. The selected interfaces have a strong foundation in social science and psychology literature and can be applied to citizen science practices as well as other online communities. Results showed that while all four approaches increased accuracy, there were differences based on the types of consensus that existed before collaboration. Such differences highlight the usefulness of different forms of feedback during collaboration for increasing data accuracy of identification and furthering users' expertise on identification tasks. We found that anonymised and goal-directed free text comments posted on social learning interfaces were most effective in improving data accuracy as well as creating opportunities for peer-learning, particularly where the species identification task was more difficult. This study has significant implications for extending the practice of citizen science across formal and informal learning environments and reaching out to a variety of users.
This paper tackles the problem of automatically labelling sentiment-bearing topics with descriptive sentence labels. We propose two approaches to the problem, one extractive and the other abstractive. Both approaches rely on a novel mechanism to automatically learn the relevance of each sentence in a corpus to sentiment-bearing topics extracted from that corpus. The extractive approach uses a sentence ranking algorithm for label selection which for the first time jointly optimises topic–sentence relevance as well as aspect–sentiment co-coverage. The abstractive approach instead addresses aspect–sentiment co-coverage by using sentence fusion to generate a sentential label that includes relevant content from multiple sentences. To our knowledge, we are the first to study the problem of labelling sentiment-bearing topics. Our experimental results on three real-world datasets show that both the extractive and abstractive approaches outperform four strong baselines in terms of facilitating topic understanding and interpretation. In addition, when comparing extractive and abstractive labels, our evaluation shows that our best performing abstractive method is able to provide more topic information coverage in fewer words, at the cost of generating less grammatical labels than the extractive method. We conclude that abstractive methods can effectively synthesise the rich information contained in sentiment-bearing topics.
We introduce the task of historical text summarisation, where documents in historical forms of a language are summarised in the corresponding modern language. This is a fundamentally important routine to historians and digital humanities researchers but has never been automated. We compile a high-quality gold-standard text summarisation dataset, which consists of historical German and Chinese news from hundreds of years ago summarised in modern German or Chinese. Based on cross-lingual transfer learning techniques, we propose a summarisation model that can be trained even with no cross-lingual (historical to modern) parallel data, and further benchmark it against state-of-the-art algorithms. We report automatic and human evaluations that distinguish the historic to modern language summarisation task from standard cross-lingual summarisation (i.e., modern to modern language), highlight the distinctness and value of our dataset, and demonstrate that our transfer learning approach outperforms standard cross-lingual benchmarks on this task.
Widespread concern over declines in pollinating insects has led to numerous recommendations of which “pollinator-friendly” plants to grow and help turn urban environments into valuable habitat for such important wildlife. Whilst communicated widely by organisations and readily taken up by gardeners, the provenance, accuracy, specificity and timeliness of such recommendations remain unclear. Here we use data (6429 records) gathered through a UK-wide citizen science programme (BeeWatch) to determine food plant use by the nations’ bumblebee species, and show that much of the plant use recorded does not reflect practitioner recommendations: correlation between the practitioners’ bumblebee-friendly plant list (376 plants compiled from 14 different sources) and BeeWatch records (334 plants) was low (r = 0.57), and only marginally higher than the correlation between BeeWatch records and the practitioners’ pollinator-friendly plant list (465 plants from 9 different sources; r = 0.52). We found pollinator-friendly plant lists to lack independence (correlation between practitioners’ bumblebee-friendly and pollinator-friendly lists: r = 0.75), appropriateness and precision, thus failing to recognise the non-binary nature of food-plant preference (bumblebees used many plants, but only in small quantities, e.g. lavender—the most popular plant in the BeeWatch database—constituted, at most, only 11% of records for any one bumblebee species) and stark differences therein among species and pollinator groups. We call for the provision and use of up-to-date dynamic planting recommendations driven by live (citizen science) data, with the possibility to specify pollinator species or group, to powerfully support transformative personal learning journeys and pollinator-friendly management of garden spaces.
Identifying private gardens in the U.K. as key sites of environmental engagement, we look at how a longer-term online citizen science programme facilitated the development of new and personal attachments of nature. These were visible through new or renewed interest in wildlife-friendly gardening practices and attitudinal shifts in a large proportion of its participants. Qualitative and quantitative data, collected via interviews, focus groups, surveys and logging of user behaviours, revealed that cultivating a fascination with species identification was key to both ‘helping nature’ and wider learning, with the programme creating a space where scientific and non-scientific knowledge could co-exist and reinforce one another.
The system transforms raw telemetric data into engaging and informative blog texts readily understood by all.