This study demonstrates the first realization of wireless strain, temperature and crack growth sensing within 3D-printed metallic structures using standard electromagnetic inspection hardware. This establishes a path toward need-based maintenance for parts operating in harsh environments driven by accurate, real-time damage assessments instead of relying on regularly scheduled maintenance teardowns. To this end, we encapsulate and embed magnetoelastic and thermomagnetic materials during additive manufacturing. Mechanical and thermal stimuli affect the magnetic permeability of the embedded materials, which modulates the flux through a coil placed on or near the part's surface. We demonstrate strain sensing accurate to +/-27x10-6 and temperature sensing accurate to +/-0.75 oC. We highlight these sensors' capabilities by detecting the onset of plasticity and fatigue-driven crack growth thousands of cycles before critical failure.
This paper explores new methods for locating the sources used to write a text, by fine-tuning a variety of language models to rerank candidate sources. After retrieving candidates sources using a baseline BM25 retrieval model, a variety of reranking methods are tested to see how effective they are at the task of source attribution. We conduct experiments on two datasets, English Wikipedia and medieval Arabic historical writing, and employ a variety of retrieval and generation based reranking models. In particular, we seek to understand how the degree of supervision required affects the performance of various reranking models. We find that semisupervised methods can be nearly as effective as fully supervised methods while avoiding potentially costly span-level annotation of the target and source documents.
The Greek and Latin classics, like many other ancient texts, have been widely translated into a variety of languages over the past two millennia. Although many digital editions and libraries contain one or two translations for a given text, about one hundred translations of the Iliad and twenty of Herodotus, for example, exist in English alone. Aligning the corpus of classical texts and translations at the sentence and word level would provide a valuable resource for studying translation theory, digital humanities, and natural language processing (NLP). Precise and faithful sentence alignment via computational methods, however, remains a challenging problem. Current alignment methods tend to have poor coverage and recall since their primary aim is to extract single sentence pairs for training machine translation systems. This paper evaluates and examines the limits of such state-of-the-art models for cross-language sentence embedding and alignment of ancient Greek and Latin texts with translations into English, French, German, and Persian. We release evaluation data for Plato’s Crito , manually annotated at the word and sentence level, and larger test datasets based on coarser structural metadata for Thucydides (Greek) and Lucretius (Latin). Testing LASER and LaBSE for sentence embedding and nearest-neighbor retrieval and Vecalign for sentence alignment, we found best results using LaBSE-Vecalign. LaBSE worked surprisingly well on ancient Greek, most probably because it had been merged with modern Greek data in its training. Both LASER-Vecalign and LaBSE-Vecalign did best when there were many ground-truth one-to-one alignments between source and target sentences, and when the order of sentences in the source was preserved in the translation. However, these conditions are often not present in the kinds of literary and free translation we wish to study, nor in editions with multiple translations, extensive commentary, or other paratext. We perform book-level and chapter-level error analysis to inform the development of a software pipeline that can be deployed on the vast corpus of translations of ancient texts.
Have digital tools and methods accelerated the rate of scholarly production over the last 20 years? If so, has this acceleration been beneficial for scholarship? This article considers examples of accelerated historical scholarship as well as calls for a “slow history.” Through an analysis of the author's own experiences with the digital humanities, it examines the advantages and disadvantages of digital technologies in the field of history. It concludes that online resources and digital technologies have expanded the archive for the historian and created new ways to reach other specialists and the general public. Nevertheless, historical scholarship must still rely on carefully crafted, well-argued prose whose production cannot be accelerated by new digital technologies, although recent developments in the field of artificial intelligence may ultimately challenge this situation. In recent decades, the field (or, at times, discipline) of digital humanities (DH) has revolutionized the scholarly profession and beyond—and with good reason. Seen at times as a democratizing force, DH has led to the creation of an increasing number of open- access databases and scholarly publications, the launching of massive archival digitization initiatives, and the development of numerous digital tools that help streamline the work of the academic researcher, student, and educator. In many ways, then, its benefits are manifest. Yet, recent years have also begun to reveal numerous problems that could influence various aspects of our trade as well as what—and how—information will be available in the future. This article discusses some of the advantages and disadvantages of DH and invites the reader to reflect on what we can do to help mitigate these problems. Exciting new modes of digital scholarship have emerged in recent years, providing us with expanded windows onto the past. This process has been accelerated by somewhat democratized ways of digitizing and analyzing source material. A main issue of contemporary knowledge production using digitized sources is how power can so easily be reinscribed into access to archives. The choice to digitize collections, even the existence of collections themselves, creates a great opportunity for research but also runs the risk of reinforcing the privilege and worldviews that have shaped and continue to shape the very processes of digitization and digitalization. Drawing on examples of Western and non-Western digital scholarship, this article argues that, although the digital facilitates greater public knowledge of collections, when it comes to decolonizing our research subjects, it also introduces significant layers of complexity. This article advances an analysis of the development and state of critical digital humanities. It posits two modalities for this approach to digital humanities (DH). The first is a modality of inward-looking, functional self-critique that comprises a rethinking of computational genesis stories, logics and methods, institutions and infrastructures, and digital capitalism, and the second is an outward-looking critique best understood as a form of situated sociopolitical engagement that embraces epistemic and social justice projects that are decolonial, anti-racist, inclusive, collaborative, and multilingual. Through these analyses, the article offers a vision of critical digital humanities in its mission to critique the ideologies, social inequities, and epistemological hierarchies that are built into technological products and computational logics and that are concomitantly fostered by knowledge- creation industries of universities, corporations, governments, and the GLAM[R] sector. In this way, the article shows how critical digital humanities helps us to envision the role that DH can play in processes of recovery, reparations, emancipation, and community-building. Drawing upon over 20 years as Editor-in-Chief of H-France, I argue that the scholarly profession, established in Cold War era, pre-digital institutions, has only begun to adapt to the transformations introduced by the global digital humanities. A generational shift is currently underway as younger scholars more natively adept with digital technologies use their skills and forms of new media to press for changes in hiring and tenure practices, to demand greater progress on diversity, equity, and inclusion (DEI) issues, and to insist that the academy confront the collapse of academic positions in the humanities and provide training for and recognition of alternative career paths. I call upon professional organizations to undertake difficult conversations and take leadership in reshaping professional organizations for a post–Cold War, digital age, especially in terms of funding priorities. Scholarly organizations will best gain influence through collaboration.
Brain-computer interfaces (BCI) are an important mode of alternative and augmentative communication for many people. Unlike keyboards, many BCI systems do not display even the 26 letters of English at one time, let alone all the symbols in more complex systems. Using language models to make character-level predictions, therefore, can greatly speed up BCI typing (Ghosh and Kristensson, 2017). While most existing BCI systems employ character n-gram models or no LM at all, this paper adapts several wordpiece-level Transformer LMs to make character predictions and evaluates them on typing tasks. GPT-2 fares best on clean text, but different LMs react differently to noisy histories. We further analyze the effect of character positions in a word and context lengths.
A range of different courts and tribunals is involved in housing work, and an of understanding this provides the foundation to an integrated PRS enforcement approach, taking into account civil and criminal procedures and possible remedies. All investigations should be thorough and ensure that good evidence is collected that proves a particular set of offences in accordance with the relevant legislation. LAs need clear approaches to civil or criminal penalties which may have consequences for subsequent decisions and actions, including any publicity pertaining to the case. All decision-making should sit within the LA framework of law enforcement.
Most NLP approaches to entity linking and coreference resolution focus on retrieving similar mentions using sparse or dense text representations. The common "Wikification" task, for instance, retrieves candidate Wikipedia articles for each entity mention. For many domains, such as bibliographic citations, authority lists with extensive textual descriptions for each entity are lacking and ambiguous named entities mostly occur in the context of other named entities. Unlike prior work, therefore, we seek to leverage the information that can be gained from looking at association networks of individuals derived from textual evidence in order to disambiguate names. We combine BERT-based mention representations with a variety of graph induction strategies and experiment with supervised and unsupervised cluster inference methods. We experiment with data consisting of lists of names from two domains: bibliographic citations from CrossRef and chains of transmission (isnads) from classical Arabic histories. We find that in-domain language model pretraining can significantly improve mention representations, especially for larger corpora, and that the availability of bibliographic information, such as publication venue or title, can also increase performance on this task. We also present a novel supervised cluster inference model which gives competitive performance for little computational effort, making it ideal for situations where individuals must be identified without relying on an exhaustive authority list.
Imbalanced classification problems are extremely common in natural language processing and are solved using a variety of resampling and filtering techniques, which often involve making decisions on how to select training data or decide which test examples should be labeled by the model. We examine the tradeoffs in model performance involved in choices of training sample and filter training and test data in heavily imbalanced token classification task and examine the relationship between the magnitude of these tradeoffs and the base rate of the phenomenon of interest. In experiments on sequence tagging to detect rare phenomena in English and Arabic texts, we find that different methods of selecting training data bring tradeoffs in effectiveness and efficiency. We also see that in highly imbalanced cases, filtering test data using first-pass retrieval models is as important for model performance as selecting training data. The base rate of a rare positive class has a clear effect on the magnitude of the changes in performance caused by the selection of training or test data. As the base rate increases, the differences brought about by those choices decreases.
This paper describes our experiences in building a live collaborative programming environment on top of the JavaScript version of the Croquet shared experience platform. Croquet provides a clean substrate for building real-time collaborative applications. We created an application framework that supports live programming, and used that framework to build the Greenlight collaborative application, then in turn, modified it to do live programming experiments. The environment allows multiple users to modify the running application from within, with changes taking effect immediately. The experiment was inspired by earlier work including Douglas Engelbart’s oN-Line System (NLS) and the Kansas system in Self. Analogically, the system is like the Smalltalk environment made collaborative. In this paper we explain the Croquet architecture, its library and framework, and the Greenlight application used to make the live programming environment. The standard version of Greenlight is available at https://croquet.io/greenlight, and the modified demo system is available at https://croquet.io/scripting.
Growing concern with online misinformation has encouraged NLP research on fact verification. Since writers often base their assertions on structured data, we focus here on verifying textual statements given evidence in tables. Starting from the Table Parsing (TAPAS) model developed for question answering (Herzig et al., 2020), we find that modeling table structure improves a language model pre-trained on unstructured text. Pre-training language models on English Wikipedia table data further improves performance. Pre-training on a question answering task with column-level cell rank information achieves the best performance. With improved pre-training and cell embeddings, this approach outperforms the state-of-the-art Numerically-aware Graph Neural Network table fact verification model (GNN-TabFact), increasing statement classification accuracy from 72.2% to 73.9% even without modeling numerical information. Incorporating numerical information with cell rankings and pre-training on a question-answering task increases accuracy to 76%. We further analyze accuracy on statements implicating single rows or multiple rows and columns of tables, on different numerical reasoning subtasks, and on generalizing to detecting errors in statements derived from the ToTTo table-to-text generation dataset.
We explore the task of quotability identification, in which, given a document, we aim to identify which of its passages are the most quotable, i.e. the most likely to be directly quoted by later derived documents. We approach quotability identification as a passage ranking problem and evaluate how well both feature-based and BERT-based (Devlin et al., 2019) models rank the passages in a given document by their predicted quotability. We explore this problem through evaluations on five datasets that span multiple languages (English, Latin) and genres of literature (e.g. poetry, plays, novels) and whose corresponding derived documents are of multiple types (news, journal articles). Our experiments confirm the relatively strong performance of BERT-based models on this task, with the best model, a RoBERTA sequential sentence tagger, achieving an average rho of 0.35 and NDCG@1, 5, 50 of 0.26, 0.31 and 0.40, respectively, across all five datasets.
Code-switching has long interested linguists, with computational work in particular focusing on speech and social media data (Sitaram et al., 2019). This paper contrasts these informal instances of code-switching to its appearance in more formal registers, by examining the mixture of languages in the Deutsches Textarchiv (DTA), a corpus of 1406 primarily German books from the 17th to 19th centuries. We automatically annotate and manually inspect spans of six embedded languages (Latin, French, English, Italian, Spanish, and Greek) in the corpus. We quantitatively analyze the differences between code-switching patterns in these books and those in more typically studied speech and social media corpora. Furthermore, we address the practical task of predicting code-switching from features of the matrix language alone in the DTA corpus. Such classifiers can help reduce errors when optical character recognition or speech transcription is applied to a large corpus with rare embedded languages.
We present our work on automatically detecting isnads, the chains of authorities for a re-port that serve as citations in hadith and other classical Arabic texts. We experiment with both sequence labeling methods for identifying isnads in a single pass and a hybrid “retrieve-and-tag” approach, in which a retrieval model first identifies portions of the text that are likely to contain start points for isnads, then a sequence labeling model identifies the exact starting locations within these much smaller retrieved text chunks. We find that the usefulness of full-document sequence to sequence models is limited due to memory limitations and the ineffectiveness of such models at modeling very long documents. We conclude by sketching future improvements on the tagging task and more in-depth analysis of the people and relationships involved in the social network that influenced the evolution of the written tradition over time.
Language models have broad adoption in predictive typing tasks. When the typing history contains numerous errors, as in open-vocabulary predictive typing with brain-computer interface (BCI) systems, we observe significant performance degradation in both n-gram and recurrent neural network language models trained on clean text. In evaluations of ranking character predictions, training recurrent LMs on noisy text makes them much more robust to noisy histories, even when the error model is misspecified. We also propose an effective strategy for combining evidence from multiple ambiguous histories of BCI electroencephalogram measurements.
The open-closed boundary (OCB) defines a region of significant transformation in Earth's protective magnetic shield. Principle among these changes is the transition of magnetic field lines from having two foot points, one in each hemisphere, to one foot point at Earth, the other mapping to the solar wind. Charged particles in the solar wind are able to follow these open field lines into Earth's upper atmosphere. The OCB also defines the polar cap boundary. Being able to identify and track the OCB allows study of several components of the geomagnetic system. Among them are the electrodynamics of the geomagnetic field and the reconnection balance between the dayside and nightside of the geomagnetic field. Furthermore, the OCB can provide insights into the precipitation of energetic protons into the ionosphere. Using the Tsyganenko model of the geomagnetic field (T96), we demonstrate a diurnal fluctuation which we call the Universal Time (UT) effect of the OCB. This UT effect is independent of all other inputs. We anticipate this UT effect to have important consequences in modeling the OCB and other polar cap-associated structures, especially polar cap absorption events that adversely affect high-frequency radio wave propagation in polar regions.
An abstract is not available for this content so a preview has been provided. Please use the Get access link above for information on how to access this content.
This paper proposes a model of information cascades as directed spanning trees (DSTs) over observed documents. In addition, we propose a contrastive training procedure that exploits partial temporal ordering of node infections in lieu of labeled training links. This combination of model and unsupervised training makes it possible to improve on models that use infection times alone and to exploit arbitrary features of the nodes and of the text content of messages in information cascades. With only basic node and time lag features similar to previous models, the DST model achieves performance with unsupervised training comparable to strong baselines on a blog network inference task. Unsupervised training with additional content features achieves significantly better results, reaching half the accuracy of a fully supervised model.
We propose a novel approach to OCR post-correction that exploits repeated texts in large corpora both as a source of noisy target outputs for unsupervised training and as a source of evidence when decoding. A sequence-to-sequence model with attention is applied for single-input correction, and a new decoder with multi-input attention averaging is developed to search for consensus among multiple sequences. We design two ways of training the correction model without human annotation, either training to match noisily observed textual variants or bootstrapping from a uniform error model. On two corpora of historical newspapers and books, we show that these unsupervised techniques cut the character and word error rates nearly in half on single inputs and, with the addition of multi-input decoding, can rival supervised methods.