Introduction. Clinical and Translational Science Award (CTSA) programs must document their scholars' research impact, but assembling each scholar's record by hand takes staff an estimated 15 hours and does not scale to a full cohort. An artificial intelligence (AI) agent could serve as a tool to gather scholar data across platforms and disciplines. Methods. We built a human-in-the-loop AI agent that assembles a dossier of sourced evidence for each scholar and drafts one-sentence Translational Science Benefits Model (TSBM) impact summaries for staff review. We evaluated it in the impact-reporting workflow of one CTSA hub across 10 career-development (KL2/K12) scholars. Two evaluation staff independently coded all 507 findings as accept, edit, or reject; the primary measure was the unanimous usable rate, defined as the share both accepted or edited. Results. Both reviewers accepted or edited 81.7
We describe the UIC-AIHealth4All system for ArchEHR-QA 2026, a shared task on grounded question answering from electronic health records. We participated in Subtasks 2 (evidence identification), 3 (answer generation), and 4 (answer-evidence alignment). For Subtasks 2 and 3, we propose an answer-first pipeline in which the model generates candidate answers citing specific note sentences before classifying the full evidence set, exploiting the asymmetry between judging relevance in the abstract versus relative to a generated answer. For Subtask 4, we apply self-consistency voting over five independent model calls, retaining links by vote threshold. Our pipeline ranked third on evidence identification (Strict Micro F1 62.90), ninth on answer generation (Overall 31.90), and fifth on answer-evidence alignment (F1 79.81). A post-hoc linguistic analysis of 45 stylistic features reveals that model outputs remain 3.2 Flesch-Kincaid grade levels harder to read than clinician-authored references despite matching their word and sentence counts, suggesting readability warrants explicit optimization in clinical NLP systems. Code and prompts are available at https://github.com/mo-arvan/archehr-qa-2026-uic-aihealth4all.
Reproducibility remains a fundamental challenge for human evaluation in Natural Language Processing (NLP), particularly due to the inherent subjectivity and variability of human judgments. This paper presents a reproduction study of the human evaluation protocol introduced by Hosking and Lapata (2021), which assesses semantic preservation in paraphrase generation models. By faithfully reproducing the original experiment with careful adaptation and applying the Quantified Reproducibility Assessment framework (Belz and Thomson, 2024a; Belz, 2022), we demonstrate strong agreement with the original findings, confirming the semantic preservation ranking among four paraphrase models. Our analyses reveal moderate inter-annotator agreement and low variability in key results, underscoring a good degree of reproducibility despite practical deviations in participant recruitment and platform. These findings highlight the feasibility and challenges of reproducing human evaluation studies in NLP. We discuss implications for improving methodological rigor, transparent reporting, and standardized protocols to bolster reproducibility in future human evaluations. The data and analysis scripts are publicly available to support ongoing community efforts toward reproducible evaluation in NLP and beyond.
Virtual assistants have become fixtures in everyday settings, but most research focuses on their development rather than their use following deployment. To facilitate study of their use in office settings, we introduce OfficeDial , a multimodal dataset containing audio recordings, transcriptions, eye tracking data, and screen recordings from conversations between humans and virtual assistants in office environments. Conversations are paired with physical and behavioral measures of cognitive load. We study the associations between verbal behavior and noise level and reveal key relationships between verbal redundancy, disfluency, and noise level. We make our new dataset available to interested researchers to inspire further exploration.
It might reasonably be expected that running multiple experiments for the same task using the same data and model would yield very similar results.Recent research has, however, shown this not to be the case for many NLP experiments.In this paper, we report extensive coordinated work by two NLP groups to run the training and testing pipeline for three neural text simplification models under varying experimental conditions, including different random seeds, run-time environments, and dependency versions, yielding a large number of results for each of the three models using the same data and train/dev/test set splits.From one perspective, these results can be interpreted as shedding light on the reproducibility of evaluation results for the three NTS models, and we present an indepth analysis of the variation observed for different combinations of experimental conditions.From another perspective, the results raise the question of whether the averaged score should be considered the 'true' result for each model.
Reproducibility is a key aspect for scientific advancement across disciplines, and reducing barriers for open science is a focus area for the theme of Interspeech 2023. Availability of source code is one of the indicators that facilitates reproducibility. However, less is known about the rates of reproducibility at Interspeech conferences in comparison to other conferences in the field. In order to fill this gap, we have surveyed 27,717 papers at seven conferences across speech and language processing disciplines. We find that despite having a close number of accepted papers to the other conferences, Interspeech has up to 40% less source code availability. In addition to reporting the difficulties we have encountered during our research, we also provide recommendations and possible directions to increase reproducibility for further studies.
We report our efforts in identifying a set of previous human evaluations in NLP that would be suitable for a coordinated study examining what makes human evaluations in NLP more/less reproducible. We present our results and findings, which include that just 13% of papers had (i) sufficiently low barriers to reproduction, and (ii) enough obtainable information, to be considered for reproduction, and that all but one of the experiments we selected for reproduction was discovered to have flaws that made the meaningfulness of conducting a reproduction questionable. As a result, we had to change our coordinated study design from a reproduce approach to a standardise-then-reproduce-twice approach. Our overall (negative) finding that the great majority of human evaluations in NLP is not repeatable and/or not reproducible and/or too flawed to justify reproduction, paints a dire picture, but presents an opportunity for a rethink about how to design and report human evaluations in NLP.
This paper presents a partial reproduction study of Data-to-text Generation with Macro Planning by Puduppully et al. (2021). This work was conducted as part of the ReproHum project, a multi-lab effort to reproduce the results of NLP papers incorporating human evaluations. We follow the same instructions provided by the authors and the ReproHum team to the best of our abilities. We collect preference ratings for the following evaluation criteria in order: conciseness, coherence, and grammaticality. Our results are highly correlated with the original experiment. Nonetheless, we believe the presented results are insufficent to conclude that the Macro system proposed and developed by the original paper is superior compared to other systems. We suspect combining our results with the three other reproductions of this paper through the ReproHum project will paint a clearer picture. Overall, we hope that our work is a step towards a more transparent and reproducible research landscape.
This study examines the effects of noise and the use of an Intelligent Virtual Assistant (IVA) on the task performance and workload of office workers. Data were collected from forty-eight adults across varied office task scenarios (i.e., sending an email, setting up a timer/reminder, and searching for a phone number/address) and noise types (i.e., silence, non-verbal noise, and verbal noise). The baseline for this study is measured without the use of an IVA. Significant differences in performance and workload were found on both objective and subjective measures. In particular, verbal noise emerged as the primary factor affecting performance using an IVA. Task performance was dependent on the task scenario and noise type. Subjective ratings found that participants preferred to use IVA for less complex tasks. Future work can focus more on the effects of tasks, demographics, and learning curves. Furthermore, this work can help guide IVA system designers by highlighting factors affecting performance.
With increasing advancements in natural language processing, intelligent virtual assistants (IVAs), such as Amazon’s Alexa, Apple’s Siri, and Microsoft’s Cortana are rapidly getting integrated into our daily lives. People interact with these command-based software agents to play music, set up alarms, control their smart homes, in-vehicle systems, and much more (Jeong & Liu, 2018). Since this technology is now available on most laptops and desktop computers, more people are using it in the workplace. One of the primary benefits of IVAs for office workers is that it allows them to multitask efficiently while their hands or eyes are busy (Seranmadevi et al., 2022). IVAs can also be beneficial in scenarios where computer input devices are inaccessible, allowing users to skip manual tasks and reduce the number of steps required for human information processing (Wickens & Carswell, 2021). Despite all the benefits of IVAs, factors such as background noise are involved in the efficiency of IVAs. The noise interrupts the speech recognition process of IVAs and may cause the system to hardly recognize and decode the user’s spoken keyword (Palconit et al., 2019). This system failure reduces not only user satisfaction but also the usefulness of the system (Schnelle-Walka, 2010). The objective of this study is to investigate the effect of an IVA in a noisy environment on office workers’ task performance and workload by comparing the outcomes to the traditional way that uses manual operations without the aid of IVA. The IVA platform used in the study was Cortana application (version 3.3.6), built into Microsoft Windows 10. A 2 (system) × 3 (noise level) within-subject design was used. The two systems were System A (performing scenarios with IVA) and System B (performing scenarios using only a keyboard and a computer mouse with no IVA support). The three noise levels were silent, non-verbal noise, and verbal noise. Forty-eight native and non-native English speakers (18 females; age M=26.1; SD=8.72) with various ethnic backgrounds participated in the study. Data were collected while they performed three office task scenarios, each with two sub-scenarios (See Table 1). Table 1. Task scenarios Scenarios Sub-scenarios S1: Sending an email S1-1: Send an email to [a fictional name] requesting feedback on a document sent before S1-2: Send an email to [a fictional name] to follow up on the previous email. S2: Setting up a timer/reminder S2-1: Set up a time for 20 minutes. S2-2: Set up a reminder to attend a meeting. S3: Running an Internet search S3-1: Look for the location of the nearest UPS store. S3-2: Look up the phone number of a university student center.
The availability of source code has been put forward as one of the most critical factors for improving the reproducibility of scientific research. This work studies trends in source code availability at major computational linguistics conferences, namely, ACL, EMNLP, LREC, NAACL, and COLING. We observe positive trends, especially in conferences that actively promote reproducibility. We follow this by conducting a reproducibility study of eight papers published in EMNLP 2021, finding that source code releases leave much to be desired. Moving forward, we suggest all conferences require self-contained artifacts and provide a venue to evaluate such artifacts at the time of publication. Authors can include small-scale experiments and explicit scripts to generate each result to improve the reproducibility of their work.
John Kelleher合作论文数School of Computing,
Dublin Institute of Technology,1