This study presents an audio stimulus corpus consisting of 444 Western popular music drum and bass patterns called the Lucerne Groove Library. Common requirements for stimuli in music psychological research, particularly studies on the experience of groove, are outlined, followed by a description of how these criteria are addressed in the presented corpus. For example, the corpus is designed to combine ecological validity with high manipulability, facilitating the use of the corpus in a variety of experimental settings. The methods section provides a detailed account of material selection, audio creation, and measures. Ground-truth behavioural data for the stimuli were obtained through a listening experiment (e.g., participant's ratings on the urge to move in response to the stimuli, and style assignments ). Several structural data (e.g., tempo and event density) and audio features (e.g., low-frequency sub-band flux) were measured for each stimulus, and an overview over all data is provided. Potential applications of the corpus in future studies are discussed, including variables along which stimuli can be selected or manipulated. The corpus is published in different formats and with a range of accompanying data, all of which can be found in the online repository.
We introduce aligned probing, a novel interpretability framework that aligns the behavior of language models (LMs), based on their outputs, and their internal representations (internals). Using this framework, we examine over 20 OLMo, Llama, and Mistral models, bridging behavioral and internal perspectives for toxicity for the first time. Our results show that LMs strongly encode information about the toxicity level of inputs and subsequent outputs, particularly in lower layers. Focusing on how unique LMs differ offers both correlative and causal evidence that they generate less toxic output when strongly encoding information about the input toxicity. We also highlight the heterogeneity of toxicity, as model behavior and internals vary across unique attributes such as Threat. Finally, four case studies analyzing detoxification, multi-prompt evaluations, model quantization, and pre-training dynamics underline the practical impact of aligned probing with further concrete insights. Our findings contribute to a more holistic understanding of LMs, both within and beyond the context of toxicity.
This paper offers a critical reassessment of claims that scientific progress is best understood through the disruptiveness of new research. Park, Leahey and Funk in Nature (2023) have re-opened the debate by presenting results using the citation-based 'CD index' to assess the extent to which individual academic publications are consolidating or disruptive. Analyzing Park et al. as a focal point of these claims, we challenge the adequacy of this approach to capture both genuine scientific disruption and scientific progress, particularly within the social sciences. Drawing on philosophy and sociology of science, we show that scientific progress is predominantly cumulative rather than disruptive, and that papers' high disruptiveness scores may often reflect phenomena such as pseudo-novelty or fragmentation rather than true epistemic breakthroughs. Our analysis demonstrates that in fields marked by intellectual pluralism and weak paradigmatic consensus, apparent disruptiveness may be an artifact of scholarly practices rather than an indication of substantive innovation. Hence, measures of disruptiveness appear ill-suited as a marker of scientific progress-as used in individual and collective research evaluations. Instead, we advance a constructive agenda by proposing that scientific progress is best conceptualized not as a dichotomy between cumulation and disruptiveness, but as a multi-dimensional process embracing elements of both disruption and consolidation within an overarching cumulative trajectory, whereby established knowledge is iteratively refined, rejected, or recombined in the light of new evidence or insight. By rethinking how scientific advancement is measured and cultivated, and suggesting ways to foster cumulative scientific progress, this article contributes to the theory and practice of research evaluation.
Multi-Party Computation (MPC) enables a set of parties to jointly compute a function while preserving the privacy of their inputs. Although the problem has been studied for several decades, most prior work considers the classical setting in which the set of parties is fixed throughout the protocol execution. This assumption is poorly suited for modern applications that are long-lived and in which parties may join or leave the computation dynamically. Motivated by this limitation, a growing body of recent work has introduced models of MPC with dynamic committees, including YOSO MPC, Fluid MPC, Layered MPC, and SCALES, among others. The proliferation of such models raises natural questions: how do these approaches relate to each other, and can they be unified under a common framework? We show that MPC with dynamic committees can be reduced to the design of a classical MPC protocol with a static set of parties and a general adversary structure, but where the interaction pattern is constrained to follow a fixed layered acyclic graph. Each party corresponds to a node in the graph and can send secret messages along its outgoing edges. We further demonstrate that existing dynamic-committee MPC models can be recovered as specific instantiations of this layered-graph framework.
Abstract Wearable light loggers and optical radiation dosimeters are increasingly used in chronobiology and circadian health research, yet their data often lack contextual information (e.g., sleep, activity, environmental conditions) and may be compromised by non-wear periods, compliance issues, or technical faults. To address these limitations, we conducted interviews (n = 21) and a survey (n = 16) with domain experts to distil and iteratively develop auxiliary data and quality-control strategies aimed at improving the accuracy and interpretability of wearable light measurements. From this process, we established a six-domain auxiliary data framework encompassing wear/non-wear logging, sleep monitoring, light-source context, participant behaviour, user experience, and environmental light levels. Survey responses showed strong consensus on the value of auxiliary information (importance 4.0/5), with sleep and wear-time tracking rated as the most essential additions. To support practical adoption, we provide implementation tools, including extensions to the open-source R package LightLogR, enabling streamlined integration of wearable and auxiliary data as well as systematic quality assurance and control. Experts agreed that combining contextual records with rigorous QA/QC procedures substantially improves the reliability of field-collected light-exposure data. These recommendations and tools aim to help researchers in chronobiology, wearable sensing, and health sciences maximise data quality and enhance interpretation in real-world light-exposure studies.