As generative artificial intelligence (GenAI) drives computational demands to unprecedented scales, digital hardware is approaching fundamental limits. Analog and optical systems promise orders-of-magnitude efficiency gains, but translating these to application-level gains is challenging due to the mismatch between hardware primitives and algorithmic requirements. Here, we introduce Analog Diffusion Models (ADMs) which implement diffusion inference with an implicit integration scheme, formulating each diffusion step as a fixed-point problem amenable for acceleration by efficient analog hardware. At the same time, training remains identical to that of conventional diffusion models, allowing adoption of established scalable training approaches with no additional overhead. We validate ADMs on analog hardware using three-dimensional optics with 2,304 programmable weights. On hardware, we generate two-dimensional distributions and latent-space distributions for MNIST, FashionMNIST, and ExtendedMNIST, demonstrating the feasibility of executing multi-layer diffusion processes entirely on noisy, non-traditional hardware. The current prototype reaches fixed-point convergence in 10–15 µs per diffusion step, with projections to nanosecond-scale convergence with miniaturization. In simulation, across multiple datasets, backbone architectures, and model sizes ranging from 32 million to 13 billion parameters, ADMs match the sample quality of standard methods with up to 16× fewer diffusion steps. Most importantly, they could achieve efficiency gains of more than 100× at the application level without sacrificing generation quality, 100× from hardware acceleration, and an additional 1-2× from algorithmic improvement, highlighting the multiplicative benefit of hardware–algorithm co-design. Together, these results establish ADMs as a scalable and general, hardware-aligned framework for low-latency and energy-efficient generative modeling on analog computing platforms.
This study systematically investigates two multi-fidelity strategies used to train machine-learned force fields (MLFFs)—pre-training/fine-tuning and multi-headed training—and elucidates the mechanisms underpinning their success. For pre-training and fine-tuning, we uncover a log–log linear relationship between pre-trained and fine-tuned accuracies that holds across model architectures, model sizes, and quantum-chemical methods. The success of this approach hinges on the quantity and quality of available pre-training data, and, critically, the inclusion of force labels. We demonstrate that pre-trained representations are inherently method-specific, requiring adaptation of the model backbone during fine-tuning. In contrast, multi-headed models learn method-independent backbone representations, where again the heads’ accuracies are log–log linearly related. Relative to pre-training and fine-tuning, these shared representations marginally reduce model performance in most cases. However, this trade-off is offset by practical advantages: multi-headed training extends naturally to multiple labelling methods and enables partial replacement of expensive labels with cheaper alternatives, paving the way towards cost-efficient universal MLFFs.
This package contains the interview guide and survey questions used for the paper Product Manager Practices for Delegating Work to Generative AI: "Accountability must not be delegated to non-human actors" (ICSE 2026). The preprint of the paper is available here: Product Manager Practices for Delegating Work to Generative AI: ‘Accountability must not be delegated to non-human actors’. For the most updated version of the paper please view: Mara Ulloa's Google Scholar. Demographics, Survey Questions, and Interview Guide.docx Here begin with a table showing demographics from telemetry data consenting participants. We list all survey questions asked to participants, including individual contributors (ICs) and People Managers (PMs). Please reference Figure 1 within the latest version of the paper (on Google scholar) for clarity. Finally, we include the interview guide used for the semi-structured interviews.
As AI systems are increasingly tested and deployed in open-ended and high-stakes domains, crowdworkers are often tasked with responsible AI (RAI) content work. These tasks include labeling violent content, moderating disturbing text, or simulating harmful behavior for red teaming exercises to shape AI system behaviors. While prior research efforts have highlighted the risks to worker well-being associated with RAI content work, far less attention has been paid to how these risks are communicated to workers by task designers or individuals who design and post RAI tasks. Existing transparency frameworks and guidelines, such as model cards, datasheets, and crowdworksheets, focus on documenting model information and dataset collection processes, but they overlook an important aspect of disclosing well-being risks to workers. In the absence of standard workflows or clear guidance, the consistent application of content warnings, consent flows, or other forms of well-being risk disclosure remains unclear. This study investigates how task designers approach risk disclosure in crowdsourced RAI tasks. Drawing on interviews with 23 task designers across academic and industry sectors, we examine how well-being risk is recognized, interpreted, and communicated in practice. Our findings highlight the need to support task designers in identifying and communicating risks not only to support crowdworker well-being but also to strengthen the ethical integrity and technical efficacy of AI development pipelines.
The widespread availability of large language models (LLMs) has provoked both fear and excitement in education. On one hand, there is the concern that students will offload their coursework to LLMs, limiting what they themselves learn. On the other hand, there is the hope that LLMs might serve as scalable, personalized tutors. To investigate how LLM-based explanations affect learning, we conducted two large, pre-registered experiments with U.S. adults (N = 1,818). In both studies, participants first practiced math problems with different forms of assistance and then answered new but similar test questions without any help. Experiment 1 compared the effects of providing participants with only correct answers versus correct LLM explanations, revealing that LLM explanations improved learning, especially when participants first attempted problems on their own. Experiment 2 was similar, but included incorrect LLM explanations with arithmetic errors that produced incorrect final answers. Learning gains from these incorrect explanations were higher than from receiving only the correct answer but lower than from receiving correct explanations. These results suggest that LLM explanations can foster learning gains, although their effectiveness depends not only on the capabilities of the underlying models, but also on how the technology is designed and used to support productive educational experiences.