Creative generation tasks, such as narrative writing and scientific ideation, demand both high-quality outputs and distinct responses across independent runs to maximize exploration. Multi-Agent Debate (MAD) has shown strong quality gains on factual and reasoning tasks, making it a natural candidate for creative generation. However, we find its convergence-driven design actively suppresses output diversity across independent runs, creating an inherent trade-off with creative tasks. We theoretically show that preserving diversity among agents within each debate session is a necessary condition for achieving diverse outputs across independent runs. Building on this finding, we propose Creative-MAD, which introduces two synergistic interventions to sustain agent divergence. Specifically, Cognitive Lens Assignment counters identity drift by anchoring each agent to a distinct and persistent cognitive mode, while Embedding-based Peer Selection counters majority pull by limiting each agent's context to its most semantically distant peers. Experiments across four creative benchmarks demonstrate that Creative-MAD significantly enhances both lexical and semantic diversity while maintaining MAD's output quality.
General movements (GMs) are part of the spontaneous movement repertoire and are present from early fetal life onwards up to age five months. GMs are connected to infants' neurological development and can be qualitatively assessed via the General Movement Assessment (GMA). In particular, between the age of three to five months, typically developing infants produce Fidgety Movements (FM) and their absence provides strong evidence for the presence of cerebral palsy (CP). To improve accessibility to the GMA, automated GMA solutions have been a key research area with proposed models becoming increasingly more accurate and interpretable. However, current models cannot gauge their ability to make decisions, which may lead to overconfident mistakes. To address this issue, we propose a Deep learning-based approach that not only classifies movements as fidgety or non-fidgety but also selectively abstains from classification when uncertain. Through two novel regularization losses, our model maintains a balanced coverage across the two movement types, which prevents bias toward an easy-to-classify subset of movements. We show that our proposed model learns to gauge its own confidence on movement classification, and our proposed regularization losses effectively ensure that the model maintains a similar confidence across movement types. We also show that the local movement abstentions have little impact on the video-level coverage and that relying on the most confident predictions improves the video-level performance.
Bayesian Optimization critically depends on the choice of acquisition function, but no single strategy is universally optimal; the best choice is non-stationary and problem-dependent. Existing adaptive portfolio methods often base their decisions on past function values while ignoring richer information like remaining budget or surrogate model characteristics. To address this, we introduce LMABO, a novel framework that casts a pre-trained Large Language Model (LLM) as a zero-shot, online strategist for the BO process. At each iteration, LMABO uses a structured state representation to prompt the LLM to select the most suitable acquisition function from a diverse portfolio. In an evaluation across 50 benchmark problems, LMABO demonstrates a significant performance improvement over strong static, adaptive portfolio, and other LLM-based baselines. We show that the LLM's behavior is a comprehensive strategy that adapts to real-time progress, proving its advantage stems from its ability to process and synthesize the complete optimization state into an effective, adaptive policy.
Accurate charge densities are central to electronic structure theory, but computing charge-state-dependent densities with density functional theory (DFT) remains too expensive for large-scale screening and defect workflows. We present ChargeFlow, a flow-matching refinement model that transforms a charge-conditioned superposition of atomic densities into the corresponding DFT electron density on the native periodic real space grid using a 3D U-Net velocity field. Trained on 9502 charged Materials Project-derived calculations and evaluated on an external 1671-structure benchmark spanning perovskites, charged defects, diamond defects, metal-organic frameworks, and organic crystals, ChargeFlow is not uniformly best on every in-distribution class but is strongest on problems dominated by nonlocal charge redistribution and charge-state extrapolation, improving deformation-density error from 3.62% to 3.21% and charge-response cosine similarity from 0.571 to 0.655 relative to a ResNet baseline. The predicted densities remain chemically useful under downstream analysis, yielding successful Bader partitioning on all 1671 benchmark structures and high-fidelity electrostatic potentials, which position flow matching as a practical density-refinement strategy for charged materials.
Discovering new solid-state materials requires rapidly exploring the vast space of crystal structures and locating stable regions. Generating stable materials with desired properties and compositions is extremely difficult as we search for very small isolated pockets in the exponentially many possibilities, considering elements from the periodic table and their 3D arrangements in crystal lattices. Materials discovery necessitates both optimized solution structures and diversity in the generated material structures. Existing methods struggle to explore large material spaces, outside the training space, and generate diverse samples with desired properties and requirements. We propose the Symmetry-aware Hierarchical Architecture for Flow-based Traversal (SHAFT), a novel generative model employing a hierarchical exploration strategy to efficiently exploit the symmetry of the materials space to generate crystal structures with desired properties. In particular, our model decomposes the exponentially large materials space into a hierarchy of subspaces consisting of symmetric space groups, lattice parameters, and atoms. To benchmark our approach, we first develop a novel, non-hierarchical GFlowNet for complete crystal structure generation. We then introduce SHAFT, a more expressive hierarchical model that leverages its architecture and increased model capacity to more efficiently explore the materials space. We demonstrate that SHAFT significantly outperforms the flat GFlowNet baseline and other state-of-the-art methods like CDVAE and DiffCSP in generating valid, stable, and diverse crystal structures.
Sample size determination for machine learning (ML) prediction models is challenging because conventional power analysis typically requires the predictor-outcome relationship and effect structure to be specified a priori. Nonlinear ML models learn complex prediction surfaces that do not admit straightforward analytical power calculations. We propose a framework that approximates nonlinear ML models with localized linear representations and estimates sample size requirements by evaluating statistical power across these local regions.
Multimorbidity, the co-occurrence of multiple chronic conditions within the same individual, is increasing globally. This is a challenge for the single patients, as these individuals are subject to a heavy disease and treatment burden, yet evidence on the epidemiology and consequences of multimorbidity remains underexplored. Historically, studies aiming to understand multimorbidity patterns predominantly utilized cross-sectional data, neglecting the essential temporal dynamics which shape multimorbidity progression. Other studies based their analyses on small datasets, or populations only targeting certain sectors of the healthcare system. In this study, we (1) introduce a novel two-step multimodal Variational Autoencoder-based approach for temporal disease-based clustering (i.e. discovering age-aware multimorbidity clusters); (2) provide quantitative experiments for the robustness of our approach and the extracted temporal clusters; and (3) demonstrate how the temporal disease clusters obtained from our model can provide novel understanding of the development of multiple conditions over time and thus generate new hypotheses for different stages of multimorbidity and their associations. We trained and evaluated our models on a dataset containing the entire adult population of Denmark in the period 1995-2015, focusing on individuals suffering from chronic heart disease, including 766,596 individuals.
Although lithium solid state electrolytes promise to mitigate the chemical instabilities of liquid electrolytes in today's mainstream rechargeable batteries, solid state electrolytes still suffer from dendrite formation which leads to battery degradation and short circuiting. Dendrite initiation and propagation in specific solid state electrolyte materials has been explained, at a microscopic scale, as emerging from the lithium-filling of pores within the solid state electrolytes via microcracks. At the atomistic scale, the thermodynamic instability of many solid state electrolyte materials can explain their susceptibility to crystal decomposition upon contact with the lithium anode. However, for a more complete picture of the dendrite formation mechanisms, an understanding of the dendrite initiation mechanism at the intermediate nanoscopic scale is required. This work applies a machine learning potential (DIEP) for simulating six different solid state electrolyte-lithium interfaces at 300 K and 1000 K, with model sizes ranging from 24k to 36k atoms, for durations exceeding 20 ps. Our simulations show that the lithium dendrite initiation process can have an underpinning nanoscopic mechanism, in which the crystal decomposition by direct lithium interaction leads to the clustering of lithium. The simulations also suggest a possible mechanism for the creation of voids within the solid-electrolyte interphase, which have been observed in the Li$|$Li$_6$PS$_5$Cl$|$Li interface.
The prediction of the electron density in molecules and crystals is a key pillar in the first-principles computation of their properties. Using machine learning to predict the electron density by using the atomic structure alone can save the computational cost of performing first-principles computations. While various machine learning approaches have been introduced for predicting the electron density, none of them predict the electron density for charged systems. This work extends a recent machine learning charge density model, DeepDFT, by including the charge of the structure as an input parameter into the model. We establish an input charge representation approach that successfully predicts the charged electron densities for several test cases, including charged defective perovskites, LiCoO2 supercells with multiple Li vacancies, diamond-based defects, metal-organic frameworks, and molecular crystals.
Neural backdoor attacks present a critical vulnerability in deep-learning systems. In this work, we introduce a more general form of trigger-based attacks that bypass existing defense methods. To counter this, we propose a novel defense mechanism grounded in an information-theoretic analysis of how backdoors are most efficiently encoded in the feature space. Specifically, we prove that the most information-efficient way to represent multiple trigger-based backdoors is to construct layered manifolds, where each layer corresponds to a unique backdoor trigger and the separation between layers is encoded radially (assuming the usual high-dimensional hyper-spherical data distribution). This structure enables the main classifier to maintain its original classification boundary while accessing backdoors through simple radial separations. Our defense exploits this insight by searching for perturbations that exhibit global behaviour, indicative of the presence of such layered manifolds and, therefore, backdoors. Extensive experiments using a variety of image datasets demonstrate that our method successfully identifies backdoors that are missed by state-of-the-art detection methods.
Artificial agents with the aid of large language models (LLMs) are effective in various real-world scenarios but struggle to cooperate in social dilemmas. When making decisions under the strain of selecting between long-term consequences and short-term benefits in commonly shared resources, LLM-based agents often exploit the environment, leading to early depletion. Inspired by the concept of consideration of future consequences (CFC), which is well-known in social psychology, we propose a framework to enable the ability to consider future consequences for LLM-based agents, which results in a new kind of agent that we term the CFC-Agent. We enable the CFC-Agent to act toward different levels of consideration for future consequences. Our first set of experiments, where LLM is directly asked to make decisions, shows that agents considering future consequences exhibit sustainable behaviour and achieve high common rewards for the population. Extensive experiments in complex environments showed that the CFC-Agent can manage a sequence of calls to LLM for reasoning and engaging in communication to cooperate with others to resolve the common dilemma better. Finally, our analysis showed that considering future consequences not only affects the final decision but also improves the conversations between LLM-based agents toward a better resolution of social dilemmas.
In this paper we propose a human-AI teaming framework for the optimization of expensive black-box functions. Inspired by the intrinsic difficulty of extracting expert knowledge and distilling it back into AI models and by observations of human behavior in real-world experimental design, our proposed algorithm lets the human expert take the lead in the experimental process. The human expert can use their domain expertise to its full potential, while the AI plays the role of a muse, injecting novelty and searching for areas of weakness to break the human out of over-exploitation induced by cognitive entrenchment. We validate our proposed algorithm using synthetic data and with human experts performing real-world experiments.
Typically developing infants, between the corrected age of 9-20 weeks, produce fidgety movements. These movements can be identified with the General Movement Assessment, but their identification requires trained professionals to conduct the assessment from video recordings. Since trained professionals are expensive and their demand may be higher than their availability, computer vision-based solutions have been developed to assist practitioners. However, most solutions to date treat the problem as a direct mapping from video to infant status, without modeling fidgety movements throughout the video. To address that, we propose to directly model infants' short movements and classify them as fidgety or non-fidgety. In this way, we model the explanatory factor behind the infant's status and improve model interpretability. The issue with our proposal is that labels for an infant's short movements are not available, which precludes us to train such a model. We overcome this issue with active learning. Active learning is a framework that minimizes the amount of labeled data required to train a model, by only labeling examples that are considered "informative" to the model. The assumption is that a model trained on informative examples reaches a higher performance level than a model trained with randomly selected examples. We validate our framework by modeling the movements of infants' hips on two representative cohorts: typically developing and at-risk infants. Our results show that active learning is suitable to our problem and that it works adequately even when the models are trained with labels provided by a novice annotator.
Large language models (LLMs) have recently demonstrated their impressive ability to provide context-aware responses via text. This ability could potentially be used to predict plausible solutions in sequential decision making tasks pertaining to pattern completion. For example, by observing a partial stack of cubes, LLMs can predict the correct sequence in which the remaining cubes should be stacked by extrapolating the observed patterns (e.g., cube sizes, colors or other attributes) in the partial stack. In this work, we introduce LaGR (Language-Guided Reinforcement learning), which uses this predictive ability of LLMs to propose solutions to tasks that have been partially completed by a primary reinforcement learning (RL) agent, in order to subsequently guide the latter's training. However, as RL training is generally not sample-efficient, deploying this approach would inherently imply that the LLM be repeatedly queried for solutions; a process that can be expensive and infeasible. To address this issue, we introduce SEQ (sample efficient querying), where we simultaneously train a secondary RL agent to decide when the LLM should be queried for solutions. Specifically, we use the quality of the solutions emanating from the LLM as the reward to train this agent. We show that our proposed framework LaGR-SEQ enables more efficient primary RL training, while simultaneously minimizing the number of queries to the LLM. We demonstrate our approach on a series of tasks and highlight the advantages of our approach, along with its limitations and potential future research directions.
Current lithium batteries do not fully meet the longevity and safety requirements of electric vehicles. Novel solid state lithium ion batteries could be a compelling solution to these problems. In this work we unravel some of these new materials with potentially high lithium conductivity by using a Bayesian optimization approach. This involves exploring the material space for new solid-state electrolyte materials with the objective of maximising - lithium diffusivity. The materials selected by the Bayesian optimisation algorithm are then examined using ab initio molecular dynamics to estimate their diffusion energy barrier. We establish that the materials are electronic insulators, a requirement in electrolyte materials, by computing the electronic bandgaps of each of the selected materials using a hybrid exchange method, and then examine the stability of the materials at the lithium metal anode interface by computing the crystal decomposition energies. Out of the selected materials, we find that Li3YBr6 has a reasonably low diffusion barrier, a high bandgap and is potentially the most stable material at the lithium metal interface. In addition to introducing stable and high-diffusivity solid-state electrolyte materials, our work presents a material discovery method that can be applied for a broad range of applications.
Adversarial attacks on deep models are often guaranteed to find a small and innocuous perturbation to easily alter the class label of a test input. We use a novel Targeted Manifold Manipulation (TMM) approach to direct the gradients from the genuine data manifold toward carefully planted traps during such adversarial attacks. The traps are assigned an additional class label (Trapclass) to make the attacks falling in them easily identifiable. Whilst low-perturbation budget attacks will necessarily end up in the traps, high-perturbation budget attacks may escape but only end up far away from the data manifold. Since our manifold manipulation is enforced only locally, we show that such out-of-distribution data can be easily detected by noting the absence of traps around them. Our detection algorithm, denoted as TMM-Def avoids learning a separate model for attack detection and thus remains semantically aligned with the original classifier. Further, since we manipulate the adversarial distribution, it avoids the fundamental difficulty associated with overlapping distributions of clean and attack samples for usual, unmanipulated models. We use nine state-of-the-art adversarial attacks with six well-known image datasets to evaluate our proposed defense. Our results show that the proposed method can detect similar to 99% attacks whilst also being robust to semantic-preserving, transformations, and adaptive attacks.
Dimensionless groups quantify the balance among key forces governing a system’s physical behaviour and are foundational in engineering for describing, comparing, and scaling processes. By condensing complex system interactions into single values, they provide a powerful means of abstraction. Yet, their potential to actively guide process optimisation remains largely untapped. This study presents a framework that integrates dimensionless analysis with Bayesian optimisation to enhance both process performance and interpretability. Using this combined approach, we demonstrate that optimisation conducted in the dimensionless space not only accelerates convergence towards optimal process conditions but also reveals the underlying physical balances driving system behaviour. The method thus bridges data-driven optimisation with physically grounded understanding, enabling more efficient and explainable control of complex manufacturing processes.
BACKGROUND:Prevention of suicide is a global health priority. Approximately 800,000 individuals die by suicide yearly, and for every suicide death, there are another 20 estimated suicide attempts. Large language models (LLMs) hold the potential to enhance scalable, accessible, and affordable digital services for suicide prevention and self-harm interventions. However, their use also raises clinical and ethical questions that require careful consideration. OBJECTIVE:This scoping review aims to identify emergent trends in LLM applications in the field of suicide prevention and self-harm research. In addition, it summarizes key clinical and ethical considerations relevant to this nascent area of research. METHODS:Searches were conducted in 4 databases (PsycINFO, Embase, PubMed, and IEEE Xplore) in February 2024. Eligible studies described the application of LLMs for suicide or self-harm prevention, detection, or management. English-language peer-reviewed articles and conference proceedings were included, without date restrictions. Narrative synthesis was used to synthesize study characteristics, objectives, models, data sources, proposed clinical applications, and ethical considerations. This review adhered to the PRISMA-ScR (Preferred Reporting Items for Systematic Reviews and Meta-Analyses extension for Scoping Reviews) standards. RESULTS:Of the 533 studies identified, 36 (6.8%) met the inclusion criteria. An additional 7 studies were identified through citation chaining, resulting in 43 studies for review. The studies showed a bifurcation of publication fields, with varying publication norms between computer science and mental health. While most of the studies (33/43, 77%) focused on identifying suicide risk, newer applications leveraging generative functions (eg, support, education, and training) are emerging. Social media was the most common source of LLM training data. Bidirectional Encoder Representations from Transformers (BERT) was the predominant model used, although generative pretrained transformers (GPTs) featured prominently in generative applications. Clinical LLM applications were reported in 60% (26/43) of the studies, often for suicide risk detection or as clinical assistance tools. Ethical considerations were reported in 33% (14/43) of the studies, with privacy, confidentiality, and consent strongly represented. CONCLUSIONS:This evolving research area, bridging computer science and mental health, demands a multidisciplinary approach. While open access models and datasets will likely shape the field of suicide prevention, documenting their limitations and potential biases is crucial. High-quality training data are essential for refining these models and mitigating unwanted biases. Policies that address ethical concerns-particularly those related to privacy and security when using social media data-are imperative. Limitations include high variability across disciplines in how LLMs and study methodology are reported. The emergence of generative artificial intelligence signals a shift in approach, particularly in applications related to care, support, and education, such as improved crisis care and gatekeeper training methods, clinician copilot models, and improved educational practices. Ongoing human oversight-through human-in-the-loop testing or expert external validation-is essential for responsible development and use. TRIAL REGISTRATION:OSF Registries osf.io/nckq7; https://osf.io/nckq7.
We introduce Pointer-Augmented Neural Memory (PANM), a versatile module designed to enhance neural networks' ability to process symbols and extend their capabilities to longer data sequences. PANM integrates an external neural memory utilizing novel physical addresses and pointer manipulation techniques, emulating human and computer-like symbol processing abilities. PANM facilitates operations like pointer assignment, dereferencing, and arithmetic by explicitly employing physical pointers for memory access. This module can be trained end-to-end on sequence data, empowering various sequential models, from simple recurrent networks to large language models (LLMs). Our experiments showcase PANM's exceptional length extrapolation capabilities and its enhancement of recurrent neural networks in symbol processing tasks, including algorithmic reasoning and Dyck language recognition. PANM enables Transformers to achieve up to 100% generalization accuracy in compositional learning tasks and significantly improves performance in mathematical reasoning, question answering, and machine translation. Notably, the generalization effectiveness scales with stronger backbone models, as evidenced by substantial performance gains when we test LLMs finetuned with PANM for tasks up to 10-100 times longer than the training data.
Tele Tan合作论文数School of Civil and Mechanical Engineering, Faculty of Science and Engineering, Curtin University10