We report a systematic and accurate approach for deriving the bulk free energy surface (FES), a function of temperature, polarization, and strain, from the first-principles density functional theory (DFT) of proper ferroelectrics. The core of our approach is the metadynamics algorithm that extracts the polarization dependence of the FES from all-atom molecular dynamics simulations without an a priori ansatz. The rest of the FES is derived from the metadynamics trajectories that span the relevant phase space. We demonstrate our approach in the case of lead titanate. The errors across the phase transition, due to DFT numerics, all-atom molecular dynamics, and free energy evaluation by enhanced sampling, can be systematically controlled and are of the order of 1meV/atom. The accuracy of the resulting ab initio FES is only limited by the adopted functional approximation of DFT.
Deep learning has advanced mass spectrometry data interpretation, yet most models remain feature extractors rather than unified scoring frameworks. We present pUniFind, a large-scale multimodal foundational model in proteomics that integrates open end-to-end peptide-spectrum scoring with open, zero-shot de novo sequencing. Trained on over 100 million open search-derived spectra, pUniFind aligns spectral and peptide modalities through cross-modality prediction alongside other carefully designed pretraining tasks. Consequently, further benefiting from its open scoring capability, pUniFind outperforms traditional engines across diverse datasets, notably achieving a 42.6% increase in identified peptides in immunopeptidomics. We propose two de novo sequencing workflows to support different applications. For modification-rich de novo sequencing, pUniFind identifies 60% more peptide-spectrum matches than existing de novo methods despite a 300 times larger search space. For regular de novo sequencing, pUniFind recovers an additional 38.5% of peptides, including 1,891 that map to the genome but are absent from reference proteomes. Crucially, it achieves this while preserving full fragment ion coverage and maintaining high consistency with database-search-based methods. Furthermore, a quality control module based on deep learning-derived features increases the consistency of results with RNA-Seq evidence from 65.4% to 85.0%. These results establish a unified, scalable deep learning framework for proteomic analysis, offering improved sensitivity, modification coverage, and interpretability.
The advancement of artificial intelligence toward agentic science is currently bottlenecked by the challenge of ultra-long-horizon autonomy, the ability to sustain strategic coherence and iterative correction over experimental cycles spanning days or weeks. While Large Language Models (LLMs) have demonstrated prowess in short-horizon reasoning, they are easily overwhelmed by execution details in the high-dimensional, delayed-feedback environments of real-world research, failing to consolidate sparse feedback into coherent long-term guidance. Here, we present ML-Master 2.0, an autonomous agent that masters ultra-long-horizon machine learning engineering (MLE) which is a representative microcosm of scientific discovery. By reframing context management as a process of cognitive accumulation, our approach introduces Hierarchical Cognitive Caching (HCC), a multi-tiered architecture inspired by computer systems that enables the structural differentiation of experience over time. By dynamically distilling transient execution traces into stable knowledge and cross-task wisdom, HCC allows agents to decouple immediate execution from long-term experimental strategy, effectively overcoming the scaling limits of static context windows. In evaluations on OpenAI's MLE-Bench under 24-hour budgets, ML-Master 2.0 achieves a state-of-the-art medal rate of 56.44
The paradigm of agentic science requires AI systems to conduct robust reasoning and engage in long-horizon, autonomous exploration. However, current scientific benchmarks remain confined to domain knowledge comprehension and complex reasoning, failing to evaluate the exploratory nature and procedural complexity of real-world research. In this work, we present research-oriented evaluations in theoretical and computational physics, a natural testbed with comprehensive domain knowledge, complex reasoning, and verifiable end-to-end workflows without reliance on experiments. Here we introduce PRL-Bench (Physics Research by LLMs), a benchmark designed to systematically map the capability boundaries of LLMs in executing end-to-end physics research. Constructed from 100 curated papers from the latest issues of Physical Review Letters since August 2025 and validated by domain experts, PRL-Bench covers five major theory- and computation-intensive subfields of modern physics: astrophysics, condensed matter physics, high-energy physics, quantum information, and statistical physics. Each task in the benchmark is designed to replicate the core properties of authentic scientific research, including exploration-oriented formulation, long-horizon workflows, and objective verifiability, thereby reconstructing the essential reasoning processes and research workflows of real physics research. Evaluation across frontier models shows that performance remains limited, with the best overall score below 50, revealing a pronounced gap between current LLM capabilities and the demands of real scientific research. PRL-Bench serves a reliable testbed for accessing next generation AI scientists advancing AI systems toward autonomous scientific discovery.
Large language models (LLMs) have demonstrated remarkable performance across a wide range of tasks and domains, with data playing a central role in enabling these advances. Despite this success, the preparation and effective utilization of the massive datasets required for LLM training remain major bottlenecks. In current practice, LLM training data is often constructed using ad hoc scripts, and there is still a lack of mature, agent-based data preparation systems that can automatically construct robust and reusable data workflows, thereby freeing data scientists from repetitive and error-prone engineering efforts. Moreover, once collected, datasets are often consumed largely in their entirety during training, without systematic mechanisms for data selection, mixture optimization, or reweighting. To address these limitations, we advocate two complementary research directions. First, we propose building a robust, agent-based automatic data preparation system that supports automated workflow construction and scalable data management. Second, we argue for a unified data-model interaction training system in which data is dynamically selected, mixed, and reweighted throughout the training process, enabling more efficient, adaptive, and performance-aware data utilization. Finally, we discuss the remaining challenges and outline promising directions for future research and system development.
Artificial intelligence (AI) is transforming scientific research, including proteomics. In this Perspective, we highlight key mass spectrometry (MS)-based proteomics areas where AI is driving innovation, ranging from protein identification to building AI virtual cells. These include improving peptide and protein identification and quantification; characterizing protein-protein interactions and protein complexes; advancing spatial and perturbation proteomics; integrating multi-omics data; and, ultimately, enabling AI virtual cells. Finally, we call for global collaboration among data producers, data consumers and other stakeholders to establish an AI-friendly ecosystem for MS-based proteomics, laying the foundation for transformative advancements in proteomics driven by AI.
Data-independent acquisition mass spectrometry (DIA-MS) has established itself as a cornerstone of proteomic profiling and large-scale systems biology, offering unparalleled depth and reproducibility. Current DIA analysis frameworks, however, require semi-supervised training within each run for peptide-spectrum match (PSM) re-scoring. This approach is prone to overfitting and lacks generalizability across diverse species and experimental conditions. Here, we present DIA-CLIP, a pre-trained model shifting the DIA analysis paradigm from semi-supervised training to universal cross-modal representation learning. By integrating dual-encoder contrastive learning framework with encoder-decoder architecture, DIA-CLIP establishes a unified cross-modal representation for peptides and corresponding spectral features, achieving high-precision, zero-shot PSM inference. Extensive evaluations across diverse benchmarks demonstrate that DIA-CLIP consistently outperforms state-of-the-art tools, yielding up to a 45
Artificial intelligence is increasingly capable of predicting chemical properties, generating candidate structures, and assisting experimental planning. Yet the rate-limiting step in energy and chemical innovation is no longer prediction alone: it is the conversion of computational proposals into reproducible experiments and deployable process decisions. In this Perspective, we argue that Artificial Intelligence for Science in chemistry is undergoing a decisive transition from model-centric performance improvement to executable, closed-loop research infrastructure. This transition involves three coupled layers. First, physically grounded models must connect molecular and materials structures with energetics, kinetics, experimental observations, and uncertainty. Second, design algorithms must operate in validation-aware workflows that link inverse design, mechanistic computation, experimentation, and process constraints. Third, autonomous laboratories require reusable agent infrastructure that integrates scientific software, instruments, analytical feedback, safety control, and human accountability. We discuss developments in molecular and catalyst design, reaction optimization, digital twins, and self-driving laboratories, and identify data provenance, physical execution, reliability benchmarking, and safety governance as central translational challenges. For energy and chemical engineering, the value of artificial intelligence will ultimately be determined not by the number of candidates it proposes, but by its ability to close the loop from hypothesis to validated technology.
The discovery of novel high-temperature superconductor materials holds transformative potential for a wide array of technological applications. However, the combinatorially vast chemical and configurational search space poses a significant bottleneck for both experimental and theoretical investigations. In this study, we employ the design of high-temperature ternary superhydride superconductors as a representative case to demonstrate how this challenge can be well addressed through a deep-learning-driven theoretical framework. This framework integrates high-throughput crystal structure exploration, physics-informed screening, and accurate prediction of superconducting critical temperatures. Our approach enabled the exploration of approximately 36 million ternary hydride structures across a chemical space of 29 elements, leading to the identification of 144 potential high-Tc superconductors with predicted Tc > 200 K and superior thermodynamic stability at 200 GPa. Among these, 129 compounds spanning 27 novel structural prototypes are reported for the first time, representing a significant expansion of the known structural landscape for hydride superconductors. This work not only greatly expands the known repertoire of high-Tc hydride superconductors but also establishes a scalable and efficient methodology for navigating the complex landscape of multinary hydrides.
Determining crystal structures from experimental powder X-ray diffraction data remains challenging because peak overlap, preferred orientation, and impurity phases obscure atomic arrangements. We present RealPXRD-Solver, a generative model trained on 6,250,238 theoretical structures with experiment-mimicking augmentations and a universal encoder of d-spacing--intensity fingerprints, enabling both lattice-conditioned and lattice-free inference. RealPXRD-Solver reaches a 98.3% Top-20 match rate on a 10,000-structure theoretical benchmark and achieves Top-1/Top-20 accuracies of 77.9%/91.9% on CNRS and 78.8%/92.9% on RRUFF experimental datasets, and it solved 39 previously unreported Powder Diffraction File entries.
To advance the computational simulation of cellular life, we propose a virtual yeast, an artificial intelligence (AI)-driven agent that models eukaryotic cellular behaviours by integrating multimodal biological data, mechanistic reasoning and active experimentation using Saccharomyces cerevisiae as a genetically tractable and data-rich model system. Cellular complexity is decomposed into eight function-centred modules, spanning genetic, metabolic and structural systems, each realized as a domain-specific AI tool coordinated through a large language model-based orchestration layer. Built on three data pillars, namely, mechanistic knowledge, subcellular architecture and dynamic states, the system integrates representation learning and generative modelling within a closed-loop learning pipeline that autonomously designs and executes experiments. The virtual yeast serves as both a conceptual and an operational platform to optimize biosynthetic pathways, support the generation and prioritization of hypotheses across diverse cellular processes, and accelerate target discovery. By coupling biological realism with autonomous AI reasoning, the virtual yeast establishes a generalizable blueprint for constructing virtual eukaryotic cells and advancing synthetic biology.
Data-centric training has emerged as a promising direction for improving large language models (LLMs) by optimizing not only model parameters but also the selection, composition, and weighting of training data during optimization. However, existing approaches to data selection, data mixture optimization, and data reweighting are often developed in isolated codebases with inconsistent interfaces, hindering reproducibility, fair comparison, and practical integration. In this paper, we present DataFlex, a unified data-centric dynamic training framework built upon LLaMA-Factory. DataFlex supports three major paradigms of dynamic data optimization: sample selection, domain mixture adjustment, and sample reweighting, while remaining fully compatible with the original training workflow. It provides extensible trainer abstractions and modular components, enabling a drop-in replacement for standard LLM training, and unifies key model-dependent operations such as embedding extraction, inference, and gradient computation, with support for large-scale settings including DeepSpeed ZeRO-3. We conduct comprehensive experiments across multiple data-centric methods. Dynamic data selection consistently outperforms static full-data training on MMLU across both Mistral-7B and Llama-3.2-3B. For data mixture, DoReMi and ODM improve both MMLU accuracy and corpus-level perplexity over default proportions when pretraining Qwen2.5-1.5B on SlimPajama at 6B and 30B token scales. DataFlex also achieves consistent runtime improvements over original implementations. These results demonstrate that DataFlex provides an effective, efficient, and reproducible infrastructure for data-centric dynamic training of LLMs.
We present Innovator-VL, a scientific multimodal large language model designed to advance understanding and reasoning across diverse scientific domains while maintaining excellent performance on general vision tasks. Contrary to the trend of relying on massive domain-specific pretraining and opaque pipelines, our work demonstrates that principled training design and transparent methodology can yield strong scientific intelligence with substantially reduced data requirements. (i) First, we provide a fully transparent, end-to-end reproducible training pipeline, covering data collection, cleaning, preprocessing, supervised fine-tuning, reinforcement learning, and evaluation, along with detailed optimization recipes. This facilitates systematic extension by the community. (ii) Second, Innovator-VL exhibits remarkable data efficiency, achieving competitive performance on various scientific tasks using fewer than five million curated samples without large-scale pretraining. These results highlight that effective reasoning can be achieved through principled data selection rather than indiscriminate scaling. (iii) Third, Innovator-VL demonstrates strong generalization, achieving competitive performance on general vision, multimodal reasoning, and scientific benchmarks. This indicates that scientific alignment can be integrated into a unified model without compromising general-purpose capabilities. Our practices suggest that efficient, reproducible, and high-performing scientific multimodal models can be built even without large-scale data, providing a practical foundation for future research.
Machine Learning Interatomic Potentials (MLIPs) enable accurate large-scale atomistic simulations, yet improving their expressive capacity efficiently remains challenging. Here we systematically investigate Mixture-of-Experts (MoE) and Mixture-of-Linear-Experts (MoLE) architectures within the DPA3 framework for MLIPs and analyze the effects of routing strategies and expert designs. We show that sparse activation combined with shared experts yields substantial performance gains, and that nonlinear MoE formulations outperform MoLE when shared experts are present, underscoring the importance of nonlinear expert specialization. Furthermore, element-wise routing consistently surpasses configuration-level routing, while global MoE routing often leads to numerical instability. The resulting element-wise MoE model consistently outperforms all DPA3-based baselines across the OMol25, OMat24, and OC20M benchmarks. Analysis of routing patterns reveals chemically interpretable expert specialization aligned with periodic-table trends, indicating that the model effectively captures element-specific chemical characteristics for precise interatomic modeling.
An efficient, reliable, and interpretable global solution method, the Deep learning-based algorithm for Heterogeneous Agent Models (DeepHAM), is proposed for solving high dimensional heterogeneous agent models with aggregate shocks. The state distribution is approximately represented by a set of optimal generalized moments. Deep neural networks are used to approximate the value and policy functions, and the objective is optimized over directly simulated paths. In addition to being an accurate global solver, this method has three additional features. First, it is computationally efficient in solving complex heterogeneous agent models, and it does not suffer from the curse of dimensionality. Second, it provides a general and interpretable representation of the distribution over individual states, which is crucial in addressing the classical question of whether and how heterogeneity matters in macroeconomics. Third, it solves the constrained efficiency problem as easily as it solves the competitive equilibrium, which opens up new possibilities for studying optimal monetary and fiscal policies in heterogeneous agent models with aggregate shocks.
Domain-specific intelligence demands specialized knowledge and sophisticated reasoning for problem-solving, posing significant challenges for large language models (LLMs) that struggle with knowledge hallucination and inadequate reasoning capabilities under constrained parameter budgets. Inspired by Bloom's Taxonomy in educational theory, we propose Retrieval-Augmented Reasoning Modeling (RARE), a novel paradigm that decouples knowledge storage from reasoning optimization. RARE externalizes domain knowledge to retrievable sources and internalizes domain-specific reasoning patterns during training. Specifically, by injecting retrieved knowledge into training prompts, RARE transforms learning objectives from rote memorization to contextualized reasoning application. It enables models to bypass parameter-intensive memorization and prioritize the development of higher-order cognitive processes. Our experiments demonstrate that lightweight RARE-trained models (e.g., Llama-3.1-8B) could achieve state-of-the-art performance, surpassing retrieval-augmented GPT-4 and Deepseek-R1 distilled counterparts. RARE establishes a paradigm shift where maintainable external knowledge bases synergize with compact, reasoning-optimized models, collectively driving more scalable domain-specific intelligence. Repo: https://github.com/Open-DataFlow/RARE
The convergence of large language models and agents is catalyzing a new era of scientific discovery: Agentic Science. While the scientific method is inherently iterative, existing agent frameworks are predominantly static, narrowly scoped, and lack the capacity to learn from trial and error. To bridge this gap, we present EvoMaster, a foundational evolving agent framework engineered specifically for Agentic Science at Scale. Driven by the core principle of continuous self-evolution, EvoMaster empowers agents to iteratively refine hypotheses, self-critique, and progressively accumulate knowledge across experimental cycles, faithfully mirroring human scientific inquiry. Crucially, as a domain-agnostic base harness, EvoMaster is exceptionally easy to scale up – enabling developers to build and deploy highly capable, self-evolving scientific agents for arbitrary disciplines in approximately 100 lines of code. Built upon EvoMaster, we incubated the SciMaster ecosystem across domains such as machine learning, physics, and general science. Evaluations on four authoritative benchmarks (Humanity's Last Exam, MLE-Bench Lite, BrowseComp, and FrontierScience) demonstrate that EvoMaster achieves state-of-the-art scores of 41.1
Large Atomistic Models (LAMs) have undergone remarkable progress recently, emerging as universal or fundamental representations of the potential energy surface defined by the first-principles calculations of atomistic systems. However, our understanding of the extent to which these models achieve true universality, as well as their comparative performance across different models, remains limited. This gap is largely due to the lack of comprehensive benchmarks capable of evaluating the effectiveness of LAMs as approximations to the universal potential energy surface. In this study, we introduce LAMBench, a benchmarking system designed to evaluate LAMs in terms of their generalizability, adaptability, and applicability. These attributes are crucial for deploying LAMs as ready-to-use tools across a diverse array of scientific discovery contexts. We benchmark ten state-of-the-art LAMs released prior to August 1, 2025, using LAMBench. Our findings reveal a significant gap between the current LAMs and the ideal universal potential energy surface. They also highlight the need for incorporating cross-domain training data, supporting multi-fidelity modeling, and ensuring the models' conservativeness and differentiability. As a dynamic and extensible platform, LAMBench is intended to continuously evolve, thereby facilitating the development of robust and generalizable LAMs capable of significantly advancing scientific research. The LAMBench code is open-sourced at https://github.com/deepmodeling/lambench, and an interactive leaderboard is available at https://www.aissquare.com/openlam?tab=Benchmark.