The rapid evolution of Large Language Models has catalyzed a surge in scientific idea production, yet this leap has not been accompanied by a matching advance in idea evaluation. The fundamental nature of scientific evaluation needs knowledgeable grounding, collective deliberation, and multi-criteria decision-making. However, existing idea evaluation methods often suffer from narrow knowledge horizons, flattened evaluation dimensions, and the inherent bias in LLM-as-a-Judge. To address these, we regard idea evaluation as a knowledge-grounded, multi-perspective reasoning problem and introduce , a deep innovation evaluation framework designed to emulate human-level idea assessment. We apply a heterogeneous deep knowledge search engine that retrieves and grounds dynamic evidence from diverse online sources. We further achieve review consensus with an innovation review board containing reviewers with distinct academic backgrounds, enabling a multi-dimensional decoupled evaluation across multiple metrics. We construct comprehensive datasets derived from authoritative peer-reviewed submissions to benchmark InnoEval. Experiments demonstrate that InnoEval can consistently outperform baselines in point-wise, pair-wise, and group-wise evaluation tasks, exhibiting judgment patterns and consensus highly aligned with human experts.
Artificial intelligence is rapidly entering the core workflows of scientific research. Yet reliable scientific reasoning requires access to accumulated scientific knowledge with sufficient breadth, depth, and standardization. Current AI scientists typically assemble scientific knowledge through workflow- and discipline-specific pipelines, which provide incomplete coverage, leave relations implicit, and make knowledge acquisition pathways fragmented. Here we present SciAtlas, a shared, machine-actionable cross-disciplinary scholarly knowledge infrastructure that integrates evidential, conceptual, disciplinary, expertise, and normative layers under a shared schema. SciAtlas further achieves a unified neuro-symbolic retrieval mechanism that grounds heterogeneous research objects, propagates relevance across the scholarly topology, and projects the resulting relevance field into the context required by each scientific workflow. Across three representative workflows, SciAtlas broadens trajectory reconstruction by recovering overlooked research branches, deepens opportunity discovery by uncovering underexplored bottlenecks and cross-domain insights, and strengthens innovation assessment by integrating evidence, expertise, and evaluation signals. Across three representative workflows, SciAtlas broadens trajectory reconstruction by recovering overlooked stages and branches, deepens opportunity discovery by uncovering underexplored bottlenecks and cross-domain connections, and standardizes innovation assessment by integrating evidence, expertise and evaluation signals. Extensive evaluations validate the foundational capabilities underpinning it as reusable knowledge infrastructure for knowledge-intensive scientific research.
The rapid growth of AI agent ecosystems is transforming how complex tasks are delegated and executed, creating a new challenge of identifying suitable agents for a given task. Unlike traditional tools, agent capabilities are often compositional and execution-dependent, making them difficult to assess from textual descriptions alone. However, existing research and benchmarks typically assume well-specified functionalities, controlled candidate pools, or only executable task queries, leaving realistic agent search scenarios insufficiently studied. We introduce AgentSearchBench, a large-scale benchmark for agent search in the wild, built from nearly 10,000 real-world agents across multiple providers. The benchmark formalizes agent search as retrieval and reranking problems under both executable task queries and high-level task descriptions, and evaluates relevance using execution-grounded performance signals. Experiments reveal a consistent gap between semantic similarity and actual agent performance, exposing the limitations of description-based retrieval and reranking methods. We further show that lightweight behavioral signals, including execution-aware probing, can substantially improve ranking quality, highlighting the importance of incorporating execution signals into agent discovery. Our code is available at https://github.com/Bingo-W/AgentSearchBench.
Detecting collaborative and problem-solving behaviours from digital traces to interpret students’ collaborative problem solving (CPS) competency is a long-term goal in the Artificial Intelligence in Education field. Although multimodal data and advanced models are argued to have the potential to detect complex CPS behaviours, empirical evidence on their value remains limited with some contrasting evidence. In this study, we investigated the potential of multimodal data to improve model performance in diagnosing secondary school students’ CPS subskills in authentic educational settings. In particular, text and acoustic embeddings from audio data were used in a multimodal classification model for CPS diagnosis. Both unimodal and multimodal transformer-based models outperformed traditional models in detecting CPS classes. Although the inclusion of multimodality did not improve the performance of traditional unimodal models, its integration into transformer-based models demonstrated improved performance for diagnosing social-cognitive CPS classes compared to unimodal transformer-based models. Based on the results, the paper argues that the value of multimodality and the selection of a particular modelling technique are limited to certain types of CPS indicators, affected by the complexity of the labels, and dependent on the composition of indicators in the dataset. We conclude by discussing the required nuance when considering the value of large language models and multimodality in automated CPS diagnosis, highlighting the need for human-AI complementarity, and proposing the exploration of relevant model architectures and techniques to improve CPS diagnosis.
Personalised text generation is essential for user-centric information systems, yet most evaluation methods overlook the individuality of users. We introduce PREF, a Personalised Reference-free Evaluation Framework that jointly measures general output quality and user-specific alignment without requiring gold personalised references. PREF operates in a three-step pipeline: (1) a coverage stage uses a large language model (LLM) to generate a comprehensive, query-specific guideline covering universal criteria such as factuality, coherence, and completeness; (2) a preference stage re-ranks and selectively augments these factors using the target user's profile, stated or inferred preferences, and context, producing a personalised evaluation rubric; and (3) a scoring stage applies an LLM judge to rate candidate answers against this rubric, ensuring baseline adequacy while capturing subjective priorities. This separation of coverage from preference improves robustness, transparency, and reusability, and allows smaller models to approximate the personalised quality of larger ones. Experiments on the PrefEval benchmark, including implicit preference-following tasks, show that PREF achieves higher accuracy, better calibration, and closer alignment with human judgments than strong baselines. By enabling scalable, interpretable, and user-aligned evaluation, PREF lays the groundwork for more reliable assessment and development of personalised language generation systems.
Conventional meta-learning typically involves adapting all meta-knowledge to specific tasks, which incurs high computational costs due to the adaption process. To address this limitation, we introduce a more efficient gradient-based meta-learning framework called Uncertainty-Aware Prompted Meta-Learning (UAPML). Instead of adapting the entire meta-knowledge, we introduce a meta-knowledge extraction paradigm inspired by the success of large language models. In this paradigm, we freeze the model backbone and employ task-specific prompts to extract meta-knowledge for few-shot tasks. To construct the task-specific prompts, a learnable Bayesian meta-prompt is employed to provide an ideal initialization. Through theoretical analysis, we demonstrate that the posterior uncertainty of the Bayesian meta-prompt aligns with that of the task-specific prompt, which can be used to modulate the construction of task-specific prompts. Accordingly, we propose two ways, i.e., the soft and hard way, to automatically construct task-specific prompts from the meta-prompt when dealing with new tasks. Experimental results demonstrate the efficiency of the meta-knowledge extraction paradigm and highlight the significantly reduced computational cost achieved by our UAPML framework without the degradation of performance.
Large language models (LLMs) have had a transformative impact on a variety of scientific tasks across disciplines such as biology, chemistry, medicine, and physics. However, ensuring the safety alignment of these models in scientific research remains an underexplored area, with existing benchmarks primarily focus on textual content and overlooking key scientific representations such as molecular, protein, and genomic languages. Moreover, the safety mechanisms of LLMs in scientific tasks are insufficiently studied. To address these limitations, we introduce SciSafeEval, a comprehensive benchmark designed to evaluate the safety alignment of LLMs across a range of scientific tasks. SciSafeEval spans multiple scientific languages - including textual, molecular, protein, and genomic - and covers a wide range of scientific domains. We evaluate LLMs in zero-shot, few-shot and chain-of-thought settings, and introduce a 'jailbreak' enhancement feature that challenges LLMs equipped with safety guardrails, rigorously testing their defenses against malicious intention. Our benchmark surpasses existing safety datasets in both scale and scope, providing a robust platform for assessing the safety and performance of LLMs in scientific contexts. This work aims to facilitate the responsible development and deployment of LLMs, promoting alignment with safety and ethical standards in scientific research.
Directed evolution is a widely-used strategy of protein engineering to improve protein function via mimicking natural mutation and selection. Machine learning-assisted directed evolution (MLDE) approaches aim to learn a fitness predictor, thereby efficiently searching for optimal mutants within the vast combinatorial mutation space. Since annotating mutants is both costly and labor-intensive, how to efficiently sample and utilize informative protein mutants to train the predictor is a critical problem in MLDE. Previous MLDE works just simply utilized pre-trained protein language models (PPLMs) for sampling without tailoring to the specific target protein of interest, which has not fully exploited the potential of PPLMs. In this work, we propose a novel method, the Actively-Finetuned Protein language model for Directed Evolution(AFP-DE), which leverages PPLMs to actively sample and fine-tune themselves, continuously improving the model’s sampling and overall performance through iterations, to achieve efficient directed protein evolution. Extensive experiments have shown the effectiveness of our method in generating optimal mutants with minimal annotation effort, outperforming previous works even with fewer annotated mutants, making it budget-friendly for biological experiments.
Molecular property is usually observed with a limited number of samples, and researchers have considered property prediction as a few-shot problem. One important fact that has been ignored by prior works is that each molecule can be recorded with several different properties simultaneously. To effectively utilize many-to-many correlations of molecules and properties, we propose a Graph Sampling-based Meta-learning (GS-Meta) framework for few-shot molecular property prediction. First, we construct a Molecule-Property relation Graph (MPG): molecule and properties are nodes, while property labels decide edges. Then, to utilize the topological information of MPG, we reformulate an episode in meta-learning as a subgraph of the MPG, containing a target property node, molecule nodes, and auxiliary property nodes. Third, as episodes in the form of subgraphs are no longer independent of each other, we propose to schedule the subgraph sampling process with a contrastive loss function, which considers the consistency and discrimination of subgraphs. Extensive experiments on 5 commonly-used benchmarks show GS-Meta consistently outperforms state-of-the-art methods by 5.71%-6.93% in ROC-AUC and verify the effectiveness of each proposed module. Our code is available at https://github.com/HICAI-ZJU/GS-Meta.
Personalized product search that provides users with customized search services is an important task for e-commerce platforms. This task remains a challenge when inferring users’ preferences from few records or even no records, which is also known as the few-shot or zero-shot learning problem. In this paper, we propose a Bayesian Online Meta-Learning Model (BOML), which transfers meta-knowledge, from the inference for other users’ preferences, to help to infer the current user’s interest behind her/his few or even no historical records. To extract meta-knowledge from various inference patterns, our model constructs a mixture of meta-knowledge and transfers the corresponding meta-knowledge to the specific user according to her/his records. Based on the meta-knowledge learned from other similar inferences, our proposed model searches a ranked list of products to meet users’ personalized query intents for those with few search records (i.e., few-shot learning problem) or even no search records (i.e., zero-shot learning problem). Under the records arriving sequentially setting, we propose an online variational inference algorithm to update meta-knowledge over time. Experimental results demonstrate that our proposed BOML outperforms state-of-the-art algorithms.
Product search has been receiving significant attention with the development of e-commerce . Existing works recognize the importance of personalization and focus on personalized product search. While these works have confirmed that personalization can improve the performance of product search, they all ignore the few-shot learning problems caused by personalization. Under the few-shot setting, personalized methods may suffer from the data-hungry issue. In this paper, we explore the data-hungry issue in personalized product search. We find that data-hungry issue exists under the few-shot setting caused by personalization, and degrades the performance under the few-shot setting when the input query consists of diverse intents. Furthermore, we illustrate that with such a data-hungry issue, the returned search results tend to be close to the products the user purchases most often, or the products the most users purchase in the market given the same query. The result in the further experiment confirms our conclusions.
The harmonic correction (HC) is one of the key quantities when using residual terrain modelling (RTM) for high-frequency gravity field modelling. In the RTM technique, high-frequency topographic gravitational signals are obtained through removing gravitational effects of a long-wavelength reference surface, e.g., MERIT2160. There might be points located below the reference surface. In such cases, the RTM gravity field is calculated in the non-harmonic condition, HC is therefore required. Over past decades, though various methods have been proposed to handle the HC issue for the RTM technique, most of them were focused on the HC for RTM gravity anomaly rather than for other gravity functionals, such as RTM geoid height. In practice, the HC for RTM geoid height was generally assumed to be negligible, but a detailed quantification was missing for present-day RTM computations. This might cause large errors in the regional geoid determination over rugged areas. In this study, we derive HC expressions for the RTM geoid height in the framework of the classical condensation method. The HC terms are derived under four different assumptions separately: residual masses approximated by an unlimited Bouguer plate, residual masses approximated by a limited Bouguer plate which overcomes the mass inconsistency effect, residual masses approximated by a Bouguer shell which overcomes the effect of planar approximation, and residual masses approximated by a limited Bouguer shell which overcomes the errors induced by both planar approximation and mass-inconsistency. The errors due to various approximations in HC terms are investigated through comparison among various terms. Besides, HC terms are computed using an expansion up to degree and order 2159. Our results show that HC for RTM geoid height is less 1 mm and could be ignored over $$\sim 99$$ % of continental areas, but be of great significance for regional geoid determination over mountain areas, e.g., more than 10 cm effect over very rugged areas. The validation through comparison with terrestrial measurements and a baseline solution of the RTM technique proves that the HC terms provided in this study can improve the accuracy of RTM geoid heights and are expected to be useful for applications of the RTM technique in regional and global gravity field modelling.
We develop high-order Galerkin methods with graded meshes for solving the two-dimensional reaction-diffusion problem on a rectangle. With the help of the comparison principle, we establish upper bounds for high order partial derivatives of an arbitrary order of its exact solution. According to prior information of the high order partial derivatives of the solution, we design both implicit and explicit graded meshes which lead to numerical solutions of the problem having an optimal convergence order. Numerical experiments are presented to confirm the theoretical estimate and to demonstrate the outperformance of the proposed meshes over the Shishkin mesh.
In this paper we design a modified version of FVCOM by adopting hybrid meshes and shifting the placement of velocity variables from the centroids of elements to the middle points of edges. A simplified version of geostrophic equation is solved to test the new scheme, and illustrates a nearly uniform error distribution.
The first-order and higher-order derivatives of a function can be viewed as the solutions of Volterra integral equations of the first kind. In this paper we propose a fast multiscale solver for the numerical solution of the Tikhonov regularization of the Volterra equations. In association with the special form of the kernels, the matrices resulting from the discretization by multiscale bases are sparse. Moreover, they can be truncated using proper strategies with only a minor loss of accuracy. In the best case, the number of nonzero entries of the truncated matrices is linear with respect to the dimensions of the matrices. The accuracy of the solution from the solver is analysed theoretically and verified by numerical experiments.
Shuffled belief propagation (SBP), as a sequential belief propagation (BP) algorithm, speeds up the convergence of BP decoding, and maintains the least complexity of flooding BP. However, its performance is remarkably inferior to informed dynamic scheduling (IDS) BP algorithms. The authors design an informed dynamic location method, based on the residuals of variable node log-likelihood ratio values, to reorder variable nodes of SBP to be updated. The location method significantly accelerates the convergence of SBP algorithm from two aspects: the unstable variable node with the largest residual to be updated first, and selecting the largest residual locally. Simulation results show that the proposed algorithm performs nearly the same as the best performance of IDS BP algorithms, and behaves prominently at high signal-to-noise ratios.
A fast multilevel augmentation method (MAM) was proposed recently by the same authors for solving a class of nonlinear boundary integral equations. In this paper, we develop accelerated quadrature formulas for computing the integrals involved in the MAM and approximate iteration for solving the resulting nonlinear system. Specifically, we employ a product integration scheme for computing the singular integrals which appear in the matrices involved in the MAM and introduce an approximation technique in the Newton iteration for solving the resulting nonlinear systems to avoid repeated computation in generating their Jacobian matrices. The use of these two techniques results in a modified MAM which speeds up its computation. We show that the modified MAM preserves the optimal convergence order of the original one while reducing computational costs. Numerical results are presented to demonstrate the approximation accuracy and computational efficiency of the proposed modified MAM, with a comparison to those of the original one and a known algorithm of Atkinson and Chandler.
In this paper we supplement matrix truncation strategies with the multilevel augmentation methods for solving the reformulated Hammerstein equations. The resulting numerical solutions have nearly optimal convergence order with linear order computational complexity up to a logarithmic factor with respect to the dimension of the discretization subspace. Numerical experiments on one and two dimensional equations illustrate that our algorithm gains remarkably high efficiency without losing accuracy.
We propose a fast algorithm for the solution of the nonlinear boundary integral equation resulting from a reformulation of a boundary value problem of the Laplace equation with nonlinear boundary conditions. The fast algorithm is developed by using the multilevel augmentation method (introduced recently by Chen, Wu, and Xu for general nonlinear integral equations), in conjunction with a matrix truncation strategy, and an error control technique of numerical integrations for integrals appeared in the process of solving the equation. We prove that the proposed algorithm has an optimal convergence order (up to a logarithmic factor) and a nearly linear computational complexity order (measured in the number of multiplications and functional evaluations). Numerical experiments are presented to demonstrate its approximation accuracy and computational efficiency, verifying the theoretical estimates, and to compare performance of the proposed algorithm with that of the Atkinson and Chandler algorithm.
Based on sparse grid multiscale piecewise polynomial bases, we develop a fast collocation method for solving Fredholm integral equations of the second kind with stochastic loading terms. It is proved that the proposed method preserves the optimal rate of convergence and has linear (up to a logarithmic factor) computational complexity. The reduction in computational costs come from three aspects: the adoption of the collocation principle, the use of the sparse grid for the solution space, and the employment of a truncation strategy for the coefficient matrix of the resulting linear system. Numerical experiments confirm the theoretical results on approximation accuracy and computational efficiency of the proposed method.