Large language models (LLMs) can generate syntactically valid optimization programs, yet often struggle to reliably choose an effective modeling strategy, leading to incorrect formulations and inefficient solver behavior. We propose , a strategy-aware framework that makes explicit in both data construction and post-training. SAGE builds a solver-verified multi-strategy dataset and trains a student model with supervised fine-tuning followed by Segment-Weighted GRPO using a composite reward over format compliance, correctness, and solver efficiency. Across eight benchmarks spanning synthetic and real-world settings, SAGE improves average pass@1 from 72.7 to 80.3 over the strongest open-source baseline. With multiple generations, SAGE discovers more distinct correct formulations and improves component-level diversity at pass@16 by 19-29%. At the largest scale, SAGE produces more compact constraint systems with 14.2% fewer constraints than the baseline, consistent with solver-efficient modeling. Overall, these results show that making explicit improves automated optimization modeling. Code is available at https://anonymous.4open.science/r/SAGE-F25B/.
NOx and VOCs emitted from flue gases in nonelectric industries processes constitute major atmospheric pollutants in China. The simultaneous removal of NOx and toluene using selective catalytic reduction (SCR) systems presents a great challenge. In this work, Co-doped MnPO/Ti catalysts (Co-MnPO/Ti) are synthesized for the simultaneous removal of two pollutants at low temperature. The results demonstrate that Co(0.5)-MnPO/Ti presents the highest catalytic activity, which achieved simultaneous conversions exceeding 80
Manganese-based mullite catalysts have attracted substantial attention owing to their inherent redox properties and structural stability. Nevertheless, reconciling superior catalytic efficiency with sustained durability during the oxidation of volatile organic compounds (VOCs) remains a persistent obstacle. In this study, a Mn3O4–YMn2O5 (Mn3O4–YMO) composite catalyst was synthesized utilizing a sol–gel method based on the interfacial synergy between Mn3O4 and YMn2O5, and its catalytic efficiency of toluene was systematically evaluated. The toluene oxidation efficiency of the optimal catalyst (1Mn3O4–YMO) was 48% higher than that of YMO at 300°C, and the catalyst maintained excellent stability. Various analyses confirmed the uniform coexistence of Mn3O4 and YMn2O5 phases, besides, the strong interfacial interaction between Mn3O4 and YMO promoted the migration of lattice oxygen, accompanied by an increase in reactive oxygen species, which consequently enhanced the low-temperature redox performance and the adsorption capacity for toluene and O2. Furthermore, the reaction mechanism of toluene oxidation over 1Mn3O4-YMO was investigated by in-situ diffuse reflectance infrared fourier transform spectroscopy (in-situ DRIFTS), results revealed that the oxidation of toluene proceeded via the following pathway: toluene → benzyl alcohol → benzaldehyde → benzoate → maleic anhydride → CO2 + H2O. This study would provide new insights into the design of highly efficient mullite–spinel composite catalysts for VOCs oxidation.
With the rapid advancement of Large Language Models (LLMs), the safety of LLMs has been a critical concern requiring precise assessment. Current benchmarks primarily concentrate on single-turn dialogues or a single jailbreak attack method to assess the safety. Additionally, these benchmarks have not taken into account the LLM's capability of identifying and handling unsafe information in detail. To address these issues, we propose a fine-grained benchmark SafeDialBench for evaluating the safety of LLMs across various jailbreak attacks in multi-turn dialogues. Specifically, we design a two-tier hierarchical safety taxonomy that considers 6 safety dimensions and generates more than 4000 multi-turn dialogues in both Chinese and English under 22 dialogue scenarios. We employ 7 jailbreak attack strategies, such as reference attack and purpose reverse, to enhance the dataset quality for dialogue generation. Notably, we construct an innovative assessment framework of LLMs, measuring capabilities in detecting, and handling unsafe information and maintaining consistency when facing jailbreak attacks. Experimental results across 17 LLMs reveal that Yi-34B-Chat and GLM4-9B-Chat demonstrate superior safety performance, while Llama3.1-8B-Instruct and o3-mini exhibit safety vulnerabilities.
MnOx-based catalysts have attracted considerable attention in the catalytic O3 decomposition owing to the high efficiency. Challenges of catalyst deactivation in high-humidity still remains. Herein, Ag-modified Mn-B catalysts (xAg/Mn-B) were prepared via the assistance of Mn-BTC and used for O3 decomposition. The catalytic activity evaluation results demonstrated that the modification of Ag significantly promoted the O3 decomposition efficiency of Mn-B catalysts, the conversion of 40 ppm O3 over the optimum sample (6Ag/Mn-B) reached 93% at relative humidity of 80%, its catalytic activity was 60% higher than that of Mn-B catalyst. The higher catalytic activity of 6Ag/Mn-B owing to its superior redox properties and the mobility of surface oxygen. Meanwhile, the mechanisms of water resistance over 6Ag/Mn-B and Mn-B catalysts were revealed by in-situ DRIFTS. Ag exists on the catalyst surface as highly dispersed AgOx, and water adsorption and O3 decomposition occur at different positions. Besides, the modification of Ag was able to inhibit the accumulation of intermediate product O22- from O3 decomposition on the catalyst surface in high humidity. Finally, 6Ag/Mn-B showed a long-term stable O3 conversion efficiency.
Large reasoning models (LRMs) have attracted much attention due to their exceptional performance. However, their performance mainly stems from thinking, a long Chain of Thought (CoT), which significantly increase computational overhead. To address this overthinking problem, existing work focuses on using reinforcement learning (RL) to train hybrid reasoning models that automatically decide whether to engage in thinking or not based on the complexity of the query. Unfortunately, using RL will suffer the the reward hacking problem, e.g., the model engages in thinking but is judged as not doing so, resulting in incorrect rewards.To mitigate this problem, existing works either employ supervised fine-tuning (SFT), which incurs high computational costs, or enforce uniform token limits on non-thinking responses, which yields limited mitigation of the problem.In this paper, we propose Thinking-Based Non-Thinking (TNT). It does not employ SFT, and sets different maximum token usage for responses not using thinking across various queries by leveraging information from the solution component of the responses using thinking. Experiments on five mathematical benchmarks demonstrate that TNT reduces token usage by around 50\\%$ compared to DeepSeek-R1-Distill-Qwen-1.5B/7B and DeepScaleR-1.5B, while significantly improving accuracy. In fact, TNT achieves the optimal trade-off between accuracy and efficiency among all tested methods. Additionally, the probability of reward hacking problem in TNT’s responses, which are classified as not using thinking, remains below $10\\%$ across all tested datasets.
Deep Research Agents (DRAs) aim to solve complex, long-horizon research tasks involving planning, retrieval, multimodal understanding, and report generation, yet their evaluation remains challenging due to dynamic web environments and ambiguous task definitions. We propose DR^3-Eval, a realistic and reproducible benchmark for evaluating deep research agents on multimodal, multi-file report generation. DR^3-Eval is constructed from authentic user-provided materials and paired with a per-task static research sandbox corpus that simulates open-web complexity while remaining fully verifiable, containing supportive documents, distractors, and noise. Moreover, we introduce a multi-dimensional evaluation framework measuring Information Recall, Factual Accuracy, Citation Coverage, Instruction Following, and Depth Quality, and validate its alignment with human judgments. Experiments with our developed multi-agent system DR^3-Agent based on multiple state-of-the-art language models demonstrate that DR^3-Eval is highly challenging and reveals critical failure modes in retrieval robustness and hallucination control. Our code and data are publicly available.
Herein, a strategy for combining vanadium phosphorous oxide (VPO) with Mn-based catalysts was proposed and the inhibiting mechanism of N2O formation in low-temperature selective catalytic reduction of NH3 (NH3-SCR) was investigated. The obtained composite catalysts (MnVPO) exhibited exceptional N-2 selectivity (approaching 100%) at 120-270 degrees C. 1.5MnVPO showed the highest activity, with > 90% NOx conversion at 210-270 degrees C, outperforming VPO and MnOx. The composite catalysts introduced the beta-VOPO4 phase while retaining the beta-MnO2 phase. X-ray photoelectron spectra (XPS) and hydrogen temperature-programmed reduction (H-2-TPR) results indicated that the composite catalysts reduced the proportion of lattice oxygen (1.5MnVPO: 41% vs MnOx: 76%) and redox capacity (H-2 consumption: 1.5MnVPO: 3.91 mmol & centerdot;g(-1) vs MnOx: (1)3.07 mmol & centerdot;g(-1)), inhibiting excessive NH3 oxidation. Moreover, temperature-programmed desorption (TPD) experiments revealed that 1.5MnVPO exhibited the largest areas for NH3 and NO desorption peaks in composite catalysts, corresponding to the best acid characteristics and NO adsorption capacity. In-situ diffuse reflectance Fourier transform spectroscopy (In-situ DRIFTS) studies indicated that the NH3-SCR reaction on 1.5MnVPO simultaneously proceeded via the Langmuir-Hinshelwood (L-H) and Eley-Rideal (E-R, dominant) mechanisms. Combined transient reaction and temperature-programmed surface reaction (TPSR) tests revealed that the main mechanism of N2O formation on MnOx was NH3 o(x)idation, with NO participating in and promoting this process. The above mechanisms were considerably inhibited on 1.5MnVPO. Density functional theory (DFT) calculations indicated that N2O formation on the MnVPO catalyst was inhibited by increasing barrier energy values during the deep dehydrogenation of key intermediates (HNO, NH2NO and NH4NO3).
With the rapid advancement of Large Language Models (LLMs), increasing complexity and widespread deployment of LLMs have exposed them to numerous security threats, necessitating a thorough examination of potential risks and associated mitigation methods. While existing studies focus on specific security aspects, a structured analysis integrating security considerations throughout the LLM lifecycle remains lacking. To address this gap, we propose a lifecycle-based perspective that systematically identifies security threats at each stage, connects them to appropriate defense mechanisms, and examines their broader societal and ethical impacts. Specifically, we first systematically analyze potential security threats across the LLM lifecycle, from training and deployment to operation phases. Building upon this analysis, we propose a security framework that aligns defensive strategies with identified threats at each lifecycle stage, encompassing data protection, model integrity, deployment security, and operational robustness. Furthermore, we investigate how these security challenges contribute to societal and ethical concerns, including privacy risks and potential misuse. Our analysis reveals the intricate relationships between technical vulnerabilities and their societal implications, suggesting directions for future research in LLM security. A curated list of related papers has been publicly available at a GitHub repository ( https://github.com/drivetosouth/Awesome-LLM-Safety ).
Although Mn-based catalysts exhibit excellent low-temperature NH3-SCR activity, their excessive oxidation ability inevitably induces NH3 over-oxidation and N2O formation. In this work, a phosphorus-modified MnOx/TiO2 catalyst was constructed by introducing Mn2P2O7 species through H3PO4 modification and employed for low-temperature NH3-SCR. The MnOx/TiO2 exhibited high NOx conversion, its N2 selectivity decreased from 100% to 60% with increasing temperature from 120 to 270 °C. In contrast, the MnPO/TiO2-Ar catalyst achieved over 90% NOx conversion at 180 °C while maintaining nearly 100% N2 selectivity throughout the investigated temperature range. Mechanistic investigations revealed that phosphorus modification optimized the Mn oxidation state distribution and reduced surface adsorbed oxygen species, thereby suppressing the excessive oxidation ability of MnOx. Meanwhile, the introduction of phosphorus significantly enriched both Brønsted and Lewis acid sites, promoting NH3 activation toward NH4+ and -NH2 intermediates while inhibiting the formation of highly reactive -NH species associated with N2O generation. Consequently, NH3 over-oxidation and undesired N2O formation pathways were effectively suppressed, enabling highly selective NH3-SCR performance.
We introduce JT-Safe-V2, a large language model designed to advance the safety and trustworthiness of foundation models, extending our previous JT-Safe model toward a more comprehensive safety-by-design paradigm. JT-Safe-V2 emphasizes the joint optimization of general intelligence and safety-by-design through several key innovations: enriching pre-training data with contextual world knowledge, high-certainty pre-training procedures, and safety strengthening post-training mechanisms for enterprise-oriented agentic capabilities. Building on these safety-enhanced foundation models, we propose Safe-MoMA (Safe Mixture of Models and Agents), a framework that enables traceable and efficient inference through the orchestrated deployment of multiple models and agents. Extensive evaluations demonstrate that JT-Safe-V2 achieves state-of-the-art performance across both general intelligence and safety benchmarks. Moreover, Safe-MoMA reduces inference costs by more than 30% compared to using the largest standalone model baseline while maintaining comparable performance. To facilitate future research on safety-by-design foundation models, we publicly release the post-trained JT-Safe-V2-35B model checkpoint.
A series of TiO2-supported phosphomolybdic heteropoly acid with vanadium substitution (PMoV/TiO2) was synthesized and tested for the selective catalytic reduction of NO with NH3 (NH3-SCR) at low temperature. The substitution of V could markedly enhance NH3-SCR performance. Among the prepared samples, PMoV3/TiO2 achieved over 95 % NO conversion across 180-300 degrees C, whereas the conversion of PMo/TiO2 exhibited merely 20 % at 180 degrees C. Furthermore, PMoV3/TiO2 exhibited better resistance of SO2 and H2O performance than PMo/TiO2. Characterization results showed that the V could substitute Mo in Keggin structure of HPMo and enhanced the interaction between V and Mo, which increased the content of Mo5+ and V4+. Moreover, the substitution of V enhanced the surface acidity of the catalyst, which improved the adsorption performance of NH3. Thus, the catalytic performance was improved with V substitution. Besides, in-situ DRIFTS results confirmed that V substitution was able to inhibit the adsorption of SO2 on the catalyst surface and improve the sulphur resistance of the catalyst.
The accumulation of nitrate in wastewater poses severe threats to aquatic ecosystems and public health, making efficient nitrate removal and resource recovery an urgent need in water purification. The electrocatalytic nitrate reduction reaction (NO3RR) offers a sustainable route to convert nitrate into valuable ammonia (NH3), simultaneously achieving contaminant separation and green NH3 synthesis. However, the eight-electron-nine-proton pathway makes hydrogenation highly dependent on the active hydrogen (H*). Pristine Co3O4 suffers from limited intrinsic active sites, sluggish charge transfer kinetics and insufficient H* generation, severely restricting NH3 selectivity. Herein, Mo-doped flower-like hierarchical Co3O4 (Mo-Co3O4) was fabricated via a solvothermal-calcination method. The effects of different Mo/Co molar ratios (1∶20, 1∶10, 1∶5) on structural, surface properties and NO3RR performance were systematically investigated. Results demonstrate that appropriate Mo doping markedly increases Co2+ active sites and induces abundant oxygen vacancies, while substantially enhancing in-situ H* generation, thereby achieving a bifunctional synergy of efficient NO3- activation and controllable H* supply. The optimized 0.1Mo-Co3O4 delivered an NH3 yield of 11.74 mg h-1 cm-2 and a Faradaic efficiency of 96.3% at -0.6 V vs. RHE, outperforming most reported Co3O4-based electrocatalysts. Mechanistic investigations confirm that Mo doping promotes water dissociation to accelerate H* generation; the resulting H* directly participates in hydrogenation of nitrogenous intermediates, enabling an efficient H* mediated NO3RR pathway. This work not only provides a high performance nonprecious metal electrocatalyst for nitrate-to-ammonia conversion, but also offers new insights for designing dual-functional catalysts that integrate substrate activation and hydrogenation reduction for wastewater purification and resource recovery.
Herein, focusing on the N2O formation pathways to gain a deeper understanding into the influence of SO2 on N2 selectivity in the low-temperature selective catalytic reduction (NH3-SCR) over MnV composite oxide catalysts. The NOx conversion rates of the catalysts exhibited a volcano-shaped relationship with the Mn/V molar ratio. The 2.0MnV catalyst with the best redox performance and the strongest surface acidity corresponded to the optimal NH3-SCR performance and the poorest N2 selectivity. The MnV catalyst consisted of Mn2O3 and MnV2O6 phases. The redox performance and NOx adsorption capacity of the SO2-poisoned 2.0MnV catalyst were significantly impaired, while the surface acidity and the Oads ratio decreased, corresponding to the inhibitions of NH3 oxidation and N2O formation. Transient reaction and TPSR experiments revealed that N2O formation on the 2.0MnV catalyst primarily followed via NH3 oxidation and NOx-involved NH3 oxidation pathways. NO partially counteracted the inhibitory effect of SO2 on the NH3 oxidation. Spectroscopic analysis revealed that the 2.0MnV catalyst did not significantly form N2O when undergoing the NH3-SCR reaction via the E-R and L-H mechanisms with SO2 presence at 180 °C. DFT calculations indicated that SO2 significantly increased the energy barriers required for the N2O formation via the “NH2NO pathway” at Lewis acid sites and the “NH4NO3 pathway” at Brønsted acid sites in Mn2O3(222) and MnV2O6(110) supercells. SO2 hindered the occurrence of key steps in the N2O formation, such as the dehydration of the NH4NO3 intermediate and the successive dehydrogenation of the NH2NO and NH2NO2 intermediates.
Alzheimer's disease (AD) is a progressive neurodegenerative disorder that places an increasing burden on patients, caregivers, and healthcare systems worldwide. Current disease-modifying therapies (DMTs) are limited by high costs, complex administration, and reliance on advanced biomarker infrastructure, highlighting the shortcomings of existing treatment paradigms. These limitations have sparked growing interest in gene- and nucleic acid-based interventions as upstream strategies to modify AD pathogenesis. Among these, small interfering RNA (siRNA) is especially compelling because it can be rationally programmed, directed at multiple molecular pathways, and paired with rapidly evolving delivery technologies. However, the clinical translation of siRNA therapies for AD is still constrained by challenges in brain-targeted delivery, safety, and sustained efficacy. In this review, we summarize current concepts in AD pathology, highlight recent clinical and translational advances, and critically assess emerging brain-targeted siRNA delivery platforms and their key bottlenecks. Within a precision-medicine framework, brain-targeted siRNA offers the possibility of aligning patient selection, molecular targets, and delivery strategies with biomarker-defined AD endotypes. We discuss both the therapeutic promise and the realistic limitations of siRNA-based approaches for AD, outline priorities for future development, and identify key gaps that must be addressed to enable meaningful clinical implementation.
Enhancing the ability of large language models (LLMs) to follow complex instructions is critical for their deployment in real-world applications. However, existing evaluation methods often oversimplify instruction complexity as a mere additive combination of atomic constraints, failing to adequately capture the high-dimensional complexity arising from the intricate interplay of content and format, logical workflow control, and real-world applications. This leads to a significant gap between current evaluation practices and practical demands. To bridge this gap, we introduce CCR-Bench, a novel benchmark designed to assess LLMs' adherence to complex instructions. CCR-Bench is characterized by: (1) deep entanglement of content and formatting requirements in task specifications; (2) instructions that involve intricate task decomposition, conditional reasoning, and procedural planning; and (3) evaluation samples derived entirely from real-world industrial scenarios. Extensive experiments on CCR-Bench demonstrate that even state-of-the-art models exhibit substantial performance deficiencies, clearly quantifying the gap between current LLM capabilities and the demands of realworld instruction understanding. We believe that CCR-Bench offers a more rigorous and realistic evaluation framework, advancing the development of LLMs toward the next generation of models capable of understanding and executing complex tasks in industrial applications.
Deep-research agents solve tasks through long trajectories of search, tool use, evidence inspection, and answer synthesis. Evaluation based on final answers shows whether an agent succeeds, but not which parts of the trajectory make the answer unreliable. We study span-level error localization for deep-research agents. We collect 2,790 real trajectories from two agent frameworks, three backbone models, and three benchmarks, convert raw logs into semantic spans, and annotate harmful error spans through LLM-assisted expert review. From these annotations, we build TELBench, a 1,000-instance benchmark for identifying error spans among normal exploration, failed searches, tentative hypotheses, and harmless noise. We further propose DRIFT, a claim-centric auditing framework that tracks agent claims, checks their support in trajectory evidence, and marks spans where unsupported or conflicting claims affect the answer path. Experiments across model families and auditing frameworks show that DRIFT improves span-level error localization and first-error accuracy by up to 30 percentage points. Our work provides a process-level view of reliability in deep-research agents.
Graph-based tasks in the zero-shot setting remain a significant challenge due to data scarcity and the inability of traditional Graph Neural Networks (GNNs) to generalize to unseen domains or label spaces. While recent advancements have transitioned toward leveraging Large Language Models (LLMs) as predictors to enhance GNNs, these methods often suffer from cross-modal alignment issues. A recent paradigm (i.e., Graph-R1) overcomes the aforementioned architectural dependencies by adopting a purely text-based format and utilizing LLM-based graph reasoning, showing improved zero-shot generalization. However, it employs a task-agnostic, one-size-fits-all subgraph extraction strategy, which inevitably introduces significant structural noise--irrelevant neighbors and edges--that distorts the LLMs' receptive field and leads to suboptimal predictions. To address this limitation, we introduce GraphSSR, a novel framework designed for adaptive subgraph extraction and denoising in zero-shot LLM-based graph reasoning. Specifically, we propose the SSR pipeline, which dynamically tailors subgraph extraction to specific contexts through a "Sample-Select-Reason" process, enabling the model to autonomously filter out task-irrelevant neighbors and overcome the one-size-fits-all issue. To internalize this capability, we develop SSR-SFT, a data synthesis strategy that generates high-quality SSR-style graph reasoning traces for supervised fine-tuning of LLMs. Furthermore, we propose SSR-RL, a two-stage reinforcement learning framework that explicitly regulates sampling and selection operations within the proposed SSR pipeline designed for adaptive subgraph denoising. By incorporating Authenticity-Reinforced and Denoising-Reinforced RL, we guide the model to achieve accurate predictions using parsimonious, denoised subgraphs for reasoning.
Resource Consumption Attacks (RCAs) have emerged as a significant threat to the deployment of Large Language Models (LLMs). With the integration of vision modalities, additional attack vectors exacerbate the risk of RCAs in large vision-language models (LVLMs). However, existing red-teaming studies have mainly overlooked visual inputs as a potential attack surface, resulting in insufficient mitigation strategies against RCAs in LVLMs. To address this gap, we propose RECITE (Resource Consumption Red-Teaming for LVLMs), the first approach for exploiting visual modalities to trigger unbounded RCAs red-teaming. First, we present Vision Guided Optimization, a fine-grained pixel-level optimization to obtain Output Recall Objective adversarial perturbations, which can induce repeating output. Then, we inject the perturbations into visual inputs, triggering unbounded generations to achieve the goal of RCAs. Empirical results demonstrate that RECITE increases service response latency by over 26 ↑, resulting in an additional 20% increase in GPU utilization and memory consumption. Our study reveals security vulnerabilities in LVLMs and establishes a red-teaming framework that can facilitate the development of future defenses against RCAs.