The application of formulas (e.g., physics formulas) is a fundamental human ability in solving numerical reasoning problems. Existing numerical reasoning datasets rarely explicitly state the formulas employed, as their questions often rely on implicit commonsense mathematical knowledge. To address this gap, we introduce FormulaReasoning, a new dataset specifically designed for formula-based numerical reasoning. It consists of 5,324 questions that require numerical calculations grounded in external physics formulas. We provide normalized, fine-grained annotations in both English and Chinese, including formula structures, parameter names, symbols, numerical values, and units-curated through extensive manual effort with LLM-assisted validation to ensure high quality. Additionally, we offer a consolidated formula database to serve as an external knowledge source. We analyze various reasoning approaches on FormulaReasoning, with emphasis on comparative evaluation of different architectural and methodological frameworks. Our assessment includes retrieval-augmented methods, approaches that decompose reasoning into formula generation, parameter extraction, and numerical calculation, as well as optimization techniques using preference data. We identify key challenges in formula-based numerical reasoning that require further investigation across different reasoning paradigms, highlighting opportunities for methodological advancement.
Joint logical-numerical reasoning remains a major challenge for language models, yet existing datasets rely on fixed rule sets and offer limited control over task complexity, constraining their generalizability for evaluation and training. We present LogiNumSynth, a flexible natural language problem synthesizer that synthesizes tasks requiring proficiency in joint logical reasoning (e.g., rule-based reasoning) and numerical reasoning (e.g., arithmetic computation). LogiNumSynth supports fine-grained control over reasoning world richness, logical reasoning depth, and the complexity of numerical computations, enabling flexible data synthesis across difficulty levels. We demonstrate three key contributions: (1) Synthesizer – synthesizing fully controllable joint reasoning tasks over natural language; (2) Evaluation Process Analysis – evaluating both process accuracy and answer accuracy; (3) Targeted Training – using synthesized data to enhance LLMs' reasoning performance. Experiments with multiple LLMs highlight persistent weaknesses in logical-numerical reasoning, showing that LogiNumSynth can serve as both a diagnostic tool and a source of targeted supervision for advancing integrated reasoning skills.
Dust heterogeneous chemistry substantially influences the atmosphere and has profound impacts on our environment and climate. Gas-particle partitioning on dust not only modifies species' chemical evolutions but also influences their deposition velocity. However, how does dust heterogeneous chemistry impacts acid deposition remains unexplored. In this research, we integrated dust photocatalytic mechanism into GEOS-Chem to assess its impact on sulfur removal across various regions and time spans in China. We find the photocatalytic mechanism enhances simulation performances of acid deposition. Observational validation demonstrates significant reductions in modeling bias for sulfur dioxide (SO2) dry deposition and sulfate (SO4) total deposition. Additionally, the improved model captures the declining trend of SO4 deposition over 2006-2020. We further identified two key impacts of photocatalytic chemistry: firstly, our findings indicate it enhances sulfur removal more efficiently in near-desert areas like North China Plain (NCP) than in downwind areas like Yunnan-Guizhou-Chongqing region (YGY). Lifetimes of total sulfur reduced from 3.29 to 2.46 days in NCP, and from 2.46 to 1.95 days in YGY. This discrepancy results from faster conversion of SO2 to dust-phase SO4 and larger proportions of coarse-mode particles in NCP, resulting in accelerated SO4 deposition velocity. Secondly, our results indicate that because dust photocatalytic chemistry amplifies removal of sulfur through SO4 formation and deposition, decreased dust emission resulted in enhanced sulfur lifetimes over 2006-2014. Sensitivity experiments further show higher dust concentrations accelerate pollutants' removal. These findings underscore the importance of dust heterogeneous chemistry in influencing sulfur deposition, providing scientific insights for mitigating acid deposition in China.
The application of formulas is a fundamental ability of humans when addressing numerical reasoning problems. However, existing numerical reasoning datasets seldom explicitly indicate the formulas employed during the reasoning steps. To bridge this gap, we propose a question answering dataset for formula-based numerical reasoning called FormulaQA, from junior high school physics examinations. We further conduct evaluations on LLMs with size ranging from 7B to over 100B parameters utilizing zero-shot and few-shot chain-of-thoughts methods and we explored the approach of using retrieval-augmented LLMs when providing an external formula database. We also fine-tune on smaller models with size not exceeding 2B. Our empirical findings underscore the significant potential for improvement in existing models when applied to our complex, formula-driven FormulaQA.
Key Point Analysis (KPA), the summarization of multiple arguments into a concise collection of key points, continues to be a significant and unresolved issue within the field of argument mining. Existing models adapt a two-stage pipeline of clustering arguments or generating key points for argument clusters. This approach rely on semantic similarity instead of measuring the existence of shared key points among arguments. Additionally, it only models the intra-cluster relationship among arguments, disregarding the inter-cluster relationship between arguments that do not share key points. To address these limitations, we propose a novel approach for KPA with pairwise generation and graph partitioning. Our objective is to train a generative model that can simultaneously provide a score indicating the presence of shared key point between a pair of arguments and generate the shared key point. Subsequently, to map generated redundant key points to a concise set of key points, we proceed to construct an arguments graph by considering the arguments as vertices, the generated key points as edges, and the scores as edge weights. We then propose a graph partitioning algorithm to partition all arguments sharing the same key points to the same subgraph. Notably, our experimental findings demonstrate that our proposed model surpasses previous models when evaluated on both the ArgKP and QAM datasets.
As Large Language Models (LLMs) and Retrieval Augmentation Generation (RAG) techniques have evolved, query rewriting has been widely incorporated into the RAG system for downstream tasks like open-domain QA. Many works have attempted to utilize small models with reinforcement learning rather than costly LLMs to improve query rewriting. However, current methods require annotations (e.g., labeled relevant documents or downstream answers) or predesigned rewards for feedback, which lack generalization, and fail to utilize signals tailored for query rewriting. In this paper, we propose ours, a framework for training query rewriting models free of annotations. By leveraging a publicly available reranker, ours~provides feedback aligned well with the rewriting objectives. Experimental results demonstrate that ours~can obtain better performance than baselines.
After recent gains achieved by large language models (LLMs) on numerical reasoning tasks, it has become of interest to have LLMs teach small models to improve on numerical reasoning. Instructing LLMs to generate Chains of Thought to fine-tune small models is an established approach. However, small models are passive in this line of work and may not be able to exploit the provided training data. In this paper, we propose a novel targeted training strategy to match LLM’s assistance with small models’ capacities. The small model will proactively request LLM’s assistance when it sifts out confusing training data. Then, LLM refines such data by successively revising reasoning steps and reducing question complexity before feeding the small model. Experiments show that this targeted training approach remarkably improves the performance of small models on a range of numerical reasoning datasets by 12–25
Polymer-drug conjugates (PDCs) provide possibilities for the development of multiresponsive drug delivery and release platforms utilized in cancer therapy. The delivery of Temozolomide (TMZ, a DNA methylation agent) by PDCs has been developed to improve TMZ stability under physiological conditions for the treatment of glioblastoma multiforme (GBM); however, with inefficient chemotherapeutic efficacy. In this work, we synthesized an amphiphilic triblock copolymer (P1-SNO) with four pendant functionalities, including (1) a TMZ intermediate (named MTIC) as a prodrug moiety, (2) a disulfide bond as a redox-responsive trigger to cage MTIC, (3) S-nitrosothiol as a light/heat-responsive donor of nitric oxide (NO), and (4) a poly(ethylene glycol) chain to enable self-assembly in aqueous media. P1-SNO was demonstrated to liberate MTIC in the presence of reduced glutathione and release gaseous NO upon exposure to light or heat. The in vitro results revealed a synergistic effect of released MTIC and NO on both TMZ-sensitive and TMZ-resistant GBM cells. The environment-responsive PDC system for codelivery of MTIC and NO is promising to overcome the efficacy issue in TMZ-based cancer therapy.
Numerical reasoning over hybrid data containing tables and long texts has recently received research attention from the AI community. To generate an executable reasoning program consisting of math and table operations to answer a question, state-of-the-art methods use a retriever-generator pipeline. However, their retrieval results are static, while different generation steps may rely on different sentences. To attend to the retrieved information that is relevant to each generation step, in this paper, we propose DyRRen, an extended retriever-reranker-generator framework where each generation step is enhanced by a dynamic reranking of retrieved sentences. It outperforms existing baselines on the FinQA dataset.
Recent machine reading comprehension datasets such as ReClor and LogiQA require performing logical reasoning over text. Conventional neural models are insufficient for logical reasoning, while symbolic reasoners cannot directly apply to text. To meet the challenge, we present a neural-symbolic approach which, to predict an answer, passes messages over a graph representing logical relations between text units. It incorporates an adaptive logic graph network (AdaLoGN) which adaptively infers logical relations to extend the graph and, essentially, realizes mutual and iterative reinforcement between neural and symbolic reasoning. We also implement a novel subgraph-to-node message passing mechanism to enhance context-option interaction for answering multiple-choice questions. Our approach shows promising results on ReClor and LogiQA.
Conversational Question Answering (ConvQA) is required to answer the current question, conditioned on the observable paragraph-level context and conversation history. Previous works have intensively studied history-dependent reasoning. They perceive and absorb topic-related information of prior utterances in the interactive encoding stage. It yielded significant improvement compared to history-independent reasoning. This paper further strengthens the ConvQA encoder by establishing long-distance dependency among global utterances in multi-turn conversation. We use multi-layer transformers to resolve long-distance relationships, which potentially contribute to the reweighting of attentive information in historical utterances.Experiments on QuAC show that our method obtains a substantial improvement (1%), yielding the F1 score of 73.7%. All source codes are available at https://github.com/jaytsien/GHR.
Diagram question answering is a challenging multi-modal machine learning task that focuses on answering questions according to given diagrams on specific fields. Compared to natural imaged, these diagrams have more abstract expressions and complex logical relations, which makes diagram question answering more difficult. In this paper, we propose a new approach for diagram question answering task. We add bottom-up and top-down attention to identify regions of interest to questions and use a same model to jointly train multiple choice questions and true false questions. Our approach on test dataset of official CCKS2022 textbook diagram question answering session achieves the accuracy of 58.09%.
Scenario-based question answering (SQA) has attracted an increasing research interest. Compared with the well-studied machine reading comprehension (MRC), SQA is a more challenging task: a scenario may contain not only a textual passage to read but also structured data like tables, i.e., tabular scenario based question answering (TSQA). AI applications of TSQA such as answering multiple-choice questions in high-school exams require synthesizing data in multiple cells and combining tables with texts and domain knowledge to infer answers. To support the study of this task, we construct GeoTSQA. This dataset contains 1k real questions contextualized by tabular scenarios in the geography domain. To solve the task, we extend state-of-the-art MRC methods with TTGen, a novel table-to-text generator. It generates sentences from variously synthesized tabular data and feeds the downstream MRC method with the most useful sentences. Its sentence ranking model fuses the information in the scenario, question, and domain knowledge. Our approach outperforms a variety of strong baseline methods on GeoTSQA.
A decline of surface biogenic secondary organic aerosols through the mediation of reduced anthropogenic aerosols has been recognized as an air quality co-benefit of anthropogenic emission control over the southeastern US. However, the climate impacts of this anthropogenic–biogenic interaction remain poorly understood. Here, we identified a substantial decline of summertime aerosol loading aloft over the southeastern US in recent decades through the interaction, which leads to a stronger decline in column-integrated aerosol optical depth and a greater increase in radiative fluxes over the southeastern than northeastern US, different from trends of anthropogenic emissions and near-surface aerosol loading. The anthropogenic–biogenic interaction is shown to explain more than 60% of the coherent increasing trend of 5.3 Wm−2decade−1 in clear-sky surface downward radiative fluxes. We show that current climate models fail to represent this interaction. The interaction is further projected to amplify the positive radiative forcing from emission control by 42.3% regionally over the southeastern US and globally by 5.4% in 2050 under RCP4.5 compared to 2005. This amplification effect implies greater challenges to achieving the Paris Agreement temperature targets with continuous emission control in future.
为了更好地揭示特殊地形下水蚀过程对土壤结构和有机碳含量分配的影响,选取典型南方红壤丘陵区—青原山小流域为研究区,采用核素137Cs示踪技术研究小流域侵蚀沟道内水土流失现状,分析了沟道侵蚀对土壤团聚体稳定性及有机碳含量的影响.结果表明:侵蚀沟道的坡顶处137Cs含量最高,且高于背景值,属于沉积区,而坡上、坡脚属于中度侵蚀,坡中属于轻度侵蚀;侵蚀沟道顺坡而下侵蚀过程依次表现为绝对沉积、绝对侵蚀、相对沉积和绝对侵蚀,其中植被和地形因子是主导因素;沉积区相比于侵蚀区平均质量直径(Mean Weight Diameter,MWD)和大团聚体含量(粒径≥0.25 mm)更高,侵蚀区中相对沉积的坡中有着更稳定的土壤团粒结构;沉积区各个粒径的土壤团聚体有机碳含量均高于侵蚀区,侵蚀区的土壤团聚体有机碳更趋向于均匀分配,土壤理化性质的空间差异也会影响土壤团聚体有机碳含量.侵蚀沟道中土壤侵蚀模式与传统坡面并不一致,土壤结构及相关碳组分主要受地形和植被支配下的土壤侵蚀程度影响.
The surface engineering of the apoferritin shell by means of traditional chemical modifications usually suffers from site inaccuracy and insufficient conjugation. This report describes a non-covalent method for precise modulation of the apoferritin surface without alteration of amino acid residues. A bifunctional macromolecule, structured as azide-poly(ethylene glycol)-porphyrin (termed TPA), was synthesized. TPA was observed to be able to recognize and bind apoferritin in a 12 : 1 stoichiometry with a higher binding affinity than arachidonate, thanks to the specific host-guest interaction between the pocket of each two-fold channel and the porphyrin moiety. This method allows for site-specific engineering of the apoferritin surface with on demand functionalities and optimization of drug encapsulation.
Scenario-based question answering (SQA) has attracted increasing research attention. It typically requires retrieving and integrating knowledge from multiple sources, and applying general knowledge to a specific case described by a scenario. SQA widely exists in the medical, geography, and legal domains---both in practice and in the exams. In this paper, we introduce the GeoSQA dataset. It consists of 1,981 scenarios and 4,110 multiple-choice questions in the geography domain at high school level, where diagrams (e.g., maps, charts) have been manually annotated with natural language descriptions to benefit NLP research. Benchmark results on a variety of state-of-the-art methods for question answering, textual entailment, and reading comprehension demonstrate the unique challenges presented by SQA for future research.
Current treatment of recurrent glioblastoma multiforme (GBM) demands dose-intense temozolomide (TMZ), a prodrug of 5-(3-methyltriazen-1-yl) imidazole-4-carboxamide (MTIC), based on the spontaneous hydrolysis of TMZ at basic pH. However, how to control the activity of MTIC remains unknown, which poses a particular challenge to search a reliable MTIC receptor. We reported that copper, for the first time, is found to recognize and bind MTIC in the process of TMZ degradation, which means copper can play an important role in enhancing the bioavailability of MTIC derived from TMZ. Using apoferritin as a model copper-bound protein, we studied the copper-TMZ interaction in protein and observed efficient MTIC immobilization with high binding efficiency (up to 92.9% based on original TMZ) and capacity (up to 185 MTIC moieties per protein). The system was stable against both alkaline and acidic pH and could be activated by glutathione to liberate MTIC, which paves a way to deliver a DNA-alkylating agent for both TMZ-sensitive and TMZ-resistant GBM chemotherapy. Our study provides a new insight for understanding the potential relationship between the special GBM microenvironment (specific copper accumulation) and the therapeutic effect of TMZ.
l-carnosine is a bipeptide with varieties of biomedical benefits, but rarely reported as a component in structurally defined polymers due to the unavailability of isolatable monomers. In this report, a simple method to synthesize a carnosine-derived methacrylamide in three steps is established and a neutral polymer with a narrow molecular weight distribution (M-n 7.6 kDa, D 1.2) and high structural regularity is prepared via free radical polymerization, and a carnosine-pendent cationic homopolymer is finally achieved after trifluoroacetic acid-mediated deprotection without a concern of impurities such as metals. The antioxidative activity, cytotoxicity, DNA-binding capability, and gene delivery performance of the cationic polymer are also studied.
Side-chain discotic liquid crystalline polymers (SDLCPs) with discotic (disclike) mesogens (discogens) attached as side groups through flexible spacers constitute a class of fascinating organic polymer semiconducting materials. While so far almost all reported SDLCPs belong to conventional C2 polymers based on polymerization of vinyl monomers, limiting their structure diversities, substitution density, and efficiency promotion. In this article we present the synthesis of a series of syndiotactic polymethylene SDLCPs of Pm(TPn) with discotic triphenylene (TP) side groups of variant peripheral alkoxy substituents (n = 6, 4, 10) and different length alkyl spacers (m = 3, 4, 6, 8, 10, 12) through an indirect two-step rhodium-complex-catalyzed Cl carbene polymerization pathway. The thermal properties and ordered organization structures of the precursor polymers and polymethylene SDLCPs have been systematically investigated with differential scanning calorimetry (DSC) and polarized optical microscopy (POM), especially through synchrotron radiation variable-temperature small/wide-angle X-ray scattering (SAXS/WAXS) analyses. All of the series Cl-type syndiotactic polymethylenes with high densely substituted TP side groups exhibit various hierarchical ordered columnar mesophases comparable to that of the typical side-chain C2 polymers of well-defined polyacrylates with TP discogens. Moreover, they possess a remarkably broadened temperature range of columnar structures persistent to very high temperatures in virtue of the stiff helical polymethylene backbone. This work provides a feasible route to prepare the Cl-type SDLCPs with high densely substituted functional side groups and may offer an in-depth understanding for the hierarchical organization of ordered columnar structures with significantly increased thermal stability.