
Reward modeling is essential for aligning large language models with human preferences through reinforcement learning. To provide accurate reward signals, a reward model (RM) should stimulate deep thinking and conduct interpretable reasoning before assigning a score or a judgment. Inspired by recent advances of long chain-of-thought on reasoning-intensive tasks, we hypothesize and validate that integrating reasoning into reward modeling significantly enhances RM's interpretability and performance. We introduce a new class of generative reward models, Reasoning Reward Models (ReasRMs), which formulate reward modeling as a reasoning task. We propose a reasoning-oriented training pipeline and train a family of ReasRMs, RM-R1. RM-R1 features a chain-of-rubrics (CoR) mechanism -- self-generating sample-level chat rubrics or math/code solutions, and evaluating candidate responses against them. The training of RM-R1 consists of two key stages: (1) distillation of high-quality reasoning chains and (2) reinforcement learning with verifiable rewards. Empirically, our models achieve superior performance across three reward model benchmarks on average, outperforming much larger open-weight models (e.g., INF-ORM-Llama3.1-70B) and proprietary ones (e.g., GPT-4o) by up to 4.9%. Beyond final performance, we perform thorough analyses to understand the key ingredients of successful ReasRM training.
We examine how lead venture capital (VC) investors adjust VC syndicate resource diversity to meet ventures’ evolving developmental needs. As ventures transition from early to later stages, lead VCs tend to assemble more resource-diverse syndicates to support expanding resource requirements. However, greater diversity also increases coordination costs of managing the syndicate. We argue that lead VCs balance these competing considerations, such that syndicate composition depends on attributes of both the lead VC and the syndicate that shape this trade-off. We test our conjectures using a sample of US ventures receiving VC funding between 1990 and 2019. Our findings offer an evolutionary perspective on VC syndicate composition across venture life stages. More diverse investors do not always mean better support for startups. Indeed, too much diversity in a VC syndicate can backfire. This study examines how lead venture capital (VC) investors adjust the diversity of their investment syndicates as startups grow. We find that as startups move from early to later stages, lead VCs tend to bring in more diverse partners to meet expanding needs, but they do so selectively because greater diversity also makes coordination more difficult. As a result, syndicate composition reflects a trade-off between accessing a broader range of resources and maintaining manageable collaboration. Our findings imply that both investors and entrepreneurs should be strategic in assembling their investor base, seeking broader expertise as their ventures grow while avoiding overly complex syndicates that may slow decision-making and execution.
The increasing penetration of distributed energy resources (DERs) necessitates innovative strategies, such as multi-transmission-node DER aggregation (M-DERA), to support their wide geographic aggregation for the wholesale market integration at scale. However, M-DERAs pose new challenges in estimating the nodal power proportions within the aggregation, inducing inaccurate power flow calculations in market operation tools such as unit commitment (UC). To this end, this paper proposes a novel chance-constrained UC (CCUC) model to determine system optimal operation plans with M-DERAs, in which the estimated nodal power proportions of M-DERAs, characterized by distribution factors (DFs), are considered as uncertain parameters, and power flow limits are modeled as bilinear chance constraints. A novel bounded hetero-dimensional mixture model is proposed to describe the complex distribution of DFs over multiple hetero-dimensional hyperplanes in a bounded space. With this, the bilinear chance constraints are reformulated into a scenario-based stochastic form and solved by Benders decomposition. Test results on the IEEE 24-bus and 118-bus systems show that, compared to other methods under various system operation conditions, the proposed method reduces UC costs by up to 6% and real-time economic dispatch (RTED) costs by up to 6.8%, while also lowering load shedding in RTED and transmission overloading after M-DERAs' self-dispatch which in the best case can be reduced to zero. These results validate the effectiveness of the proposed method in managing M-DERA integration while ensuring operational economics and mitigating transmission line overloading.
Going beyond simple text processing, financial auditing requires detecting semantic, structural, and numerical inconsistencies across large-scale disclosures. As financial reports are filed in XBRL, a structured XML format governed by accounting standards, auditing becomes a structured information extraction and reasoning problem involving concept alignment, taxonomy-defined relations, and cross-document consistency. Although large language models (LLMs) show promise on isolated financial tasks, their capability in professional-grade auditing remains unclear. We introduce FinAuditing, a taxonomy-aligned, structure-aware benchmark built from real XBRL filings. It contains 1,102 annotated instances averaging over 33k tokens and defines three tasks: Financial Semantic Matching (FinSM), Financial Relationship Extraction (FinRE), and Financial Mathematical Reasoning (FinMR). Evaluations of 13 state-of-the-art LLMs reveal substantial gaps in concept retrieval, taxonomy-aware relation modeling, and consistent cross-document reasoning. These findings highlight the need for realistic, structure-aware benchmarks. We release the evaluation code at https://github.com/The-FinAI/FinAuditing and the dataset at https://huggingface.co/collections/TheFinAI/finauditing. The task currently serves as the official benchmark of an ongoing public evaluation contest at https://open-finance-lab.github.io/SecureFinAI_Contest_2026/.
We study the impact of environmental, social, and governance (ESG) scores on out-of-sample portfolio gains. Our shrinkage approach accommodates investors with heterogeneous beliefs and enables us to assess the incremental value of ESG relative to market information in a data-driven manner. We find that ESG-based portfolio rules do not consistently outperform market-based strategies in terms of risk-adjusted returns. Moreover, investors concerned with ex-post ESG standing can achieve comparable goals using return-based rules alone without integrating ESG scores into their portfolio choices, suggesting that these scores are a second-order priced information. Our paper raises questions about the efficiency of ESG-driven portfolios and their long-term financial stability.