Existing legal benchmarks for Large Language Models (LLMs) primarily focus on final judgment outcomes or isolated subtasks, making it difficult to systematically evaluate the multi-step reasoning processes and reasoning quality required in real-world legal decision-making. To address this gap, we introduce MSLR, a Chinese benchmark for multi-step legal reasoning, which models and evaluates complete judicial reasoning trajectories based on the IRAC (Issue-Rule-Application-Conclusion) framework. MSLR comprises 1389 Chinese insider trading judgments issued between 2005 and 2024, with approximately 60,000 step-level reasoning paths annotated through a combination of human and automated methods, averaging 43 intermediate reasoning steps per case. Based on MSLR, we adopt two evaluation metrics, IRAC Recall and LLM Score, to provide a fine-grained assessment of legal reasoning capabilities. We systematically evaluate 19 mainstream general-purpose, legal-domain, and reasoning-oriented models. The results show that current models exhibit overall limited performance on multi-step legal reasoning tasks; even the best-performing o1-mini model fails to exceed 75% IRAC Recall. In addition, we report the performance of mainstream models on structured annotation tasks and legal judgment reasoning tasks. The MSLR dataset and benchmark construction code are publicly available at https://github.com/yuwenhan07/MSLR-Bench.
Drawing on the spillover-crossover model, we investigate how workplace telepressure after hours (WTA) influences employees’ behaviors at home and crosses over to affect their spouses’ well-being, including relationship satisfaction and family emotional exhaustion. We suggest that work-related rumination serves as a key mechanism, and that core self-evaluations (CSE) moderate these effects. Using a three-wave, multisource design with 227 dual-earner couples, we found that WTA was indirectly and negatively linked to family undermining through problem-solving pondering, which in turn improved spouses’ relationship satisfaction and decreased their family emotional exhaustion. This positive spillover-crossover effect was only present among employees with high CSE. Conversely, WTA was indirectly and positively associated with family undermining through affective rumination, which then lowered spouses’ relationship satisfaction and increased their family emotional exhaustion. This negative spillover-crossover effect was not significant when employees had high CSE. Overall, this study reveals new insights into how telepressure affects both employees and their spouses, emphasizing the crucial role of personal resources in managing spillover and crossover effects of telepressure.
While Large Language Models (LLMs) have demonstrated impressive general capabilities, their direct application in the legal domain is often hindered by a lack of precise domain knowledge and complexity of performing rigorous multi-step judicial reasoning. To address this gap, we present LegalOne, a family of foundational models specifically tailored for the Chinese legal domain. LegalOne is developed through a comprehensive three-phase pipeline designed to master legal reasoning. First, during mid-training phase, we propose Plasticity-Adjusted Sampling (PAS) to address the challenge of domain adaptation. This perplexity-based scheduler strikes a balance between the acquisition of new knowledge and the retention of original capabilities, effectively establishing a robust legal foundation. Second, during supervised fine-tuning, we employ Legal Agentic CoT Distillation (LEAD) to distill explicit reasoning from raw legal texts. Unlike naive distillation, LEAD utilizes an agentic workflow to convert complex judicial processes into structured reasoning trajectories, thereby enforcing factual grounding and logical rigor. Finally, we implement a Curriculum Reinforcement Learning (RL) strategy. Through a progressive reinforcement process spanning memorization, understanding, and reasoning, LegalOne evolves from simple pattern matching to autonomous and reliable legal reasoning. Experimental results demonstrate that LegalOne achieves state-of-the-art performance across a wide range of legal tasks, surpassing general-purpose LLMs with vastly larger parameter counts through enhanced knowledge density and efficiency. We publicly release the LegalOne weights and the LegalKit evaluation framework to advance the field of Legal AI, paving the way for deploying trustworthy and interpretable foundation models in high-stakes judicial applications.
Legal dispute mediation plays a crucial role in resolving civil disputes, yet its empirical study is limited by privacy constraints and complex multivariate interactions. To address this limitation, we present AgentMediation, the first LLM-based agent framework for simulating dispute mediation. It simulates realistic mediation processes grounded in real-world disputes and enables controlled experimentation on key variables such as disputant strategies, dispute causes, and mediator expertise. Our empirical analysis reveals patterns consistent with sociological theories, including Group Polarization and Surface-level Consensus. As a comprehensive and extensible platform, AgentMediation paves the way for deeper integration of social science and AI in legal research.
This paper pioneers the application of chain-of-thought (CoT) prompting in large language models (LLMs) for financial forecasting and portfolio optimization. Leveraging anonymized financial statements from 608 Chinese A-share listed companies (2010-2023), we benchmark ChatGPT 4.0 against human analysts in earnings direction prediction. The CoT-enhanced model achieved superior accuracy (64.35% vs. 58.37%) and generated a 17.14% alpha with a Sharpe ratio of 1.5959 in backtests, demonstrating LLMs' capacity to automate financial reasoning while reducing reliance on specialized expertise. Our findings bridge AI and quantitative finance by validating LLMs' cross-domain adaptability using purely numerical data, offering practical implications for AI-driven investment decision systems.