
耶鲁大学1701年创立时的原名为Collegiate School,1718年为感谢伊利胡·耶鲁先生的捐助,将校名改为耶鲁学院,1887年,耶鲁学院改名为耶鲁大学,而耶鲁学院成为大学的本科部名称。耶鲁学院实行住宿学院制,由12所学院组成。这种制度始于1933年,由一位非常赞赏牛津大学和剑桥大学类似制度的校友Edward S. Harkness捐款建立。每所学院对学生都有一套完备的支持体系,包括院长(Master)、学监(Dean)、驻院学者和Fellow。每所学院都有不同的建筑风格,但都包括庭院和完备的设施。虽然每所学院都设有自己的讨论课程(但向所有学生开放通选),举办自己的社交活动和院长茶会(Master’s Tea),耶鲁学生仍然积极参与全校的学习和社会活动。耶鲁大学的住宿学院都以校史中的著名人物或者校友命名,而故意避免以捐款者命名。
The Freshwater Lake model (FLake) has recently been integrated into the land surface scheme of ERA5 and subsequently downscaled to ERA5-Land, enabling the investigation of global-scale lake thermodynamics at an enhanced spatial resolution. Although ERA5-Land's capability in simulating lake surface water temperature (Ts) has been recognized, its performance in reproducing other critical lake variables remains inadequately explored. This study comprehensively evaluated ERA5-Land atmospheric forcings and lake products against long-term observations at Lake Taihu, a large, subtropical shallow lake. Results show that ERA5-Land captures the surface microclimates reasonably well at Lake Taihu across diurnal to annual scales. Satisfactory performance is also achieved in simulating monthly and annual latent heat fluxes (λE) with a positive bias of 1.0 W m−2 (or 1% of the annual mean). However, water stratification simulated by ERA5-Land is excessively strong in spring and summer. ERA5-Land systematically overestimates downward shortwave radiation (+16 W m−2) and underestimates downward longwave radiation (−13 W m−2) due to cloud cover underestimation. These intrinsic radiative biases only introduce a marginal error in simulated Ts and λE due to the compensation effect between downward shortwave and longwave components. Our proposed correction algorithm can reduce monthly radiative biases both for an independent period at Lake Taihu and at geographically independent Poyang Lake. Notably, despite lacking lake-specific parameter optimization, the default ERA5-Land configuration still outperforms two other FLake simulations using satellite-calibrated parameters in reproducing the multi-year mean Ts. Furthermore, Flake model in ERA5-Land simulates an identical lake warming trend (0.26 °C decade−1 from 1981 to 2020) and a comparable evaporation changing rate (2.9 W m−2 decade−1 from 1979 to 2013) to other independent simulations with tuned parameters, highlighting its reliability for investigating long-term trends in lake thermodynamics.
AI agents may soon become capable of autonomously completing valuable, long-horizon tasks in diverse domains. Current benchmarks either do not measure real-world tasks, or are not sufficiently difficult to meaningfully measure frontier models. To this end, we present Terminal-Bench 2.0: a carefully curated hard benchmark composed of 89 tasks in computer terminal environments inspired by problems from real workflows. Each task features a unique environment, human-written solution, and comprehensive tests for verification. We show that frontier models and agents score less than 65% on the benchmark and conduct an error analysis to identify areas for model and agent improvement. We publish the dataset and evaluation harness to assist developers and researchers in future work at tbench.ai.
With the growing adoption of large language model agents in persistent real-world roles, they naturally encounter continuous streams of tasks. A key limitation, however, is their failure to learn from the accumulated interaction history, forcing them to discard valuable insights and repeat past errors. We propose ReasoningBank, a novel memory framework that distills generalizable reasoning strategies from an agent's self-judged successful and failed experiences. At test time, an agent retrieves relevant memories from ReasoningBank to inform its interaction and then integrates new learnings back, enabling it to become more capable over time. Building on this powerful experience learner, we further introduce memory-aware test-time scaling (MaTTS), which accelerates and diversifies this learning process by scaling up the agent's interaction experience. By allocating more compute to each task, the agent generates abundant, diverse experiences that provide rich contrastive signals for synthesizing higher-quality memory. The better memory in turn guides more effective scaling, establishing a powerful synergy between memory and test-time scaling. Across web browsing and software engineering benchmarks, ReasoningBank consistently outperforms existing memory mechanisms that store raw trajectories or only successful task routines, improving both effectiveness and efficiency; MaTTS further amplifies these gains. These findings establish memory-driven experience scaling as a new scaling dimension, enabling agents to self-evolve with emergent behaviors naturally arise. Our code can be found at https://github.com/google-research/reasoning-bank.
Training large language models (LLMs) for complex reasoning via Reinforcement Learning with Verifiable Rewards (RLVR) is effective but limited by reliance on costly, domain-specific supervision. We explore Reinforcement Learning from Internal Feedback (RLIF), a framework that enables LLMs to learn from intrinsic signals without external rewards or labeled data. We propose Intuitor, an RLIF method that uses a model's own confidence—termed self-certainty—as its sole reward signal. Intuitor replaces external rewards in Group Relative Policy Optimization (GRPO) with self-certainty scores, enabling fully unsupervised learning. Experiments demonstrate that Intuitor matches GRPO's performance on mathematical benchmarks while achieving better generalization to out-of-domain tasks like code generation, without requiring gold solutions or test cases. Our findings show that intrinsic model signals can drive effective learning across domains, offering a scalable alternative to RLVR for autonomous AI systems where verifiable rewards are unavailable. Code is available at https://github.com/sunblaze-ucb/Intuitor
Rhizomes, horizontal underground stems, play fundamental roles in plant persistence and perennial growth by enabling clonal propagation, resource storage, and stress resilience. Despite their ecological and agronomic importance across plant lineages, the genetic and developmental regulation of rhizomes remains poorly characterized. Here, we synthesize findings from in vitro induction studies, in vivo developmental analyses, quantitative trait loci (QTL) mapping, comparative transcriptomics, and limited functional studies to evaluate current knowledge and highlight outstanding questions in rhizome biology. Results from both in vitro and whole-plant studies show that phytohormones, particularly auxin, cytokinin, and gibberellin, are central regulators of rhizome initiation and growth, with effects mediated in a context-dependent manner through interactions with environmental and developmental cues. Across rhizomatous species, traits such as rhizome initiation, branching, and elongation are often polygenic, although comparatively simpler genetic architectures associated with repeated rhizome evolution have been documented in emerging model systems like Mimulus. Transcriptomic analyses further highlight hormone signaling, stress response, and carbohydrate metabolism pathways as key regulatory components. However, few genes have been functionally validated, underscoring the need for tractable systems for genetic dissection. Perennial Mimulus species are proposed as promising models for rhizome research due to their experimental accessibility, ecological relevance, and established genomic resources. Integrated approaches leveraging fine-mapping, near-isogenic lines, multi-omics, and gene editing are poised to accelerate discovery of causal loci and regulatory networks underlying rhizome development, thereby clarifying the genetic and developmental bases of rhizome traits underlying their repeated evolution, with broader implications for perenniality, environmental responses, and crop improvement.