We study how firms exposed to concentrated algorithmic infrastructure are repriced when U.S. export controls alter access to that infrastructure. We construct an Algorithmic Dependence Index (ADI) from 89,965 corporate 10-K filings measuring exposure to cloud platforms, AI accelerators, and semiconductor supply chains. Around nineteen policy events between 2018 and 2025, a one-standard-deviation increase in ADI predicts -0.201% lower three-day abnormal returns (t = -2.39); on the options-listed subsample the effect strengthens to -0.299%. Infrastructure providers are insulated. The mechanism is regime uncertainty rather than directional restriction: loosening events hurt high-ADI firms as much as tightening events, and implied volatility rises for dependent firms even when headlines signal relaxation.
Corporate environmental disclosure is now abundant, but the ability to check it against physical reality is not. Tools for assessing climate and environmental claims have largely compared text against text; they rarely confront a firm's claim with the observed state of the land. We present LLEO (Large Language Models for Earth Observation), an agentic system in which a language model orchestrates satellite Earth-observation tools in Google Earth Engine to ground site-level corporate environmental claims in Sentinel-2 and Sentinel-5P evidence. Across 1,245 claims from 214 European companies' 2024 sustainability reports, with companion 2023 and 2025 corpora, satellite land-cover change corroborates 26.1% of analysed claims and contradicts 7.5%, while 66.3% remain under-determined by optical imagery alone; a contradiction flags a site for review, not misreporting. Verifiability is structurally uneven across claim types, and the binding constraint is geolocation rather than imagery: only 9% of claims disclose precise coordinates, and geocoding failures dominate every other failure mode. The margin that matters is verifiable disclosure, not more disclosure, with site coordinates its cheapest missing input.
Long-form legal reasoning remains a key challenge for large language models (LLMs) in spite of recent advances in test-time scaling. To address this, we introduce ***LEXam***, a novel benchmark derived from 340 law exams spanning 116 law school courses across a range of subjects and degree levels. The dataset comprises 4,886 law exam questions in English and German, including 2,841 long-form, open-ended questions and 2,045 multiple-choice questions. Besides reference answers, the open questions are also accompanied by explicit guidance outlining the expected legal reasoning approach such as issue spotting, rule recall, or rule application. Our evaluation on both open-ended and multiple-choice questions present significant challenges for current LLMs; in particular, they notably struggle with open questions that require structured, multi-step legal reasoning. Moreover, our results underscore the effectiveness of the dataset in differentiating between models with varying capabilities. Deploying an ensemble LLM-as-a-Judge paradigm with rigorous human expert validation, we demonstrate how model-generated reasoning steps can be evaluated consistently and accurately, closely aligning with human expert assessments. Our evaluation setup provides a scalable method to assess legal reasoning quality beyond simple accuracy metrics. Anonymous repository: [this URL](https://anonymous.4open.science/r/LEXam-anonymous-12EB).
IntroductionHuman–AI interaction is commonly framed as a problem of uniformly minimizing uncertainty across the joint system. We challenge this assumption by proposing Asymmetric Uncertainty Regulation (AUR), a dynamical framework in which stable and adaptive collaboration requires directional rather than symmetric uncertainty regulation.MethodsHuman–AI systems are modeled as coupled entropy dynamics. The artificial subsystem rapidly contracts predictive entropy through statistical inference and optimization, whereas the human subsystem maintains bounded but persistently non-zero entropy. System stability is characterized by a positive stability margin that prevents both symmetric entropy collapse and runaway instability. The framework is conceptually operationalized through illustrative applications in nutritional counseling, clinical diagnosis under novelty, and criminal investigation.ResultsThe model yields empirically testable signatures, including pronounced timescale separation between AI and human entropy trajectories, a non-zero human entropy plateau, and bounded fluctuations whose variance increases near the stability boundary. Systems that enforce symmetric entropy minimization are predicted to become rigid or fragile under distributional shift. By contrast, systems operating within the AUR regime are predicted to maintain predictive reliability while preserving exploratory variability, contextual adaptation, and value-sensitive human judgment.DiscussionAUR reframes uncertainty as a structured resource rather than a deficiency to be uniformly eliminated. The illustrative applications suggest how reliable AI prediction can be combined with human flexibility and normative judgment, although they constitute conceptual operationalizations rather than empirical validations. By preserving an asymmetry between AI certainty and bounded human uncertainty, AUR offers a principled foundation for designing human–AI systems that augment rather than erode human agency.
Financial sentiment analysis models increasingly drive automated trading and risk assessment, yet their vulnerability to adversarial manipulation remains poorly understood. We demonstrate that subtle, human-imperceptible textual changes can systematically fool leading sentiment classifiers. Using GPT-4o to generate semantically equivalent paraphrases, optimized via embedding similarity ratios, we attack FinBERT and FinGPT across three financial datasets. Our method alters predictions in 20%-54% of cases, reducing accuracy by 10-26 percentage points. We identify three critical vulnerabilities: difficulty with numbers lacking directional cues, misinterpretation of double negatives, and oversensitivity to trigger words. These findings reveal security risks in automated financial systems and underscore the need for more robust models.
Tracking financial investments in climate adaptation is a complex and expertise-intensive task, particularly for Early Warning Systems (EWS), which lack standardized financial reporting across multilateral development banks (MDBs) and funds. To address this challenge, we introduce an LLM-based agentic AI system that integrates contextual retrieval, fine-tuning, and multi-step reasoning to extract relevant financial data, classify investments, and ensure compliance with funding guidelines. Our study focuses on a real-world application: tracking EWS investments in the Climate Risk and Early Warning Systems (CREWS) Fund. We analyze 25 MDB project documents and evaluate multiple AI-driven classification methods, including zero-shot and few-shot learning, fine-tuned transformer-based classifiers, chain-of-thought (CoT) prompting, and an agent-based retrieval-augmented generation (RAG) approach. Our results show that the agent-based RAG approach significantly outperforms other methods, achieving 87% accuracy, 89% precision, and 83% recall. Additionally, we contribute a benchmark dataset and expert-annotated corpus, providing a valuable resource for future research in AI-driven financial tracking and climate finance transparency.
PDFs are the second-most used document type on the internet (after HTML). Yet, existing QA datasets commonly start from text sources or only address specific domains. In this paper, we present pdfQA, a multi-domain 2K human-annotated (real-pdfQA) and 2K synthetic dataset (syn-pdfQA) differentiating QA pairs in ten complexity dimensions (e.g., file type, source modality, source position, answer type). We apply and evaluate quality and difficulty filters on both datasets, obtaining valid and challenging QA pairs. We answer the questions with open-source LLMs, revealing existing challenges that correlate with our complexity dimensions. pdfQA presents a basis for end-to-end QA pipeline evaluation, testing diverse skill sets and local optimizations (e.g., in information retrieval or parsing).
This paper examines the effects of an educational program on Sustainable Finance Literacy (SFL) and its influence on sustainable investment decisions. Through a randomized controlled trial and an incentivized choice experiment, we found that our SFL program significantly improves literacy. The program also increased the probability of investing in a highly sustainable fund by 6 percentage points on the extensive margin and decreased allocations between 3.2% and 2.7% for the less sustainable funds on the intensive margin. Among participants who already held pro-sustainability attitudes, the treatment additionally led to more investments in the highly sustainable fund on the intensive margin. Higher SFL further led to more critical sustainability assessments of mid-tier funds and reduced tendencies to chase past high returns.
The European Commission adopted revised European Sustainability Reporting Standards (ESRS) on 3 July 2026, following EFRAG's December 2025 technical advice. Using the datapoint schedule in that technical advice, which reduces required datapoints, when material, by 61% overall and ESRS E4 biodiversity datapoints by 78%, we assess the nature-related disclosures of STOXX Europe 600 companies over FY2023-FY2025 (801 TNFD and 1,047 ESRS firm-year observations) with a retrieval-augmented generation pipeline that extends the ASKNATURE system. Three findings bear on the proposal. First, mandatory reporting delivers: ESRS disclosures carry a specificity premium of +0.13 on the unit scale (p < 0.01) and an enforceability premium of +0.16 (p < 0.01) over voluntary TNFD reporting, and ESRS E1+E2 evidence depth rises from 0.785 to 0.805 (FY2025 partial) with narrowing variance under the CSRD. Second, mandatory reporting has limits: it produces no commitment premium, the roughly 4:1 positive-to-negative sentiment skew in nature-impact reporting survives the mandate, and dual TNFD+ESRS reporters show no disclosure premium over single-framework reporters. Third, under proportional datapoint retention the simplification would erase 57% of observable ESRS evidence depth on average and 78% of E4 biodiversity evidence, collapsing the cross-sector differentiation on which institutional nature-risk screening relies. The improving climate-disclosure trend under the CSRD offers no assurance that biodiversity disclosure would survive the simplified regime. Preserving core quantitative E4 requirements, in particular the E4-5 KPIs and site-level granularity, is necessary to keep European nature reporting from reverting to the voluntary cheap-talk baseline we document.
Financial news media shapes trillion-dollar climate investment decisions, yet discourse in this elite domain remains underexplored. We analyze two decades of climate-related articles (2000-2023) from Dow Jones Newswire using an Actor-Frame-Argument (AFA) pipeline that extracts who speaks, how issues are framed, and which arguments are deployed. We validate extractions against 2,000 human-annotated articles using a Decompositional Verification Framework that evaluates completeness, faithfulness, coherence, and relevance. Our longitudinal analysis uncovers a structural transformation: pre-2015 coverage emphasized risk and regulatory burden; post-Paris Agreement, discourse shifted toward economic opportunity and innovation, with financial institutions becoming dominant voices. Methodologically, we provide a replicable paradigm for longitudinal media analysis with LLMs; substantively, we reveal how financial elites have internalized and reframed the climate crisis across two decades.
Automated fact-checking (AFC) systems retrieve evidence and predict claim veracity, yet evaluations omit simple baselines, systems are developed for a single benchmark and cannot be trusted to generalise across domains. No prior work cross-evaluates the full two-stage retrieve-then-verify pipeline across diverse datasets, complementing retrieval-only studies (Thakur et al., 2021) and single-stage benchmarking studies (Calamai et al., 2025). We benchmark nine models, ranging from random and sparse baselines to fine-tuned transformers, zero-shot LLMs, and the two highest-ranked systems from the AVeriTeC 2025 shared task, across four datasets spanning scientific, open-web, and climate domains. Three findings stand out: (1) on ClimateCheck claim-only and fine-tuned models outperform zero-shot LLM and top-performing AVeriTeC 2025 systems, highlighting that noisy evidence can degrade veracity prediction; (2) system rankings are strongly domain- and metric-dependent: the best model on SciFact (macro-F1 0.70) drops to 0.31 on ClimateCheck, while the AVeriTeC 2025 winner and runner-up swap rankings based on evaluation metrics and datasets; (3) replacing retrieved evidence with gold annotations improves veracity accuracy by 14-22 points across models, confirming retrieval remains primary bottleneck. We release code, pre-processed datasets, and all results to support reproducible AFC research.
The emerging paradigm of AI co-scientists focuses on tasks characterized by repeatable verification, where agents explore search spaces in 'guess and check' loops. This paradigm does not extend to problems where repeated evaluation is impossible and ground truth is established by the consensus synthesis of theory and existing evidence. We evaluate a Gemini-based AI environment designed to support collaborative scientific assessment, integrated into a standard scientific workflow. In collaboration with a diverse group of 13 scientists working in the field of climate science, we tested the system on a complex topic: the stability of the Atlantic Meridional Overturning Circulation (AMOC). Our results show that AI can accelerate the scientific workflow. The group produced a comprehensive synthesis of 79 papers through 104 revision cycles in just over 46 person-hours. AI contribution was significant: most AI-generated content was retained in the report. AI also helped maintain logical consistency and presentation quality. However, expert additions were crucial to ensure its acceptability: less than half of the report was produced by AI. Furthermore, substantial oversight was required to expand and elevate the content to rigorous scientific standards.
Variation in human annotation (i.e., disagreements) is common in NLP, often reflecting important information like task subjectivity and sample ambiguity. Modeling this variation is important for applications that are sensitive to such information. Although RLVR-style reasoning (Reinforcement Learning with Verifiable Rewards) has improved Large Language Model (LLM) performance on many tasks, it remains unclear whether such reasoning enables LLMs to capture informative variation in human annotation. In this work, we evaluate the influence of different reasoning settings on LLM disagreement modeling. We systematically evaluate each reasoning setting across model sizes, distribution expression methods, and steering methods, resulting in 60 experimental setups across 3 tasks. Surprisingly, our results show that RLVR-style reasoning degrades performance in disagreement modeling, while naive Chain-of-Thought (CoT) reasoning improves the performance of RLHF LLMs (RL from human feedback). These findings underscore the potential risk of replacing human annotators with reasoning LLMs, especially when disagreements are important.
Do transformer-era text-processing technologies expand sell-side coverage of disclosure-intensive firms, or do they deepen productivity within existing analyst-firm relationships? Using SEC filings matched to I/B/E/S from 2015 to 2025, we exploit predetermined disclosure intensity, a measure of where such tools should be valuable rather than of individual analyst use. The evidence rejects broad coverage expansion: firm-year analyst counts on disclosure-intensive firms are flat, new analyst-firm pair formation declines, and entrant shares are similar across exposure groups. The gains are concentrated within existing relationships, consistent with incumbent deepening. Within analyst-firm pairs, scaled forecast error falls by $0.59$ per standard deviation of disclosure intensity, roughly $59\%$ of the pre-period within-pair median, and the decline survives analyst-by-period fixed effects. Direct reallocation evidence is suggestive only. Broker access and revision timing are supportive only. The pattern is consistent with text-processing technologies raising incumbent productivity without expanding sell-side coverage.
Conventional information retrieval is concerned with identifying the relevance of texts for a given query. Yet, the conventional definition of relevance is dominated by aspects of similarity in texts, leaving unobserved whether the text is truly useful for addressing the query. For instance, when answering whether Paris is larger than Berlin, texts about Paris being in France are relevant (lexical/semantic similarity), but not useful. In this paper, we introduce UsefulBench, a domain-specific dataset curated by three professional analysts labeling whether a text is connected to a query (relevance) or holds practical value in responding to it (usefulness). We show that classic similarity-based information retrieval aligns more strongly with relevance. While LLM-based systems can counteract this bias, we find that domain-specific problems require a high degree of expertise, which current LLMs do not fully incorporate. We explore approaches to (partially) overcome this challenge. However, UsefulBench presents a dataset challenge for targeted information retrieval systems.
BackgroundWhile heat-related mortality is well-documented, the full economic burden of non-fatal illness remains underexplored. This gap hinders evidence-based health planning in a warming climate. We aimed to quantify the current and projected financial burden of heat-related hospital admissions in Switzerland.MethodsWe linked daily hospital admissions (1998-2022) from six Swiss cantons, representing 60% of the national total, to temperature records using a Distributed Lag Non-Linear Model (DLNM) meta-analysis. Costs were estimated via the Swiss Diagnosis-Related Groups (SwissDRG) tariff system across different disease categories and age groups. Future climate projections were generated using CH2018 climate simulations with Shared Socioeconomic Pathway (SSP) scenarios (SSP1-Representative Concentration Pathway (RCP)2.6, SSP2-RCP4.5, SSP5-RCP8.5).ResultsHistorically, extreme heat significantly increased hospital admissions in major Swiss regions, with endocrine/metabolic disorders showing the highest relative risk (RR 2.02). The elderly (75+) and children were the most vulnerable populations. The average annual cost of these direct hospitalizations across the six cantons studied was Swiss Francs (CHF) 20.6 million (2013-2022), notably a conservative estimate representing only a fraction of the total economic burden. Future projections show this burden escalating sharply. Under a high-emissions pathway (SSP5-RCP8.5), these direct costs are projected to increase 2.5-fold by the 2060s, with costs for the elderly quintupling. Critically, even with aggressive mitigation (SSP1-RCP2.6), costs are still projected to triple compared to the baseline. This is driven primarily by demographic aging, with climate change acting as a significant amplifier, responsible for 15-30% of the projected cost increases for the elderly.ConclusionsOur findings from major Swiss regions reveal a substantial and growing financial burden on the healthcare system. Given that these figures represent a lower bound, the true costs are likely much higher. This evidence underscores the urgent need for nationally coordinated adaptation policies to protect public health and ensure healthcare sustainability, even as mitigation efforts continue.
The extent to which firms are adapting and building resilience to environmental change is crucial information for financial institutions, regulators and governments. While corporates’ physical climate risk exposure of their assets to environmental change can be calculated using models, additional information is needed to evaluate their vulnerability to physical climate change, how well they are adapting and broader alignment with societal adaptation and resilience (A&R) goals. This paper empirically evaluates the extent of A&R-related information in current corporate sustainability reports to provide such insights. We build on established sustainability disclosure frameworks and develop an A&R disclosure framework that we combine with the latest advances in large language models to assess S&P 500 company sustainability reports. We prove that corporate A&R information in sustainability reports is lacking, particularly around risks, metrics and targets, underlining the need to consider other data sources when assessing firm-level risks and contributions to societal A&R goals.
LLMs can solve complex tasks by generating long, multi-step reasoning chains. Test-time scaling (TTS) can further improve LLM performance by sampling multiple variants of intermediate reasoning steps, verifying their correctness, and strategically choosing the best steps for continuation. However, existing verification approaches, such as Process Reward Models (PRMs), are computationally expensive, limited to specific domains, and require large-scale human or model-generated annotations. We propose a lightweight alternative for step-level reasoning verification based on probing the internal states of LLMs. We train a transformer-based probe that uses the internal states of the frozen LLM to estimate the credibility of its reasoning steps during generation. Annotation can be generated either by another larger LLM (e.g., DeepSeek-R1) or in a self-supervised manner by the original model itself. The probes are both effective and lightweight, containing fewer than 10M parameters. Across multiple domains, including mathematics, planning, and general knowledge question answering, our probes match or even exceed the performance of PRMs that are up to 810× larger. Our findings suggest that the internal states of LLMs encode their confidence in reasoning processes and can serve as reliable signals for reasoning step verification, offering a promising direction towards scalable and generalizable TTS and introspective LLMs.
This article introduces a novel firm-level green innovation measure using ClimateBERT and GPT-3 to analyze earnings call transcripts. The measure captures a broader range of innovative activities beyond patents and distinguishes between invention and adoption dimensions of such efforts. These dimensions reveal distinctive geographical and industry patterns as well as key determinants. Firms engaged in green innovation, including many from carbon-intensive sectors, exhibit lower expected returns than their industry peers, likely due to improved environmental performance and ability to hedge against transition risk. The effects extend to non-patentable inventions and the widespread adoption of green innovations.