Navigating the complex landscape of corporate climate disclosures and their real impacts is crucial for managing climate-related financial risks. However, current disclosures oftentimes suffer from imprecision, inaccuracy, and greenwashing. We introduce ClimateBert CTI, a deep learning algorithm, to identify climate-related cheap talk in MSCI World index firms’ annual reports. We find that only targeted climate engagement is associated with less cheap talk. Voluntary climate disclosures are associated with more cheap talk. Moreover, cheap talk correlates with increased negative news coverage and higher emissions growth. Hence, cheap talk helps assess climate initiatives’ effectiveness and anticipate reputation and transition risk exposure.
This paper introduces a novel approach to enhance Large Language Models (LLMs) with expert knowledge to automate the analysis of corporate sustainability reports by benchmarking them against the Task Force for Climate-Related Financial Disclosures (TCFD) recommendations. Corporate sustainability reports are crucial in assessing organizations' environmental and social risks and impacts. However, analyzing these reports' vast amounts of information makes human analysis often too costly. As a result, only a few entities worldwide have the resources to analyze these reports, which could lead to a lack of transparency. While AI-powered tools can automatically analyze the data, they are prone to inaccuracies as they lack domain-specific expertise. This paper introduces a novel approach to enhance LLMs with expert knowledge to automate the analysis of corporate sustainability reports. We christen our tool CHATREPORT, and apply it in a first use case to assess corporate climate risk disclosures following the TCFD recommendations. CHATREPORT results from collaborating with experts in climate science, finance, economic policy, and computer science, demonstrating how domain experts can be involved in developing AI tools. We make our prompt templates, generated data, and scores available to the public to encourage transparency.
Deep learning approaches have revolutionized Natural Language Processing (NLP) and textual analysis in the last decade. The performance of the popular chatbot ChatGPT is an impressive demonstration of what deep learning is capable of. However, this remarkable progress with the potential to significantly reduce measurement error has apparently not yet reached accounting and finance. We review a large corpus of accounting and finance literature that uses textual analysis and observe that rule-based and traditional machine learning approaches still dominate. We then compare the performance of rule-based, traditional machine learning, and deep learning approaches on four datasets and find deep learning approaches to consistently perform best. However, we also find that for simpler topics requiring less context, rule-based approaches can produce decent results. These findings argue for more evaluation and increased, but not universal, use of deep learning for textual analysis in accounting and finance.
Large language models (LLMs) have significantly transformed the landscape of artificial intelligence by demonstrating their ability in generating human-like text across diverse topics. However, despite their impressive capabilities, LLMs lack recent information and often employ imprecise language, which can be detrimental in domains where accuracy is crucial, such as climate change. In this study, we make use of recent ideas to harness the potential of LLMs by viewing them as agents that access multiple sources, including databases containing recent and precise information about organizations, institutions, and companies. We demonstrate the effectiveness of our method through a prototype agent that retrieves emission data from ClimateWatch (https://www.climatewatchdata.org/) and leverages general Google search. By integrating these resources with LLMs, our approach overcomes the limitations associated with imprecise language and delivers more reliable and accurate information in the critical domain of climate change. This work paves the way for future advancements in LLMs and their application in domains where precision is of paramount importance.
Large Language Models have made remarkable progress in question-answering tasks, but challenges like hallucination and outdated information persist. These issues are especially critical in domains like climate change, where timely access to reliable information is vital. One solution is granting these models access to external, scientifically accurate sources to enhance their knowledge and reliability. Here, we enhance GPT-4 by providing access to the Sixth Assessment Report of the Intergovernmental Panel on Climate Change (IPCC AR6), the most comprehensive, up-to-date, and reliable source in this domain (refer to the ’Data Availability’ section). We present our conversational AI prototype, available at www.chatclimate.ai , and demonstrate its ability to answer challenging questions in three different setups: (1) GPT-4, (2) ChatClimate, which relies exclusively on IPCC AR6 reports, and (3) Hybrid ChatClimate, which utilizes IPCC AR6 reports with in-house GPT-4 knowledge. The evaluation of answers by experts show that the hybrid ChatClimate AI assistant provide more accurate responses, highlighting the effectiveness of our solution.
To transition to a green economy, environmental claims made by companies must be reliable, comparable, and verifiable. To analyze such claims at scale, automated methods are needed to detect them in the first place. However, there exist no datasets or models for this. Thus, this paper introduces the task of environmental claim detection. To accompany the task, we release an expert-annotated dataset and models trained on this dataset. We preview one potential application of such models: We detect environmental claims made in quarterly earning calls and find that the number of environmental claims has steadily increased since the Paris Agreement in 2015.
In the face of climate change, are companies really taking substantial steps toward more sustainable operations? A comprehensive answer lies in the dense, information-rich landscape of corporate sustainability reports. However, the sheer volume and complexity of these reports make human analysis very costly. Therefore, only a few entities worldwide have the resources to analyze these reports at scale, which leads to a lack of transparency in sustainability reporting. Empowering stakeholders with LLM-based automatic analysis tools can be a promising way to democratize sustainability report analysis. However, developing such tools is challenging due to (1) the hallucination of LLMs and (2) the inefficiency of bringing domain experts into the AI development loop. In this paper, we ChatReport, a novel LLM-based system to automate the analysis of corporate sustainability reports, addressing existing challenges by (1) making the answers traceable to reduce the harm of hallucination and (2) actively involving domain experts in the development loop. We make our methodology, annotated datasets, and generated analyses of 1015 reports publicly available.
Corporate climate disclosures are considered an essential prerequisite to managing climate-related financial risks. At the same time, current disclosures are imprecise, inaccurate, and greenwashing-prone. We introduce a deep learning approach to enable comprehensive climate disclosure analyses by fine-tuning the \climatebert model. From 14,584 annual reports of the MSCI World index firms from 2010 to 2020, we extract the amount of cheap talk, defined as the share of precise versus imprecise climate commitments. We then test various hypotheses by linking three different climate initiatives, namely the Task Force on Climate-Related Financial Disclosure, the Science-Based Targets Initiative, and the Climate Action 100+, to the economic channels of signaling, credibility, and active engagement. In particular, we ask whether these initiatives decrease cheap talk by disciplining companies in how they define and disclose actionable climate commitments in their annual reports.
Corporate climate disclosures are considered an essential prerequisite to managing climate-related financial risks. At the same time, current disclosures are imprecise, inaccurate, and greenwashing-prone. We introduce a deep learning approach to enable comprehensive climate disclosure analyses by fine-tuning the ClimateBert model. From 14,584 annual reports of the MSCI World index firms from 2010 to 2020, we extract the amount of cheap talk, defined as the share of precise versus imprecise climate commitments. We then test various hypotheses by linking three different climate initiatives, namely the Task Force on Climate-Related Financial Disclosure, the Science-Based Targets Initiative, and the Climate Action 100+, to the economic channels of signaling, credibility, and active engagement. In particular, we ask whether these initiatives decrease cheap talk by disciplining companies in how they define and disclose actionable climate commitments in their annual reports.
The climate impact of AI, and NLP research in particular, has become a serious issue given the enormous amount of energy that is increasingly being used for training and running computational models. Consequently, increasing focus is placed on efficient NLP. However, this important initiative lacks simple guidelines that would allow for systematic climate reporting of NLP research. We argue that this deficiency is one of the reasons why very few publications in NLP report key figures that would allow a more thorough examination of environmental impact. As a remedy, we propose a climate performance model card with the primary purpose of being practically usable with only limited information about experiments and the underlying computer hardware. We describe why this step is essential to increase awareness about the environmental impact of NLP research and, thereby, paving the way for more thorough discussions.
Disclosure of climate-related financial risks greatly helps investors assess companies’ preparedness for climate change. Voluntary disclosures such as those based on the recommendations of the Task Force for Climate-related Financial Disclosures (TCFD) are being hailed as an effective measure for better climate risk management. We ask whether this expectation is justified. We do so by training ClimateBERT, a deep neural language model fine-tuned based on the language model BERT. In analyzing the disclosures of TCFD-supporting firms, ClimateBERT comes to the sobering conclusion that the firms’ TCFD support is mostly cheap talk and that firms cherry-pick to report primarily non-material climate risk information.
The ongoing shift towards sustainable asset management fuels the demand for stock indices with improved ESG profiles. Such indices exhibit different performance metrics and risk characteristics compared to their conventional parent indices. At first sight, it might be tempting to assume that their improved ESG profiles cause the main differences. We introduce Monte Carlo permutation tests and apply them to two S&P 500 ESG indices and find, however, that claims about ESG ratings impacting index performance metrics are not supported by our results. We observe that the better ESG profiles of these indices go hand in hand with exposures to small or value companies, depending on the index construction. Thus, it is essential for investors to adequately assess the methodologies used in constructing ESG indices to make expedient investment and benchmark decisions.
Over the recent years, large pretrained language models (LM) have revolutionized the field of natural language processing (NLP). However, while pretraining on general language has been shown to work very well for common language, it has been observed that niche language poses problems. In particular, climate-related texts include specific language that common LMs can not represent accurately. We argue that this shortcoming of today's LMs limits the applicability of modern NLP to the broad field of text processing of climate-related texts. As a remedy, we propose CLIMATEBERT, a transformer-based language model that is further pretrained on over 2 million paragraphs of climate-related texts, crawled from various sources such as common news, research articles, and climate reporting of companies. We find that CLIMATEBERT leads to a 48% improvement on a masked language model objective which, in turn, leads to lowering error rates by 3.57% to 35.71% for various climate-related downstream tasks like text classification, sentiment analysis, and fact-checking.
Index providers increasingly offer sustainable stock indices based on ESG (Environmental, Social, and Governance) ratings of firms. The performance of such indices with ESG tilts is driven by the impact of the applied weighting methodology and by the ESG firm ratings. In this paper, we focus on the S&P 500 ESG Factor Weighted Index, which outperforms the conventional S&P 500 index in terms of raw returns. Based on a simulation analysis, we find the ESG ratings to contribute little to the index return. Rather, the used weighting methodology mainly causes this outperformance which is connected to exposures to the well-established size, investment and momentum factors. Interestingly, the S&P 500 Equal Weighted Index shows comparable exposures to these factors. The weighting methodologies of both indices result in a relative high weighting of smaller stocks with similar return characteristics. Thus, our key finding is that the weighting methodologies of sustainable indices can be a main return driver, which has to be taken into account by investors evaluating the risk and return profile of stock indices with ESG tilts.