
ABSTRACT Recent headlines indicate that the Big Four accounting, auditing, and professional services firms (Big Four) have invested billions in artificial intelligence (AI), as they digitally transform their firms. As part of that analysis, this paper reviews several large language model (LLM) applications that have been developed and considers some of the implications of those applications, such as the democratization of knowledge and the potential impact on barriers to entry. Because LLMs are early in their life cycle and there is limited information about those applications, I take a demonstration‐based field work approach, gathering information from firm demonstrations and then building some prototypes in LLMs that simulate the demonstrated capabilities.
ABSTRACT This study aims to enhance bankruptcy prediction models by employing the gradient‐boosting TreeNet algorithm. This research assesses the predictive accuracy of TreeNet in bankruptcy classification using a high‐dimensional approach. Specifically, it categorizes bankruptcy as nonbankrupt, Chapter 7, or Chapter 11 categories. This research utilizes a large dataset comprised of 76,069 firm‐year observations for the years between 1991 and 2019. The findings highlight the high classification accuracy of TreeNet, particularly in predicting Chapter 7 cases, and its robust performance across different time horizons. The TreeNet prediction model is a valuable tool for decision‐making and risk management in finance, auditing, and policymaking sectors.
Tax authorities face growing volumes of filings and payments and must manage procedural non-compliance (e.g., late filing, late payment and the accumulation of tax arrears) with limited administrative capacity. Many existing machine learning (ML) and artificial intelligence (AI) applications in tax administration rely on binary outcomes, which limits severity-based prioritisation and the targeting of low-cost interventions. This study develops a multiclass prediction model for administrative tax compliance severity using Slovakia's public tax reliability index, which classifies companies into three categories based on regulator-defined administrative criteria. Using only financial statement ratios and governance indicators, we evaluate nine classifiers and five resampling techniques for class imbalance. Gradient boosting models (XGBoost and CatBoost) perform best, reaching an OvR AUC-ROC above 96% for 1-year forecasts, with modest declines for 2- and 3-year horizons. SHAP explanations indicate that smaller boards and indicators consistent with liquidity constraints and tax-payment pressure are associated with higher-severity administrative classes. The proposed workflow offers a transferable framework for multiclass, long-horizon compliance risk prediction and can support proactive case management (e.g., targeted reminders and payment facilitation, including payment plans, and debt prioritisation) in advance of the regulator's semi-annual updates; it may also provide researchers with a potential early-warning label of administrative compliance frictions that could be examined in relation to financial distress.
Pictures are a powerful medium to communicate complex and emotive messages. In particular, the human face expresses corporate culture including diversity and equal opportunity. However, despite the recent visual turn in accounting and finance, quantitative research on diversity in photos is scant because automated solutions for identifying and classifying human faces were not readily available. This paper seeks to bridge this gap by tailoring automated large-sample facial analysis from the recent computing literature into the accounting literature. Our automated model identifies and classifies faces with sufficient accuracy and precision to draw reliable inferences, and this model is made available for future research. We use the resulting quantitative dataset to analyse intermodal discourse in the annual report, asking the question: Do cover photos augment textual diversity disclosure, or are they PR window-dressing? Results suggest that the decision to publish faces on the annual report cover is associated with an integrated reporting strategy and high-quality diversity disclosure, consistent with pictures augmenting textual disclosures. Gender and ethnic diversity of faces in cover photos tell a different story, tending towards PR window-dressing. Methods and findings from this paper may be of interest to researchers, government and policy makers involved in diversity research and regulation.
Genetic programming (GP) is used to obtain multiperiod bankruptcy prediction models, as well as to perform a prior feature selection process for these models. Given the controversy in the field of bankruptcy prediction about the need to include (or not) variables from the economic environment as input information for the prediction models, an analysis is carried out to check whether the impact that the economic environment undoubtedly has on the firms can be captured using only the financial variables of the firm as explanatory variables. To this end, the analysis includes a study of the correlation between the estimates of the prediction models and certain economic indicators. The results confirm the possibility of capturing the evolution of the economic environment using only financial information as input, as strong correlations are shown between the predictions of the models and important economic indicators over a very long postlearning period (2008-2020) and varied in terms of the economic environment (crisis, recovery, COVID, etc.).
Crypto assets have experienced significant growth in recent years, attracting substantial investments from institutional entities and individual investors alike. This surge in popularity necessitates sophisticated strategies to optimize returns. Concurrently, advancements in machine learning have revolutionized the forecasting of crypto asset returns, facilitating algorithmic trading. Leveraging robust algorithms, this approach enables comprehensive market exploration and capitalizes on escalating computational capabilities. This manuscript presents a comparative analysis of neural networks, genetic algorithms, and fuzzy logic, framed within the ordered weighted average (OWA) operator paradigm. These methods are integrated with deep learning and quantum computing principles to predict price movements in crypto assets and other financial indices. Our findings indicate that the quantum genetic algorithm excels in accurately forecasting asset price trends, whereas the quantum fuzzy approach exhibits comparatively lower precision in predicting cryptocurrency price fluctuations. The empirical analysis employs high-frequency data sampled at 10-, 30-, and 60-min intervals from October 2021 to February 2023. The dataset encompasses 11 cryptocurrencies (e.g., Bitcoin and Ethereum), 10 fan tokens, 10 NFTs, and nine reference financial indices (including Gold, WTI Oil, S&P 500, and Euro Stoxx 60). The implications of this research extend to the development of advanced algorithmic trading strategies, offering valuable tools for market participants and stakeholders in the financial sector. The methodologies discussed herein provide versatile and quantitative frameworks for analyzing diverse financial markets, highlighting their potential to enhance decision-making and improve investment outcomes.
The rise of cryptocurrencies has generated significant interest from the public and investors due to their decentralized nature, advanced security features, and potential for high returns. This research uses K-Means clustering and Inverse Covariance Clustering (ICC) to optimize cryptocurrency portfolios by addressing market dynamics and traditional portfolio management limitations. The study involved three phases: collecting daily price data from the top 100 cryptocurrencies from January 2018 to January 2024, performing calculations to identify cryptocurrencies through clustering methods, and constructing and dynamically optimizing investment portfolios from early 2022 to early 2024. We evaluate the constructed portfolios against the Cryptocurrency Benchmark Index (CRIX) using metrics like the Sharpe and Treynor ratios. Results show that both clustering methods can create efficient portfolios, but their effectiveness varies with dataset characteristics and investor objectives. K-Means produces more diversified portfolios, while ICC yields lower volatility portfolios, with ICC generally outperforming K-Means compared to the CRIX index. The findings highlight the potential of clustering methods in enhancing cryptocurrency portfolio selection and suggest the need for further research on real-world applications and advanced techniques tailored for the cryptocurrency market.
The complexity of International Financial Reporting Standards (IFRS) challenges accounting professionals to navigate intricate judgment calls and estimations. This paper tackles a pressing question: Can OpenAI's ChatGPT (Version GPT-4) serve as a reliable artificial intelligence (AI) advisory tool to interpret and apply IFRS standards in real-world scenarios? The importance of this inquiry lies in the potential of generative AI to revolutionize financial reporting by enhancing accuracy, efficiency, and decision-making speed, which are critical demands in today's globalized financial environment. Through an experimental design employing practical case studies, this research evaluates GPT-4's performance under three prompting strategies: zero shot (ZS), few shot (FS), and chain of thought (CoT). This research examines the ability of AI to address judgment-driven, complex IFRS problems, expanding the scope of prior studies that primarily relied on theoretical exams or professional certification tests. Our findings reveal that GPT-4 can consistently identify the correct IFRS standard and produce professionally usable guidance, exhibiting strong potential. ZS proved fastest and most practical for a first advisory pass, FS delivered more structured and accounting-like answers but required greater preparation, and CoT generated the richest explanations at the expense of efficiency. Across all strategies, expert review remained necessary in areas involving item and measurement choices, contract integration, or business-model interpretation. This study efforts to advance the dialogue on AI's role in accounting and lays a foundation for future research exploring its broader implications in accounting decision-making. With insights into GPT-4's strengths and constraints, this study emphasizes its role as a transformative, yet supplementary, tool in advancing IFRS compliance and reporting standards.
To address the limitations of existing product concept design (PCD) methods in the rapidly changing market environments, this study proposes a PCD method using e-commerce product data and artificial intelligence techniques. First, data of competing e-commerce products are acquired from an e-commerce platform. Second, monthly sales of products are categorized and selected as the indicator for evaluating product concepts (PCs). Third, Doc2Vec is used to vectorize the product description to obtain the semantic representation of PCs, and a machine learning-based PC evaluation model is built using the concept vector as features. Finally, a PC element library is built based on Word2Vec, and the tabu search algorithm is applied to identify the optimal combination of concept elements, determining the most favorable combination of PCs for the new product. Results indicate that the PC evaluation model based on multilayer perceptron achieves an average accuracy of 85.62% in predicting the quartiles of sales in the case of middle-aged and elderly home products, with the area under the receiver operating characteristic curve ranging from 0.96 to 0.99. The proposed PCD method can produce novel PCs with good market potential and a high degree of automation, improving the time efficiency and quality of PCD.
This article examines the decision-making processes in open innovation labs (OI-labs) in government. Through a qualitative single case study, we explore how the use of causal and effectual reasoning, as dichotomous logics, evolves over time and is manifested in the form of organizational practices to tackle temporal, relational, and cultural complexity. The findings reveal three episodes: the conceptualizing of the lab (predominantly causation), the building of the lab (predominantly effectuation), and the sustaining of the lab (hybrid causation-effectuation). Moreover, shifts in the logic are aimed at addressing different types of complexity, and over time, a hybrid logic emerges.
There is an increasing interest in financial text mining tasks. Significant progress has been made by using deep learning-based models on a generic corpus, which also shows reasonable results on financial text mining tasks such as financial sentiment analysis. However, financial sentiment analysis is still demanding work because of the insufficiency of labeled data for the financial domain and its specialized language. General-purpose deep learning methods are not as effective mainly due to specialized language used in the financial context. In this study, we focus on enhancing the performance of financial text mining tasks by improving the existing pretrained language models via NLP transfer learning. Pretrained language models demand a small quantity of labeled samples, and they could be enhanced to a greater extent by training them on domain-specific corpora instead. We propose an enhanced model FinSentiment, which incorporates enhanced versions of a number of recently proposed pretrained models, such as BERT, XLNet, RoBERTa, GPT, Llama, and T5, to better perform across NLP tasks in financial domain by training these models on financial domain corpora. The corresponding finance-specific models in FinSentiment are called Fin-BERT, Fin-XLNet, Fin-RoBERTa, Fin-GPT, Fin-Llama, and Fin-T5, respectively. We also propose variants of these models jointly trained over financial domain and general corpora. Our finance-specific FinSentiment models, in general, show the best performance across three financial sentiment analysis datasets, even when only a subpart of these models is fine-tuned with a smaller training set. Our results exhibit enhancement for each tested performance criteria on the existing results for these datasets. Extensive experimental results demonstrate the effectiveness and robustness of especially RoBERTa pretrained on financial corpora. Overall, we show that NLP transfer learning techniques are favorable solutions to financial sentiment analysis tasks. Our source code has been deposited at https://github.com/seferlab/finsentiment .
Machine learning dominates automated property valuation, yet comprehensive comparisons of predictive models remain scarce. This study compares 28 rent prediction models using 79,735 Belgian residential rental properties from 2022. Predictive performance is evaluated with traditional and alternative metrics for train data, test data, and across deciles. The results confirm that tree-based ensemble models outperform others, with stacking and averaging yielding superior results at a higher computational cost. Furthermore, middle-range rents show better predictive accuracy than extremes. Traditional and alternative metrics provide consistent findings. These insights aid real estate stakeholders seeking to enhance their expert systems for real estate price modeling.
Data bias is a critical challenge in machine learning applications within the financial and insurance sectors, as it can lead to misleading risk assessments and inaccurate predictive models. A prevalent source of bias in real-world datasets is the imbalanced distribution of classes, which is particularly problematic in fraud detection, credit risk assessment, and claim prediction. Traditional approaches to handling imbalanced data often rely on undersampling or oversampling techniques. However, these methods may generate unrealistic minority class samples or fail to perform effectively when dealing with extreme class imbalances. In this paper, we propose a configurable technique based on the underbagging method, integrated with a classifier for highly imbalanced datasets. Our approach is designed to enhance the predictive accuracy of the minority class while maintaining robust performance for the majority class. We incorporate our methodology into a classification ensemble framework and evaluate its effectiveness by comparing it against 100 combinations of 10 different oversampling and undersampling techniques applied to 10 different machine learning algorithms. The evaluation is conducted on two highly imbalanced real-world datasets: one related to auto insurance claims and another focused on credit card fraud detection. Our statistical analysis demonstrates that Balanced Underbagged Ensemble achieves superior classification performance in terms of recall for both classes, regardless of the base machine learning model used within the ensemble. Furthermore, our method finds an optimal balance between classification performance and computational efficiency.
This paper presents an innovative approach to comprehensively and systematically evaluate manual journal entries (MJEs) and enhance the control procedures in auditing. The proposed approach combines quantitative and qualitative information to develop various Key Risk Indicators (KRIs) that provide insights into potential risks associated with MJEs. The approach incorporates textual analytics into traditional quantitative measures. Using the data obtained from a multinational company, the application of the proposed testing approach demonstrates its effectiveness in identifying potential high-risk MJEs and improving the company's journal entry testing and monitoring procedures. The findings contribute to current audit practices by offering a more efficient and comprehensive method for evaluating MJEs.
This paper describes some experimentation with the evolving ability of large language models to generate sentiment estimates. We find that current models seem to equal or even exceed the ability of human annotators in a case study of single sentiment sentences. In addition, using the large language models, we were able to identify a small number of sentences in the data set, where it appears that the annotator made errors in assessing the sentiment. Unfortunately, analysis of the LLM results also illustrates apparent cognitive biases in the LLM behavior. Those effects appear to include an “ostrich effect” and a “no one is good enough” effect cognitive bias in LLM sentiment estimates.
This paper examines the prediction of IPO withdrawal using machine learning methods (lasso and random forest) and conventional regression (logit). The dataset comprises 2444 US first-time IPOs from 1997 to 2014. Results show that random forest outperforms both logit and lasso in in-sample and cross-sectional out-of-sample predictions when the training and test sets are drawn from the same time period. However, when models are trained on past data and tested on future observations, all models fail to accurately predict IPO withdrawal. This failure is attributed to concept drift-a change in the relationship between predictors and IPO withdrawal over time. I show that concept drift occurs at multiple points in time, affects various predictors, and persists even when accounting for economic shocks, institutional changes, or different prediction horizons. These findings suggest that the generalizability of previous results on IPO withdrawal is limited, as the relationship between various predictors and IPO withdrawal seems to vary across time periods.
Class imbalance remains a persistent challenge in predictive modeling, often leading to biased machine learning outcomes that disproportionately favor the majority class. This study investigates the effectiveness of advanced resampling techniques-both undersampling and oversampling-across two large and highly imbalanced datasets involving credit and loan default prediction. In addition to evaluating established oversampling techniques, the study introduces and validates a novel resampling approach, Deep Adaptive Resampling Technique (DART). Each technique is assessed using a consistent suite of classifiers, including logistic regression, gradient descent, na & iuml;ve Bayes, random forest, CatBoost, and artificial neural networks. The results reveal that K-MeansSMOTE and NearMiss outperform other resampling strategies in oversampling and undersampling, respectively, by achieving balanced trade-offs in precision, recall, F1 score, AUC, and Matthews correlation coefficient. Notably, DART demonstrates exceptional performance across both datasets, achieving nearly perfect classification scores across all metrics, suggesting strong generalizability and robustness. The study further analyzes the strengths and limitations of each resampling technique and emphasizes the importance of metric selection when evaluating imbalanced datasets. By integrating empirical evaluation with theoretical insights, this research contributes to the growing body of literature on imbalanced learning and offers practical guidance for selecting appropriate resampling strategies. These findings have broader implications for domains such as finance, healthcare, and fraud detection, where class imbalance is common. Overall, the study affirms the value of hybrid and adaptive resampling methods in building more accurate and generalizable predictive models.
This study investigates the daily price patterns and behavioral similarities among cryptocurrencies, focusing on two key research questions: (1) Do cryptocurrency prices vary consistently throughout the day? (2) Can cryptocurrencies be meaningfully grouped based on their behavioral patterns? Using Gaussian mixture models (GMMs), we analyze the opening, closing, high, and low prices of a broad range of cryptocurrencies. The findings reveal that while opening prices exhibit uniform patterns, closing, high, and low prices show more complex, multi-component behaviors, reflecting diverse market dynamics throughout the day. Consensus clustering identifies four distinct cryptocurrency clusters, each demonstrating unique price behaviors, challenging the notion of cryptocurrencies as a homogeneous group. The results suggest that cryptocurrencies behave as differentiated financial products, influenced by factors such as volatility, adoption, and technology. These findings contribute to the understanding of cryptocurrency market dynamics and have implications for investment strategies, risk management, and regulatory approaches.
This paper evaluates the use of supervised machine learning to automatically identify going concern–modified audit reports. Models based on two different classifiers—logistic regression and extreme gradient boosting—achieve strong classification performance for this task. The same classifiers, along with naïve Bayes, also demonstrate strong performance in the ancillary task of identifying audit report pages in financial reports. These results have practical implications, including the application of the presented methods for timely accounting information retrieval for users, automated peer comparison for auditors, or as a data extraction method for researchers, particularly in settings with limited audit data availability.
The objective of our study is to demonstrate the feasibility of predicting chief executive officer (CEO) compensation by exploring various machine learning methods. In our analysis, we examine six models: k -nearest neighbors, random forest, decision tree, extra trees, extreme gradient boosting, and support vector machines regressors. We find that XGBoost, random forest, and extra trees regressors exhibit the highest predictive power with the lowest error. Decision tree feature importance analyses identify firm size, CEO age, tangibility, cash holding, and Tobin's Q as key factors in predicting CEO compensation. We also conduct ordinary least squares regressions and find that the significance levels of the coefficients are comparable to the feature importances from the machine learning analysis. Our feature-grouping analysis shows that firms' financial performance and economic characteristics play the most significant role in determining CEO compensation, followed by board characteristics. The analysis further indicates that the predictive power of random forest and extra trees is stronger for forecasting next year's compensation than for predicting the current year's. These findings are valuable for compensation consultants and stakeholders involved in benchmarking decisions.