We compare different models on estimating the probability of default over a time horizon, considering censored data. Models were fitted from a survival analysis perspective, considering both classic models (Cox Proportional Hazard) with penalized variations and ensemble machine learning methods (boosting and bagging). Using a dataset of credit card refinancing operations, we assess accuracy performance in both out-of-sample and out-of-time observations. We assess well-established metrics, such as the concordance index, and measures that indicate calibration power (integrated Brier score and time-dependent dynamic AUC). Results show that a boosting approach with component-wise regression as base learner outperform other models for short term operations (36 months), in contrast to longer term transactions (60 months), where Cox Proportional Hazard (with and without penalization) depicted better results.
The exponential growth of the agricultural industry in the Amazon region has brought about notable economic advancements. However, this growth has substantially cost the region’s ecosystems, manifesting in increased deforestation and biodiversity degradation within the Amazon forest. This article is dedicated to assessing the eco-efficiency of agricultural production in Amazon Biome municipalities. It places particular emphasis on identifying critical determinants through the utilization of the classic Data Envelopment Analysis (DEA) model for efficiency computation, super-efficiency models for distinctive characterization, bootstrap computational techniques for robust resampling, and the Malmquist index for calculating annual eco-efficiency indices of each Decision-Making Unit (DMU). An exploration of the correlation between efficiency and meteorological attributes of the municipalities is conducted. The findings of this study reveal the following significant points: Eco-efficient municipalities within the Amazon Biome can serve as benchmarks for other DMUs striving to attain optimal input–output levels, most municipalities in the Amazon Biome operate close to the productive frontier due to the prevalent technology employed in their agricultural activities, the nature of the technological frontier’s return suggests that small and large DMUs possess eco-efficiency potential, and the current dataset does not yield conclusive evidence regarding a direct correlation between the variables. Leveraging this information, strategic pathways can be formulated to drive economic development in tandem with the sustainability of Amazon Biome municipalities. These strategies promise to foster social, economic, and environmental benefits for the populace while providing valuable insights to inform future research within this thematic domain.
Portfolio optimization relies in three main elements: the utility function, the use of multivariate dependence measure between assets and the returns of the assets. The first problem is well known in finance since one can use several different types of portfolio utility functions depending on her objective function and constraints. The two remaining problems are relevant topics of research: the measure and forecast of the dependence between assets and the forecast of the returns of assets. In this paper, we present a new method with artificial intelligence to estimate the variance and covariance matrices of financial time series data accounting for small sample size, large number of assets and presences of outliers. We compare results using different approaches, e.g., Minimum covariance determinant (MCD), Minimum Volume Ellipsoid (MVE), Minimum Regularized Covariance Determinant (MRCD) and Orthogonalized Gnanadesikan-Kettenring (OGK). The proposed approach is robust under several financial stylized facts and presents good performance with respect to the cumulative returns in the US markets (NASDAQ Stock Exchange) for 95 assets during 2015 to the beginning of 2019. These study contributes to fill in the literature gap in covariance estimation for small sample size, large number of assets and presences of outliers and contributes to investors to making better portfolio management.
We use regularized machine learning models to forecast Brazilian power electricity consumption for short and medium terms. We compare our models to benchmark specifications such as Random Walk and Autoregressive Integrated Moving Average. Our results show that machine learning methods, especially Random Forest and Lasso Lars, give more accurate forecasts for all horizons. Random Forest and Lasso Lars managed to keep up with the trend and the seasonality for various time horizons. The gain in predicting PEC using machine learning models relative to the benchmarks is considerably higher for the very short-term. Machine learning variable selection further shows that lagged consumption values are extremely important for very short-term forecasting due to the series high autocorrelation. Other variables such as weather and calendar variables are important for longer time horizons.
This paper discusses the application of ensemble techniques for the prediction of time series, presenting an in-depth review of the main techniques and algorithms used by the recent literature, with emphasis on the bootstrap aggregation (bagging) and boosting approaches. We also discuss the theoretical foundations of the ensemble-based models, presenting measures of model stability and the main aggregation methods to combine the forecasts of the individual models, as well as recommendations for future developments for related research agendas.
We investigated happiness at work reflected in the indexes presented by the 100 best companies to work for in the United States, as well as its impact on Return on Assets (ROA). Indicators used by Glassdoor were collected, as well as financial data from the Eikon. We performed a confirmatory factor analysis (CFA) to obtain the Happiness at Work General Index. The results were applied to regression to understand the relationship between the index and the ROA. Happiness at Work General Index and ROA are positively related, providing evidence that happiness at work positively impacts the operational performance of companies.
This study used real data from a Brazilian financial institution on transactions involving Consumer Direct Credit (CDC), granted to clients residing in the Distrito Federal (DF), to construct credit scoring models via Logistic Regression and Geographically Weighted Logistic Regression (GWLR) techniques. The aims were: to verify whether the factors that influence credit risk differ according to the borrower’s geographic location; to compare the set of models estimated via GWLR with the global model estimated via Logistic Regression, in terms of predictive power and financial losses for the institution; and to verify the viability of using the GWLR technique to develop credit scoring models. The metrics used to compare the models developed via the two techniques were the AICc informational criterion, the accuracy of the models, the percentage of false positives, the sum of the value of false positive debt, and the expected monetary value of portfolio default compared with the monetary value of defaults observed. The models estimated for each region in the DF were distinct in their variables and coefficients (parameters), with it being concluded that credit risk was influenced differently in each region in the study. The Logistic Regression and GWLR methodologies presented very close results, in terms of predictive power and financial losses for the institution, and the study demonstrated viability in using the GWLR technique to develop credit scoring models for the target population in the study.
Human motion analysis provides useful information for the diagnosis and recovery assessment of people suffering from pathologies, such as those affecting the way of walking, i.e., gait. With recent developments in deep learning, state-of-the-art performance can now be achieved using a single 2D-RGB-camera-based gait analysis system, offering an objective assessment of gait-related pathologies. Such systems provide a valuable complement/alternative to the current standard practice of subjective assessment. Most 2D-RGB-camera-based gait analysis approaches rely on compact gait representations, such as the gait energy image, which summarize the characteristics of a walking sequence into one single image. However, such compact representations do not fully capture the temporal information and dependencies between successive gait movements. This limitation is addressed by proposing a spatiotemporal deep learning approach that uses a selection of key frames to represent a gait cycle. Convolutional and recurrent deep neural networks were combined, processing each gait cycle as a collection of silhouette key frames, allowing the system to learn temporal patterns among the spatial features extracted at individual time instants. Trained with gait sequences from the GAIT-IT dataset, the proposed system is able to improve gait pathology classification accuracy, outperforming state-of-the-art solutions and achieving improved generalization on cross-dataset tests.
This study aimed to verify whether the use of support vector regression (SVR) makes the portfolio's return exceed the market. For such proposal, SVR was applied for 15 different kernel functions to select the best stocks for each quarter, calculating the quarterly portfolio return and cumulative return along the period. Subsequently, the returns of these portfolios were compared with the returns of a market benchmark. White's (2000) test was applied to avoid the data-snooping effect in assessing the statistical significance of the portfolios developed by the training strategies. The portfolio selected by SVR with inverse multiquadric kernel presented the highest cumulative return of 374.40% and a value at risk (VaR) of −6.87%. The results of this study corroborate the superiority hypothesis of the innovative method of SVR in the formation of portfolios, thus constituting a robust predictive method capable to cope with high dimensionality interactions.
Resumo Este artigo especifica e valida um modelo derivado de abordagens teóricas na literatura para mensurar as capacidades do Estado, especificamente do governo federal brasileiro. Dados coletados por Survey foram analisados usando a técnica de modelagem de equações estruturais (MEE). Os achados indicam que as características weberianas da burocracia ainda são uma referência útil para estudos sobre a capacidade estatal, uma vez que o nível de profissionalização e de habilidades dos burocratas apresentaram efeito positivo e estatisticamente significativo sobre o desempenho percebido do Estado. No que diz respeito à autonomia burocrática, os achados indicam que seu efeito no desempenho do Estado é indireto, mediado pela profissionalização. Ao contrário das previsões teóricas, não encontramos efeitos diretos significativos entre os relacionamentos da burocracia com atores não estatais e desempenho do Estado nem entre este e a dotação de recursos organizacionais. O artigo contribui para a literatura ao utilizar dados obtidos diretamente dos burocratas, ao desenvolver e validar um modelo replicável que relaciona as diferentes dimensões do conceito de capacidades estatais e ao utilizar a MEE para estimar os efeitos das dimensões do conceito sobre os resultados da ação estatal.
This paper uses particle filter to estimate daily volatility in the Brazilian financial stocks market and obtain an optimal allocation of assets via Monte Carlo approach. Our volatility model outperforms the Kalman filter besides overcoming non-additivity and non-Gaussian disturbance pattern. The historical statistics use an optimist Black-Litterman priori view to systematise our analysis in a rolling window. Our proposed method has better out-of-sample metrics than Markowitz, Naive (equal assets weight) and Bovespa Index benchmark.
Several pathologies can alter the way people walk, i.e., their gait. Gait analysis can be used to detect such alterations and, therefore, help diagnose certain pathologies or assess people's health and recovery. Simple vision-based systems have a considerable potential in this area, as they allow the capture of gait in unconstrained environments, such as at home or in a clinic, while the required computations can be done remotely. State-of-the-art vision-based systems for gait analysis use deep learning strategies, thus requiring a large amount of data for training. However, to the best of our knowledge, the largest publicly available pathological gait dataset contains only 10 subjects, simulating five types of gait. This paper presents a new dataset, GAIT-IT, captured from 21 subjects simulating five types of gait, at two severity levels. The dataset is recorded in a professional studio, making the sequences free of background camouflage, variations in illumination and other visual artifacts. The dataset is used to train a novel automatic gait analysis system. Compared to the state-of-the-art, the proposed system achieves a drastic reduction in the number of trainable parameters, memory requirements and execution times, while the classification accuracy is on par with the state-of-the-art. Recognizing the importance of remote healthcare, the proposed automatic gait analysis system is integrated with a prototype web application. This prototype is presently hosted in a private network, and after further tests and development it will allow people to upload a video of them walking and execute a web service that classifies their gait. The web application has a user-friendly interface usable by healthcare professionals or by laypersons. The application also makes an association between the identified type of gait and potential gait pathologies that exhibit the identified characteristics.
In this chapter, we analyzed the impacts of the COVID-19 outbreak in Brazil, one of the most severely affected countries by the pandemic in the world. We discussed the overall characteristics of the pathogen (SARS-CoV-2) and its diffusion in Brazil over time and its main differences compared to other common infectious diseases in Brazil, such as Dengue and Zika. We then performed quantitative experiments to forecast the trend of confirmed cases and deaths from COVID-19 using a machine learning method for both country and state levels. We also estimated the instantaneous reproducing number over time for each Brazilian state and grouped them into clusters, aiming at identifying the heterogeneity between the Brazilian States. We also analyzed the magnitude of underreporting of COVID-19 cases in Brazil's most affected states during the early stages of the pandemic. Moreover, we summarized the impacts of the pandemic on Brazil's economic activities and social life, analyzing the dynamics of macroeconomic indicators, the effectiveness of the emergency financial assistance granted by the Brazilian government, and the relative level of adherence to social distancing in the Brazilian states since the beginning of lockdown measures. Finally, we presented the main challenges for Brazil after the pandemic fades away, focusing on the economic recovery and future public policies.
This paper analyzes the factor zoo, which has theoretical and empirical implications for finance, from a machine learning perspective. More specifically, we discuss feature selection in the context of deep neural network models to predict the stock price direction. We investigated a set of 124 technical analysis indicators used as explanatory variables in the recent literature and specialized trading websites. We applied three feature selection methods to shrink the feature set aiming to eliminate redundant information from similar indicators. Using daily data from stocks of seven global market indexes between 2008 and 2019, we tested neural networks with different settings of hidden layers and dropout rates. We compared various classification metrics, taking into account profitability and transaction costs levels to analyze economic gains. The results show that the variables were not uniformly chosen by the feature selection algorithms and that the out -ofsample accuracy rate of the prediction converged to two values - besides the 50% accuracy value that would suggest market efficiency, a "strange attractor"of 65% accuracy also was achieved consistently. We also found that the profitability of the strategies did not manage to significantly outperform the Buy -and -Hold strategy, even showing fairly large negative values for some hyperparameter combinations.
ABSTRACT The objective of this study is to propose a methodology that, using multiple decreases, in addition to classified by actuarial profile and source of social security costs, calculates actuarially fair and balanced rates for unscheduled collective costing benefits from Defined Contribution (DC) pension plans. There are no studies in Brazil about costing rates for benefits not scheduled in pension plans of the DC modality. Any institution that pays collective cost social security benefits must determine an actuarial rate that is not insufficient, generating a financial imbalance in the fund, nor excessive, compromising the participant’s income. This work is the first study on costing rates for collective costing benefits from pension plans with DC modalities. Actuarially fair rates are obtained considering multiple decreases and equalizing the present value of contributions and the present value of pension and disability benefits, classified by actuarial profile and source of social security cost. The specific balance rate is determined for each source of social security costs and is obtained considering the actuarially fair rates for each actuarial profile. The general balance rate is obtained by the marginal contribution of each specific balance rate. The proposed methodology was used to calculate the rates of unscheduled benefits with collective costing in DC modality plans. The proposed methodology estimated that the legal changes, resulting from Constitutional Amendment 103/2019, indirectly increased by more than 4% the general balance rate of the unscheduled benefits of the Supplementary Social Security Foundation of the Federal Public Servant of the Executive Branch of the Federal Government (FUNPRESP-Exe).
The study specified and validated a model derived from theoretical approaches in the literature to measure state capacities. Survey data were analyzed using structural equation modeling (SEM). The findings indicate that bureaucracy’s Weberian features are still a valuable reference for state capacity studies since bureaucrats’ professionalization and skills presented a positive and statistically significant effect on the state’s perceived performance. Concerning bureaucratic autonomy, the findings indicate that its effect on state performance is indirect, mediated by professionalization. Unlike theoretical predictions, we found no significant direct effects of bureaucracy relationships with non-state actors and of organizational resources for state performance. The article contributes to the literature by using data obtained directly from bureaucrats, developing and validating a replicable model that relates the different dimensions of the concept, and using SEM to estimate the effects of each of the concept’s dimensions of state capacity on the outcomes of State action.
Abstract: The objective of this study is to develop a quantitative tool, based on Machine Learning and Geomarketing to identify business opportunities and contribute to the strategic process of local choice of franchises’ network selecting regions that have a high demand forecast and a lack of product supply. In addition, we conducted a qualitative analysis of the selected business places based on defined criteria. This prediction is given by constructing a consumption pattern, defined by a classifier, based on the characteristics of the reserved rights. Initially, for a better understanding on this subject, a theoretical background was made covering the main concepts about Geomarketing and Machine Learning and its applications. After that for a demonstration of the results, we opted for the application of the method for the market of fine chocolates (Cacau-Show) in the Distrito Federal. The main databases used in this paper were Pesquisa de Orçamentos Familiares and from Instituto Brasileiro de Estatística e Geografia (IBGE). As a result, the Standardized Spend was obtained, which indicates the requirement for each Censitar Sector, as georeferenced information of the competition, containing 44 stores that have as their main product of fine chocolate, and as digital meshes of the Federal District. The crossing is available for the elaboration of a map that facilitates the identification of the business opportunities for the market of fine chocolates in the Distrito Federal, Brazil.
In this paper we present the RMCriteria package to support the decision making system in the R software environment for statistical computing and graphics. The RMCriteria is a support decision system that implements all PROMETHEE family methods and also graphical and sensitivity analysis assisting the analysts to make their own decision process and also facilitating the analysis of the decisions.
In this article, we studied 13 parametric CAViaR models for 27 stock's indices concerning the bias-variance dilemma, providing an empirical golden rule for choosing the CAViaR structure over unknown information distributional features of financial data. Our findings pointed out that the adaptive model should be chosen when no prior information is available since it presented the smallest MSE in 23 of 27 assets. Furthermore, we also noted that in most cases, the CAViaR models overestimate the validation value-at-risk. This might not be troublesome from a regulators' point of view, since firms and financial institutions that would use those models will likely overestimate risk and hence adopt more conservative politics. However, from the firm's point of view, this means that they will likely operate in a suboptimal risk regime which.