We compare different models on estimating the probability of default over a time horizon, considering censored data. Models were fitted from a survival analysis perspective, considering both classic models (Cox Proportional Hazard) with penalized variations and ensemble machine learning methods (boosting and bagging). Using a dataset of credit card refinancing operations, we assess accuracy performance in both out-of-sample and out-of-time observations. We assess well-established metrics, such as the concordance index, and measures that indicate calibration power (integrated Brier score and time-dependent dynamic AUC). Results show that a boosting approach with component-wise regression as base learner outperform other models for short term operations (36 months), in contrast to longer term transactions (60 months), where Cox Proportional Hazard (with and without penalization) depicted better results.
This study introduces a machine learning competing risks survival analysis model aiming at exploring the Probability of Default component of credit risk. Due to modelling of a cumulative probability of default over time, the model is applicable to assess Lifetime Expected Credit Loss under the International Financial Reporting Standard (IFRS) 9 regulation for financial institutions. Whilst most credit models focus on the default event itself, in many loan transactions, there is a competing event affecting risk: the possibility of the borrower prepay their debt before maturity. In this case, credit risk ceases to exist. We derive a statistical model that supports handling competing risks (credit risk and prepayment risk) in a machine learning survival analysis setup. As there is no available implemented computer package or library, we build the computational algorithm with subdistribution hazards using boosting as an ensemble method. Results of the model are generated using a dataset of credit card refinancing operations of a US financial institution. We observe, comparing different survival analysis techniques, that ComponentWise Gradient Boosting (CWGB) models showed better performance on both scenarios (subdistribution hazards and cause-specific models), closely followed by cause specific Cox Proportional Hazards, and that Gradient Boosting Survival was outperformed in all comparisons. The derived model is useful to address the guidelines of the IFRS 9 for credit risk, taking into account the context of lifetime credit exposure.
The literature has few studies on the seasonality of tuberculosis (TB) in the southern hemisphere, entailing the fill of this knowledge gap. This study aims to analyze whether TB incidence in Brazilian capitals and the Federal District is seasonal. This is an ecological study of a time series (2001-2019) of TB cases, conducted with 26 capitals and the Federal District. The Ministry of Health database, with 516,524 TB cases, was used. Capitals and the Federal District were divided into five groups based on social indicators, disease burden, and the Koppen climate classification. The seasonal variation of TB notifications and group amplitude were evaluated. We found TB seasonality in Brazil with a 1% significance in all capital groups (Stability assumption and Krusall-Wallis tests, p < 0.01). In the combined seasonality test, capital groups A, D, and E showed seasonality, whereas groups B and C, its probability. Our findings showed that health service supply and/or demand - rather than climate - may be the most relevant underlying factor in TB seasonality. It is challenging to raise the other seasonal factors underlying TB seasonality in tropical regions in the Southern Hemisphere.
Existe uma limitação de trabalhos na literatura acerca da sazonalidade da tuberculose (TB) no hemisfério sul, o que torna necessário o preenchimento dessa lacuna de conhecimento para a região. O estudo objetiva analisar se existe sazonalidade da incidência de TB nas capitais brasileiras do Brasil e no Distrito Federal, por meio de um estudo ecológico de série temporal (2001-2019) dos casos da doença. Utilizou-se a base de 516.524 casos de TB do Ministério da Saúde. As capitais e o Distrito Federal foram distribuídos em cinco grupos, com base em indicadores sociais, carga da doença e classificação climática de Koppen. Avaliou-se a variação sazonal das notificações de TB e a amplitude sazonal por grupo. Identificou-se a presença da sazonalidade da TB no Brasil ao nível de significância de 1% em todos os grupos de capitais (teste de estabilidade assumida e Krusall-Wallis, p < 0,01) e, no teste combinado de sazonalidade, os grupos A, D e E de capitais mostraram presença de sazonalidade; e, provavelmente presentes, os grupos B e C. Os achados mostraram que é um desafio levantar os fatores sazonais subjacentes à sazonalidade da TB nas regiões tropicais do Hemisfério Sul: o clima pode não ser o fator subjacente mais relevante encontrado na sazonalidade da TB, mas sim a oferta e/ou procura por serviços de saúde.
BCB), que disponibiliza dados entre
This paper analyzes the factor zoo, which has theoretical and empirical implications for finance, from a machine learning perspective. More specifically, we discuss feature selection in the context of deep neural network models to predict the stock price direction. We investigated a set of 124 technical analysis indicators used as explanatory variables in the recent literature and specialized trading websites. We applied three feature selection methods to shrink the feature set aiming to eliminate redundant information from similar indicators. Using daily data from stocks of seven global market indexes between 2008 and 2019, we tested neural networks with different settings of hidden layers and dropout rates. We compared various classification metrics, taking into account profitability and transaction costs levels to analyze economic gains. The results show that the variables were not uniformly chosen by the feature selection algorithms and that the out -ofsample accuracy rate of the prediction converged to two values - besides the 50% accuracy value that would suggest market efficiency, a "strange attractor"of 65% accuracy also was achieved consistently. We also found that the profitability of the strategies did not manage to significantly outperform the Buy -and -Hold strategy, even showing fairly large negative values for some hyperparameter combinations.
This work aimed to reproduce the methodology of Carl Benedikt Frey and Michael Osborne of 2017 for estimating the automation probabilities of occupations in Brazil. These estimates are potentially important for professionals and policymakers because they can guide the career of a worker, as well as define priority courses that educational institutions should offer in order to maximize employment opportunities in the country. We consulted the opinion of 69 scholars and professionals that are experts in machine learning to ground the estimation the automation probability of Brazilian occupations. The findings indicate that a large part of the occupations can be automated in the next years. In addition, it can be seen that these professions with a higher risk of automation show a trend of growth over time, which may result in a high level of unemployment in the coming years if professionals and the government do not prepare for this scenario.
In this paper, we replicated the method applied by Frey and Osborne to investigate the automation probability of jobs in Brazil, using data from Brazilian labor market administrative records between 1986 and 2017. We categorized each job listed on Brazil's Occupational Classification System into Job Zones based on its technical qualifications requirements and estimated the future demand for workers of each Job Zone from 2018 to 2046. To estimate the probability of automation for each occupation, we first collected the expert opinion of 69 specialists on artificial intelligence and then applied a Gaussian process using as input the text description of each occupation. The results showed that in 2017 55% of all formally employed workers in Brazil are in jobs with high or very high risk of automation, a value consistent with similar works found in the literature for other countries. The findings of this paper can aid policy makers to anticipate potential increased unemployment for occupations with high risk of automation, anticipate future transformations of the Brazilian labor market, and consequently support the planning of economic and social interventions.