Using novel data from a survey with 2,055 Brazilians (aged 16-26 years old), this article explores youth motives to voluntarily join military careers. Our results found that non-White men were more likely to enter the military. Economic affluence negatively influenced prospects for police careers while showing a non-significant positive effect on the armed forces. We also found that individuals with more conservative views and values are more likely to pursue roles in the military police or the armed forces. Interviewees with strong convictions about the negative effects of the COVID-19 pandemic on the job market also showed a higher likelihood of joining military careers. We interpret these findings in a context of extreme job insecurity and progressive militarization of Brazilian bureaucracy during Bolsonaro's administration (2018-2022). This study offers new evidence of youths' inclination to pursue a military career.
The COVID-19 pandemic has generated an unprecedented amount of epidemiological data. Yet, concerns regarding the validity and reliability of the information reported by health surveillance systems have emerged worldwide. In this paper, we develop a novel approach to evaluating data integrity by combining the Newcomb-Benford Law with outlier methods. We demonstrate the advantages of our framework using a case study from China. To ensure more robust findings, we employ multiple diagnostic procedures, including three conformity estimates, four goodness-of-fit tests, and two distance measures (Cook and Mahalanobis). To promote transparency, we have made all computational scripts publicly available. Our findings indicates a significant deviation in the distribution of new deaths from the theoretical expectations of Benford's Law. Importantly, these results remain accurate even when considering alternative model specifications and conducting various statistical tests. Furthermore, the procedures developed here are easily applicable in other areas of knowledge and can be scaled to assess data quality in both the public and private sectors.
In this paper, we critically reevaluate Koch and Okamura’s (2020) conclusions on the conformity of Chinese COVID-19 data with Benford’s Law. Building on Figueiredo et al. (2022), we adopt a framework that combines multiple tests, including Chi-square, Kolmogorov-Smirnov, Euclidean Distance, Mean Absolute Deviation, Distortion Factor, and Mantissa Distribution. The primary rationale behind employing multiple tests is to enhance the robustness of our inference. The main finding of the study indicates that COVID-19 infections in China do not adhere to the distribution expected under Benford’s Law, nor does it align with the figures observed in the U.S. and Italy. The usefulness of deviations from Benford’s Law in detecting misreported or fraudulent data remains controversial. However, addressing this question requires a more careful statistical analysis than what is presented in the Koch and Okamura (2020) paper. By employing a combination of several tests using fully transparent procedures, we establish a more reliable approach to evaluating conformity to the Newcomb-Benford Law in applied research.
Background: In Brazil, the Ministry of Health (MH) monitors leprosy using 15 indicators, with the aim of implementing and evaluating evidence-based public policies. However, an excessive number of variables can complicate the definition of objectives and verification of epidemiological goals. Methods: In this paper, we develop the Global Leprosy Assessment Index (GLAI), a composite measure that integrates two key dimensions for the control the disease: epidemiological and operational. Using a confirmatory factor analysis model to examine 2020 state-level data, we have standardized GLAI to a range of 0 to 1. Results: Higher values within this range indicate a greater severity of the disease. The mean value of the GLAI was 0.67, with a standard deviation of 0.22. Roraima has the highest value, followed by Paraíba with 0.88 while Tocantins records the lowest value of the indicator, followed by Mato Grosso with 0.14. The epidemiological and operational indicators have a positive but statistically insignificant correlation (r = 0.25; p-value = 0.20). Conclusions: The development of evidence-based public policies depends on the availability of valid and reliable indicators. The GLAI presented in this paper is easily reproducible and can be used to monitor the disease with disaggregated information. Furthermore, the GLAI has the potential to serve as a more robust parameter for evaluating the impact of actions designed to eradicate leprosy in Brazil.
Resumo Como inferir causalidade a partir de dados observacionais? Este artigo apresenta uma introdução intuitiva ao pareamento, técnica estatística útil para identificar relações causais em desenhos de pesquisa não experimentais. Metodologicamente, apresentamos as principais características do matching a partir de três exemplos: a) o efeito das escolas militares sobre aprendizagem; b) o impacto do Bolsa Família sobre a propensão em votar no Partido dos Trabalhadores; e c) a influência do gênero sobre o desempenho eleitoral. Mostramos a implementação computacional no R e explicamos a interpretação substantiva dos resultados. Com o objetivo de aumentar o potencial didático da pesquisa, disponibilizamos todos os materiais de replicação, o que facilita que estudantes e profissionais utilizem os dados e scripts em suas atividades de estudo e trabalho. Com este artigo, esperamos difundir o uso de técnicas quase-experimentais nas Ciências Sociais e incentivar a replicabilidade como estratégia de ensino de análise de dados.
Resumo Este artigo analisa o impacto da posição em que as questões são apresentadas sobre o desempenho dos estudantes no Exame Nacional do Ensino Médio (Enem) no Brasil em 2016. A partir de uma amostra de 4.427.790 casos, calculamos o índice de acerto por questão para os diferentes cadernos de prova da área de Matemática e suas Tecnologias. Os resultados indicam a presença do efeito fadiga na prova do Enem 2016, ou seja, a ordem de apresentação das questões afeta a proporção de respostas corretas, que diminui à medida que o item é apresentado mais próximo do final da prova. As evidências exploratórias também sugerem que o efeito fadiga se manifesta tanto em estudantes de baixo quanto de alto desempenho. Por exemplo, a posição do item reduziu o índice de acerto em até 18 pontos percentuais, controlando pelo nível de desempenho. Este artigo faz a primeira avaliação empírica do efeito fadiga no Enem e os resultados representam uma contribuição para a literatura sobre influências não cognitivas em avaliação e são úteis para fundamentar estudos mais sistemáticos sobre o impacto do efeito fadiga em testes padronizados de larga escala, inclusive para além do caso específico analisado. Ao final, sugerimos medidas que podem mitigar esse efeito no Enem.
Background Claims of inconsistency in epidemiological data have emerged for both developed and developing countries during the COVID-19 pandemic. Methods In this paper, we apply first-digit Newcomb-Benford Law (NBL) and Kullback-Leibler Divergence (KLD) to evaluate COVID-19 records reliability in all 20 Latin American countries. We replicate country-level aggregate information from Our World in Data. Results We find that official reports do not follow NBL’s theoretical expectations ( n = 978; chi-square = 78.95; KS = 4.33, MD = 2.18; mantissa = .54; MAD = .02; DF = 12.75). KLD estimates indicate high divergence among countries, including some outliers. Conclusions This paper provides evidence that recorded COVID-19 cases in Latin America do not conform overall to NBL, which is a useful tool for detecting data manipulation. Our study suggests that further investigations should be made into surveillance systems that exhibit higher deviation from the theoretical distribution and divergence from other similar countries.
Abstract This paper analyzes the impact of the position of questions on students’ performance on the National Secondary Education Examination (Enem) in 2016. From a sample of 4,427,790 cases, we calculated the hit rate per question for the different workbooks in the Mathematics and its Technologies test. The results indicate presence of the fatigue effect on the 2016 Enem, that is, the order in which the questions are presented affects the proportion of correct answers, which is diminished as an item is presented closer to the end of the test. The exploratory evidence also suggests that the fatigue effect is manifested in students of both low and high performance. For example, the position of an item reduced the hit rate up to 18%, controlling for performance level. This paper conducts the first empirical evaluation of the fatigue effect during the Enem. The results contribute to the literature on the non-cognitive influences in evaluation, being useful to substantiate more systematic studies on the fatigue effect’s impact on large-scale standardized tests, beyond the case analyzed. At the end, we suggest measures that can mitigate this effect during the Enem.
ObjectiveThis paper studies the integrity of the vote counting system in Brazil.MethodWe analyze data from the Superior Electoral Court (TSE) for the 2018 Brazilian presidential election to assess suspicious vote count patterns deploying five techniques commonly used to detect fraud: a) the second-digit Benford's law test; b) the last digit mean; c) frequency analysis of last digits 0 and 5; d) correlation between the percentage of votes and the turnout rate; and e) resampled Kernel density of the proportion of votes.ResultsThe results show that the second-digit distributions for the three most voted candidates – Jair Messias Bolsonaro (PSL), Fernando Haddad (PT), and Ciro Gomes (PDT) – conform to Benford's law. We also find that last digit means and last digit frequency are within normal parameters, indicating no irregularities. Similarly, the fingerprint plot indicates a correlation coefficient that is consistent with the theoretical expectation of a fair election. The resampled Kernel density suggests that the vote count was performed without statistically significant distortions. These results are robust at different levels of data aggregation (polling station and municipality).ConclusionThe joint application of digit-focused tests, regression-based techniques, and patterns in the distribution of vote-shares provide a more reliable method for detecting anomalous cases. Relying on this unified framework, we find no evidence of electoral fraud in the 2018 Brazilian presidential election. These results advance our current understanding of statistical forensics tools and may be easily replicated to examine electoral integrity in other countries.
O p-valor pode ser definido como uma probabilidade que informa o nível de incompatibilidade dos dados observados com um modelo teórico esperado. Por essa razão, atua como um dos principais parâmetros de significância estatística de pesquisas empíricas. Contudo, a utilização incorreta, associada a problemas como viés de publicação e ausência de padrões específicos de reprodutibilidade, tem gerado problemas em áreas do conhecimento. Objetivo do trabalho é discutir aspectos ligados à importância do p-valor nas pesquisas empíricas. O estudo é teórico-reflexivo, baseado nas recomendações da Associação Americana de Estatística sobre a correta interpretação do p-valor. Além disso, discutimos o papel da significância estatística a partir de uma perspectiva empírica. Em particular, o p-valor: (1) não informa a probabilidade de que a hipótese nula é verdadeira; (2) não indica que os resultados foram produzidos aleatoriamente; (3) não estima o tamanho do efeito observado; (4) não mensura a importância substantiva dos resultados; (5) nunca deve ser interpretado sozinho; (6) não deve ser interpretado quando os pressupostos de seu cálculo forem violados e (7) não pode ser interpretado quando se trabalha com a população. A discussão crítica sobre a utilização de testes de significância é sinal de maturidade estatística. Contudo, os pesquisadores não podem decidir sobre como utilizar o p-valor antes de compreenderem integralmente o seu papel na pesquisa empírica.
Objetivo: Este artigo estima o impacto das medidas de distanciamento social sobre a incidência de COVID-19 a partir de uma perspectiva multissetorial. Métodos: O desenho de pesquisa utiliza um modelo de regressão em painel para analisar a relação entre restrições de mobilidade em diferentes setores econômicos e a dinâmica longitudinal da doença nos estados do Brasil. Resultados: Os principais resultados indicam que apenas os coeficientes das variáveis que representam os setores de restaurantes (p-valor < 0,05), compras (p-valor < 0,05) e transporte (p-valor < 0,001) obtiveram significância estatística. Em especial, o transporte (std= -0,674) é a variável que mais influencia a variação do número de casos de COVID-19. Conclusões: As evidências reportadas nesta pesquisa podem auxiliar o processo de tomada de decisão dos gestores governamentais a respeito da eficácia de intervenções não farmacológicas como instrumento para reduzir a disseminação da COVID-19.
RESUMO Este artigo analisa a relação entre o saneamento básico e a disseminação da COVID-19 nas capitais brasileiras. Para tanto, estima-se o Índice de Acesso ao Saneamento Básico pela redução das dimensões cobertura do saneamento e qualidade da gestão, obtidas por dados disponíveis no Sistema Nacional de Informação sobre Saneamento. Em seguida, aferiu-se o nível de associação entre saneamento e taxas de incidência e mortalidade da doença em todas as capitais brasileiras entre março e setembro de 2020. Os resultados sugerem que Curitiba (0,824), Campo Grande (0,808) e Goiânia (0,794) lideram o ranking de acesso ao saneamento básico. Além disso, as evidências apontam para uma correlação negativa entre saneamento e taxas de incidência e mortalidade por COVID-19. Contudo, a significância estatística das estimativas varia em função do tempo. Esses achados estão alinhados com a literatura internacional, que identifica o acesso ao saneamento como uma medida chave de profilaxia de doenças infecciosas.
Apesar da crescente oferta de dados em formato de painel, ainda são raros os estudos no Brasil que combinam as dimensões transversal e longitudinal na mesma análise. Para se ter uma ideia, numa amostra de mais de 7 mil artigos publicados entre 2000 e 2018 em periódicos de CPRI, apenas 45 casos citavam técnicas específicas para lidar com observações de unidades espaciais (países, estados, pessoas) repetidas em intervalos regulares do tempo (anos, meses, dias). Diante dos benefícios inferenciais que este tipo de perspectiva pode proporcionar e da escassez de pesquisas sobre o tema, este artigo apresenta uma introdução à regressão de painel. Metodologicamente, sintetizamos as principais recomendações da literatura e mostramos a implementação no R Statistical com o pacote plm,indo desde a seleção de modelos até o tratamento dos dados e apresentação de resultados. Para aumentar o potencial pedagógico do trabalho, disponibilizamos os dados originais e scripts computacionais. Com este artigo esperamos difundir a utilização de análises longitudinais na pesquisa empírica em CPRI no Brasil.
The beta regression has been received considerable attention in the last decade because of its applications to proportional data in several fields. We study the variability of coronavirus death rates in the first wave of twenty European countries using the beta regression with two systematic components for the mean and dispersion parameters. We prove empirically that the population density, proportion of urban population, hospital beds per 100 thousand and running time explain the variability of the COVID-19 death rates in the first wave of these countries.
This article analyzes the relationship between basic sanitation and the spread of COVID-19 in Brazilian state capitals. For that, the Basic Sanitation Access Index is estimated based on the reduction in the dimensions of sanitation coverage and management quality, obtained from data available in the National Sanitation Information System. Then, the level of association between sanitation and disease incidence and mortality rates in all Brazilian capitals between March and September 2020 is measured. The results suggest that Curitiba (0.824), Campo Grande (0.808), and Goiania (0.794) lead the ranking of access to basic sanitation. Also, evidence points to a negative correlation between sanitation and COVID-19 incidence and mortality rates. However, the statistical significance of the estimates varies with time. These findings are in line with the international literature, which identifies access to sanitation as a key measure of infectious disease prophylaxis.
RESUMO Introducao: E se a minha variavel resposta for categorica binaria? Este artigo apresenta uma introducao intuitiva a regressao logistica, tecnica estatistica mais adequada para lidar com variaveis dependentes dicotomicas. Materiais e Metodos: estimamos o efeito dos escândalos de corrupcao sobre a chance de reeleicao de candidatos concorrentes a deputado federal no Brasil a partir dos dados de Castro e Nunes (2014). Em particular, mostramos a implementacao computacional no R e explicamos a interpretacao substantiva dos resultados. Resultados: disponibilizamos todos os materiais de replicacao, permitindo que estudantes e profissionais utilizem os procedimentos discutidos aqui em suas atividades de estudo e pesquisa. Discussao: esperamos incentivar o uso da regressao logistica e difundir a replicabilidade como ferramenta de ensino de analise de dados. PALAVRAS-CHAVE: regressao; regressao logistica; replicacao; metodos quantitativos; transparencia.
The intentional killing of one human being by its own kind is considered the worst of the crimes. Therefore, homicide prevention is a major concern for policy makers in both developing and developed countries. We propose regression modeling for the homicide rates in Brazil along with appropriately chosen distributions for these responses that are in agreement with the restriction of values to the unit interval. We adopt the beta and simplex regression models with systematic components for the mean and dispersion parameters to explain the homicide rates in 27 state capitals of Brazil from the following explanatory variables: time, Gini coefficient, municipal human development index (MHDI), illiteracy and poverty rates. We employ standard likelihood techniques, perform influence and residual analysis and calculate goodness-of-fit statistics to select the best regression to explain homicides rates in these capitals. We perform the computations in the R package. The main results suggest the following: the mean homicide rate is increasing over time; there is a negative correlation between MHDI and murder rate; the poverty has a quite small negative impact on the mean homicide rates in the beta regression. The Gini coefficient and the illiteracy and poverty rates explain the dispersion of the homicide rates.
Abstract We employ Newcomb–Benford law (NBL) to evaluate the reliability of COVID-19 figures in Brazil. Using official data from February 25 to September 15, we apply a first digit test for a national aggregate dataset of total cases and cumulative deaths. We find strong evidence that Brazilian reports do not conform to the NBL theoretical expectations. These results are robust to different goodness of fit (chi-square, mean absolute deviation and distortion factor) and data sources (John Hopkins University and Our World in Data). Despite the growing appreciation for evidence-based-policymaking, which requires valid and reliable data, we show that the Brazilian epidemiological surveillance system fails to provide trustful data under the NBL assumption on the COVID-19 epidemic.
ABSTRACT Introduction: What if my response variable is binary categorical? This paper provides an intuitive introduction to logistic regression, the most appropriate statistical technique to deal with dichotomous dependent variables. Materials and Methods: we estimate the effect of corruption scandals on the chance of reelection of candidates running for the Brazilian Chamber of Deputies using data from Castro and Nunes (2014). Specifically, we show the computational implementation in R and we explain the substantive interpretation of the results. Results: we share replication materials which quickly enables students and professionals to use the procedures presented here for their studying and research activities. Discussion: we hope to facilitate the use of logistic regression and to spread replication as a data analysis teaching tool.
In response to the COVID-19 pandemic, governments worldwide have implemented social distancing policies with different levels of both enforcement and compliance. We conducted an interrupted time series analysis to estimate the impact of lockdowns on reducing the number of cases and deaths due to COVID-19 in Brazil. Official daily data was collected for four city capitals before and after their respective policies interventions based on a 14 days observation window. We estimated a segmented linear regression to evaluate the effectiveness of lockdown measures on COVID-19 incidence and mortality. The initial number of new cases and new deaths had a positive trend prior to policy change. After lockdown, a statistically significant decrease in new confirmed cases was found in all state capitals. We also found evidence that lockdown measures were likely to reverse the trend of new daily deaths due to COVID-19. In São Luís, we observed a reduction of 37.85% while in Fortaleza the decrease was 33.4% on the average difference in daily deaths if the lockdown had not been implemented. Similarly, the intervention diminished mortality in Recife by 21.76% and Belém by 16.77%. Social distancing policies can be useful tools in flattening the epidemic curve.