The scarcity of primary data and challenges in analyzing secondary data hinder comprehensive mental health monitoring and the development of evidence-based policies. In Brazil, although multiple health information systems capture critical mental health data, these systems lack integration, robust analysis, and systematic presentation of mental health indicators. This study introduces the Digital Public Mental Health Dataset, a comprehensive, transparent, and reproducible data resource that consolidates data from existing Brazilian health information systems. This dataset enables detailed tracking of mental health indicators, supports future research, and informs public health interventions to address mental health disparities.
Abstract Objective: Hospital admissions data is an important source of information for health surveillance and health policy makers. The Brazilian Health Ministry offers access to this data in scattered files using a legacy format, justifying the effort of this work to present this data in a unique, coherent and comprehensible format and infrastructure, suitable to its big data dimension and importance. Data Description: The produced dataset covers hospital admissions in the Brazilian Universal Health System as a whole, keeping the characteristics of the original files and additional new variables with valuable enrichments to understand the data better.
Multivariate time series find extensive applications in conjunction with machine learning methodologies for scenario forecasting across various domains. Nevertheless, certain domains exhibit inherent complexities and diversities, which detrimentally impact the predictive efficacy of global models. This ongoing study introduces a Subset Modeling Framework designed to acknowledge the inherent diversity within a domain’s multivariate space. Comparative assessments between subset models and global models are conducted in terms of performance, revealing compelling findings and suggesting the potential for further exploration and refinement of this novel framework.
Objectives : This article presents the process of extraction and treatment of two datasets from the General Ombudsman of the Brazilian Unified Health System (OUVSUS). The resulting datasets allow the analysis of manifestation characteristics and sociodemographic profile of the citizens that performed these manifestations. Data description: The first dataset depicts the characteristics of the manifestations registered by the General Ombudsman. Each row represents an individual manifestation and contains information such as the registration date, classification, input channel, and subject, among others. The second dataset is constituted of sociodemographic information for each citizen that performed a manifestation, and characteristics such as sexual orientation, race, age, and geographic location of the citizen are presented, among others.
Climate trends and weather indicators are used in several research fields due to their importance in statistical modeling, frequently used as covariates. Usually, climate indicators are available as grid files with different spatial and time resolutions. The availability of a time series of climate indicators compatible with administrative boundaries is scattered in Brazil, not fully available for several years, and produced with diverse methodologies. In this paper, we propose time series of climate indicators for the Brazilian municipalities produced using zonal statistics derived from the ERA5-Land reanalysis indicators. As a result, we present datasets with zonal statistics of climate indicators with daily data, covering the period from 1950 to 2022.
Introduction Epidemiology is considered both a field of research and a methodological approach within the broader health sciences. It aims to understand health-related events’ causes and effects and provide the evidence necessary to prevent disease and implement effective control and prevention strategies. One of the main focuses of epidemiology is identifying the determinant factors in the health situation of populations since health-related anomalies are not randomly distributed among people. This understanding brings up the necessity of considering each place’s particularities and observing the regularity of diseases in a population context. Methods We present the Contextual-Compositional Approach (CCA) for the discovery of associations between Health Indicators (HI) and Health Determinants (HD) for neonatal mortality rate monitoring in situations of anomalies. CCA uses time series concepts, anomaly detection, and data distribution between classes for studying HD under expected conditions and comparing them to the anomaly conditions indicated by the anomaly detection in the HI. CCA is evaluated using a neonatal mortality database in health facilities in Rio de Janeiro, Brazil. Results The results show that CCA can reveal essential associations between the health condition and the population’s social, economic, and cultural characteristics on different scales. Conclusion CCA stands out because it is easy to apply and understand, requiring little computational resources and parameters.
OBJECTIVES:The control chart is a classic statistical technique in epidemiology for identifying trends, patterns, or alerts. One meaningful use is monitoring and tracking Infant Mortality Rates, which is a priority both domestically and for the World Health Organization, as it reflects the effectiveness of public policies and the progress of nations. This study aims to evaluate the applicability and performance of this technique in Brazilian cities with different population sizes using infant mortality data. RESULTS:In this article, we evaluate the effectiveness of the statistical process control chart in the context of Brazilian cities. We present three categories of city groups, divided based on population size and classified according to the quality of the analyses when subjected to the control method: consistent, interpretable, and inconsistent. In cities with a large population, the data in these contexts show a lower noise level and reliable results. However, in intermediate and small-sized cities, the technique becomes limited in detecting deviations from expected behaviors, resulting in reduced reliability of the generated patterns and alerts.
Objectives Surveillance of infant and fetal deaths is of paramount importance in thinking about government strategies to reduce these rates, provide greater visibility of these mortality figures in the country, enable the adoption of prevention measures, as well as contribute to a better record of deaths. Data description The dataset comprises fetal, neonatal, early neonatal, late neonatal, and perinatal Mortality Rates of Brazilian municipalities with their respective information, between 2010 to 2020, aggregated by epidemiological week.
Objectives The National Registry of Healthcare Facilities is a system with the registry of every healthcare facility in Brazil with information on the capacity building and healthcare workforce regarding its public or private nature. Despite being publicly available, it can only be accessed in separated disjoint tables, with different primary units of analysis. The objective is to offer an interoperable dataset containing monthly data from 2005 to 2021 with information on healthcare facilities, including their physical and human resources, services and teams, enriched with municipal information. Data description Database with historical data and geographic information for each health facility in Brazil. It is composed by 5 distinct tables, organized according to combinations of time, space, and types of resources, services and teams. This database opens up a range of possibilities for research topics, from case studies in a single health facility and period, analysis of a group of health facilities with characteristics of interest, to a broader study using the entire dataset and aggregated data by municipality. Furthermore, the fact that there is a row for each health facility/month/year facilitates the integration with other datasets from the Brazilian healthcare system. In addition to being a potential object of study in the health area, the dataset is also convenient in data science, especially for studies focused on time series.
The Data Science Platform Applied to Health (PCDaS) is a research and technological development project that aims to develop and apply novel data analysis methods to public health data. It fills a technological gap between the variety of data sources available in legacy and unstandardized formats and the current needs and possibilities of Data Science applications to consume and explore data for the benefit of the Brazilian Health System. PCDaS provides democratic access to health-related datasets and information by requiring fewer technological abilities from its users while maintaining a continuously updated stack of technologies. As a data ecosystem, our primary goal is to provide secure and remote access to health data, technological tools, and a robust infrastructure provided by our platform to process and analyze a large amount of data that generally demand computational power often unavailable to researchers. The infrastructure consists of multi-region on-premise and cloud servers prepared to deal with the heavy analysis of Big Data from anywhere from multiple users simultaneously. Providing secure and remote access to health databases, whether in their original form or processed, is a daily breakthrough for a public health researcher. Knowing that there is a place where they can access integrated data in a standard format makes the research process much more manageable. To ensure quality, our data engineering and governance teams process these data sources following a gold standard based on cross-tables provided by the Health Ministry (the TabNET system) and decoding the original variables into meaningful names provided by the sources. It is very relevant to emphasize the comprehensive documentation of metadata, attributes, and the ETL (Extract, Transform, Load) process for databases. Every part of these steps is described in detail on the PCDaS website, ensuring the comprehension and reproducibility of the process. These features ensure that PCDaS users can effectively leverage the platform’s resources and capabilities, enabling them to conduct research, perform data analysis, and collaborate within a secure and supportive environment to contribute to the Brazilian Health System.
Abstract Wastewater Based Epidemiology (WBE) supports sanitary surveillance enablingan early identification of viral spread. The procedure involves the genome copy(GC) of SARS-CoV-2 capture from sewage samples infected by symptomatic andasymptomatic people. During the pandemic of COVID-19 in Brazil(2020/2021)the WBE studies followed different guidelines still incipient, revealinglow concern for a common agreed-upon procedure. As a result, when compilingthe available WBE data for the training of Artificial Intelligence (AI) models forCOVID-19 number of cases prediction, we found few quantity, obtained throughdifferent adopted procedures and difficult to co-related to extra information. Thelack of a common WBE procedure makes it hard to build useful predictiveMachine Learning (ML) models. In this context, guidelines that link ML andWBE are explored here. We aim at raising an alert highlighting the relevance onthe design of useful strategies to join WBE and IA. The proposal aims atstandardization and consistency without any detriment to the initial objective ofsurveillance. The approach includes processes related to: the definition of samplecollection approach, sampling frequency, related information, physical andtemporal sample characteristics, laboratory methods for genes amplification anddetection and results dissemination.
Resumo A investigação analisou a tendência da mortalidade por HIV/Aids segundo características sociodemográficas nos estados brasileiros entre 2000 e 2018. Estudo ecológico de série temporal das taxas padronizadas de mortalidade por Aids geral, por sexo, faixa etária, estado civil e raça/cor. Foi utilizado o modelo linear generalizado de Prais-Winsten. Os resultados do estudo evidenciaram que os estados com as maiores taxas foram Rio Grande do Sul, Rio de Janeiro, São Paulo e Santa Catarina. A tendência foi crescente nas regiões Norte e Nordeste. Os homens tiveram taxas mais elevadas quando comparados às mulheres e à população geral. Quanto às faixas etárias, as mais avançadas mostraram tendência a crescimento. A análise de acordo com o estado civil evidenciou taxas mais elevadas entre os não casados e tendência a crescimento concentrada nesta população. De acordo com raça/cor, identificou-se que os negros apresentaram maiores taxas, exceto no Paraná, e a tendência foi majoritariamente crescente. A mortalidade por HIV/Aids apresenta tendências distintas segundo as características sociodemográficas, verificando-se necessidade de ações de prevenção e cuidado aos homens, adultos, idosos, não casados e negros em vista de mudança no perfil da mortalidade.
This investigation analyzed the trend of HIV/AIDS mortality by sociodemographic characteristics in the Brazilian states from 2000 to 2018. This is an ecological study of time-series of standardized rates of mortality from AIDS overall, by gender, age group, marital status, and ethnicity/skin color, employing the Prais-Winsten generalized linear model. The results showed that the states with the highest rates were Rio Grande do Sul, Rio de Janeiro, São Paulo, and Santa Catarina. The trend was increasing in the North and Northeast. Men had higher rates than women and the general population. The most advanced age groups showed a growing trend. The analysis by marital status showed higher and growing rates among the unmarried. Blacks had higher rates, except for Paraná, with a mainly increasing trend. Mortality due to HIV/AIDS had different trends by sociodemographic characteristics, with a need for preventive and care actions for men, adults, older adults, unmarried, and black people due to the change in the mortality profile.
Our objective is to describe the differences in the sampling plans of the two editions of the Brazilian National Health Survey (PNS 2013 and 2019) and to evaluate how the changes affected the coefficient of variation (CV) and the design effect (Deff) of some estimated indicators. Variables from different parts of the questionnaire were analyzed to cover proportions with different magnitudes. The prevalence of obesity was included in the analysis since anthropometry measurement in the 2019 survey was performed in a subsample. The value of the point estimate, CV, and the Deff were calculated for each indicator, considering the stratification of the primary sampling units, the weighting of the sampling units, and the clustering effect. The CV and the Deff were lower in the 2019 estimates for most indicators. Concerning the questionnaire indicators of all household members, the Deffs were high and reached values greater than 18 for having a health insurance plan. Regarding the indicators of the individual questionnaire, for the prevalence of obesity, the Deff ranged from 2.7 to 4.2, in 2013, and from 2.7 to 10.2, in 2019. The prevalence of hypertension and diabetes per Federative Unit had a higher CV and lower Deff. Expanding the sample size to meet the diverse health objectives and the high Deff are significant challenges for developing probabilistic household-based national survey. New probabilistic sampling strategies should be considered to reduce costs and clustering effects.
Resumo Introdução O termo “big data” no ambiente acadêmico tem deixado de ser uma novidade, tornando-se mais comum em publicações científicas e em editais de fomento à pesquisa, levando a uma revisão profunda da ciência que se faz e se ensina. Objetivo Refletir sobre as possíveis mudanças que as ciências de dados podem provocar nas áreas de estudos populacionais e de saúde. Método Para fomentar esta reflexão, artigos científicos selecionados da área de big data em saúde e demografia foram contrastados com livros e outras produções científicas. Resultados Argumenta-se que o volume dos dados não é a característica mais promissora de big data para estudos populacionais e de saúde, mas a complexidade dos dados e a possibilidade de integração com estudos convencionais por meio de equipes interdisciplinares são promissoras. Conclusão No âmbito do setor de saúde e de estudos populacionais, as possibilidades da integração dos novos métodos de ciência de dados aos métodos tradicionais de pesquisa são amplas, incluindo um novo ferramental para a análise, monitoramento, predição de eventos (casos) e situações de saúde-doença na população e para o estudo dos determinantes socioambientais e demográficos.
Objectives Neonatal mortality is a global public health problem, and the efforts to reduce child mortality is one of the goals of the 2030 Agenda for Sustainable Development, launched in 2015 by the United Nations. The availability of historical neonatal mortality rates (NMR) data in Brazilian municipalities is crucial to evaluate trends at local, regional and national level, identifying gaps and vulnerable territories. Therefore, the objective of this article is to offer an integrated dataset containing monthly data in a historical series from 1996 to 2017 with information on all births, neonatal deaths, and NMR (total, early and late components) enriched with information related to the municipality. Data description It is a dataset of historical data with information on the number of births, the number of neonatal deaths, the neonatal mortality rate (including early and late), and geographic information for each month (between January 1996 and December 2017) and Brazilian municipality.
Due to its impact, COVID-19 has been stressing the academy to search for curing, mitigating, or controlling it. It is believed that under-reporting is a relevant factor in determining the actual mortality rate and, if not considered, can cause significant misinformation. Therefore, this work aims to estimate the under-reporting of cases and deaths of COVID-19 in Brazilian states using data from the InfoGripe. InfoGripe targets notifications of Severe Acute Respiratory Infection (SARI). The methodology is based on the combination of data analytics (event detection methods) and time series modeling (inertia and novelty concepts) over hospitalized SARI cases. The estimate of real cases of the disease, called novelty, is calculated by comparing the difference in SARI cases in 2020 (after COVID-19) with the total expected cases in recent years (2016–2019). The expected cases are derived from a seasonal exponential moving average. The results show that under-reporting rates vary significantly between states and that there are no general patterns for states in the same region in Brazil. The states of Minas Gerais and Mato Grosso have the highest rates of under-reporting of cases. The rate of under-reporting of deaths is high in the Rio Grande do Sul and the Minas Gerais. This work can be highlighted for the combination of data analytics and time series modeling. Our calculation of under-reporting rates based on SARI is conservative and better characterized by deaths than for cases.
In data analysis, the mining of frequent patterns plays an important role in the discovery of associations and correlations between data. During this process, it is common to produce thousands of association rules (ARs), making the study of each one arduous. This problem weakens the process of finding useful information. There is a scientific effort to develop approaches capable of filtering interesting patterns, balancing the number of ARs produced with the goal of not being trivial and known by specialists. However, even when such approaches are adopted, the number of produced ARs can still be high. This work contributes by presenting Divergent Association Rules Approach (DARA), a novel approach for obtaining ARs that presents themselves in divergence with the data distribution. DARA is applied right after traditional approaches to filtering interesting patterns. To validate our approach, we studied the dataset related to the occurrence of malaria in the Brazilian Legal Amazon. The discovered patterns highlight that ARs brought relevant insights from the data. This article contributes both in the medical and computer science fields since this novel computational approach enabled new findings regarding malaria in Brazil.
Objectives Malaria is an infectious disease that annually presents around 200,000 cases in Brazil. The availability of data on malaria is crucial for enabling and supporting studies that can promote actions to prevent it. Therefore, the goal of this paper is to contribute to such studies by offering an integrated dataset containing data on reported and suspected cases of malaria in the Brazilian Legal Amazon comprising the period from the years 2009 to 2019. Data description This paper presents a dataset with all medical records of patients who were tested for malaria in the Brazilian Legal Amazon from 2009 to 2019. The dataset has 40 attributes and 22,923,977 records of suspected cases of malaria. Around 12% of the data correspond to confirmed cases of malaria. The attributes include data regarding the notifications, examinations, as well as personal patient information, which are organized into health regions.