A method based on Bayesian structural time series is proposed to predict healthcare usage trends and to test for changes in the series levels during or after an abnormal year, such as that of the 2020 COVID-19 pandemic. Our method can also serve to calculate correction factors for frequency count data that can be integrated in a preprocessing step before undertaking a cross-sectional statistical analysis, and, in this way, the impact of a shock can be eliminated. Here, adjustments are derived for a large private health insurer in Spain from estimates of average healthcare usage. Median claims rate levels in 2020 were 15 % down on 2019 figures, but rose in 2021 and 2022, when the rate was 11 % and 8 % higher than in 2019, respectively. Once the shock correction is incorporated in the preprocessing step, our approach is shown to outperform traditional time series techniques. Healthcare insurance usage in Spain did not fully go back to normal levels (assuming that pre-pandemic values represent normality) in 2022, with the exception of some patient groups and specific medical services. Our method can be implemented in other areas of risk analysis when frequency counts are exposed to shocks and it allows estimating the difference in claims volume between real figures and those estimated, had the shock not occurred.
Seasonal infectious-disease interventions are commonly evaluated with interrupted time-series or pre–post designs that align epidemics by calendar week. When epidemic onset, speed or peak timing differs between seasons, such comparisons confound a shift in epidemic phase with a change in disease burden. We propose a Bayesian causal count model in which season-specific affine transformations map calendar time to a latent epidemic clock, and intervention effects are estimated on that clock rather than on the calendar. The alignment is a model component rather than a preprocessing step, so uncertainty about epidemic timing propagates into every causal contrast. The model uses a negative-binomial observation distribution, hierarchical area, season and area-season effects, a shrunk Fourier epidemic curve, and a continuous programme-intensity exposure. Posterior g-computation yields prevented cases, prevented fractions, peak attenuation and epidemic displacement, under both a controlled contrast and a dynamic contrast that propagates disease history within each arm. A two-tier simulation study evaluates bias, root mean squared error, interval coverage and parameter recovery under stable timing, epidemic-clock variation, intensity-dependent ascertainment and area-level confounding. We illustrate the framework using open Catalan primary-care surveillance and respiratory syncytial virus immunisation data, with explicit attention to the overlap in programme intensity that identifies the effect.
ABSTRACT Gender-based violence refers to violence directed against a person because of that person’s gender or violence that affects persons of a particular gender disproportionately. It is estimated that 30% of women worldwide have suffered either physical and/or sexual violence in their lifetime. Primary Health Care could be one of the ideal places for the detection of these situations, but most of the cases remain undetected as the victims often decline to seek for medical care after suffering an event. This work shows that public primary health care system in Catalonia might be registering only around 50% of the cases currently, and it will take more than 20 years to see the whole picture of the phenomenon, and the situation could be the same in countries with similar socioeconomic contexts. We found in previous studies that gender-based violence cases are severely underregistered from the public health and judicial perspectives, on the basis of qualitative analyses and survey data. Furthermore, we propose a statistical modelling approach able to estimate the actual burden of this issue accurately. Our results show that awareness training campaigns focused on primary healthcare professionals are very effective in reducing the underreporting issue but should be conducted repeatedly and not only once.
Abstract Organisms‐related data often appear as counts. The Poisson distribution is the most popular choice for modelling count data, but this distribution assumes equidispersion, which is usually not satisfied in real‐world data. Deviations from the Poisson assumption lead to discrete‐valued distributions that can fit over‐ and/or underdispersion. Although models for count data with over‐dispersion have been widely considered in the literature, models for underdispersion—the opposite phenomenon—have received less attention because underdispersion is relatively common only in certain research fields, including ecology. The Good distribution is a flexible option for modelling count data with over‐dispersion or underdispersion, although no R packages are available so far offering functionalities such as calculating quantiles, probabilities, etc., of a Good distribution or providing a method for modelling a Good‐distributed output based on a number of potential predictors. This paper presents the R package good, which computes the standard probabilistic functions, generates random samples from a population following a Good distribution and estimates the Good regression.
"Response to Giraudo, Ricceri and Rosso (2022)." Communications in Statistics - Simulation and Computation, ahead-of-print(ahead-of-print), p. 1
Introducción: Basar procesos de toma de decisiones en datos que contienen errores e imprecisiones es inevitable en muchos de contextos por diferentes razones. La situación derivada de la pandemia mundial de COVID-19 es un claro ejemplo, donde los datos proporcionados por fuentes oficiales no siempre fueron fiables debido a problemas de recopilación de datos y a la alta proporción de casos asintomáticos. Objetivos: Cuantificar la gravedad de la información errónea en una serie temporal y reconstruir la evolución más probable del proceso, así como una discusión sobre los métodos estadísticos más adecuados para obtener predicciones en este contexto. Métodos: Se propone el uso de un modelo autoregresivo con heterocedasticidad condicional y estimación de los parámetros mediante Bayesian synthetic likelihood. Resultados: Solo alrededor del 51% de los casos de COVID-19 en el período 23 de febrero de 2020 al 27 de febrero de 2022 se notificaron en España, observándose también diferencias relevantes en la intensidad del subregistro entre comunidades autónomas. Conclusión: La metodología propuesta proporciona a los tomadores de decisiones en salud pública una valiosa herramienta para mejorar la evaluación de la evolución de una enfermedad bajo diferentes escenarios, ya que permite generar predicciones realistas en este contexto.
Imputation of missing values is a strategy for handling non-responses in surveys or data loss in measurement processes, which may be more effective than ignoring them. When the variable represents a count, the literature dealing with this issue is scarce. Likewise, if problems of over- or under-dispersion are observed, generalisations of the Poisson distribution are recommended for carrying out imputation. In order to assess the performance of various regression models in the imputation of a discrete variable compared to classical counting models, this work presents a comprehensive simulation study considering a variety of scenarios and real data. To do so we compared the results of estimations using only complete data, and using imputations based on the Poisson, negative binomial, Hermite, and COMPoisson distributions, and the ZIP and ZINB models for excesses of zeros. The results of this work reveal that the COMPoisson distribution provides in general better results in any dispersion scenario, especially when the amount of missing information is large. When the variable presenting missing values is a count, the most widely used method is to assume that a classical Poisson model is the best alternative to impute the missing counts; however, in real-life research this assumption is not always correct, and it is common to find count variables exhibiting overdispersion or underdispersion, for which the Poisson model is no longer the best to use in imputation. In several of the scenarios considered the performance of the methods analysed differs, something which indicates that it is important to analyse dispersion and the possible presence of excess zeros before deciding on the imputation method to use. The COMPoisson model performs well as it is flexible regarding the handling of counts with characteristics of over- and under-dispersion, as well as with equidispersion.
Introduction: It is well known that work has a great influence on the well-being of workers. In the aftermath of the COVID-19 pandemic, it seems evident that work organization, in particular, plays a key role to face and control a pandemic. Consequently, it is essential to establish specific and sustainable tools to further study the relationship between work organization and workers' health. The aim of this paper is to describe the study design and baseline data of the OTS PANEL ("OTS" stands for "Work Organization and Health" in Spanish). Methods: Panel -type cohort study to be carried out annually applying an online self-administered questionnaire. Work organization and health indicators and their corresponding questions were selected through a multistep process carried out by a team composed by professionals of different disciplines. The sample is composed of n = 1824 salaried workers, aged 25-64, residing in Spain. Results: Mean response time was 17.4 +/- 7 min (median 15.8). 84.6 % of the indicators had percentages of missing values lower than 3 %, with labor market insecurity being the highest (5.8 %). We compute 39 indicators in which, except for a few cases, women and manual workers show consistently worse results. Conclusions: OTS PANEL can represent a valuable information source in Spain to contribute to generate solid evidence for research and for decision-making to improve the living and health conditions of the working population.
Introduction: Basing decision-making processes on data containing errors and inaccuracies is unavoidable in many situations. The COVID-19 pandemic related data is a clear example, where the information provided by official sources was often unreliable due to data collection mechanisms and the amount of asymptomatic cases. Objectives: To estimate the amount of misreported data in a time series and reconstructing the most probable evolution of the process and provides a discussion on the more appropriate statistical methods able to yield reliable forecasts in this context. Methods: The usage of a model based on autoregressive conditional heteroskedastic time series is proposed, estimating the parameters by Bayesian synthetic likelihood. Results: Only around 51% of the cases of COVID-19 in the period from February 23rd, 2020 to February 27th, 2022 were observed in Spain, also detecting remarkable differences in the reporting issues between Autonomous communities. Conclusion: The presented method allows generating realistic predictions under different possible scenarios, and therefore it represents a valuable tool for policy makers in order to improve the evaluation of the evolution of a situation.
Background:Short-term forecasts of infectious disease burden can contribute to situational awareness and aid capacity planning. Based on best practice in other fields and recent insights in infectious disease epidemiology, one can maximise the predictive performance of such forecasts if multiple models are combined into an ensemble. Here, we report on the performance of ensembles in predicting COVID-19 cases and deaths across Europe between 08 March 2021 and 07 March 2022. Methods:We used open-source tools to develop a public European COVID-19 Forecast Hub. We invited groups globally to contribute weekly forecasts for COVID-19 cases and deaths reported by a standardised source for 32 countries over the next 1-4 weeks. Teams submitted forecasts from March 2021 using standardised quantiles of the predictive distribution. Each week we created an ensemble forecast, where each predictive quantile was calculated as the equally-weighted average (initially the mean and then from 26th July the median) of all individual models' predictive quantiles. We measured the performance of each model using the relative Weighted Interval Score (WIS), comparing models' forecast accuracy relative to all other models. We retrospectively explored alternative methods for ensemble forecasts, including weighted averages based on models' past predictive performance. Results:Over 52 weeks, we collected forecasts from 48 unique models. We evaluated 29 models' forecast scores in comparison to the ensemble model. We found a weekly ensemble had a consistently strong performance across countries over time. Across all horizons and locations, the ensemble performed better on relative WIS than 83% of participating models' forecasts of incident cases (with a total N=886 predictions from 23 unique models), and 91% of participating models' forecasts of deaths (N=763 predictions from 20 models). Across a 1-4 week time horizon, ensemble performance declined with longer forecast periods when forecasting cases, but remained stable over 4 weeks for incident death forecasts. In every forecast across 32 countries, the ensemble outperformed most contributing models when forecasting either cases or deaths, frequently outperforming all of its individual component models. Among several choices of ensemble methods we found that the most influential and best choice was to use a median average of models instead of using the mean, regardless of methods of weighting component forecast models. Conclusions:Our results support the use of combining forecasts from individual models into an ensemble in order to improve predictive performance across epidemiological targets and populations during infectious disease epidemics. Our findings further suggest that median ensemble methods yield better predictive performance more than ones based on means. Our findings also highlight that forecast consumers should place more weight on incident death forecasts than incident case forecasts at forecast horizons greater than 2 weeks. Funding:AA, BH, BL, LWa, MMa, PP, SV funded by National Institutes of Health (NIH) Grant 1R01GM109718, NSF BIG DATA Grant IIS-1633028, NSF Grant No.: OAC-1916805, NSF Expeditions in Computing Grant CCF-1918656, CCF-1917819, NSF RAPID CNS-2028004, NSF RAPID OAC-2027541, US Centers for Disease Control and Prevention 75D30119C05935, a grant from Google, University of Virginia Strategic Investment Fund award number SIF160, Defense Threat Reduction Agency (DTRA) under Contract No. HDTRA1-19-D-0007, and respectively Virginia Dept of Health Grant VDH-21-501-0141, VDH-21-501-0143, VDH-21-501-0147, VDH-21-501-0145, VDH-21-501-0146, VDH-21-501-0142, VDH-21-501-0148. AF, AMa, GL funded by SMIGE - Modelli statistici inferenziali per governare l'epidemia, FISR 2020-Covid-19 I Fase, FISR2020IP-00156, Codice Progetto: PRJ-0695. AM, BK, FD, FR, JK, JN, JZ, KN, MG, MR, MS, RB funded by Ministry of Science and Higher Education of Poland with grant 28/WFSN/2021 to the University of Warsaw. BRe, CPe, JLAz funded by Ministerio de Sanidad/ISCIII. BT, PG funded by PERISCOPE European H2020 project, contract number 101016233. CP, DL, EA, MC, SA funded by European Commission - Directorate-General for Communications Networks, Content and Technology through the contract LC-01485746, and Ministerio de Ciencia, Innovacion y Universidades and FEDER, with the project PGC2018-095456-B-I00. DE., MGu funded by Spanish Ministry of Health / REACT-UE (FEDER). DO, GF, IMi, LC funded by Laboratory Directed Research and Development program of Los Alamos National Laboratory (LANL) under project number 20200700ER. DS, ELR, GG, NGR, NW, YW funded by National Institutes of General Medical Sciences (R35GM119582; the content is solely the responsibility of the authors and does not necessarily represent the official views of NIGMS or the National Institutes of Health). FB, FP funded by InPresa, Lombardy Region, Italy. HG, KS funded by European Centre for Disease Prevention and Control. IV funded by Agencia de Qualitat i Avaluacio Sanitaries de Catalunya (AQuAS) through contract 2021-021OE. JDe, SMo, VP funded by Netzwerk Universitatsmedizin (NUM) project egePan (01KX2021). JPB, SH, TH funded by Federal Ministry of Education and Research (BMBF; grant 05M18SIA). KH, MSc, YKh funded by Project SaxoCOV, funded by the German Free State of Saxony. Presentation of data, model results and simulations also funded by the NFDI4Health Task Force COVID-19 (https://www.nfdi4health.de/task-force-covid-19-2) within the framework of a DFG-project (LO-342/17-1). LP, VE funded by Mathematical and Statistical modelling project (MUNI/A/1615/2020), Online platform for real-time monitoring, analysis and management of epidemic situations (MUNI/11/02202001/2020); VE also supported by RECETOX research infrastructure (Ministry of Education, Youth and Sports of the Czech Republic: LM2018121), the CETOCOEN EXCELLENCE (CZ.02.1.01/0.0/0.0/17-043/0009632), RECETOX RI project (CZ.02.1.01/0.0/0.0/16-013/0001761). NIB funded by Health Protection Research Unit (grant code NIHR200908). SAb, SF funded by Wellcome Trust (210758/Z/18/Z).
Background The problem of dealing with misreported data is very common in a wide range of contexts for different reasons. The current situation caused by the Covid-19 worldwide pandemic is a clear example, where the data provided by official sources were not always reliable due to data collection issues and to the high proportion of asymptomatic cases. In this work, a flexible framework is proposed, with the objective of quantifying the severity of misreporting in a time series and reconstructing the most likely evolution of the process. Methods The performance of Bayesian Synthetic Likelihood to estimate the parameters of a model based on AutoRegressive Conditional Heteroskedastic time series capable of dealing with misreported information and to reconstruct the most likely evolution of the phenomenon is assessed through a comprehensive simulation study and illustrated by reconstructing the weekly Covid-19 incidence in each Spanish Autonomous Community. Results Only around 51% of the Covid-19 cases in the period 2020/02/23–2022/02/27 were reported in Spain, showing relevant differences in the severity of underreporting across the regions. Conclusions The proposed methodology provides public health decision-makers with a valuable tool in order to improve the assessment of a disease evolution under different scenarios.
Epidemiological mathematical models have been proved crucial in supporting the decision-making of the health authorities during the COVID-19 pandemic. In this context, this work presents two contributions. The first one is a methodology to integrate different data sources into a single time series that provides realistic COVID-19 incidence rates considering both the reported and unreported cases in Spain and Comunidad de Madrid. The second contribution is a novel ensemble forecast model that uses as input the predictions of three different COVID-19 forecasts models. These approaches have been used to provide forecast predictions in the scope of PredCov project, supporting both the Spanish and the European Union -via the European Centre for Disease Prevention and Control-health authorities. The output generated by the ensemble model provides a combined -and more accurate-prediction of the COVID-19 incidence. This work includes a description of both contributions and discusses the results provided by them.
The aim of this study is to describe the discordance between the self-perceived risk and actual risk of HIV among young men who have sex with men (YMSM) and its associated factors. An online, cross-sectional study was conducted with 405 men recruited from an Argentinian NGO in 2017. Risk discordance (RD) was defined as the expression of the underestimation of risk, that is, as a lower self-perception of HIV risk, as measured with the Perceived Risk of HIV Scale, than the current risk of HIV infection, as measured by the HIV Incidence Risk Index. Multivariate logistic regression models were used to analyze the associations between the RD and the explanatory variables. High HIV risk was detected in 251 (62%), while 106 (26.2%) showed high self-perceived risk. RD was found in 230 (56.8%) YMSM. The predictors that increased RD were consistent condom use with casual partners (aOR = 3.8 [CI 95:1.5–11.0]), the use of Growler to meet partners (aOR = 10.38 [CI 95:161–121.94]), frequenting gay bars (aOR = 1.9 [95% CI:1.1–3.5]) and using LSD (aOR = 5.44 [CI 95:1.32–30.29]). Underestimation of HIV risk in YMSM is associated with standard HIV risk behavior and modulated by psychosocial aspects. Thus, prevention campaigns aimed at YMSM should include these factors, even though clinical practice does not. Health professionals should reconsider adapting their instruments to measure the risk of HIV in YMSM. It is unknown what score should be used for targeting high-risk YMSM, so more research is needed to fill this gap. Further research is needed to assess what score should be used for targeting high-risk in YMSM.
To estimate the prospective relationships between exposure to psychosocial risks dimensions included in the COPSOQ-Istas21 and the deterioration of general and mental health and sleep problems among workers residing in Spain.Cohort whose baseline corresponds to the 2016 Psychosocial Risks Survey with a new measurement after one year.Social capital and interpersonal relations and leadership dimensions, as well as work̶life conflict, were related to all health variables. Dimensions of work organization and job contents did it especially with the mental health, the quantitative demands with the general health and the emotional ones with the mental health. The dimensions related to job insecurity did not show relationships with health.The results obtained reinforce the role of the COPSOQ-Istas21 as a useful instrument for the evaluation and prevention of psychosocial risks at work.
Background Bed occupancy in the ICU is a major constraint to in-patient care during COVID-19 pandemic. Diagnoses of acute respiratory infection (ARI) by general practitioners have not previously been investigated as an early warning indicator of ICU occupancy. Methods A population-based central health care system registry in the autonomous community of Catalonia, Spain, was used to analyze all diagnoses of ARI related to COVID-19 established by general practitioners and the number of occupied ICU beds in all hospitals from Catalonia between March 26, 2020 and January 20, 2021. The primary outcome was the cross-correlation between the series of COVID-19-related ARI cases and ICU bed occupancy taking into account the effect of bank holidays and weekends. Recalculations were later implemented until March 27, 2022. Findings Weekly average incidence of ARI diagnoses increased from 252.7 per 100,000 in August, 2020 to 496.5 in October, 2020 (294.2 in November, 2020), while the average number of ICU beds occupied by COVID-19-infected patients rose from 1.7 per 100,000 to 3.5 in the same period (6.9 in November, 2020). The incidence of ARI detected in the primary care setting anticipated hospital occupancy of ICUs, with a maximum correlation of 17.3 days in advance (95% confidence interval 15.9 to 18.9). Interpretation COVID-19-related ARI cases may be a novel warning sign of ICU occupancy with a delay of over two weeks, a latency window period for establishing restrictions on social contacts and mobility to mitigate the propagation of COVID-19. Monitoring ARI cases would enable immediate adoption of measures to prevent ICU saturation in future waves.
Abstract Background The main goal of this work is to estimate the actual number of cases of Covid-19 in Spain in the period 01-31-2020/06-01-2020 by Autonomous Communities. Based on these estimates, this work allows us to accurately re-estimate the lethality of the disease in Spain, taking into account unreported cases. Methods A hierarchical Bayesian model recently proposed in the literature has been adapted to model the actual number of Covid-19 cases in Spain. Results The results of this work show that the real load of Covid-19 in Spain in the period considered is well above the data registered by the public health system. Specifically, the model estimates show that, cumulatively until June 1st, 2020, there were 2 425 930 cases of Covid-19 in Spain with characteristics similar to those reported (95% credibility interval: 2 148 261 2 813 864), from which were actually registered only 518 664. Conclusions Considering the results obtained from the second wave of the Spanish seroprevalence study, which estimates 2 350 324 cases of Covid-19 produced in Spain, in the period of time considered, it can be seen that the estimates provided by the model are quite good. This work clearly shows the key importance of having good quality data to optimize decision-making in the critical context of dealing with a pandemic.
Objective: To explore the decisional process of people living with human immunodeficiency virus (HIV) currently enrolled in antiretroviral clinical trials. Method: Cross-sectional retrospective study. Outcome variables were reasons to participate, perceived decisional role (Control Preference Scale), the Decisional Conflict Scale and the Decisional Regret Scale. Descriptive statistics were calculated, and associations among these variables and with sociodemographic and clinical characteristics were analyzed with non-parametric techniques. Results: Main reasons to participate were gratitude towards Fundaci?n Huesped (47%), the doctor?s recommendation (32%), and perceived difficulty to access treatment in a public hospital (28%). Most patients thought that they made their decision alone (54.8%) or collaboratively with the physician (43%). Decisional conflict was low, with only some conflict in the support subscale (median = 16.67). Education was the only significant correlate of the total decisional conflict score (higher in less educated patients; p = 0.018), whereas education, recent diagnosis, living alone, lower age, being man and doctor?s recommendation to go to Fundaci?n Hu?sped related to higher conflict in different subscales. Nobody regretted to participate. Conclusions: The decision making regarding participation in HIV trials, from the perspective of participants, was made respecting their autonomy and with very low decisional conflict. Currently, patients show no signs of regret. However, even in this favorable context, results highlight the necessity of enhancing the decision support in more vulnerable patients (e.g., less educated, recently diagnosed or with less social support), thus warranting equity in the quality of the decision making process. ? 2020 SESPAS. Published by Elsevier Espana, S.L.U. This is an open access article under the CC BY-NC-ND license (http://creativecommons.org/licenses/by-nc-nd/4.0/).
The problem of dealing with misreported data is very common in a wide range of contexts and for different reasons.This has been and still is an important issue for data analysts and statisticians as not accounting for it could led to biased estimates and conclusions, and in many cases that would have implications in a posterior decision making process, as we all have seen in the current worldwide Covid-19 pandemic.In the last few years, many approaches have been proposed in the literature to accomodate data presenting this issue, especially in the fields of epidemiology and public health but also in other areas as social science.In this work, a comprehensive review of the recently proposed methods based on mixture models for longitudinal data (correlated and uncorrelated) is presented and several examples of application are discussed, including several approaches to the burden of Covid-19 infection cases in Spain and different approaches to deal with underreported registries of human papillomavirus infections and genital warts in Catalunya.