In many real-world scenarios, especially those involving privacy constraints or data summarization, data are available only in aggregated forms, such as histograms or frequency tables. This work introduces a novel Bayesian method for inferring the underlying population distribution by fitting a mixture model to binned data. While we focus on mixtures of normal distributions, the framework is flexible and can be extended to other distributional families. We place a prior distribution on the number of mixture components, accommodating both finite and countably infinite mixtures, and perform inference using reversible jump MCMC. The proposed approach demonstrates strong performance on large-scale data, showcasing the potential of nonparametric Bayesian modeling in practical applications. Furthermore, we extend the method to model multiple histograms simultaneously and cluster them using the Dirichlet process. This enables information sharing across populations and provides a principled posterior probability to assess homogeneity between groups. Some theoretical results supporting the performance of our proposed methodology are also discussed.
Time-dependent regionalization, or spatially restricted grouping, is a significant area of research focused on understanding the evolution of spatial clusters over time. In this study we adopt a probabilistic approach to regionalization, conceptualizing it as a random partition of geographic space at each time point, with the sequence of spatial partitions exhibiting time dependency. This methodology facilitates inference regarding the temporal dynamics of clusters. We employ a product partition prior for the random partitions at each time point, introducing temporal correlation among partitions through the temporal structure associated with prior cohesions. To explore partition search space effectively and ensure spatially constrained clustering, we utilize random spanning trees. This research is motivated by a pertinent applied problem: the identification of spatial and temporal patterns associated with mosquito-borne diseases. Given the overdispersion inherent in this type of data, we propose a spatiotemporal Poisson mixture model in which both mean and dispersion parameters vary according to spatiotemporal covariates. We apply the proposed model to analyze weekly reported cases of dengue from 2018 to 2023 in the Southeast region of Brazil. Additionally, we assess modeling performance using simulated data. Results indicate that our model is competitive in analyzing the temporal evolution of spatial clustering.
We introduce a computationally efficient Bayesian nonparametric model for spatio-temporal data clustering. The model is applied to the logarithm of weekly PM2.5 concentrations to cluster monitoring sites with similar pollution levels taking into account spatial dependence and temporal variations. We propose an Autoregressive Product Partition Model (ARPPM) prior to model spatial random partitions over time, which can integrate past observations and covariates. The associated MCMC algorithm can be executed in parallel, making it suitable for large longitudinal datasets. We validate the approach using air quality data from Lombardy, demonstrating its ability to identify meaningful clusters reflecting both spatial and temporal pollution trends.
More than half of the world's population is exposed to mosquito-borne diseases, leading to millions of cases and hundreds of thousands of deaths every year. Analyzing this type of data is complex and poses several interesting challenges, mainly due to the usually vast geographic area involved, the peculiar temporal behavior, and the potential correlation between infections. Motivation for this work stems from the analysis of tropical disease data, namely, the number of cases of dengue and chikungunya, for the 145 microregions in Southeast Brazil from 2018 to 2022. As a contribution to the literature on multivariate disease data, we develop a flexible Bayesian multivariate spatio-temporal model where temporal dependence is defined for areal clusters. The model features a prior distribution for the random partition of areal data that incorporates neighboring information. It also incorporates an autoregressive structure and terms related to seasonal patterns into temporal components that are disease- and cluster-specific. Furthermore, it considers a multivariate directed acyclic graph autoregressive structure to accommodate spatial and inter-disease dependence. We explore the properties of the model through simulation studies and show results that prove our proposal compares well to competing alternatives. Finally, we apply the model to the motivating dataset with a twofold goal: finding clusters of areas with similar temporal trends for some of the diseases and exploring the existence of correlation between two diseases transmitted by the same mosquito.
Spatially constrained clustering is an important field of research, particularly when it involves changes over time. Partitioning a map is not simple since there is a vast number of possible partitions within the search space. In spatio-temporal clustering, this task becomes even more difficult, as we must consider sequences of partitions. Motivated by these challenges, we introduce a Bayesian model for time-dependent sequences of spatial random partitions by proposing a prior distribution based on product partition models that correlates partitions. Additionally, we employ random spanning trees to facilitate the exploration of the partition search space and to guarantee spatially constrained clustering. This work is motivated by a relevant applied problem: identifying spatial and temporal patterns of mosquito-borne diseases. Given the overdispersion present in this type of data, we introduce a spatio-temporal Poisson mixture model in which mean and dispersion parameters vary according to spatio-temporal covariates. The proposed model is applied to analyze the number of dengue cases reported weekly from 2018 to 2023 in the Southeast region of Brazil. We also evaluate model performance using simulated data. Overall, the proposed model has proven to be a competitive approach for analyzing the temporal evolution of spatial clustering.
Overweight and obesity in adults are known to be associated with increased risk of metabolic and cardiovascular diseases. Obesity has now reached epidemic proportions, increasingly affecting children. Therefore, it is important to understand if this condition persists from early life to childhood and if different patterns can be detected to inform intervention policies. Our motivating application is a study of temporal patterns of obesity in children from South Eastern Asia. Our main focus is on clustering obesity patterns after adjusting for the effect of baseline information. Specifically, we consider a joint model for height and weight over time. Measurements are taken every six months from birth. To allow for data-driven clustering of trajectories, we assume a vector autoregressive sampling model with a dependent logit stick-breaking prior. Simulation studies show good performance of the proposed model to capture overall growth patterns, as compared to other alternatives. We also fit the model to the motivating dataset, and discuss the results, in particular highlighting cluster differences. We have found four large clusters, corresponding to children sub-groups, though two of them are similar in terms of both height and weight at each time point. We provide interpretation of these clusters in terms of combinations of predictors.
Incomplete covariate vectors are known to be problematic for estimation and inferences on model parameters, but their impact on prediction performance is less understood. We develop an imputation-free method that builds on a random partition model admitting variable-dimension covariates. Cluster-specific response models further incorporate covariates via linear predictors, facilitating estimation of smooth prediction surfaces with relatively few clusters. We exploit marginalization techniques of Gaussian kernels to analytically project response distributions according to any pattern of missing covariates, yielding a local regression with internally consistent uncertainty propagation that utilizes only one set of coefficients per cluster. Aggressive shrinkage of these coefficients regulates uncertainty due to missing covariates. The method allows in- and out-of-sample prediction for any missingness pattern, even if the pattern in a new subject's incomplete covariate vector was not seen in the training data. We develop an MCMC algorithm for posterior sampling that improves a computationally expensive update for latent cluster allocation. Finally, we demonstrate the model's effectiveness for nonlinear point and density prediction under various circumstances by comparing with other recent methods for regression of variable dimensions on synthetic and real data.
Hydrosalpinx is a fluid occlusion and distension of the fallopian tubes, often resulting from pelvic inflammatory disease, which reduces the success of artificial reproductive technologies (ARTs) by 50%. Tubal factors account for approximately 25% of infertility cases, but their underlying molecular mechanisms and functional impact on other reproductive tissues remain poorly understood. This proteomic profiling study applied sequential window acquisition of all theoretical fragment ion spectra mass spectrometry (SWATH-MS) to study hydrosalpinx cyst fluid and pre- and post-salpingectomy endometrial fluid. Among the 967 proteins identified, we found 19 and 17 candidate biomarkers for hydrosalpinx in pre- and post-salpingectomy endometrial fluid, respectively. Salpingectomy significantly affected 76 endometrial proteins, providing insights into the enhanced immune response and inflammation present prior to intervention, and enhanced coagulation cascades and wound healing processes occurring one month after intervention. These findings confirmed that salpingectomy reverses the hydrosalpinx-related functional impairments in the endometrium and set a foundation for further biomarker validation and the development of less-invasive diagnostic strategies for hydrosalpinx.
OBJECTIVE:The aim of our study was to assess if the addition of PRGF to healthy human sperm affects its motility and vitality.METHODS:This was a prospective study, with 44 sperm donors on whom sperm analysis was performed. Nine mL of blood was collected and PRGF was obtained using PRGF-Endoret® technology. The influence of different dilutions of PRGF (5%, 10%, 20%, 40%) applied to 15 sperm donors was compared, and sperm motility was assessed after 30 minutes. In the second part of the study, 29 sperm donors were studied to analyze the influence of 20% dilution of PRGF at 15, 30 and 45 minutes in fresh and thawed sperm samples. Motility was assessed after the addition of PRGF and after analysis each aliquot was frozen. After thawing, concentration and motility were assessed at the same time periods.RESULTS:There were no differences in sperm motility in fresh samples between dilutions of PRGF when assessed 30 minutes after administration, nor between them, nor when compared to the control group immediately prior to treatment. No trend was observed between motility and PRGF dilution in linear regression analysis. There were no significant differences in thawed samples.CONCLUSIONS:The administration of 20% PRGF dilution had no effect on sperm motility compared to samples without PRGF. In addition, there was no change in sperm vitality when comparing samples with and without PRGF. More studies focusing on subnormal sperm samples, analyzing different PRGF concentrations and increasing the number of study variables are needed.
More than half of the world's population is exposed to the risk of mosquito-borne diseases, which leads to millions of cases and hundreds of thousands of deaths every year. Analyzing this type of data is often complex and poses several interesting challenges, mainly due to the vast geographic area, the peculiar temporal behavior, and the potential correlation between infections. Motivation stems from the analysis of tropical diseases data, namely, the number of cases of two arboviruses, dengue and chikungunya, transmitted by the same mosquito, for all the 145 microregions in Southeast Brazil from 2018 to 2022. As a contribution to the literature on multivariate disease data, we develop a flexible Bayesian multivariate spatio-temporal model where temporal dependence is defined for areal clusters. The model features a prior distribution for the random partition of areal data that incorporates neighboring information, thus encouraging maps with few contiguous clusters and discouraging clusters with disconnected areas. The model also incorporates an autoregressive structure and terms related to seasonal patterns into temporal components that are disease and cluster-specific. It also considers a multivariate directed acyclic graph autoregressive structure to accommodate spatial and inter-disease dependence, facilitating the interpretation of spatial correlation. We explore properties of the model by way of simulation studies and show results that prove our proposal compares well to competing alternatives. Finally, we apply the model to the motivating dataset with a twofold goal: clustering areas where the temporal trend of certain diseases are similar, and exploring the potential existence of temporal and/or spatial correlation between two diseases transmitted by the same mosquito.
Model-based clustering is a powerful tool that is often used to discover hidden structure in data by grouping observational units that exhibit similar response values. Recently, clustering methods have been developed that permit incorporating an ``initial'' partition informed by expert opinion. Then, using some similarity criteria, partitions different from the initial one are down weighted, i.e. they are assigned reduced probabilities. These methods represent an exciting new direction of method development in clustering techniques. We add to this literature a method that very flexibly permits assigning varying levels of uncertainty to any subset of the partition. This is particularly useful in practice as there is rarely clear prior information with regards to the entire partition. Our approach is not based on partition penalties but considers individual allocation probabilities for each unit (e.g., locally weighted prior information). We illustrate the gains in prior specification flexibility via simulation studies and an application to a dataset concerning spatio-temporal evolution of ${\rm PM}_{10}$ measurements in Germany.
STUDY QUESTION:In lesbian couples, is shared motherhood IVF (SMI) associated with an increase in perinatal complications compared with artificial insemination with donor sperm (AID)?SUMMARY ANSWER:Singleton pregnancies in SMI and AID had very similar outcomes, except for a non-significant increase in the rate of preeclampsia/hypertension (PE/HT) in SMI (recipient's age-adjusted odds ratio (OR) = 1.9, 95% CI = 0.7-5.2; P = 0.19), but twin SMI pregnancies had a much higher frequency of PE/HT than AID twins (recipient's age-adjusted OR = 21.7, 95% CI = 2.8-289.4; P = 0.01).WHAT IS KNOWN ALREADY:Oocyte donation (OD) pregnancies are associated with an increase in perinatal complications, in particular, preterm delivery and low birth weight, and PE/HT. However, it is unclear to what extent these complications are due to OD process or to the conditions why OD was performed, such as advanced age and underlying health conditions. Unfortunately, the literature concerning perinatal outcomes in SMI is scarce.STUDY DESIGN, SIZE, DURATION:Retrospective study involving 660 SMI cycles (299 pregnancies) and 4349 AID cycles (949 pregnancies) assisted over a 10-year period.PARTICIPANTS/MATERIALS, SETTING, METHODS:All cycles fulfilling the inclusion criteria performed in lesbian couples seeking fertility treatment in 17 Spanish clinics of the same group. Pregnancy rates of SMI and AID cycles were compared. Perinatal outcomes were compared: gestational length, newborn weight, preterm and low birth rates, PE/HT rates, cesarean section rates, perinatal mortality, and newborn malformations.MAIN RESULTS AND THE ROLE OF CHANCE:Pregnancy rates were higher in SMI than in AID (45.3% versus 21.8%, P < 0.001). There was a non-significant trend to higher multiple rate in AID (4.7% versus 8.5%, P = 0.08). In single pregnancies, there were no differences between SMI and AID in gestational age (278 days (268-285) versus 279 (272-284), P = 0.24), preterm rate (8.3% versus 7.3%, P = 0.80), preterm <28 weeks (0.6% versus 0.4%, P = 1.00), newborn weight (3195 g (2915-3620) versus 3270 g (2980-3600), P = 0.296), low birth rate (6.4% versus 6.4%, P = 1.00), extremely low birth weight (0.6% versus 0.5%, P = 1.00), and the distribution of newborns by weight groups. Cesarean section rate, newborn malformation rate, and perinatal mortality were also similar in SMI and AID. Additionally, there was non-significant trend in hypertensive disorders to an increase in PE/HT among SMI (recipient's age-adjusted OR = 1.9, 95% CI = 0.7-5.2). Overall, perinatal data are consistent with what is reported in the general population. In twin pregnancies, the aforementioned perinatal parameters were also very similar in SMI and AID. However, SMI twin pregnancies had a very high risk of PE/HT when compared with AID (recipient's age-adjusted OR = 21.7, 95% CI = 2.8-289.4, P = 0.01).LIMITATIONS, REASONS FOR CAUTION:Our data regarding the pregnancy course were obtained from information registered in the delivery report as well as from what was reported by the patients themselves, so a certain degree of inaccuracy cannot be ruled out. Additionally, in some parameters, there was up to 10% of data missing. However, since the methodology of reporting was the same in SMI and AID groups, one should not expect a differential reporting bias. It cannot be ruled out that the risk of PE/HT in simple gestations would be significant in a larger study. Additionally, in the SMI group allocation to the transfer of 2 embryos was not randomized so some bias is possible.WIDER IMPLICATIONS OF THE FINDINGS:SMI, if single embryo transfer is performed, seems to be is a safe procedure. Double embryo transfer should not be performed in SMI. Our data suggest that the majority of complications in OD could be related more with recipient status than with OD itself, since with SMI (performed in women without fertility problems) the perinatal complications were much lower than usually described in OD.STUDY FUNDING/COMPETING INTEREST(S):No external funding was received. The authors declare that they have no conflict of interest.TRIAL REGISTRATION NUMBER:N/A.
Spatio-temporal areal data can be seen as a collection of time series which are spatially correlated, according to a specific neighbouring structure. Motivated by a dataset on mobile phone usage in the Metropolitan area of Milan, Italy, we propose a semi-parametric hierarchical Bayesian model allowing for time-varying as well as spatial model-based clustering. Our approach incorporates the notion of regimes that describe changing patterns over work and night hours as well as weekdays/weekends. Changes across regimes are considered by means of temporal changepoint components that allow for different hierarchical structures specified across time points. The changepoints might occur within fixed time windows over the day. The model features a novel random partition prior that incorporates the desired spatial features and encourages co-clustering based on areal proximity. We explore properties of the model by way of extensive simulation studies from which we collect valuable information. Finally, we discuss the application to the motivating data, where the main goal is to spatially cluster population patterns of mobile phone usage.
Regression is one of the most fundamental statistical inference problems. A broad definition of regression problems is as estimation of the distribution of an outcome using a family of probability models indexed by covariates. Despite the ubiquitous nature of regression problems and the abundance of related methods and results there is a surprising gap in the literature. There are no well established methods for regression with a varying dimension covariate vectors, despite the common occurrence of such problems. In this paper we review some recent related papers proposing varying dimension regression by way of random partitions.
Research question: Is magnetic-activated cell sorting (MACS) a safe semen sample processing technique for newborns and mothers when used for semen processing prior to intracytoplasmic sperm injection (ICSI) cycles?Design: This retrospective multicentre cohort study involved patients undergoing ICSI cycles with either donor or autologous oocytes from January 2008 to February 2020. They were divided into two groups: those who underwent standard semen preparation (reference group) and those who had an added MACS procedure (MACS group). A total of 25,356 deliveries were assessed in the case of cycles using donor oocytes, and 19,703 deliveries from cycles using autologous oocytes. Of these, 20,439 and 15,917, respectively, were singleton deliveries. Obstetric and perinatal outcomes were retrospectively assessed. All means, rates and incidences were computed per live newborn in each study group.Results: There were no significant differences between the main obstetric and perinatal morbidities affecting the mothers' and newborns' well-being between groups using either donated or autologous oocytes. There was a significant increase in the incidence of gestational anaemia in both subpopulations (donor oocytes P = 0.01; autologous oocytes P < 0.001). However, this incidence was within the estimated prevalence for gestational anaemia in the general population. There was a statistically significant decrease in preterm (P = 0.02) and very preterm (P = 0.01) birth rates in the MACS group in cycles using donor oocytes.Conclusions: The use of MACS during semen preparation before ICSI using either donor or autologous oocytes appears to be safe for the mothers' and newborns' well-being during pregnancy and birth. Nevertheless, a close follow-up of these parameters in the future is advised, especially concerning anaemia, in order to detect even smaller effect sizes.
Abstract Study question Does the use of platelet growth factors instilled into thin endometrium in oestrogen-replacement treatment cycles improve endometrial thickness? Summary answer The use of growth factors in thin endometrium improves endometrial growth What is known already Thin endometrium impairs implantation rates and IVF outcomes. To date there is no effective treatment to improve the proliferation. PRGF-Endoret is used in multiple clinical indications. The autologous collection of growth factors (PRGF), applied to the area to be treated, produces a controlled release of multiple molecules with regenerative capacity in the focus of the lesion. Study design, size, duration Prospective randomised clinical trial to determine the effects of PRGF-ENDORET therapy in patients with thin endometrium undergoing oestrogen treatment for embryo transfer. Twenty-two patients were included from March 2018 to May 2021. At the end of the trial, a complementary retrospective study was conducted analysing all the treatments of these patients without live birth(21) before and after participating in the study until December 2022. 15 embryotransfer in study group and 27 in control group. Participants/materials, setting, methods Patients of IVI Bilbao clinic with endometrium less than or equal to 5mm after 10 days of treatment. PRGF was prepared according to the PRGF-Endoret technique using the medical device developed by BTI Biotechnology Institute. The study group received three instillations of PRGF Endoret using a kitazato IUI cannula. The first instillation was performed with in situ activated coagulum and the 2nd and 3rd instillations with PRGF-rich supernatants that were previously stored at -20 °C. Main results and the role of chance The clinical trial showed a significant increase in endometrial thickness 1.3 ± 0.67 mm compared to the control group of 0.58 ± 0.51 mm. Prior to treatment the mean endometrial thickness was 4.29 ± 0.88 in the control group and 4.44 ± 0.40 in the study group. After randomisation the mean was 4.87 ± 0.76 in the control group and 5.74 ± 0.87 in the study group. 2 pregnancies only in the study group and 1 live birth. The retrospective study shows 0.7mm increase in mean endometrial thickness in the post-PRGF cycles in the study group patients. Of the 15 transfers that were performed a posteriori, there were implantation rate of 40% and a 20% live birth rate, 40% biochemical miscarriage and 20% clinical miscarriage, with a significant increase of 0.94mm compared to the endometrium of the instillation cycle. In the control group of patients who received a posteriori factors there was also a significant increase in thickness of 1.6mm compared to the study cycle with an implantation rate of 58.33% and a 33% live birth rate, 58.33 biochemical miscarriage and 25% clinical miscarriage. Comparing all the cycles of these patients before and after PRGF, no statistical significance was found, although there was an increase of 0.29mm. Limitations, reasons for caution The number of patients included was lower than initially planned, but once the statistical power was adjusted, its safety and efficacy was demonstrated. More studies should be done to demonstrate when to instil, how many times, when transfer, but it is clear that it is safe and effective. Wider implications of the findings The use of factors has been shown to increase endometrial thickness. It could be thought to open the door to favouring endometrial receptivity and therefore live birth. Trial registration number 2016-001716-38
The product partition model (PPM) is widely used for detecting multiple change points. Because changes in different parameters may occur at different times, the PPM fails to identify which parameters experienced the changes. To solve this limitation, we introduce a multipartition model to detect multiple change points occurring in several parameters. It assumes that changes experienced by each parameter generate a different random partition along the time axis, which facilitates identifying those parameters that changed and the time when they do so. We apply our model to detect multiple change points in Normal means and variances. Simulations and data illustrations show that the proposed model is competitive and enriches the analysis of change point problems.