To compare ART cycle outcomes based on source of LH activity in mixed COH protocols Single center retrospective cohort study evaluating all ‘mixed protocol’ (using both FSH and LH receptor stimulation) COH cycles with planned egg retrieval, carried out from October 2022 to January 2023, during the hMG (menopur) shortage. The default protocol at our center uses hMG for LH activity and during this time, dhCG was substituted only for those patients without access, hence providing a native experiment. Outcomes were evaluated per retrieval and included gonadotropin doses, oocytes retrieved, oocyte maturity, positive serum hCG, biochemical and clinical pregnancy rates. These were compared using independent sample t-tests and Mann-Whitney U tests, as appropriate. Generalized estimating equation modeling was used to adjust for patient age, antral follicle count (AFC), anti-Mullerian hormone (AMH), and number of prior IVF cycles. 740 dhCG + FSH cycles were compared to 3,056 hMG + FSH cycles. The groups were similar in BMI, infertility diagnosis, baseline FSH, and gravidity. Patients using dhCG were younger (35.3 years vs 36.1, p=<0.001), with fewer prior IVF cycles (0.88 vs 1.2, p=<0.001), and higher AMH (3.2 vs 2.9, p=0.016) and higher AFC (16.7 vs 14.5, p=0<0.001). The dhCG + FSH group had a higher total number of oocytes retrieved (17.9 vs 15.0, p=<0.001), and number of M2 oocytes (12.3 vs 10.0, p=<0.001). These differences remained significant after adjustment. No difference was noted in duration of stimulation (13.97 vs 14.03d, p=0.51), number of M1 oocytes retrieved, rate of positive serum hCG per fresh transfer, biochemical pregnancy rate, and clinical pregnancy rate. An hMG shortage requiring replacement with dHCG for mixed COH ART protocols in a subset of patients provided a naturally-occurring opportunity to compare these medications. Patients stimulated with dhCG had a similar clinical pregnancy rate and more total oocytes and M2 oocytes retrieved, even after adjusting for confounding variables.
To evaluate the performance of the endometrial receptivity assay (ERA) as a clinical diagnostic test. The control arm of the Synchrony trial served as an a priori non-selection prospective cohort study, since the ERA was performed in both the control and intervention groups, but no intervention was made in the control group. Using these data, the sensitivity, specificity, negative and positive predictive values (NPV, PPV) and other markers were calculated. The primary outcome was sustained implantation of euploid blastocysts (clinical pregnancy with fetal heart rate). A sensitivity analysis was also conducted, comparing patients with a nonreceptive ERA defined as recommending ≥ 24 hours change in progesterone exposure, with changes ≤12 hours considered normal. There were no differences in the clinical outcomes between the control and interventions arms of the clinical trial, nor in sensitivity analyses (Table 1). The sensitivity and specificity of the ERA to predict failure of sustained embryo implantation were poor (55% and 49% respectively). The ERA had a PPV of 74 and 78% in both models, but extremely poor NPV of 29%. Positive and negative likelihood ratios were also low in both analyses. Receiver operating characteristic (ROC) results demonstrated an area under curve (AUC) of 0.52.Tabled 1Receptive vs Nonreceptive1 (Primary Analysis)Receptive or Nonreceptive2 (Sensitivity Analysis)Clinical OutcomeReceptive (n=208)Nonreceptive (n=178)P valueReceptive(n=120)Nonreceptive(n=266)P valuePositive hCG (%)167 (80.3)140 (78.7)0.78799 (82.5)208 (78.2)0.404Sustained implantation (%)155 (74.5)126 (70.8)0.48094 (78.3)187 (70.3)0.129Clinical pregnancy loss (%)25 (12)16 (9)0.42518 (15)23 (8.6)0.090Live birth130 (62.5)109 (61.2)0.88176 (63.3)163 (61.3)0.786Statistical OutcomeReceptive vs Nonreceptive1 (Primary Analysis)Receptive or Nonreceptive2 (Sensitivity Analysis)Sensitivity55.2%33.5%Specificity49.5%75.2%Positive predictive value74.5%78.3%Negative predictive value29.2%29.7%Positive likelihood ratio1.091.35Negative likelihood ratio0.910.88ROC AUC0.520.541Nonreceptive: any result recommending ≥12 hours progesterone change2Nonreceptive: any result recommending ≥24 hours progesterone change Open table in a new tab 1Nonreceptive: any result recommending ≥12 hours progesterone change 2Nonreceptive: any result recommending ≥24 hours progesterone change ERA thresholds of both 12- and 24- hour displacement demonstrated poor diagnostic performance for predicting failure to achieve sustained implantation, with a negative predictive value of only 29%. Furthermore, likelihood ratios demonstrated the test to be highly unlikely to change clinical management or outcomes.
The endometrial receptivity assay (ERA) is a diagnostic tool designed to identify the optimal window of implantation on the basis of the transcriptomic signature of 238 genes expressed in the receptive phase of the endometrium ( 1 Igenomix. https://www.igenomix.com/our-services.era/#why-eraDate accessed: July 17, 2023 Google Scholar ). The assay purports to identify whether the endometrium is receptive (optimal time for embryo transfer and implantation), prereceptive (additional progesterone exposure time recommended), or postreceptive (less progesterone exposure time recommended). After its introduction to the market, ERA utilization was rapid and widespread in reproductive medicine ( 2 Díaz-Gimeno P. Horcajadas J.A. Martínez-Conejero J.A. Esteban F.J. Alamá P. Pellicer A. et al. A genomic diagnostic tool for human endometrial receptivity based on the transcriptomic signature. Fertil Steril. 2011; 95: 50-60 Abstract Full Text Full Text PDF PubMed Scopus (447) Google Scholar ). However, the clinical effectiveness of the ERA has been questioned recently by several studies. The Synchrony trial by Doyle et al. ( 3 Doyle N. Jahandideh S. Hill M. Widra E. Levy M. Devine K. Effect of timing by endometrial receptivity testing vs standard timing of frozen embryo transfer on live birth in patients undergoing in vitro fertilization. JAMA. 2022; 328: 2117-2125 Crossref PubMed Scopus (18) Google Scholar ) reported no difference in implantation and live birth in women undergoing vitrified euploid embryo transfer randomized to transfer timing on the basis of ERA testing results vs. standard progesterone exposure. The objective of this study was to evaluate the performance of the ERA as a clinical diagnostic test on the basis of an a priori analysis of the data from the Synchrony trial. Letter to “Evaluating the Endometrial Receptivity Assay: a nested diagnostic accuracy study within the Synchrony randomized clinical trial.”Fertility and SterilityPreviewWe have reviewed the research letter by Chae-Kim et al. (1) who assessed endometrial receptivity analysis (ERA) on the basis of the partial reanalysis of their Synchrony randomized clinical trial (RCT) data (2), and we commend them for evaluating this important topic. However, there are several concerns that we would like to raise. Full-Text PDF (In)Accuracy of the endometrial receptivity assay in the general fertility populationFertility and SterilityVol. 120Issue 6PreviewThe study by Chae-Kim et al. (1) in this month’s Fertility and Sterility seeks to characterize the diagnostic accuracy of the endometrial receptivity assay (ERA) in predicting implantation failure after frozen embryo transfer (FET). Their group used data from the control arm of the 2022 Synchrony trial, in which 386 participants underwent ERA testing followed by euploid FET (2). Individuals in the control arm of the Synchrony trial all underwent a standardized FET protocol after 123 (±3) hours of progesterone exposure because patients and providers were double blinded to the results of ERA testing. Full-Text PDF
Abstract Study question What is the expected improvement in pregnancy rates using an artificial intelligence (AI) model for embryo ranking compared to manual grading systems? Summary answer A large-scale retrospective bootstrapped analysis shows that use of an AI model for embryo ranking can improve pregnancy rates compared to manual grading. What is known already Embryo evaluation is one of the most important steps of an in vitro fertilization (IVF) procedure. Recently, artificial intelligence (AI) models have been developed to automate embryo analysis and reduce the subjectivity of manual grading. While models are often evaluated in terms of classification accuracy or area under the curve (AUC), a more relevant metric is improvement in pregnancy rates. Here we evaluate a previously developed model using a large-scale bootstrapped analysis of virtual patient pregnancy rates and compare its performance to manual grading. Study design, size, duration Historical, de-identified images of transferred blastocyst-stage embryos and manual morphology grades were collected from 11 IVF clinics in the United States for cycles started between 2015-2020. Images were captured on day 5, 6, or 7 using the inverted microscope prior to biopsy or freeze. A total of 1,776 test set images from 3-fold cross validation were used for this analysis. Participants/materials, setting, methods Embryos were matched by age, PGT status, and race to create 16 distinct categories. Virtual patient panels were created within each category using a random selection of 3-5 embryos. Embryos were re-used across different panels, but each individual panel was unique. Three different manual ranking systems were created incorporating the morphology grade and day of image capture. The AI and one randomly chosen manual ranking system independently selected a top embryo for each panel. Main results and the role of chance On average, 105,263 unique virtual patient panels were constructed from the 1,776 embryos. Within these panels, the AI model and manual ranking system selected different top embryos from each other in 27,860 cases, or 26% of the time. The average pregnancy rate of the top-ranked embryo using manual grading was 53.1%, and the average pregnancy rate of the top-ranked embryo using the AI model was 59.4%. The average pregnancy rate improvement from using the AI model was 6.3%, with a standard deviation of 0.2% measured across 10 repetitions of the simulation with different random seeds. Limitations, reasons for caution The primary limitation is the retrospective nature of this study. Also, this bootstrapped panel study relied on recorded manual morphology grades at the time of embryo transfer or freeze rather than on the actual selection of the top embryo in each panel by an embryologist. Wider implications of the findings Our results demonstrate the potential of using an AI model for embryo ranking in terms of improved pregnancy rates. Results from this large-scale bootstrapped retrospective analysis will help inform the design of future clinical validation studies. Trial registration number not applicable
Objective: To perform a series of analyses characterizing an artificial intelligence (AI) model for ranking blastocyst-stage embryos. The primary objective was to evaluate the benefit of the model for predicting clinical pregnancy, whereas the secondary objective was to identify limitations that may impact clinical use. Design: Retrospective study. Setting: Consortium of 11 assisted reproductive technology centers in the United States. Patient(s): Static images of 5,923 transferred blastocysts and 2,614 nontransferred aneuploid blastocysts. Intervention(s): None. Main Outcome Measure(s): Prediction of clinical pregnancy (fetal heartbeat). Result(s): The area under the curve of the AI model ranged from 0.6 to 0.7 and outperformed manual morphology grading overall and on a per-site basis. A bootstrapped study predicted improved pregnancy rates between +5% and +12% per site using AI compared with manual grading using an inverted microscope. One site that used a low-magnification stereo zoom microscope did not show predicted improvement with the AI. Visualization techniques and attribution algorithms revealed that the features learned by the AI model largely overlap with the features of manual grading systems. Two sources of bias relating to the type of microscope and presence of embryo holding micropipettes were identified and mitigated. The analysis of AI scores in relation to pregnancy rates showed that score differences of >= 0.1 (10%) correspond with improved pregnancy rates, whereas score differences of <0.1 may not be clinically meaningful. Conclusion(s): This study demonstrates the potential of AI for ranking blastocyst stage embryos and highlights potential limitations related to image quality, bias, and granularity of scores. (C) 2021 by American Society for Reproductive Medicine.
Patient success during in vitro fertilization (IVF) cycles varies based on several factors. The most commonly used measure of success is the cumulative live birth rate (CLBR), which incorporates outcomes from fresh and frozen-thawed embryo transfers. A number of counseling tools are available that can estimate the CLBR for a patient who is considering IVF. However, a limitation of such tools is that they typically do not account for unused frozen embryos. There is a need for new methodologies to predict CLBR that account for all embryos, especially with the increased prevalence of embryo banking.
Importance:Endometrial receptivity testing is purported to improve live birth following frozen embryo transfer by identifying the optimal embryo transfer time for an individual patient; however, data are conflicting. Objective:To compare live birth from single euploid frozen embryo transfer according to endometrial receptivity testing vs standardized timing. Design, Setting, and Participants:Double-blind, randomized clinical trial at 30 sites within a multicenter private fertility practice in the Eastern US. Enrollment was from May 2018 to September 2020; follow-up concluded in August 2021. Participants underwent in vitro fertilization, preimplantation genetic testing for aneuploidy, endometrial receptivity testing, and frozen embryo transfer. Those with euploid blastocyst(s) and an informative receptivity result were randomized. Exclusion criteria included recurrent pregnancy loss, recurrent implantation failure, surgically aspirated sperm, donor egg(s), and unmitigated anatomic uterine cavity defects. Interventions:The intervention group (n = 381) underwent receptivity-timed frozen embryo transfer, with adjusted duration of progesterone exposure prior to transfer, if indicated by receptivity testing. The control group (n = 386) underwent transfer at standard timing, regardless of receptivity test results. Main Outcomes and Measures:The primary outcome was live birth. There were 3 secondary outcomes, including biochemical pregnancy and clinical pregnancy. Results:Among 767 participants who were randomized (mean age, 35 years), 755 (98%) completed the trial. All randomized participants were analyzed. The primary outcome of live birth occurred in 58.5% of transfers (223 of 381) in the intervention group vs 61.9% of transfers (239 of 386) in the control group (difference, -3.4% [95% CI, -10.3% to 3.5%]; rate ratio [RR], 0.95 [95% CI, 0.79 to 1.13]; P = .38). There were no significant differences in the intervention vs the control group for the prespecified secondary outcomes, including biochemical pregnancy rate (77.2% vs 79.5%, respectively; difference, -2.3% [95% CI, -8.2% to 3.5%]; RR, 0.97 [95% CI, 0.83 to 1.14]; P = .48) and clinical pregnancy rate (68.8% vs 72.8%, respectively; difference, -4.0% [95% CI, -10.4% to 2.4%]; RR, 0.94 [95% CI, 0.80 to 1.12]; P = .25). There were no reported adverse events. Conclusions and Relevance:Among patients for whom in vitro fertilization yielded a euploid blastocyst, the use of receptivity testing to guide the timing of frozen embryo transfer, compared with standard timing for transfer, did not significantly improve the rate of live birth. The findings do not support routine use of receptivity testing to guide the timing of embryo transfer during in vitro fertilization. Trial Registration:ClinicalTrials.gov Identifier: NCT03558399.
Obesity rates in the United States have increased over the past several decades. Obesity has many known adverse consequences on fertility and obstetrical outcomes. This study aimed to determine whether body mass index (BMI) adversely affects pregnancy outcomes following donor egg IVF embryo transfer, by analyzing donor oocyte recipient cycles using paired sibling oocytes.
Abstract Study question What is the sensitivity of an embryo-grading artificial intelligence (AI) model to different focal planes and how do we obtain consistent scores across focal planes? Summary answer Test-time augmentation and ensemble modeling reduce sensitivity of the AI model to different focal planes while maintaining performance. What is known already When prioritizing embryos for transfer, embryologists assess the 3D morphological features under a microscope, by zooming up and down, and assign a score that reflects the embryo quality. In comparison, some AI-based embryo grading models typically take one 2D focal plane of an embryo and output a score based on that focal plane. AI models such as convolutional neural networks (CNNs) are known to be sensitive to perturbations in its input. In order to reduce sensitivity and generalization error and thus improve predictive performance, techniques such as ensemble learning and test-time augmentation can be used. Study design, size, duration Historical, de-identified images of blastocyst-stage embryos were collected from 11 IVF clinics in the United States for cycles between 2015-2020. 5,100 blastocysts were matched to pregnancy outcomes as determined by fetal heartbeat. 2,900 blastocysts were matched to aneuploid PGT-A results and added to the negative training group to reduce selection bias. Data was split to 70% for training and 30% for testing. A set of 10 embryos were used for focal plane sensitivity. Participants/materials, setting, methods A single model (ResNet18), a three-model (ResNet18), and a six-model (ResNet18 and EfficientNet-b1) ensemble with and without test-time augmentation were trained to rank embryos according to their likelihood of reaching clinical pregnancy. Test-time augmentation involved taking the average scores from 4 flipped and rotated copies of the original input image. Manual grades were mapped to numeric scores for comparison. The AUC was used to evaluate the ability of the models to rank embryos. Main results and the role of chance Focal plane sensitivity was calculated as the range, or difference between the maximum and minimum score, for an embryo at different focal planes. Between 12 and 100 focal plane images were available for each of the 10 embryos. On average, the focal plane range was 0.26 for the single model, 0.22 for the single model with test-time augmentation, 0.14 for a 3-model ensemble with test-time augmentation, and 0.11 for a 6-model ensemble with test-time augmentation. Test-time augmentation on the single model reduced the range by 17%; whereas ensemble modeling with test-time augmentation reduced the range by 46% for the 3-model ensemble and 60% for the 6-model ensemble. Reduction in range did not compromise performance. The AUC for the test set for all embryos was 0.73 for the single model, 0.74 for the single model with test-time augmentation, 0.75 for the three-model ensemble with test-time augmentation and 0.74 for the six-model ensemble with test-time augmentation. All models outperformed manual grading, which was estimated to have an AUC of 0.67 for all embryos. Limitations, reasons for caution Our analysis on focal plane sensitivity was limited to a small sample size of 10 embryos, so more samples will be needed to confirm our findings. Wider implications of the findings Test-time augmentation and ensemble techniques can be used to reduce sensitivity while maintaining model performance. By reducing sensitivity to different focal planes, an AI model can produce one reliable score for a single embryo as is done currently in practice with manual grading. Trial registration number not applicable
The mean maternal and paternal age have increased over the past several decades (1, 2). While studies have examined the impact of advanced maternal age, there is limited, conflicting data regarding advanced paternal age. This disparity makes it difficult to counsel couples undergoing assisted reproductive technology (ART). A recent small study using paired sibling oocyte recipients concluded that paternal age ≥45 years had a negative effect on clinical pregnancy rates. We sought to determine, in a large cohort, whether paternal age >50 impacts pregnancy outcomes by analyzing donor oocyte recipient cycles using paired sibling oocytes. We performed a retrospective cohort study from 2016-2019 at a large fertility center. Cycles where the same donor's oocytes were split between at least two couples with sperm from disparate paternal ages, were paired and categorized based on paternal age < 45 and ≥ 45. Sensitivity analyses were also performed using 50 years of age as a threshold, paternal age as a continuous variable, and greater than efficiency curves. All patients with uterine factor, male factor, surgically obtained sperm, or donor sperm were excluded. Primary outcome was live birth per embryo transfer. Secondary outcomes included clinical pregnancy and miscarriage. GEE analysis was performed to account for cycle repeats and maternal age. 1255 donor oocytes cycles were assessed. Paternal age as a continuous variable was not associated with live birth or secondary outcomes. The database was initially analyzed by paired sibling oocytes from men < 45 years-old and ≥ 45 years-old accounting for 547 split donor oocyte recipients. There were no significant differences in live birth (54.6% vs 54.4%, p=1.00), clinical pregnancy, implantation, or miscarriage (Table 2). We then performed the paired sibling oocyte analysis by paternal age < 50 and ≥ 50. A total of 236 split donor oocyte recipients were included. Group A (< 50 years-old) consisted of 149 recipients, and group B (≥ 50 years-old) included 87 recipients. Both groups were similar in terms of BMI, type of sperm used, gravity, and ethnicity; however, recipient age was significantly lower in group A, 41 years-old, compared to group B, 46 years-old (Table 1). There were no differences in live birth (47.3% vs 51.1%, p= 0.66), clinical pregnancy (66.2% vs 64.8%, p= 0.93), or miscarriage (18.9% vs 19.6%, p= 0.39) (Table 2). After further adjusting for confounding variables, the results remained unchanged. In this large study uniquely controlling for oocyte quality utilizing donor oocyte sibling pairs, increased paternal age, even past the age of 50, was not associated with pregnancy outcomes. These data may reassure couples undergoing ART in the setting of advanced paternal age.
To develop an interpretable machine learning model for individualized gonadotropin dose selection during controlled ovarian stimulation. Historical, de-identified electronic medical record (EMR) data was collected from 4 IVF clinics in the United States. Records were filtered for autologous, non-canceled IVF retrievals, resulting in 7,977 cycles started between 2014 and 2020. A multiple linear regression model was developed with cross validation and recursive feature elimination to predict the number of eggs and mature (MII) eggs retrieved using baseline parameters available prior to start of treatment. The predictor variables were then used to create a patient similarity model based on K nearest neighbors (KNN), an interpretable machine learning technique. After identifying the best performing distance metrics, neighbor weights, and number of neighbors, the model was used to predict the number of eggs and MII eggs retrieved by calculating the weighted average from the set of K neighbors most similar to the patient of interest. The performance of the KNN model was compared to linear regression in terms of R-squared (R2) and mean absolute error (MAE). The KNN model was then used to (a) query the K most similar patients, and (b) identify the optimal gonadotropin dose in terms of highest number of MII eggs retrieved. We developed linear regression and KNN models using patient age, BMI, diagnosis, AMH, AFC, number of previous IVF cycles, and parity. KNN achieved highest performance using the Manhattan distance, 50-80 similar patients, and distance-based neighbor weighting. The KNN model outperformed linear regression for eggs retrieved (R2: 0.43 vs. 0.39, MAE: 4.84 vs. 4.98) and for MII eggs retrieved (R2: 0.39 vs. 0.35, MAE: 4.01 vs. 4.11). We then investigated the application of these models for gonadotropin dose selection. Linear models indicated that gonadotropin dose is negatively correlated with MII eggs, which may in part reflect that poor-prognosis patients are prescribed higher doses. In contrast, the KNN model showed that 22% of patients had a concave dose response curve, in which there was an optimal dose that maximized the number of MII eggs. We developed a patient similarity model using K nearest neighbors. The model showed better accuracy than linear regression for predicting eggs and MII eggs retrieved, and allowed the evaluation of which starting dose maximized the number of MII eggs retrieved, which is not possible with a linear model. Future work will optimize techniques for matching similar patients and extend the modeling for protocol selection.
ObjectiveUse of frozen sperm in non-male factor infertility is often needed in donor cycles. Most studies to date examining outcomes of fresh vs frozen sperm are unable to control for oocyte quality. Studies examining sibling oocytes represent a unique model to control for oocyte quality. A recent small study using this model found worse outcomes in the frozen group. We sought to evaluate, in a large cohort, if fresh and frozen ejaculated sperm are associated with similar pregnancy outcomes by analyzing paired donor egg recipient (DER) cycles.Materials and MethodsRetrospective cohort study from 2016-2019 at a large fertility center. Patients who underwent DER cycles where oocytes were split between two couples and one couple used fresh sperm and the other used frozen sperm were included. All patients with uterine factor, male factor or surgically obtained sperm were excluded. Primary outcome was Ongoing pregnancy/Live birth rate (OPR). Secondary outcome included clinical pregnancy rate (CPR) and miscarriage rate. GEE analysis was performed to control for confounding factors and donors providing oocytes to both study cohorts.Results1255 donor oocytes cycles were screened. A total of 205 unique oocytes donors were identified with oocytes inseminated with discrepant sperm in different recipient cycles. There were 698 recipient transfer cycles, 405 fresh and 293 frozen. Cohorts were similar in baseline characteristics (table 1). There were no differences in OPR/LBR with fresh vs frozen sperm (53.6% vs 55.6%, p=0.7) or clinical pregnancies (66.4% vs 63.5%, p=0.4). Spontaneous miscarriage (<20 weeks) was significantly higher in the fresh cohort (12.3% vs 6.1%, p=0.01).ConclusionsIn this large study uniquely controlling for oocyte quality, there are no differences in live birth rate when fresh or frozen sperm was utilized on the same donor oocytes. This type of comparison is important as it helps control as much as possible the oocyte, thus isolating the discrepant sperm state as a determinant of outcome. There was a significant increase in miscarriage rate when fresh sperm was used.Impact StatementTabled 1FreshFrozenp valuen405293Recipient age (years) (mean (SD))41.93 (4.92)42.47 (4.16)0.13Recipient BMI (kg/m2) (mean (SD))26.70 (5.06)25.99 (5.15)0.07Gravidity (mean (SD))2.81 (1.75)2.38 (1.59)0.001Cycle outcome (%)CPR269 (66.4)186 ( 63.5)0.40OPR/LBR217 (53.6)163 (55.6)0.70SAB50 (12.3)18 (6.1)0.01 Open table in a new tab ObjectiveUse of frozen sperm in non-male factor infertility is often needed in donor cycles. Most studies to date examining outcomes of fresh vs frozen sperm are unable to control for oocyte quality. Studies examining sibling oocytes represent a unique model to control for oocyte quality. A recent small study using this model found worse outcomes in the frozen group. We sought to evaluate, in a large cohort, if fresh and frozen ejaculated sperm are associated with similar pregnancy outcomes by analyzing paired donor egg recipient (DER) cycles. Use of frozen sperm in non-male factor infertility is often needed in donor cycles. Most studies to date examining outcomes of fresh vs frozen sperm are unable to control for oocyte quality. Studies examining sibling oocytes represent a unique model to control for oocyte quality. A recent small study using this model found worse outcomes in the frozen group. We sought to evaluate, in a large cohort, if fresh and frozen ejaculated sperm are associated with similar pregnancy outcomes by analyzing paired donor egg recipient (DER) cycles. Materials and MethodsRetrospective cohort study from 2016-2019 at a large fertility center. Patients who underwent DER cycles where oocytes were split between two couples and one couple used fresh sperm and the other used frozen sperm were included. All patients with uterine factor, male factor or surgically obtained sperm were excluded. Primary outcome was Ongoing pregnancy/Live birth rate (OPR). Secondary outcome included clinical pregnancy rate (CPR) and miscarriage rate. GEE analysis was performed to control for confounding factors and donors providing oocytes to both study cohorts. Retrospective cohort study from 2016-2019 at a large fertility center. Patients who underwent DER cycles where oocytes were split between two couples and one couple used fresh sperm and the other used frozen sperm were included. All patients with uterine factor, male factor or surgically obtained sperm were excluded. Primary outcome was Ongoing pregnancy/Live birth rate (OPR). Secondary outcome included clinical pregnancy rate (CPR) and miscarriage rate. GEE analysis was performed to control for confounding factors and donors providing oocytes to both study cohorts. Results1255 donor oocytes cycles were screened. A total of 205 unique oocytes donors were identified with oocytes inseminated with discrepant sperm in different recipient cycles. There were 698 recipient transfer cycles, 405 fresh and 293 frozen. Cohorts were similar in baseline characteristics (table 1). There were no differences in OPR/LBR with fresh vs frozen sperm (53.6% vs 55.6%, p=0.7) or clinical pregnancies (66.4% vs 63.5%, p=0.4). Spontaneous miscarriage (<20 weeks) was significantly higher in the fresh cohort (12.3% vs 6.1%, p=0.01). 1255 donor oocytes cycles were screened. A total of 205 unique oocytes donors were identified with oocytes inseminated with discrepant sperm in different recipient cycles. There were 698 recipient transfer cycles, 405 fresh and 293 frozen. Cohorts were similar in baseline characteristics (table 1). There were no differences in OPR/LBR with fresh vs frozen sperm (53.6% vs 55.6%, p=0.7) or clinical pregnancies (66.4% vs 63.5%, p=0.4). Spontaneous miscarriage (<20 weeks) was significantly higher in the fresh cohort (12.3% vs 6.1%, p=0.01). ConclusionsIn this large study uniquely controlling for oocyte quality, there are no differences in live birth rate when fresh or frozen sperm was utilized on the same donor oocytes. This type of comparison is important as it helps control as much as possible the oocyte, thus isolating the discrepant sperm state as a determinant of outcome. There was a significant increase in miscarriage rate when fresh sperm was used. In this large study uniquely controlling for oocyte quality, there are no differences in live birth rate when fresh or frozen sperm was utilized on the same donor oocytes. This type of comparison is important as it helps control as much as possible the oocyte, thus isolating the discrepant sperm state as a determinant of outcome. There was a significant increase in miscarriage rate when fresh sperm was used.
To identify and reduce potential sources of bias when training deep learning models for analyzing images of human embryos. Historical, de-identified images of blastocyst-stage embryos were collected from 11 IVF clinics in the United States between 2015-2020. Each laboratory captured a single image using their existing inverted microscope, stereo zoom microscope, or time-lapse microscope. Approximately 8,000 images were matched to positive clinical pregnancy, negative clinical pregnancy, or PGT-A aneuploid result. We trained a series of deep convolutional neural networks (CNNs) to rank embryo images according to their likelihood of having a positive or negative outcome. Experiments were performed using different techniques for combining images from the clinical sites, including naive and balanced methods. For each experiment, the aggregated data was split into 70% train and 30% test. The area under the receiver operating curve (AUC) was used for evaluating the ability of the models to rank embryos according to their likelihood of achieving a positive outcome. Total and per-clinic AUCs, as well as total and per-clinic inference probabilities, were evaluated for each experiment to identify and reduce potential sources of bias. Using a naive approach for combining data together from all clinics achieved the highest total AUC for the test set (0.75) but also the lowest per-clinic AUC (0.51). Investigation of this discrepancy revealed two strong sources of bias, which artificially inflated the total AUC and significantly limited per-clinic performance. The biases included the unique optical signature of each type of microscope, and the presence of foreign objects, such as a holding micropipette present in the image. If a certain optical signature or foreign artifact appeared more in the positive training class compared to the negative training class, the CNN models were found to learn these biases and give a higher score to those images regardless of the embryo morphology. With these insights, a new dataset was prepared that balanced the ratios of positive-to-negative samples for each type of microscope and for each group containing foreign objects. This provided a non-inflated total AUC of 0.72 and significantly raised the lowest per-site AUC from 0.51 to 0.61. There has been significant recent interest in using deep learning for analyzing images of embryos at the blastocyst stage. The black-box nature of deep learning models such as CNNs makes it difficult to recognize when potential sources of bias have been introduced during the training process. We performed a series of experiments that identified and reduced two sources of bias, and improved per-clinic performance of the CNN. Future work will continue to search for other sources of bias and address them accordingly.
Research question: Ovarian stimulation during IVF cycles involves close monitoring of oestradiol, progesterone and ultrasound measurements of follicle growth. In contrast to blood draws, sampling saliva is less invasive. Here, a blind validation is presented of a novel saliva-based oestradiol and progesterone assay carried out in samples collected in independent IVF clinics. Design: Concurrent serum and saliva samples were collected from 324 patients at six large independent IVF laboratories. Saliva samples were frozen and run blinded. A further 18 patients had samples collected more frequently around the time of HCG trigger. Saliva samples were analysed using an immunoassay developed with Salimetrics LLC. Results: In total, 652 pairs of saliva and serum oestradiol were evaluated, with correlation coefficients ranging from 0.68 to 0.91. In the European clinics, a further 237 of saliva and serum progesterone samples were evaluated; however, the correlations were generally poorer, ranging from -0.02 to 0.22. In the patients collected more frequently, five out of 18 patients (27.8%) showed an immediate decrease in oestradiol after trigger. When progesterone samples were assessed after trigger, eight out of 18 (44.4%) showed a continued rise. Conclusions: Salivary oestradiol hormone testing correlates well to serum-based assessment, whereas progesterone values, around the time of trigger, are not consistent from patient to patient. ABSTRACT Research question: Ovarian stimulation during IVF cycles involves close monitoring of oestradiol, progesterone and ultrasound measurements of follicle growth. In contrast to blood draws, sampling saliva is less invasive. Here, a blind validation is presented of a novel saliva-based oestradiol and progesterone assay carried out in samples collected in independent IVF clinics. Design: Concurrent serum and saliva samples were collected from 324 patients at six large independent IVF laboratories. Saliva samples were frozen and run blinded. A further 18 patients had samples collected more frequently around the time of HCG trigger. Saliva samples were analysed using an immunoassay developed with Salimetrics LLC. Results: In total, 652 pairs of saliva and serum oestradiol were evaluated, with correlation coefficients ranging from 0.68 to 0.91. In the European clinics, a further 237 of saliva and serum progesterone samples were evaluated; however, the correlations were generally poorer, ranging from -0.02 to 0.22. In the patients collected more frequently, five out of 18 patients (27.8%) showed an immediate decrease in oestradiol after trigger. When progesterone samples were assessed after trigger, eight out of 18 (44.4%) showed a continued rise. Conclusions: Salivary oestradiol hormone testing correlates well to serum-based assessment, whereas progesterone values, around the time of trigger, are not consistent from patient to patient.
To develop a generalizable model for ranking embryos at the blastocyst stage that is applicable to any type of microscope, day of image capture, and cycle type. Historical, de-identified images of blastocyst-stage embryos were collected from 11 IVF clinics in the United States for cycles started between 2015-2020. Each laboratory captured a single image using their existing inverted microscope, stereo zoom microscope, or time-lapse incubation system. Images were captured on day 5, 6 or 7 prior to transfer, biopsy, or freeze. 5,100 blastocysts from fresh transfers, frozen transfers, and frozen-euploid transfers were matched to clinical pregnancy outcomes as determined by fetal heartbeat. An additional 2,900 blastocysts were matched to aneuploid (abnormal) PGT-A results. Aneuploid embryos were added to the negative training group to reduce selection bias. Data was split to 70% for training and 30% for testing. We trained a deep convolutional neural network (CNN) to rank embryo images according to their likelihood of reaching clinical pregnancy. A shallow model architecture (Resnet-18) with dropout was used. Performance was optimized using data augmentation, a custom weighted sampling technique, and hyperparameter tuning. Scores were personalized to each patient by incorporating age and donor egg status. Manual morphology grades were mapped to a numeric score for comparison. The area under the receiver operating curve (AUC) was used for evaluating the ability of the models to rank embryos. Bootstrapped analysis was performed using combinations of 2-4 euploid embryos to compare pregnancy rates of the top-ranked embryo for the CNN compared to manual grading. The CNN model AUC on the test set was 0.72 for all embryos (including transferred embryos and non-transferred aneuploids), 0.65 for fresh and frozen non-PGT transfers, and 0.63 for euploid transfers. For euploid transfers, the CNN model AUC outperformed manual grading overall (+7%), by clinical site (ranging from +5% to +11% per site), and by day of image capture (+7% for day-5, +8% for days-6/7). Bootstrapped analysis predicted improved pregnancy rates on first transfer of between +3% to +8% per site using the CNN model compared to manual grading. We developed a deep learning-based model for ranking embryos at the blastocyst-stage. Previous studies using deep learning have been application-specific, using images from a single type of microscope (e.g. a time-lapse instrument), or captured at a specific day (e.g. day 5), or from a specific cycle type (e.g. non-PGT cycles). With access to a large and diverse dataset, we developed a generalizable model that is broadly applicable and outperforms manual grading when analyzed in aggregate and by laboratory and day of image capture. Future work will focus on expanding our training dataset and performing clinical studies to validate performance.
STUDY QUESTIONDo donor oocyte recipients benefit from preimplantation genetic testing for aneuploidy (PGT-A)?SUMMARY ANSWERPGT-A did not improve the likelihood of live birth for recipients of vitrified donor oocytes, but it did avoid embryo transfer in cycles with no euploid embryos.WHAT IS KNOWN ALREADYRelative to slow freeze, oocyte vitrification has led to increased live birth from cryopreserved oocytes and has led to widespread use of this technology in donor egg IVF programs. However, oocyte cryopreservation has the potential to disrupt the meiotic spindle leading to abnormal segregation of chromosome during meiosis II and ultimately increased aneuploidy in resultant embryos. Therefore, PGT-A might have benefits in vitrified donor egg cycles. In contrast, embryos derived from young donor oocytes are expected to be predominantly euploid, and trophectoderm biopsy may have a negative effect relative to transfer without biopsy.STUDY DESIGN, SIZE, DURATIONThis is a paired cohort study analyzing donor oocyte-recipient cycles with or without PGT-A performed from 2012 to 2018 at 47 US IVF centers.PARTICIPANTS/MATERIALS, SETTING, METHODSVitrified donor oocyte cycles were analyzed for live birth as the main outcome measure. Outcomes from donors whose oocytes were used by at least two separate recipient couples, one couple using PGT-A (study group) and one using embryos without PGT-A (control group), were compared. Generalized estimating equation models controlled for confounders and nested for individual donors contributing to both PGT-A and non-PGT-A cohorts, enabling a single donor to serve as her own control.MAIN RESULTS AND THE ROLE OF CHANCEIn total, 1291 initiated recipient cycles from 223 donors were analyzed, including 262 cycles with and 1029 without PGT-A. The median aneuploidy rate per recipient was 25%. Forty-three percent of PGT-A cycles had only euploid embryos, whereas only 12.7% of cycles had no euploid embryos. On average 1.09 embryos were transferred in the PGT-A group compared to 1.38 in the group without PGT-A (P < 0.01). Live birth occurred in 53.8% of cycles with PGT-A versus 55.8% without PGT-A (P = 0.44). Similar findings persisted in cumulative live birth from per recipient cycle.LIMITATIONS, REASONS FOR CAUTIONPooled clinical data from 47 IVF clinics introduced PGT-A heterogeneity as genetic testing were performed using different embryology laboratories, PGT-A companies and testing platforms.WIDER IMPLICATIONS OF THE FINDINGSPGT-A testing in donor oocyte-recipient cycles does not improve the chance for live birth nor decrease the risk for miscarriage in the first transfer cycle but does increase cost and time for the patient. Further studies are required to test if our findings can be applied to the young infertility patient population using autologous oocytes.STUDY FUNDING/COMPETING INTEREST(S)No external funding was used for this study. There are no conflicts of interest to declare.TRIAL REGISTRATION NUMBERN/A.