The Flare protocol is often used for poor responders, as it is thought that stimulating endogenous production of FSH with a low dose of Lupron can increase the responsiveness of the ovaries. However, studies have reported conflicting results, and it remains unclear whether Flare provides better outcomes for poor responders compared to Antagonist.
While numerous studies have evaluated the relationship between serum progesterone concentration (P) and pregnancy outcomes in programmed frozen embryo transfer (FET) cycles with vaginal (PV) progesterone replacement, P is higher with intramuscular progesterone (IMP) relative to PV, and it remains unknown whether P correlates with success following programmed FET with IMP. The current study aimed to evaluate the relationship between serum progesterone (P) and live birth rates (LBR) in women receiving IMP in programmed FET cycles. The study included all single, euploid FET cycles at a single academic center from 2014 to 2022. All patients underwent programmed endometrial preparation cycle with IMP replacement. IMP was started at 50mg once daily (QD) and was increased if P was <18 ng/ml at any time during monitoring. The dose was initially increased to 75mg QD and further increased to 100mg QD if P remained <18 ng/ml. If a patient required a dose increase in a prior cycle, then a higher dose was started in the subsequent cycle(s). Serum P was obtained the day prior to FET, two days after FET and at time of pregnancy test. The primary outcome was live birth following euploid FET. Mixed effects logistic regression model was performed to account for multiple cycles in the same patient for the primary outcome of live birth. A total of 4731 cycles from 3343 unique patients were analyzed. The mean serum P and interquartile ranges for the day prior to FET, two days after FET and at the time of pregnancy test were 26.57 ng/ml (19.7 - 32.1), 28.95 ng/ml (22.4 - 34.6), and 28.97 ng/ml (22.7 - 34.3), respectively. After adjusting for patient age, body mass index, endometrial thickness, embryo grade, embryo development, number of prior cycles and multiple cycles within the same patient, there was no statistically significant association between serum progesterone and live birth rate at each time interval with p values of 0.617, 0.844, and 0.712 for P the day prior to FET, two days after FET and at the time of pregnancy test, respectively. To our knowledge, this is the largest study evaluating the relationship between P and LBR in programmed FET cycles utilizing IMP replacement. While the study was limited by the varying doses of IMP utilized, it indicates that P with IMP was generally sufficient and was not associated with LBR. Therefore, routine monitoring prior to and following FET may not be indicated when IMP is used. However, given the retrospective nature of this study, additional prospective studies should be performed to further elucidate the relationship between P and pregnancy outcomes in FET cycles using IMP for luteal support.
Objective: To investigate the association between the number of oocytes retrieved and the numbers of fertilized oocytes and blastocysts and cumulative and primary transfer live birth rates (LBRs). Design: Retrospective study. Setting: Retrieval cycles and linked embryo transfers from the Society for Assisted Reproductive Technology Clinic Outcome Reporting System. Patient(s): Patients in the United States undergoing autologous in vitro fertilization cycles from 2014 to 2019 (n 1/4 402,411 cycles). Intervention(s): None. Main Outcome Measure(s): Normally fertilized oocytes, blastocysts, and cumulative and primary transfer LBRs. Result(s): There was a strong positive linear correlation between oocytes and fertilized oocytes and between oocytes and blastocysts. The cumulative LBR increased rapidly with the number of oocytes retrieved to approximately 16-20 oocytes, at which point it continued to increase but with diminishing returns. The increasing trend of the cumulative LBR was observed when stratifying patients by age and antimueurollerian hormone and after controlling for confounding variables using multivariate logistic regression. The primary transfer LBR also increased with the number of oocytes to approximately 16-20 oocytes, at which point it plateaued but did not decline. Conclusion(s): A higher number of oocytes retrieved improves the cumulative LBR without impairing the primary transfer LBR. This suggests that ovarian stimulation strategies should aim to safely maximize the number of oocytes retrieved. (Fertil Sterile 2023;119:762-9. (c) 2023 by American Society for Reproductive Medicine.) El resumen esta disponible en Espanol al final del articulo.
Objective: To use causal inference to investigate whether the flare or antagonist protocol is better for poor responders going through controlled ovarian stimulation. Design: A retrospective study.Setting: Retrieval cycles from the Society for Assisted Reproductive Technology Clinic Outcomes Reporting System.Patients: Patients in the United States underwent autologous in vitro fertilization cycles from 2014 to 2019 using either the flare or antagonist protocol.Intervention: Not applicable.Main Outcome Measure: Primary outcomes included oocytes retrieved, fertilized oocytes (2PNs), blastocysts, the cumulative live birth rate (CLBR), and cycle cancelation rate.Results: After propensity score matching, patients with a predicted poor response (antimullerian hormone, <0.5) on their first in vitro fertilization cycle had similar outcomes on the antagonist protocol (CLBR of 14.2%, 95% confidence intervals [CIs]: 13.6%, 14.8%) compared with flare (CLBR of 13.6%, 95% CIs: 12.4%, 14.8%). We evaluated patients undergoing a second cycle after having a poor response (<4 oocytes retrieved) on their first cycle. Patients in the antagonist-to-antagonist group had a similar change in outcomes between the first and second cycles (average CLBR improvement of 13.9%, 95% CIs: 12.1%, 15.6%) compared with the antagonist-to-flare group (average CLBR improvement of 14.4%, 95% CIs: 10.9%, 18.3%). In addition, patients in the flare-to-antagonist group had a similar change in outcomes between the first and second cycles (average CLBR improvement of 10.4%, 95% CIs: 6.6%, 14.5%) compared with the flare-to-flare group (average CLBR improvement of 9.0%, 95% CIs: 5.1%, 13.4%). Conclusion: Poor responders have similar outcomes on an antagonist protocol compared with a flare protocol for both the first and second cycles. (Fertil Steril (R) 2023;120:289-96. (C)2023 by American Society for Reproductive Medicine.)
The daily workload for an embryology laboratory varies depending on the number of egg retrievals performed each day and how many eggs were retrieved from each patient. An artificial intelligence (AI) model to predict upcoming embryology workload could assist with staffing and identify which days are likely to necessitate a higher resource allocation for procedures such as intracytoplasmic sperm injection (ICSI) and preimplantation genetic testing (PGT).
To evaluate the integration of two independently-developed artificial intelligence (AI) tools: (1) automated follicle measurements from MyCycleClarity, and (2) predictions of number of eggs retrieved using Alife Health's Stim AssistTM. MyCycleClarity uses AI to automatically count and measure follicles from 3D ultrasound images, and was previously trained on 91,782 follicles in 19,776 ovaries. The Alife Stim AssistTM Trigger Tool uses a linear regression model to predict the number of eggs retrieved based on a patient's individual follicle sizes and estradiol level on the day of trigger, and was previously trained on 26,179 cycles. In the current study, electronic medical record data from 553 patients from one US clinic was collected. Data from this clinic contained both manually counted human follicle measurements as well as AI follicle measurements from MyCycleClarity. On the day of trigger, there were 82 cycles with human follicle measurements and 186 cycles with AI follicle measurements, and 25 cycles with both human and AI follicle measurements. To assess the integration of these two machine learning tools, we evaluated the accuracy of Stim AssistTM using human follicle measurements compared to using AI follicle measurements. The linear regression model coefficients from the Stim AssistTM Trigger Tool showed that follicles size 7-25mm were significantly associated with the number of eggs retrieved (p < 0.05). Follicles size 14-17mm measured at trigger had the strongest association with egg outcomes. Across all patients at the test clinic on the day of trigger, MyCycleClarity counted more small follicles (<10 mm) compared to human measurements (9.5 ± 9.4 vs 0.8 ± 0.9); however, it counted a similar number of large follicles (>10 mm) (14.1 ± 10.5 vs 11.7 ± 7.7). On the set of 25 cycles with both AI and human follicle measurements, the Stim AssistTM Trigger Tool had a mean absolute prediction error of 3.30 eggs using the AI follicle measurements and 3.84 eggs using the human follicle measurements. In an analysis of patients with both AI-counted follicles from MyCycleClarity and human-counted follicles, the Alife's Stim AssistTM tool had slightly more accurate predictions using AI counted follicles compared to human measurements. This is likely because MyCycleClarity more thoroughly counted follicles, especially the small ones <10mm, on the day of trigger.
Embryo evaluation is a critical step of in vitro fertilization (IVF). Here, we sought to develop an AI model that can automate the Gardner scale morphology grading that is performed routinely in labs, including degree of expansion (3,4,5,6), inner cell mass (ICM) grade (A,B,C), and trophectoderm (TE) grade (A,B,C). Historical, de-identified images of blastocyst-stage embryos and manual morphology grades were collected from multiple IVF clinics in the US for cycles between 2015-2020. Images were captured on day 5, 6, or 7 using the inverted microscope prior to biopsy or freeze. The dataset contains 9,478 images. A separate test dataset of 50 images was collected from an independent IVF clinic, including manual morphology grades given by 6-10 embryologists each year for 4 years. Convolutional neural networks (CNNs) were trained independently for each morphological component. First, the images were sorted into 3 ICM grades (A,B, or C), and an ensemble of 2 CNNs (ResNet and EfficientNet) were trained to predict the ICM grade. This process was then repeated independently for TE and expansion. The final model for predicting the morphological grade consisted of 6 CNNs. After training and validation, the model was evaluated on an independent test dataset. The ICM, TE, and expansion deep learning models reached training and validation accuracies of approximately 80%.. Visual inspection of images with prediction errors revealed issues with image quality and inconsistent labeling between embryologists. The independent test dataset was used to evaluate consensus agreement between a group of embryologists and the AI model. For expansion, the embryologists agreed unanimously on the expansion grade 12% of the time, showed majority (>50%) consensus 100% of the time, and the AI model agreed with the embryologist-consensus 88% of the time. For ICM, the embryologists agreed unanimously on the ICM grade 4% of the time, showed majority (>50%) consensus 94% of the time, and the AI model agreed with the embryologist-consensus 60% of the time. For TE, the embryologists agreed unanimously on the TE grade 0% of the time, showed majority (>50%) consensus 98% of the time, and the AI model agreed with the embryologist-consensus 84% of the time. The most common AI prediction errors were A-to-B or B-to-A, but never A-to-C or C-to-A. After combining all three categories (expansion, ICM, and TE), the average rate at which individual embryologists agree with the common consensus is 43%, while the ratio for the AI model is 46%. While the subjectivity of ground-truth labels poses a challenge, automated morphology grading of blastocyst-stage embryos can be achieved with deep learning at human-level accuracy.
To evaluate the performance of a machine learning model for ranking blastocyst stage embryos for transfer using a double-blinded randomized comparative reader study. In previous work, a machine learning model was developed that predicts the likelihood of clinical pregnancy using embryo morphology grades assigned by embryologists using Gardner classification and day of development (5, 6, or 7). This model was trained on data from over 12,000 single-blastocyst transfer cycles from multiple U.S. IVF clinics performed between 2014 to 2021. To independently test the model, a retrospective, double-blinded, randomized comparative reader study was performed. The study included data from 438 single-blastocyst transfers from 10 different IVF clinics in the U.S. that were not part of previous model development or testing. Using this data, a large set of 1,257 simulated, or virtual, patient panels were created. Each virtual patient panel included between 2-5 embryos that were matched by age (18 - 29, 30 - 34, 35 - 37, ≥38), race (white, non-white, and unknown) and PGT-status (untested or euploid transfers). A group of 5 embryologists (readers) with varying levels of experience were then asked to select their top embryo for transfer for each virtual patient panel (control arm), and the machine learning model was also used to select a top embryo for transfer from each patient panel (treatment arm). The clinical pregnancy rates (CPR) of the top-selected embryos were calculated and compared using a comparison of proportions (clinical pregnancy rates), using a 2-sided type-1 error rate of 5% (α=0.05). The average CPR of the control arm (embryos selected by embryologists) was 61.0% (individual rates of 58.9%, 59.6%, 61.5%, 61.6%, and 63.3%), and the CPR of the treatment arm (embryos selected by machine learning model) was 62.1% (demonstrating non-inferiority with p<.001). In 35% of cases there was inter-embryologist variability in the top embryo selected for transfer, and when there were 3 or more embryos to choose from the variability increased to 44%. When all 5 readers agreed (65% of the time), the machine learning model also selected the same top embryo in nearly all cases (99% of the time), showing very high concordance with the group consensus. The machine learning model was non-inferior to manual embryo selection overall. When there was group consensus on the top embryo for selection, the machine learning model agreed in nearly all cases.
To develop a comprehensive model for predicting the probability of live birth prior to the start of progesterone during artificial frozen embryo transfer (FET) cycles that can identify at-risk cycles and set expectations.
Prior studies have estimated the efficiency of PGT-A using published implantation and aneuploidy rates [1]. Here we developed a more rigorous approach that incorporates a novel methodology for imputing the likelihood of failed untested transfers resulting from aneuploidy.
Dilated cardiomyopathy (DCM) is characterized by reduced cardiac output, as well as thinning and enlargement of left ventricular chambers. These characteristics eventually lead to heart failure. Current standards of care do not target the underlying molecular mechanisms associated with genetic forms of heart failure, driving a need to develop novel therapeutics for DCM. To identify candidate therapeutics, we developed an in vitro DCM model using induced pluripotent stem cell-derived cardiomyocytes (iPSC-CMs) deficient in B-cell lymphoma 2 (BCL2)-associated athanogene 3 (BAG3). With these BAG3-deficient iPSC-CMs, we identified cardioprotective drugs using a phenotypic screen and deep learning. From a library of 5500 bioactive compounds and siRNA validation, we found that inhibiting histone deacetylase 6 (HDAC6) was cardioprotective at the sarcomere level. We translated this finding to a BAG3 cardiomyocyte-knockout (BAG3(cKO)) mouse model of DCM, showing that inhibiting HDAC6 with two isoform-selective inhibitors (tubastatin A and a novel inhibitor TYA-018) protected heart function. In BAG3(cKO) and BAG3E455K mice, HDAC6 inhibitors improved left ventricular ejection fraction and reduced left ventricular diameter at diastole and systole. In BAG3(cKO) mice, TYA-018 protected against sarcomere damage and reduced Nppb expression. Based on integrated transcriptomics and proteomics and mitochondrial function analysis, TYA-018 also enhanced energetics in these mice by increasing expression of targets associated with fatty acid metabolism, protein metabolism, and oxidative phosphorylation. Our results demonstrate the power of combining iPSC-CMs with phenotypic screening and deep learning to accelerate drug discovery, and they support developing novel therapies that address underlying mechanisms associated with heart disease.
Abstract Study question What is the combined expected benefit of using machine learning algorithms to optimize the starting gonadotropin dose and day of trigger during ovarian stimulation? Summary answer Patients who had an optimal starting dose and optimal day of trigger had significantly improved outcomes compared to propensity matched patients who did not. What is known already Choosing the starting dose of follicle-stimulating hormones (FSH) and deciding when to inject the final trigger shot are two critical decisions made during an ovarian stimulation protocol. Although studies have investigated the effect of these decisions on patient outcomes, in practice, they remain subjective and can vary significantly across providers. Recently, machine learning techniques to support these decisions have been investigated, providing evidence that following model recommendations can improve outcomes. However, the combined effect of multiple clinical decision support tools on patient outcomes has not been studied. Study design, size, duration We performed a retrospective analysis of patients undergoing autologous, non-cancelled IVF cycles from 2014 - 2020 (n = 15,522) at three different IVF clinics in the United States. The primary outcomes were the average number of MIIs, 2PNs, and usable blastocysts in relation to total doses of FSH. Participants/materials, setting, methods To select the optimal starting FSH dose, a K-nearest neighbor model identified 100 similar patients and a dose response curve was created by plotting the number of MIIs retrieved relative to starting FSH across all neighbors. To select the optimal trigger day, linear regression models used daily follicles sizes and estradiol levels to predict MIIs retrieved today versus tomorrow, and a trigger day was identified by looking at day-by-day predicted MII trends. Main results and the role of chance Across all cycles, 27% were given the recommended optimal starting FSH dose. 51% of patients were triggered earlier and 13% were triggered later than the recommendation. Combining both algorithms, 11% of patients were given both the optimal starting dose and optimal trigger day, while the remaining 89% of patients had cycles that did not follow both recommendations. Patients following both model recommendations had on average 3.2 more MIIs, 2.3 more 2PNs, and 1.2 more usable blastocysts, using 730 IU’s less of total FSH, compared to propensity-matched patients with cycles that did not match both recommendations Limitations, reasons for caution The primary limitation is the retrospective nature of this study. Clinicians did not use either decision support tool when planning patients’ ovarian stimulation protocol. Further, we did not differentiate between different protocol or medication types in our analyses, which will be the focus of future work. Wider implications of the findings Our results suggest that following the combined recommendations of two clinical decision support tools can improve outcomes and reduce total FSH used in ovarian stimulation. Future work will include continuing to increase the diversity of our dataset and performing validation studies to show improved outcomes with model use. Trial registration number NA
To investigate whether particular changes in gonadotropin dose or stimulation protocols can improve ovarian stimulation outcomes for patients undergoing successive retrieval cycles.
Current success estimators for IVF do not take into account unused embryos and may under-predict cumulative live birth rate (CLBR) for some patients. We aimed to develop a new approach that considers predicted outcomes of all embryos.
Research question: Can we develop an interpretable machine learning model that optimizes starting gonadotrophin dose selection in terms of mature oocytes (metaphase II [MII]), fertilized oocytes (2 pronuclear [2PN]) and usable blastocysts?Design: This was a retrospective study of patients undergoing autologous IVF cycles from 2014 to 2020 (n = 18,591) in three assisted reproductive technology centres in the USA. For each patient cycle, an individual dose-response curve was generated from the 100 most similar patients identified using a K-nearest neighbours model. Patients were labelled as dose-responsive if their dose-response curve showed a region that maximized MII oocytes, and flat -responsive otherwise.Results: Analysis of the dose-response curves showed that 30% of cycles were dose-responsive and 64% were flat-responsive. After propensity score matching, patients in the dose-responsive group who received an optimal starting dose of FSH had on average 1.5 more MII oocytes, 1.2 more 2PN embryos and 0.6 more usable blastocysts using 10 IU less of starting FSH and 195 IU less of total FSH compared with patients given non-optimal doses. In the flat-responsive group, patients who received a low starting dose of FSH had on average 0.3 more MII oocytes, 0.3 more 2PN embryos and 0.2 more usable blastocysts using 149 IU less of starting FSH and 1375 IU less of total FSH compared with patients with a high starting dose.Conclusions: This study demonstrates retrospectively that using a machine learning model for selecting starting FSH can achieve optimal laboratory outcomes while reducing the amount of starting and total FSH used.
Abstract Study question What is the expected improvement in pregnancy rates using an artificial intelligence (AI) model for embryo ranking compared to manual grading systems? Summary answer A large-scale retrospective bootstrapped analysis shows that use of an AI model for embryo ranking can improve pregnancy rates compared to manual grading. What is known already Embryo evaluation is one of the most important steps of an in vitro fertilization (IVF) procedure. Recently, artificial intelligence (AI) models have been developed to automate embryo analysis and reduce the subjectivity of manual grading. While models are often evaluated in terms of classification accuracy or area under the curve (AUC), a more relevant metric is improvement in pregnancy rates. Here we evaluate a previously developed model using a large-scale bootstrapped analysis of virtual patient pregnancy rates and compare its performance to manual grading. Study design, size, duration Historical, de-identified images of transferred blastocyst-stage embryos and manual morphology grades were collected from 11 IVF clinics in the United States for cycles started between 2015-2020. Images were captured on day 5, 6, or 7 using the inverted microscope prior to biopsy or freeze. A total of 1,776 test set images from 3-fold cross validation were used for this analysis. Participants/materials, setting, methods Embryos were matched by age, PGT status, and race to create 16 distinct categories. Virtual patient panels were created within each category using a random selection of 3-5 embryos. Embryos were re-used across different panels, but each individual panel was unique. Three different manual ranking systems were created incorporating the morphology grade and day of image capture. The AI and one randomly chosen manual ranking system independently selected a top embryo for each panel. Main results and the role of chance On average, 105,263 unique virtual patient panels were constructed from the 1,776 embryos. Within these panels, the AI model and manual ranking system selected different top embryos from each other in 27,860 cases, or 26% of the time. The average pregnancy rate of the top-ranked embryo using manual grading was 53.1%, and the average pregnancy rate of the top-ranked embryo using the AI model was 59.4%. The average pregnancy rate improvement from using the AI model was 6.3%, with a standard deviation of 0.2% measured across 10 repetitions of the simulation with different random seeds. Limitations, reasons for caution The primary limitation is the retrospective nature of this study. Also, this bootstrapped panel study relied on recorded manual morphology grades at the time of embryo transfer or freeze rather than on the actual selection of the top embryo in each panel by an embryologist. Wider implications of the findings Our results demonstrate the potential of using an AI model for embryo ranking in terms of improved pregnancy rates. Results from this large-scale bootstrapped retrospective analysis will help inform the design of future clinical validation studies. Trial registration number not applicable
Patient success during in vitro fertilization (IVF) cycles varies based on several factors. The most commonly used measure of success is the cumulative live birth rate (CLBR), which incorporates outcomes from fresh and frozen-thawed embryo transfers. A number of counseling tools are available that can estimate the CLBR for a patient who is considering IVF. However, a limitation of such tools is that they typically do not account for unused frozen embryos. There is a need for new methodologies to predict CLBR that account for all embryos, especially with the increased prevalence of embryo banking.
Objective: To perform a series of analyses characterizing an artificial intelligence (AI) model for ranking blastocyst-stage embryos. The primary objective was to evaluate the benefit of the model for predicting clinical pregnancy, whereas the secondary objective was to identify limitations that may impact clinical use. Design: Retrospective study. Setting: Consortium of 11 assisted reproductive technology centers in the United States. Patient(s): Static images of 5,923 transferred blastocysts and 2,614 nontransferred aneuploid blastocysts. Intervention(s): None. Main Outcome Measure(s): Prediction of clinical pregnancy (fetal heartbeat). Result(s): The area under the curve of the AI model ranged from 0.6 to 0.7 and outperformed manual morphology grading overall and on a per-site basis. A bootstrapped study predicted improved pregnancy rates between +5% and +12% per site using AI compared with manual grading using an inverted microscope. One site that used a low-magnification stereo zoom microscope did not show predicted improvement with the AI. Visualization techniques and attribution algorithms revealed that the features learned by the AI model largely overlap with the features of manual grading systems. Two sources of bias relating to the type of microscope and presence of embryo holding micropipettes were identified and mitigated. The analysis of AI scores in relation to pregnancy rates showed that score differences of >= 0.1 (10%) correspond with improved pregnancy rates, whereas score differences of <0.1 may not be clinically meaningful. Conclusion(s): This study demonstrates the potential of AI for ranking blastocyst stage embryos and highlights potential limitations related to image quality, bias, and granularity of scores. (C) 2021 by American Society for Reproductive Medicine.
Abstract Study question What is the sensitivity of an embryo-grading artificial intelligence (AI) model to different focal planes and how do we obtain consistent scores across focal planes? Summary answer Test-time augmentation and ensemble modeling reduce sensitivity of the AI model to different focal planes while maintaining performance. What is known already When prioritizing embryos for transfer, embryologists assess the 3D morphological features under a microscope, by zooming up and down, and assign a score that reflects the embryo quality. In comparison, some AI-based embryo grading models typically take one 2D focal plane of an embryo and output a score based on that focal plane. AI models such as convolutional neural networks (CNNs) are known to be sensitive to perturbations in its input. In order to reduce sensitivity and generalization error and thus improve predictive performance, techniques such as ensemble learning and test-time augmentation can be used. Study design, size, duration Historical, de-identified images of blastocyst-stage embryos were collected from 11 IVF clinics in the United States for cycles between 2015-2020. 5,100 blastocysts were matched to pregnancy outcomes as determined by fetal heartbeat. 2,900 blastocysts were matched to aneuploid PGT-A results and added to the negative training group to reduce selection bias. Data was split to 70% for training and 30% for testing. A set of 10 embryos were used for focal plane sensitivity. Participants/materials, setting, methods A single model (ResNet18), a three-model (ResNet18), and a six-model (ResNet18 and EfficientNet-b1) ensemble with and without test-time augmentation were trained to rank embryos according to their likelihood of reaching clinical pregnancy. Test-time augmentation involved taking the average scores from 4 flipped and rotated copies of the original input image. Manual grades were mapped to numeric scores for comparison. The AUC was used to evaluate the ability of the models to rank embryos. Main results and the role of chance Focal plane sensitivity was calculated as the range, or difference between the maximum and minimum score, for an embryo at different focal planes. Between 12 and 100 focal plane images were available for each of the 10 embryos. On average, the focal plane range was 0.26 for the single model, 0.22 for the single model with test-time augmentation, 0.14 for a 3-model ensemble with test-time augmentation, and 0.11 for a 6-model ensemble with test-time augmentation. Test-time augmentation on the single model reduced the range by 17%; whereas ensemble modeling with test-time augmentation reduced the range by 46% for the 3-model ensemble and 60% for the 6-model ensemble. Reduction in range did not compromise performance. The AUC for the test set for all embryos was 0.73 for the single model, 0.74 for the single model with test-time augmentation, 0.75 for the three-model ensemble with test-time augmentation and 0.74 for the six-model ensemble with test-time augmentation. All models outperformed manual grading, which was estimated to have an AUC of 0.67 for all embryos. Limitations, reasons for caution Our analysis on focal plane sensitivity was limited to a small sample size of 10 embryos, so more samples will be needed to confirm our findings. Wider implications of the findings Test-time augmentation and ensemble techniques can be used to reduce sensitivity while maintaining model performance. By reducing sensitivity to different focal planes, an AI model can produce one reliable score for a single embryo as is done currently in practice with manual grading. Trial registration number not applicable
Abstract Study question What is the expected benefit of using a machine learning model for predicting the optimal starting dose of gonadotropin during ovarian stimulation? Summary answer Patients who had an optimal starting gonadotropin dose had improved outcomes and used significantly less total FSH compared to propensity matched patients who did not. What is known already The relationship between the starting dose of follicle-stimulating hormones (FSH) and ovarian response is complex. In general, too little starting FSH may lead to inadequate follicle recruitment, while too much may lead to excessive response. In completed cycles, there exists conflicting evidence of whether higher doses are beneficial or detrimental to the number of oocytes retrieved. The field of assisted reproduction has begun to apply machine learning techniques to clinical decision support for ovarian stimulation, but no studies have specifically investigated optimizing starting FSH dose selection. Study design, size, duration We performed a retrospective analysis of patients undergoing autologous, non-cancelled IVF cycles from 2014 - 2020 (n = 18,591) at three different IVF clinics in the United States. The primary outcomes were the average number of MIIs, 2PNs, and usable blastocysts in relation to starting and total doses of FSH. Participants/materials, setting, methods A K-nearest neighbor similarity model was trained on all cycles and used to identify the 100 most similar patients to a patient-of-interest using age, BMI, baseline anti-mullerian hormone (AMH), and baseline antral follicle count (AFC). For each patient, a patient-specific dose response curve was created by fitting a constrained second order polynomial to the number of MII oocytes relative to the starting dose of FSH across all of the neighbors. Main results and the role of chance For each patient, their individual dose response curve was used to determine if there was an optimal dose that maximizes the prediction of MIIs (called dose-responsive patients), or if the dose response curve shows no optimal dose (called non-responsive patients). 30% of cycles were identified as dose-responsive, 64% were identified as non-responsive, and 6% were inconclusive and excluded from analysis. Dose-responsive patients who received an optimal starting dose had, on average, 1.5 more MIIs, 1.0 more 2PNs, and 0.5 more usable blastocysts using 10 IU’s less of starting FSH and 195 IU’s less of total FSH compared to propensity-matched patients with non-optimal doses. Non-responsive patients who received a low starting dose had, on average, 0.3 more MIIs, 0.4 more 2PNs, and 0.3 more usable blastocysts using 150 IU’s less of FSH and 1375 IU’s less of total FSH compared to propensity-matched patients with a high starting dose. Limitations, reasons for caution The primary limitation is the retrospective nature of this study. Further, our calculations of starting FSH combined the contribution of pure FSH plus the FSH component of FSH/LH medication, rather than evaluating each separately. We also did not differentiate between types of protocols, as the majority were antagonist cycles. Wider implications of the findings Our results suggest that a patient similarity model for selecting starting FSH can help increase MII outcomes while reducing the amount of FSH given to a patient. Future work will include continuing to increase the diversity of our dataset and performing validation studies to show improved outcomes with model use. Trial registration number Not applicable