The objective of this review was to evaluate the efficacy of three promising technologies for assessment of ploidy status in IVF embryos [i.e. preimplantation genetic testing for aneuploidy (PGT-A)]: artificial intelligence (AI), non-invasive PGT-A (niPGT-A) and metabolomics. Publications where >80% correlation with blastocyst biopsies could be demonstrated in ≥50 cycles were prioritized. AI was found to classify the chance of an embryo implanting with an average area under the curve (AUC) of 0.7. AI is thus a superior selection method compared with morphological selection alone, but is still inferior to invasive PGT-A. Some niPGT-A studies have up to 100% concordance with PGT-A, but a multicentre study showed 78% concordance due to maternal contamination, which can improve with specific changes in culture conditions. niPGT-A has thus improved significantly and has the potential to reach 100% with PGT-A if the issue of maternal contamination is solved; however, >30% of euploid embryos never implant. Finally, metabolomics is the least developed technique of the three, but some preliminary data show >90% concordance with implantation and with PGT-A without changing culture conditions. Metabolomics thus has the potential to identify euploid embryos that, metabolically, are incapable of implanting. A combination of two or all of these approaches is possible.
To develop an image-based artificial intelligence (AI) algorithm for combined morphological and genetic assessment of embryo quality. Two AI algorithms, one trained on images of embryos with pre-implantation genetic testing for aneuploidies (PGT-A) outcomes (genetics AI)1, and another trained on images of embryos with pregnancy outcomes (viability AI)2, were combined to create a single AI score to assess embryo quality (EQ score). The optimum ratio of genetics AI score to viability AI score for identifying embryos that were both euploid and of high morphological quality was 2.4:1. The EQ score was assessed for its ability to identify embryos of 3 Gardner-based quality levels: ≥ expansion grade 3 euploid embryos, ≥ 3BB euploid embryos, and ≥ 3AA euploid embryos. Performance was evaluated on a blind test set of 1474 embryo images using ROC-AUC and simulated cohort ranking3. The test set was balanced for morphology as follows: 25% ≥ 3AA, 40% ≥ 3BB, 75% ≥ expansion Grade 3, 25% expansion grades 1-2. Two independent blind test sets of 943 and 664 embryos were used to validate EQ performance based on Gardner and ASEBIR grading, respectively (not balanced). Finally, the EQ score was compared to genetics and viability AI scores alone for its ability to identify euploid embryos (test set of 936 embryo images) and pregnancy outcomes (test set of 479 embryo images), respectively. The EQ score demonstrated high predictive ability for identifying ≥ expansion grade 3 euploid embryos, ≥ 3BB euploid embryos, and ≥ 3AA euploid embryos on 2 datasets, with ROC-AUC values up to 0.772, 0.814, and 0.921, respectively. The EQ score was also able to predict embryo quality according to ASEBIR grading, with ROC-AUC values of 0.716 for ≥ Grade B euploid embryos and 0.814 for ≥ Grade A euploid embryos. Accuracy on the balanced test set was 73-74% for each quality level. Ranking analyses showed the probability of selecting a good quality embryo as the top one in each cohort was 58%, 66%, and 76% for ≥ 3AA euploid, ≥ 3BB euploid, and ≥ expansion grade 3 euploid embryos, respectively. This increased to >95% in each case for the probability of identifying at least 1 good quality embryo in the top-3 ranked embryos in each cohort. The EQ score outperformed both genetics and viability AI scores alone for identification of good quality embryos. It showed a similar probability of selecting a euploid embryo to the genetics AI score alone (83% versus 81%, respectively), and was at least as good at identifying embryos that led to a pregnancy as the viability AI score alone (7% reduction in transfers relative to Gardner-based ranking versus 5.8%, respectively). The EQ score is highly predictive for identifying euploid embryos with high morphological quality. The combined score was comparable to individual genetics and viability AI scores for predicting PGT-A and pregnancy outcomes, respectively.
PDF file - 178K, Nutlin-3a induces cytostatic responses in sarcoma patient material.
PDF file - 67K, TP53 pathway alterations do not mediate sarcoma cytostatic response to Nutlin-3a.
PDF file - 43K, Induction of GADD45A and BBC3 is associated with Nutlin-3a induced apoptosis.
Background and Aims: Artificial intelligence (AI) is being increasingly used for non-invasive evaluation of embryo quality during IVF. Previous studies described development of AI for selecting embryos likely to be euploid (genetics AI), or likely to lead to clinical pregnancy (viability AI), based on analysis of images of blastocysts on day 5 of development. The aim of this study was to determine if a combination of these AI scores could be used to effectively evaluate both outcomes. Method: 936 embryo images with pre-implantation genetic testing for aneuploidies (PGT-A) outcomes, and 479 embryo images with clinical pregnancy outcomes, were retrospectively obtained from 12 IVF clinics in 5 countries. Performance was evaluated for each AI score alone, and the average score of the two AIs. The ability to select euploid or viable embryos was evaluated using ROC-AUC analyses, and a simulated cohort ranking method reported in the literature. Results: The average score of the two AIs was generally as effective at selecting euploid embryos as the genetics AI, and just as effective at selecting viable embryos as the viability AI. Results for both analyses are presented below. Conclusion: An AI score that can evaluate both embryo ploidy and viability simultaneously is useful for selecting preferred embryos for analysis or transfer. These results suggest that it is feasible to generate a single score for evaluating overall embryo quality using a non-invasive approach.
To develop an image-based artificial intelligence (AI) algorithm for automated embryo classification. Two AI algorithms for analyzing embryo morphology were developed previously – the first was trained on pre-implantation genetic testing for aneuploidies (PGT-A) outcomes (genetics AI)1, and the second was trained on clinical pregnancy outcomes (fetal heartbeat at first ultrasound scan) (viability AI)2. Scores from both AIs were shown to correlate with features indicative of embryo quality to differing extents (expansion grade, inner cell mass grade, and trophectoderm grade). Given these observations, the scores were combined to define a series of Morphology Categories as follows: Good Morphology Category = combined AI score of 9.0-10.0 (consisting of the genetics AI score and the viability AI score in a ratio of 2.4:1), Poor Morphology Category = viability AI score of 0.0-4.0, Fair Morphology Category = all other embryos. The Morphology Categories were evaluated for correlation with the US-based SART and EU-based ASEBIR classification systems, using a US dataset of 1764 embryos and a Spanish dataset of 483 embryos, respectively. The US dataset was balanced for morphology according to Gardner Grade as follows: 25% ≥ 3AA, 40% ≥ 3BB, 75% ≥ expansion Grade 3, and 25% expansion grades 1-2. The Spanish dataset was balanced for morphology using ASEBIR grading as follows: 1/3 Grade A, 1/3 Grade B, 1/3 Grade C (no Grade D embryos were available). The proportion of SART Good, Fair, and Poor embryos, and ASEBIR Grade A, B, and C embryos, was calculated for each Morphology Category. The Morphology Categories were effective for classifying embryos of high and low quality according to SART and ASEBIR systems. The Good Morphology Category consisted of 65% SART Good embryos (87% SART Good + Fair embryos), and 77% ASEBIR Grade A embryos, whereas the Poor Morphology Category consisted of 82% SART Poor embryos, and 67% ASEBIR Grade C embryos. The proportion of high-quality embryos in the Poor Morphology Category was only 3% for both SART and ASEBIR systems, and the proportion of low-quality embryos in the Good Morphology Category was 13% and 1% for SART and ASEBIR, respectively. The Fair Morphology Category contained a more heterogeneous mix of embryo qualities, although the predominant embryo quality was the intermediate quality in both cases (50% SART Fair and 44% ASEBIR Grade B). This finding was not unexpected, as even manual grading of these intermediate quality embryos is difficult, and grading demonstrates a high level of inconsistency between embryologists and different IVF clinics. The AI-based Morphology Categories showed a correlation with both US-based SART quality categories and EU-based ASEBIR quality categories, demonstrating applicability across multiple demographics.
Background and Aims: Embryo selection is critical in determining IVF success yet continues to be challenging due to the subjectivity of morphology grading methods, especially when grading fair/average quality embryos. Improving embryo selection could optimise implantation rates and minimise financial/emotional burden on patients. Artificial Intelligence (AI) algorithms represent promising, non-invasive methods of standardising embryo grading and potentially increasing IVF success rates. This study assessed whether an AI algorithm (Life Whisperer Viability) for evaluating the likelihood of clinical pregnancy improves time to pregnancy (TTP) when compared to or combined with standard morphology grading. Method: 305 de-identified 2D images of day 5 blastocysts (121 fresh/184 frozen) with matched clinical pregnancy outcomes (fetal heartbeat at first scan) from women who underwent IVF treatment from 2020-2023 were retrospectively assessed. All images were taken prior to transfer/freezing. TTP was assessed using a simulated cohort ranking method, with TTP being defined as the average number of transfers needed to obtain a clinical pregnancy. Results: A positive linear correlation of LWV scores with pregnancy outcomes was observed (p<0.001). ROC-AUC results indicate that LWV is selecting embryos leading to pregnancy at least as well, if not better, than Gardner morphology grading (0.641 vs 0.624), with further improvement observed when LWV and Gardner grading were combined. The TTP analysis showed a 7.3% reduction in TTP when using LWV over Gardner grading. Combined use of LWV+Gardner grading reduced TTP by up to 10.8%, with the largest improvement (5.3%) seen in the frozen group, where there was a higher distribution of average quality embryos. Conclusion: LWV showed improved embryo rankingand reduction in the estimated average number oftransfers needed to achieveclinical pregnancy. Furthermore, evaluation of TTP supports the combined use of LWV+Gardner grading, showing that they work synergistically to further improve ranking performance when selecting average quality embryos.
Background and Aims: Embryologist evaluation of embryos is critical for ensuring successful pregnancy outcomes. Standard, manual evaluation is variable, subjective, and time-consuming. The aim of this study was to evaluate whether an artificial intelligence (AI) algorithm can standardize and improve embryo evaluation during IVF. Method: 20 images of blastocyst-stage embryos on day 5 of in vitro development were selected to represent a range of morphological qualities. All embryos had been transferred and the clinical pregnancy outcome was known for each embryo based on detection of fetal heartbeat at first ultrasound scan (∼7-9 weeks gestation). 50% of embryos in the dataset resulted in pregnancy. 158 embryologists made a total of 236 attempts at providing their evaluation of the morphological quality of the 20 embryo images using the Gardner system. The embryologist-assigned grades were then used to generate their prediction of whether that embryo would lead to pregnancy or not (≥ 3BB indicated a pregnancy prediction, and <3BB indicated a non-pregnancy prediction). The same 20 embryo images were also assessed by a previously developed viability AI algorithm for evaluating the likelihood of clinical pregnancy based on embryo images. An AI score of ≥5.0/10 indicated a pregnancy prediction, and <5.0/10 indicated a non-pregnancy prediction. The AI algorithm provided the same score for each embryo image regardless of how many times the analysis was performed. Results: The AI algorithm correctly predicted pregnancy outcome for 14/20 embryo images (70%). Embryologists also correctly predicted 14/20 images in 14/236 attempts (6%), and in 1 attempt correctly predicted 15/20 images. In the remaining 221 attempts (94%) embryologists correctly predicted between 6-13 images, representing a range of accuracies from 30-75%. Conclusion: This study demonstrates the inherent variability and lack of objectivity associated with an embryologist’s evaluation of embryos. It highlights the benefits of accurate AI algorithms for standardizing embryo assessment
Research question: Can better methods be developed to evaluate the performance and characteristics of an artificial intelligence model for evaluating the likelihood of clinical pregnancy based on analysis of day-5 blastocyst-stage embryos, such that performance evaluation more closely reflects clinical use in IVF procedures, and correlations with known features of embryo quality are identified?Design: De-identified images were provided retrospectively or collected prospectively by IVF clinics using the artificial intelligence model in clinical practice. A total of 9359 images were provided by 18 IVF clinics across six countries, from 4709 women who underwent IVF between 2011 and 2021. Main outcome measures included clinical pregnancy outcome (fetal heartbeat at first ultrasound scan), embryo morphology score, and/or pre-implantation genetic testing for aneuploidy (PGT-A) results.Results: A positive linear correlation of artificial intelligence scores with pregnancy outcomes was found, and up to a 12.2% reduction in time to pregnancy (TTP) was observed when comparing the artificial intelligence model with standard morphological grading methods using a novel simulated cohort ranking method. Artificial intelligence scores were significantly correlated with known morphological features of embryo quality based on the Gardner score, and with previously unknown morphological features associated with embryo ploidy status, including chromosomal abnormalities indicative of severity when considering embryos for transfer during IVF.Conclusion: Improved methods for evaluating artificial intelligence for embryo selection were developed, and advantages of the artificial intelligence model over current grading approaches were highlighted, strongly supporting the use of the artificial intelligence model in a clinical setting.
Artificial intelligence (AI) is a tool thought to revolutionize the field of reproductive medicine in the years to come. Specifically, machine learning (ML), which is a subset of AI methods used to detect patterns and make predictions based on large datasets, has been used in ART to predict implantation outcome, embryo transfer strategies or adverse outcomes for IVF treatments (for a review see, Wang et al., 2019). However, ML methods have not yet been widely adopted in medically assisted reproduction (MAR). This slow uptake could be due to a lack of in-depth interdisciplinary communication between ML experts and clinicians/embryologists, as well as some uncertainty on how ML models could be generalized for different populations. The February edition of the ESHRE Journal Club discussed a paper from Yland et al. (2022) where the authors compared different ML algorithms to predict the chance of pregnancy in couples actively trying to conceive (without undergoing MAR). By using data from the Pregnancy Study Online (PRESTO), a web-based prospective cohort study of 4133 couples in North America (Wise et al., 2015), the authors analysed 163 potential variables and used supervised ML classification algorithms to identify the strength of each variable in predicting pregnancy outcomes of couples across the fertility spectrum (infertile: <12 menstrual cycles; subfertile: within 6 menstrual cycles; and fecundability: the average probability of pregnancy per menstrual cycle). The authors were able to predict pregnancy with a discrimination as high as 71.2% and identify the variables that most consistently predicted conception. These variables included lifestyle and reproductive characteristics such as age, BMI, history of infertility, daily use of vitamins or folic acid, intercourse timing, diet quality and reduced stress. ESHRE Journal Club discussion focused on promoting interdisciplinarity among ML and reproductive medicine specialists. It hosted 50 participants on Twitter together with experts Michelle Perugini, Vajira Thambawita and authors Jennifer Yland, Yannis Paschalidis and Lauren Wise. The discussion resulted in 902 tweets and around 1 million impressions over a 24-h period.
To investigate whether a non-invasive, deep learning AI algorithm trained on static images of oocytes, denuded prior to ICSI, can predict whether oocytes will develop into a usable blastocyst.
To investigate if a non-invasive AI algorithm developed to evaluate the likely genetic status of embryos at transfer is predictive of live birth.
To determine if a non-invasive AI algorithm for evaluating the likelihood of embryo euploidy (genetics AI) improves selection of viable embryos when used in combination with an AI for evaluating the likelihood of clinical pregnancy (viability AI).
Analysis of clinical data suggests inherent errors in the classification of Day 5 blastocyst images, where viable embryos are wrongly classified non-viable based on a negative pregnancy outcome. A novel AI technique (UDC) was used to identify and remove mis-classified data to obtain a cleaned dataset which improves AI performance and reduces misleading reporting of AI accuracy. Retrospective analysis in private reproductive technology programs. We assessed ∼5,500 static 2D images of Day 5 blastocysts with known clinical pregnancy outcomes. Clinical analysis considered patients under 35 years because they are likely to contain more mis-classified non-viable embryos with patient factors preventing a pregnancy. A novel AI technique (UDC) which identifies incorrectly classified (labeled) data, was used to identify viable embryos incorrectly classified as non-viable. We compared the performance of AI trained using the original embryo dataset and a new cleaned dataset, by assessing accuracy on both an uncleaned and cleaned blind test dataset. Patients <35 that did not achieve a pregnancy had a higher rate (63.6%) of patient factors (e.g. endometriosis) compared with patients ≥35 (49.1%). For patients <35, 49.2% of embryos transferred did not lead to a pregnancy, despite only 17% of these being deemed non-viable by traditional morphological grading. This indicates that there are many examples of embryos deemed non-viable that are likely viable, but did not result in a pregnancy. These mis-classified cases are deemed poor quality data. Applying the UDC to the images identified a significant proportion of embryos suspected to be viable but labeled as non-viable. We removed mis-classified nonviable data to create a clean AI training dataset, and a clean test dataset which is used to report the performance of the AI. Cleaning the training data improved overall AI performance from 59.7% to 61.1%, as measured on an unclean test dataset. There was a large accuracy increase in the (correct) viable class from 76.8% to 80.6%, and a drop in the (misclassified) non-viable class from 37.3% to 35.4%. When measuring the AI performance of the same model on the cleaned test dataset with mis-classified data removed, we found that the original AI accuracy was under-reported, and the true performance overall was 77.1%. For the non-viable class of embryos the under-reported accuracy was even more pronounced, consistent with a larger amount of poor quality data in this class, and the true performance was actually 58.8%. These data suggest that in the class of embryos deemed non-viable due to a negative pregnancy outcome, there are many embryos that are viable and just wrongly classified. The UDC is a unique technique that is effective at identifying these mis-classified cases, which when removed from the AI training datasets results in improved AI performance and enables the true reporting of AI performance. This also calls into question whether it is even possible to achieve the high accuracy (above 90%) reported by others in the literature when embryo viability data is inherently poor quality.
To establish evidence for superior ranking of blastocysts using artificial intelligence (AI) that assesses viability based on Day 5 embryo images. A simulated cohort study was performed to establish a measure of Time-to-Pregnancy (TTP), using AI to rank embryos within cohorts then calculating how many transfers are needed before a successful clinical pregnancy occurs. Retrospective analysis in private reproductive technology programs. An AI model (Life Whisperer) for classifying Day 5 embryo images in terms of viability was developed by training and testing on datasets totaling 3,900 images with a known pregnancy outcome, sourced from 16 clinics across 5 countries. A simulated cohort study was designed whereby retrospective embryo images from a blind test dataset with known pregnancy outcomes were randomized into 116 groups. Each group represented a simulated patients' cohort of embryos, using cohort sizes based on a known clinical distribution. The AI was used to rank embryos in each cohort from most to least likely to be viable. We defined a new measure, TTP, as the position of the first embryo in the ranked cohort to give a positive pregnancy outcome. If the first embryo in the cohort resulted in a positive pregnancy, the TTP for that cohort was 1; if the first embryo was negative but the second embryo was positive, the TTP was 2, etc. A lower TTP was interpreted as a superior ranking outcome. Mean TTP value of the AI ranking was compared to the result expected from random chance, since all embryos in the dataset were already chosen by an embryologist and transferred. The 3,900 embryo images were randomized into 116 simulated cohorts 1,000 times, and the entire set of cohorts were used to provide a bootstrapped statistical analysis. A mean TTP value of 1.506 and standard error of 0.003 was observed for the AI model, compared to a mean TTP of 1.750 and standard error of 0.004 for ranking based on random chance. The differences in mean TTP distribution were modeled as an asymmetric Laplace distribution. Overall, these results translated to a 13.6% improvement in TTP using AI compared to that of random chance, with statistical significance. An AI model trained on clinical pregnancy data showed superior ranking ability and a shorter TTP compared with random chance, for simulated cohorts of transferred embryos. Out-of-pocket expenses for IVF are estimated at $19,000 for the first cycle and $7,000 for each additional cycle (Wu, et al. 2014). Given this, the reduction in TTP means a potential cost saving of $1,200 per patient and clinic revenue increase of $4,600 per treatment through the ability to service more higher value first cycle patients. In the USA with 300,000 annual IVF cycles, AI could achieve total patient savings of $360M, and $1.38B increase in revenue for IVF clinics. Globally with over 2.5M cycles, AI could achieve global patient savings of $3B and $13.8B increase in revenue for IVF clinics.