To develop an image-based artificial intelligence (AI) algorithm for combined morphological and genetic assessment of embryo quality. Two AI algorithms, one trained on images of embryos with pre-implantation genetic testing for aneuploidies (PGT-A) outcomes (genetics AI)1, and another trained on images of embryos with pregnancy outcomes (viability AI)2, were combined to create a single AI score to assess embryo quality (EQ score). The optimum ratio of genetics AI score to viability AI score for identifying embryos that were both euploid and of high morphological quality was 2.4:1. The EQ score was assessed for its ability to identify embryos of 3 Gardner-based quality levels: ≥ expansion grade 3 euploid embryos, ≥ 3BB euploid embryos, and ≥ 3AA euploid embryos. Performance was evaluated on a blind test set of 1474 embryo images using ROC-AUC and simulated cohort ranking3. The test set was balanced for morphology as follows: 25% ≥ 3AA, 40% ≥ 3BB, 75% ≥ expansion Grade 3, 25% expansion grades 1-2. Two independent blind test sets of 943 and 664 embryos were used to validate EQ performance based on Gardner and ASEBIR grading, respectively (not balanced). Finally, the EQ score was compared to genetics and viability AI scores alone for its ability to identify euploid embryos (test set of 936 embryo images) and pregnancy outcomes (test set of 479 embryo images), respectively. The EQ score demonstrated high predictive ability for identifying ≥ expansion grade 3 euploid embryos, ≥ 3BB euploid embryos, and ≥ 3AA euploid embryos on 2 datasets, with ROC-AUC values up to 0.772, 0.814, and 0.921, respectively. The EQ score was also able to predict embryo quality according to ASEBIR grading, with ROC-AUC values of 0.716 for ≥ Grade B euploid embryos and 0.814 for ≥ Grade A euploid embryos. Accuracy on the balanced test set was 73-74% for each quality level. Ranking analyses showed the probability of selecting a good quality embryo as the top one in each cohort was 58%, 66%, and 76% for ≥ 3AA euploid, ≥ 3BB euploid, and ≥ expansion grade 3 euploid embryos, respectively. This increased to >95% in each case for the probability of identifying at least 1 good quality embryo in the top-3 ranked embryos in each cohort. The EQ score outperformed both genetics and viability AI scores alone for identification of good quality embryos. It showed a similar probability of selecting a euploid embryo to the genetics AI score alone (83% versus 81%, respectively), and was at least as good at identifying embryos that led to a pregnancy as the viability AI score alone (7% reduction in transfers relative to Gardner-based ranking versus 5.8%, respectively). The EQ score is highly predictive for identifying euploid embryos with high morphological quality. The combined score was comparable to individual genetics and viability AI scores for predicting PGT-A and pregnancy outcomes, respectively.
It is reported that ∼25% of IVF cycles in some countries utilize PGT-A for aneuploidy detection and/or sex selection. Clinical reasons for sex selection include family balancing and avoidance of sex-linked disease, although ethical concerns prohibit this practice in many countries. Sex identification during IVF is currently performed using invasive PGT-A screening. This study shows it may be possible to utilize non-invasive techniques for evaluating the likelihood of embryo sex prior to transfer.
Medical datasets inherently contain errors from subjective or inaccurate test results, or from confounding biological complexities. It is difficult for medical experts to detect these elusive errors manually, due to lack of contextual information, limiting data privacy regulations, and the sheer scale of data to be reviewed. Current methods for training robust artificial intelligence (AI) models on data containing mislabeled examples generally fall into one of several categories-attempting to improve the robustness of the model architecture, the regularization techniques used, the loss function used during training, or selecting a subset of data that contains cleaner labels. This last category requires the ability to efficiently detect errors either prior to or during training, either relabeling them or removing them completely. More recent progress in error detection has focused on using multi-network learning to minimize deleterious effects of errors on training, however, using many neural networks to reach a consensus on which data should be removed can be computationally intensive and inefficient. In this work, a deep-learning based algorithm was used in conjunction with a label-clustering approach to automate error detection. For dataset with synthetic label flips added, these errors were identified with an accuracy of up to 85%, while requiring up to 93% less computing resources to complete compared to a previous model consensus approach developed previously. The resulting trained AI models exhibited greater training stability and up to a 45% improvement in accuracy, from 69 to over 99% compared to the consensus approach, at least 10% improvement on using noise-robust loss functions in a binary classification problem, and a 51% improvement for multi-class classification. These results indicate that practical, automated a priori detection of errors in medical data is possible, without human oversight.
To develop an image-based artificial intelligence (AI) algorithm for automated embryo classification. Two AI algorithms for analyzing embryo morphology were developed previously – the first was trained on pre-implantation genetic testing for aneuploidies (PGT-A) outcomes (genetics AI)1, and the second was trained on clinical pregnancy outcomes (fetal heartbeat at first ultrasound scan) (viability AI)2. Scores from both AIs were shown to correlate with features indicative of embryo quality to differing extents (expansion grade, inner cell mass grade, and trophectoderm grade). Given these observations, the scores were combined to define a series of Morphology Categories as follows: Good Morphology Category = combined AI score of 9.0-10.0 (consisting of the genetics AI score and the viability AI score in a ratio of 2.4:1), Poor Morphology Category = viability AI score of 0.0-4.0, Fair Morphology Category = all other embryos. The Morphology Categories were evaluated for correlation with the US-based SART and EU-based ASEBIR classification systems, using a US dataset of 1764 embryos and a Spanish dataset of 483 embryos, respectively. The US dataset was balanced for morphology according to Gardner Grade as follows: 25% ≥ 3AA, 40% ≥ 3BB, 75% ≥ expansion Grade 3, and 25% expansion grades 1-2. The Spanish dataset was balanced for morphology using ASEBIR grading as follows: 1/3 Grade A, 1/3 Grade B, 1/3 Grade C (no Grade D embryos were available). The proportion of SART Good, Fair, and Poor embryos, and ASEBIR Grade A, B, and C embryos, was calculated for each Morphology Category. The Morphology Categories were effective for classifying embryos of high and low quality according to SART and ASEBIR systems. The Good Morphology Category consisted of 65% SART Good embryos (87% SART Good + Fair embryos), and 77% ASEBIR Grade A embryos, whereas the Poor Morphology Category consisted of 82% SART Poor embryos, and 67% ASEBIR Grade C embryos. The proportion of high-quality embryos in the Poor Morphology Category was only 3% for both SART and ASEBIR systems, and the proportion of low-quality embryos in the Good Morphology Category was 13% and 1% for SART and ASEBIR, respectively. The Fair Morphology Category contained a more heterogeneous mix of embryo qualities, although the predominant embryo quality was the intermediate quality in both cases (50% SART Fair and 44% ASEBIR Grade B). This finding was not unexpected, as even manual grading of these intermediate quality embryos is difficult, and grading demonstrates a high level of inconsistency between embryologists and different IVF clinics. The AI-based Morphology Categories showed a correlation with both US-based SART quality categories and EU-based ASEBIR quality categories, demonstrating applicability across multiple demographics.
Abstract Study question What is the effect of time-point on performance of a non-invasive artificial intelligence (AI) algorithm for evaluating embryo genetic status? Summary answer While predictive ability was maintained across different time-points on day 5, optimal performance for ranking and selecting euploid embryos was observed at 120 hours post-fertilization. What is known already Studies have shown that it is possible to develop computer vision-based AI algorithms capable of predicting embryo ploidy status using single images of blastocyst-stage embryos. The genetic status of embryos is linked to morphokinetic development, with aneuploidy generally resulting in earlier arrest. Given the dynamic nature of embryo development, it might be expected that the time-point selected for analysis could influence AI performance. The key questions remaining to be answered are to what extent is AI analysis affected by expansion grade, and what does this mean for selection of a time-point for evaluation? Study design, size, duration 2,683 images of day 5 blastocyst-stage embryos with matched ploidy outcomes from pre-implantation genetic testing for aneuploidies (PGT-A) were provided by 10 IVF clinics in the USA, Australia, Malaysia, and India. A subset of 182 embryos had images provided at 110, 115, and 120 hour time points (GERI and EmbryoScope time lapse systems). Participants/materials, setting, methods Images were analysed by a previously developed AI algorithm which evaluates the likelihood of an embryo being euploid according to PGT-A (Diakiw et al, 2022. Hum Reprod, Jul 30;37(8):1746-1759). Evaluation was performed on embryos of each expansion grade, and at three time-points on day 5. Correlations were assessed using Chi-squared test for trend, with pair-wise comparisons conducted using Student’s t-test. Performance was evaluated using ROC-AUC, and a simulated cohort ranking analysis method. Main results and the role of chance AI scores positively correlated with expansion grade, and expansion grade likewise correlated with an increasing proportion of euploid embryos. AI scores also increased over time on day 5, consistent with continued embryo expansion. Scores for grade 4 (expanded) embryos increased more than for grade 5 (hatching) embryos (+2.2-fold and +0.8-fold, respectively), indicative of continued expansion becoming limited at later stages. Despite the valid association of AI scores with expansion grade, results showed the AI could predict euploidy even amongst embryos of the same expansion grade (ROC-AUC ranging from 0.61−0.69). While predictive ability was maintained at each time-point on day 5, ROC-AUC values were highest at 120 hours (0.64, 0.64, and 0.68 for 110, 115, and 120 hours, respectively). Simulated cohort ranking analyses also showed that the AI performed best at 120 hours, selecting a euploid embryo as the top-ranked embryo in 77.1% of patient cohorts (71.4%, 75.1%, and 77.1% for 110, 115, and 120 hours, respectively). These results suggest that regardless of expansion grade, all embryos should be assessed at the same time-point on day 5, preferably closer to the 120 hour time-point. Limitations, reasons for caution The time-point analyses were conducted on a relatively small dataset, and therefore findings should be validated on a larger dataset including embryos of other expansion grades. The analysis was limited to day 5 post-fertilization, however it would be interesting to extend the study to embryos on day 6 also. Wider implications of the findings These results suggest that the AI is providing additional information regarding embryo genetic status, over and above that provided by known morphological parameters. The dynamic nature of AI score related to expansion is of interest as it relates to the optimal time-point for conducting analyses for selection of euploid embryos. Trial registration number Not applicable
Research question: Can better methods be developed to evaluate the performance and characteristics of an artificial intelligence model for evaluating the likelihood of clinical pregnancy based on analysis of day-5 blastocyst-stage embryos, such that performance evaluation more closely reflects clinical use in IVF procedures, and correlations with known features of embryo quality are identified?Design: De-identified images were provided retrospectively or collected prospectively by IVF clinics using the artificial intelligence model in clinical practice. A total of 9359 images were provided by 18 IVF clinics across six countries, from 4709 women who underwent IVF between 2011 and 2021. Main outcome measures included clinical pregnancy outcome (fetal heartbeat at first ultrasound scan), embryo morphology score, and/or pre-implantation genetic testing for aneuploidy (PGT-A) results.Results: A positive linear correlation of artificial intelligence scores with pregnancy outcomes was found, and up to a 12.2% reduction in time to pregnancy (TTP) was observed when comparing the artificial intelligence model with standard morphological grading methods using a novel simulated cohort ranking method. Artificial intelligence scores were significantly correlated with known morphological features of embryo quality based on the Gardner score, and with previously unknown morphological features associated with embryo ploidy status, including chromosomal abnormalities indicative of severity when considering embryos for transfer during IVF.Conclusion: Improved methods for evaluating artificial intelligence for embryo selection were developed, and advantages of the artificial intelligence model over current grading approaches were highlighted, strongly supporting the use of the artificial intelligence model in a clinical setting.
Abstract Study question Does patient age need to be explicitly factored into AI-based embryo quality assessment, or does embryo morphology alone capture the age-related decline in embryo quality? Summary answer Age-related effects on embryo quality are inherently captured in embryo morphology. AI algorithms that assess morphology correlate with expected decline in embryo quality with age. What is known already Patient age strongly correlates with genetic aneuploidy in oocytes, which results in a dramatic reduction in genetic integrity and viability of embryos with patient age1. This negative correlation ultimately leads to poorer implantation and clinical pregnancy outcomes. AI imaging tools assess the quality of embryos generally, using morphology alone2. However, it is unknown whether these morphological assessments inherently consider age-related quality factors like cytoplasmic and/or genetic competence, or whether age should be incorporated as a separate variable. The current study aimed to assess the correlation of AI-based scores with the age-related decline in embryo quality. Study design, size, duration The study used a retrospective dataset of static Day 5 blastocyst images taken using an optical light microscope with associated PGT-A or pregnancy outcomes. The dataset comprised images of 4,000 embryos sourced from 1,199 consecutive patients treated between 2011 and 2020 at five IVF clinics (USA). The study evaluated correlation of algorithms Life Whisperer Genetics and Life Whisperer Viability with patient or donor age. Data were excluded in donor cases where age was not known. Participants/materials, setting, methods 4,000 embryo images were used to report a linear correlation between proportion of euploids(%) and pregnancies(%) across six age-brackets, between 20 to 50 years old. Life Whisperer Genetics AI was applied to a blind dataset of 809 images to assess likelihood of euploidy, and Life Whisperer Viability AI applied to a dataset of 556 images to assess likelihood of pregnancy. Scores within each age-bracket were averaged and chi-squared analyses was used to assess significance. Main results and the role of chance As expected, there was a significant negative correlation between the number of euploid embryos(%) and patient/donor age on a dataset of 4,000 images (slope of -13.2±0.2), and on a blind test set of 809 embryos (slope of -11.2±0.2). The Life Whisperer Genetics AI score was then reported on the blind test set, showing a significant negative correlation with age (-0.45±0.16 with a χ2/dof value of 0.41). The significant downward trend indicates that the AI, using morphology alone, can account for the age-related impact in the genetic competence without a corresponding reduction in accuracy, and without needing additional age-related variables in its calculation. The AI was able to generalize correctly, identifying morphological signs of ploidy well, regardless of age. Regarding cytoplasmic or metabolic competence, we report on a blind dataset of 556 images that the proportion of viable embryos(%) reduces with increasing patient age, although exhibiting a peak in proportion of viable embryos in the 25-29 year bracket. Similarly, we show that Life Whisperer Viability AI scores within each age-bracket reduce with age. Our results suggest that both AI algorithms for genetic competence and metabolic competence in terms of viability take into account patient age based on morphology. Limitations, reasons for caution Although age was shown to be represented in embryo morphology, adding a separate age-related variable could be considered in future studies. However, for embryo ranking and selection for a given patient, this is likely to be of value only when comparing embryos corresponding to different donor oocytes. Wider implications of the findings As the age of the patient increases, the morphology of their embryos also changes, corresponding to a decrease in embryo quality. This justifies morphology-based embryo quality assessment, giving credence to generalizable AI that perform robust assessment of embryo quality for patients of all ages, and do not require calibration. Trial registration number Not Applicable
Abstract STUDY QUESTION Can an artificial intelligence (AI) model predict human embryo ploidy status using static images captured by optical light microscopy? SUMMARY ANSWER Results demonstrated predictive accuracy for embryo euploidy and showed a significant correlation between AI score and euploidy rate, based on assessment of images of blastocysts at Day 5 after IVF. WHAT IS KNOWN ALREADY Euploid embryos displaying the normal human chromosomal complement of 46 chromosomes are preferentially selected for transfer over aneuploid embryos (abnormal complement), as they are associated with improved clinical outcomes. Currently, evaluation of embryo genetic status is most commonly performed by preimplantation genetic testing for aneuploidy (PGT-A), which involves embryo biopsy and genetic testing. The potential for embryo damage during biopsy, and the non-uniform nature of aneuploid cells in mosaic embryos, has prompted investigation of additional, non-invasive, whole embryo methods for evaluation of embryo genetic status. STUDY DESIGN, SIZE, DURATION A total of 15 192 blastocyst-stage embryo images with associated clinical outcomes were provided by 10 different IVF clinics in the USA, India, Spain and Malaysia. The majority of data were retrospective, with two additional prospectively collected blind datasets provided by IVF clinics using the genetics AI model in clinical practice. Of these images, a total of 5050 images of embryos on Day 5 of in vitro culture were used for the development of the AI model. These Day 5 images were provided for 2438 consecutively treated women who had undergone IVF procedures in the USA between 2011 and 2020. The remaining images were used for evaluation of performance in different settings, or otherwise excluded for not matching the inclusion criteria. PARTICIPANTS/MATERIALS, SETTING, METHODS The genetics AI model was trained using static 2-dimensional optical light microscope images of Day 5 blastocysts with linked genetic metadata obtained from PGT-A. The endpoint was ploidy status (euploid or aneuploid) based on PGT-A results. Predictive accuracy was determined by evaluating sensitivity (correct prediction of euploid), specificity (correct prediction of aneuploid) and overall accuracy. The Matthew correlation coefficient and receiver-operating characteristic curves and precision-recall curves (including AUC values), were also determined. Performance was also evaluated using correlation analyses and simulated cohort studies to evaluate ranking ability for euploid enrichment. MAIN RESULTS AND THE ROLE OF CHANCE Overall accuracy for the prediction of euploidy on a blind test dataset was 65.3%, with a sensitivity of 74.6%. When the blind test dataset was cleansed of poor quality and mislabeled images, overall accuracy increased to 77.4%. This performance may be relevant to clinical situations where confounding factors, such as variability in PGT-A testing, have been accounted for. There was a significant positive correlation between AI score and the proportion of euploid embryos, with very high scoring embryos (9.0–10.0) twice as likely to be euploid than the lowest-scoring embryos (0.0–2.4). When using the genetics AI model to rank embryos in a cohort, the probability of the top-ranked embryo being euploid was 82.4%, which was 26.4% more effective than using random ranking, and ∼13–19% more effective than using the Gardner score. The probability increased to 97.0% when considering the likelihood of one of the top two ranked embryos being euploid, and the probability of both top two ranked embryos being euploid was 66.4%. Additional analyses showed that the AI model generalized well to different patient demographics and could also be used for the evaluation of Day 6 embryos and for images taken using multiple time-lapse systems. Results suggested that the AI model could potentially be used to differentiate mosaic embryos based on the level of mosaicism. LIMITATIONS, REASONS FOR CAUTION While the current investigation was performed using both retrospectively and prospectively collected data, it will be important to continue to evaluate real-world use of the genetics AI model. The endpoint described was euploidy based on the clinical outcome of PGT-A results only, so predictive accuracy for genetic status in utero or at birth was not evaluated. Rebiopsy studies of embryos using a range of PGT-A methods indicated a degree of variability in PGT-A results, which must be considered when interpreting the performance of the AI model. WIDER IMPLICATIONS OF THE FINDINGS These findings collectively support the use of this genetics AI model for the evaluation of embryo ploidy status in a clinical setting. Results can be used to aid in prioritizing and enriching for embryos that are likely to be euploid for multiple clinical purposes, including selection for transfer in the absence of alternative genetic testing methods, selection for cryopreservation for future use or selection for further confirmatory PGT-A testing, as required. STUDY FUNDING/COMPETING INTEREST(S) Life Whisperer Diagnostics is a wholly owned subsidiary of the parent company, Presagen Holdings Pty Ltd. Funding for the study was provided by Presagen with grant funding received from the South Australian Government: Research, Commercialisation, and Startup Fund (RCSF). ‘In kind’ support and embryology expertise to guide algorithm development were provided by Ovation Fertility. ‘In kind’ support in terms of computational resources provided through the Amazon Web Services (AWS) Activate Program. J.M.M.H., D.P. and M.P. are co-owners of Life Whisperer and Presagen. S.M.D., M.A.D. and T.V.N. are employees or former employees of Life Whisperer. S.M.D, J.M.M.H, M.A.D, T.V.N., D.P. and M.P. are listed as inventors of patents relating to this work, and also have stock options in the parent company Presagen. M.V. sits on the advisory board for the global distributor of the technology described in this study and also received support for attending meetings. TRIAL REGISTRATION NUMBER N/A.
Training on multiple diverse data sources is critical to ensure unbiased and generalizable AI. In healthcare, data privacy laws prohibit data from being moved outside the country of origin, preventing global medical datasets being centralized for AI training. Data-centric, cross-silo federated learning represents a pathway forward for training on distributed medical datasets. Existing approaches typically require updates to a training model to be transferred to a central server, potentially breaching data privacy laws unless the updates are sufficiently disguised or abstracted to prevent reconstruction of the dataset. Here we present a completely decentralized federated learning approach, using knowledge distillation, ensuring data privacy and protection. Each node operates independently without needing to access external data. AI accuracy using this approach is found to be comparable to centralized training, and when nodes comprise poor-quality data, which is common in healthcare, AI accuracy can exceed the performance of traditional centralized training.
To investigate whether a non-invasive, deep learning AI algorithm trained on static images of oocytes, denuded prior to ICSI, can predict whether oocytes will develop into a usable blastocyst.
To investigate if a non-invasive AI algorithm developed to evaluate the likely genetic status of embryos at transfer is predictive of live birth.
To determine if a non-invasive AI algorithm for evaluating the likelihood of embryo euploidy (genetics AI) improves selection of viable embryos when used in combination with an AI for evaluating the likelihood of clinical pregnancy (viability AI).
Abstract Study question Do AI models used to assess embryo viability (based on pregnancy outcome) also correlate with known embryo quality measures such as ploidy status? Summary answer An AI for embryo viability assessment correlated with ploidy status, and with karyotypic features of aneuploidy, supporting its use for embryo selection. What is known already One factor that can influence pregnancy success is the genetic status of the embryo. PGT-A is commonly used to test for embryo ploidy, with the aim of identifying karyotypically normal embryos (euploid embryos), for preferential transfer. There is evidence suggesting that transfer of euploid embryos produces favorable clinical outcomes over aneuploid embryos. Given the AI model was trained to evaluate clinical pregnancy, it was hypothesized that the score might also correlate with ploidy status, and with different types of aneuploidies. Little is known about morphological correlations with embryo ploidy status, so we also sought to explore this relationship. Study design, size, duration This study involved analysis of a retrospective dataset of single static Day 5 embryo (blastocyst) images with associated PGT-A results and AI viability scores. The dataset comprised images of 5,469 embryos from 2,615 consecutive patients treated at five US IVF clinics between February 2015 and April 2020. The AI was trained on thousands of Day 5 embryo images from multiple IVF laboratories in multiple countries, but was not trained on data used in this study. Participants/materials, setting, methods Average patient age was 36.2 years, and average embryo cohort size was 2.1/patient. PGT-A analysis was performed on embryos at time of evaluation. The dataset comprised 3,251 (59.4%) euploid embryos, 1,815 (33.2%) aneuploid embryos, and 403 (7.4%) mosaic embryos. The AI was retrospectively used to provide a score between 0 (predicted non-viable) and 10 (predicted viable) for each image. Correlation between the AI viability score and euploid, mosaic and aneuploid embryos was then assessed. Main results and the role of chance Results showed a statistically significant correlation between AI viability score and PGT-A outcome, consistent with a relationship between pregnancy outcome and ploidy status. The average score for euploid embryos was 8.20, which was significantly higher than the average score for aneuploid embryos of 7.80 (p < 0.0001). There was a significant linear increase in confidence score from full aneuploid embryos, through mosaic embryos (average score 7.97), to full euploid embryos (mosaic threshold of 20–80%). High mosaic embryos tended to have a lower average score (7.60) than low mosaic embryos (7.96), consistent with correlation of viability (pregnancy outcome) with the degree of mosaicism. AI viability score also correlated with ploidy features believed to affect pregnancy outcomes. Trisomic changes had higher average scores than monosomic changes. Segmental changes had higher average scores than full gain or loss. The AI score differentiated euploid from aneuploid status more efficiently in embryos with poorer morphology than those with good morphology. Whilst there was an evident correlation between pregnancy outcome and ploidy status, the AI was only weakly predictive of euploidy, with an accuracy of 57.3% using an AI viability score threshold of 7.5/10.This suggests pregnancy-related morphological features are somewhat correlated with embryo ploidy, but not completely. Limitations, reasons for caution The PGT-A technique is held to have some limitations for evaluating ploidy status, therefore it would be of benefit to perform additional confirmatory studies on independent datasets. It would be of interest to conduct prospective studies evaluating correlations between the AI’s evaluation of morphology and pregnancy outcome with ploidy status. Wider implications of the findings: The AI score correlated with genetic features of embryos that are known to correlate with pregnancy, which further supports the efficacy and use of AI for embryo viability assessment. The AI identified morphological features that are somewhat predictive of ploidy status, with potential application to embryos of poorer Gardner score. Trial registration number none
What are the major problems faced by embryologists at 1) Clinic level, 2) Professional level, 3) Personal level, and 4) What are their career goals? Embryologists, essential professionals of Fertility Centres, are less satisfied in many quantifiable aspects, but they love their profession and have many aspirational goals. IVF success depends in part on embryologists’ skills. The need to recognize clinical embryology as a specialty and clinical embryologists’ educational level, responsibilities, and workload have been addressed by a few national societies. However, data are lacking from the embryologists’ viewpoint at a global level about their profession. Qualitative data-analysis methods provide thick, rich descriptions of subjects’ thoughts, feelings, and lived experiences but can be time-consuming, labor-intensive, and prone to bias. A questionnaire was prepared using SurveyMonkey online software (SurveyMonkey, Inc., USA) and distributed to IVF lab professionals through embryology societies, online social media, and email databases. The questionnaire consisted of open-ended questions focused on identifying problems faced by embryologists at the clinic, in the profession, and in a personal level, as well as questions about their career outlook. The survey was active from May 2016 until February 2017. From 73 countries, 720 responses were obtained. Using natural language processing (NLP), the top 15 most frequently used keywords were identified and correlated with each other. Stronger correlation (≥0.5) between semantically similar words expressing a strong signal from each answer, and their usage was further analyzed for positive versus negative sentiment. By normalizing the frequency of positive/negative samples for each keyword as a percentage, “sentiment wheels” were produced, identifying the key concepts that respondents answered and quantifying how they felt about them. The responses received were from 80% private, 17% public and 3% other ART settings distributed all over the world. From the embryologists’ viewpoints reported and after the NLP processing it was shown that the common topics related to strong negative sentiments were: embryologists’ remuneration (0.6) at the Clinic level; certification (0.7), recognition (0.5), respect (0.5), learn (0.5) and experience (0.5) at the Professional level; and remuneration (0.7), emotional (0.5) dealing (0.5) at the Personal level. Renumeration was reported and strongly related to embryologists’ viewpoint at both the clinic and personal level in combination with the need for certification, recognition and ongoing development at the Professional level. Moreover, the NLP processing demonstrated that the common topics on career goal analysis related to strong positive sentiments were: teaching (0.7), education (0.7), and continuation (0.5) all three topics are compatible with a professional orientation open to ongoing development and practice advancement. The NLP and the manual data analysis project an image of the typical embryologist as a knowledge seeking professional who is deeply dedicated to the job but feels the need for professional development and suffers some lack of recognition and feels in some cases not fairly treated as an employee. The data obtained is limited. Only one natural language processing model was used to analyze the results. Different analysts using other methods may have different results. For these reasons, the results should be interpreted with caution. Wider implications of the findings: It is important to focus on the lab as an organization and not just a service for the patients in treatment at the moment. The NLP results ultimately obtained may help streamline professional satisfaction efforts, and guide future quality management strategies Not applicable
The detection and removal of poor-quality data in a training set is crucial to achieve high-performing AI models. In healthcare, data can be inherently poor-quality due to uncertainty or subjectivity, but as is often the case, the requirement for data privacy restricts AI practitioners from accessing raw training data, meaning manual visual verification of private patient data is not possible. Here we describe a novel method for automated identification of poor-quality data, called Untrainable Data Cleansing. This method is shown to have numerous benefits including protection of private patient data; improvement in AI generalizability; reduction in time, cost, and data needed for training; all while offering a truer reporting of AI performance itself. Additionally, results show that Untrainable Data Cleansing could be useful as a triage tool to identify difficult clinical cases that may warrant in-depth evaluation or additional testing to support a diagnosis.
Do artificial intelligence (AI) models used to assess embryo viability (based on pregnancy outcomes) also correlate with known embryo quality measures such as Gardner score? An AI for embryo viability assessment also correlated with Gardner score, further substantiating the use of AI for assessment and selection of good quality embryos. The Gardner score consists of three separate components of embryo morphology that are graded individually, then combined to give a final score describing Day 5 embryo (blastocyst) quality. Evidence suggests the Gardner score has some correlation with clinical pregnancy. We hypothesized that an AI model trained to evaluate likelihood of clinical pregnancy based on fetal heartbeat (in clinical use globally) would also correlate with components of the Gardner score itself. We also compared the ability of the AI and Gardner score to predict pregnancy outcomes. This study involved analysis of a prospectively collected dataset of single static Day 5 embryo images with associated Gardner scores and AI viability scores. The dataset comprised time-lapse images of 1,485 embryos (EmbryoScope) from 638 patients treated at a single in vitro fertilization (IVF) clinic between November 2019 and December 2020. The AI model was not trained on data from this clinic. Average patient age was 35.4 years. Embryologists manually graded each embryo using the Gardner method, then subsequently used the AI to obtain a score between 0 (predicted non-viable, unlikely to lead to a pregnancy) and 10 (predicted viable, likely to lead to a pregnancy). Correlation between the AI viability score and Gardner score was then assessed. The average AI score was significantly correlated with the three components of the Gardner score: expansion grade, inner cell mass (ICM) grade, and trophectoderm grade. Average AI score generally increased with advancing blastocyst developmental stage. Blastocysts with expansion grades of ≥ 3 are generally considered suitable for transfer. This study showed that embryos with expansion grade 3 had lower AI scores than those with grades 4-6, consistent with a reduced pregnancy rate. AI correlation with trophectoderm grade was more significant than with ICM grade, consistent with studies demonstrating that trophectoderm grade is more important than ICM in determining clinical pregnancy likelihood. The AI predicted Gardner scores of ≥ 2BB with an accuracy of 71.7% (sensitivity 75.1%, specificity 45.9%), and an AUC of 0.68. However, when used to predict pregnancy outcome, the AI performed 27.9% better than the Gardner score (accuracies of 49.8% and 39.0% respectively). Even though the AI was highly correlated with the Gardner score, the improved efficacy for predicting pregnancy suggests that a) the AI provides an advantage in standardization of scoring over the manual and subjective Gardner method, and b) the AI is likely identifying and evaluating morphological features of embryo quality that are not captured by the Gardner method. The Gardner score is not a linear score, creating challenges with setting a suitable threshold relating to the prediction of pregnancy. The 2BB treshold was chosen based on literature (Munné et al 2019) and verified by experienced embryologists. This correlative study may also require additional confirmatory studies on independent datasets. The correlation between AI scores and known features of embryo quality (Gardner score) substantiates the use of the AI for embryo assessment. The AI score provides further insight into components of the Gardner score, and may detect morphological features related to clinical pregnancy beyond those evaluated by the Gardner method. Not applicable
Abstract Study question Does embryo quality/viability change over time, suggesting the use of video for AI-based embryo quality assessment has limited benefit over single point-in-time images? Summary answer AI assessment of single static embryo images at multiple time-points indicates embryo viability is dynamic, and past viability is a limited predictor of future pregnancy. What is known already Artificial Intelligence (AI) has been applied to the problem of embryo quality (viability) assessment using either video or single static images. However, whether historical data within video provide an additional advantage over single static images of embryos (at the time of transfer) for assessing embryo viability is not known. This applies to both manual and AI-based embryo assessment. If embryo viability changes over time prior to transfer, then the implication is that the assessment of future pregnancy using historical embryo data from videos would provide limited additional value over single static images taken immediately prior to transfer. Study design, size, duration Retrospective dataset of single embryo images taken at up-to three time-points prior to transfer: Early Day 5, Late Day 5 (8 hours later), and Early Day 6 (16 hours later), with corresponding fetal heartbeat (pregnancy) outcomes. The AI assessed the viability of each embryo at its available timepoints. Viability prediction was compared with pregnancy outcome to assess viability predictiveness at each timepoint prior to transfer, and assess the variability of viability over time. Participants/materials, setting, methods Single static images of 173 embryos were taken using time-lapse incubators from a single IVF clinic. 116 embryos were viable (led to a pregnancy) and 57 were non-viable (did not lead to a pregnancy). The AI was trained on thousands of Day 5 static embryo images taken from multiple IVF laboratories and countries, but was not trained on data from this clinic. Main results and the role of chance When embryos were assessed as viable by the AI immediately prior to transfer (no delay), the AI accuracy (sensitivity) in predicting pregnancy was 88.1% (59/67) for Early Day 5, 84.8% (28/33) for Late Day 5 and 87.5% (14/16) for Early Day 6. When the delay between AI assessment and transfer is 8 hours, 16 hours and 24 hours, the the accuracy drops to 66.7% (22/33), 31.3% (5/16) and 12.5% (2/16), respectively. These results indicate that the viability of the embryo is dynamic, and therefore time series analysis, i.e. using video, may not be well suited for embryo viability assessment because past viability is not necessarily a good predictor of future viability or pregnancy outcome. The viability of the embryo immediately prior to transfer, from a single static image, is a reliable predictor of viability. This is consistent with the current clinical practice of using Gardner score end-point assessment for embryo quality. Results also suggest significant benefits from using time-lapse with AI, where AI continually assesses embryo viability over time using static images. The time point at which the embryo should be transferred to maximize pregnancy outcome is when the embryo has the greatest AI viability score. Limitations, reasons for caution Although evidence suggests past embryo viability is a limited predictor of future pregnancy, a side-by-side comparison of video versus single static image AI assessment would further verify that the historical or change in embryo development or viability has minimal impact on embryo viability assessment at the time prior to transfer. Wider implications of the findings: Time-lapse and AI can beneficially change the way embryos are assessed. Continual AI monitoring of embryos enables optimization of which embryo to transfer and when, to ultimately improve pregnancy outcomes for patients. The findings also suggest that static end-point AI assessment is sufficient for predicting embryo implantation potential. Trial registration number Not applicable
Analysis of clinical data suggests inherent errors in the classification of Day 5 blastocyst images, where viable embryos are wrongly classified non-viable based on a negative pregnancy outcome. A novel AI technique (UDC) was used to identify and remove mis-classified data to obtain a cleaned dataset which improves AI performance and reduces misleading reporting of AI accuracy. Retrospective analysis in private reproductive technology programs. We assessed ∼5,500 static 2D images of Day 5 blastocysts with known clinical pregnancy outcomes. Clinical analysis considered patients under 35 years because they are likely to contain more mis-classified non-viable embryos with patient factors preventing a pregnancy. A novel AI technique (UDC) which identifies incorrectly classified (labeled) data, was used to identify viable embryos incorrectly classified as non-viable. We compared the performance of AI trained using the original embryo dataset and a new cleaned dataset, by assessing accuracy on both an uncleaned and cleaned blind test dataset. Patients <35 that did not achieve a pregnancy had a higher rate (63.6%) of patient factors (e.g. endometriosis) compared with patients ≥35 (49.1%). For patients <35, 49.2% of embryos transferred did not lead to a pregnancy, despite only 17% of these being deemed non-viable by traditional morphological grading. This indicates that there are many examples of embryos deemed non-viable that are likely viable, but did not result in a pregnancy. These mis-classified cases are deemed poor quality data. Applying the UDC to the images identified a significant proportion of embryos suspected to be viable but labeled as non-viable. We removed mis-classified nonviable data to create a clean AI training dataset, and a clean test dataset which is used to report the performance of the AI. Cleaning the training data improved overall AI performance from 59.7% to 61.1%, as measured on an unclean test dataset. There was a large accuracy increase in the (correct) viable class from 76.8% to 80.6%, and a drop in the (misclassified) non-viable class from 37.3% to 35.4%. When measuring the AI performance of the same model on the cleaned test dataset with mis-classified data removed, we found that the original AI accuracy was under-reported, and the true performance overall was 77.1%. For the non-viable class of embryos the under-reported accuracy was even more pronounced, consistent with a larger amount of poor quality data in this class, and the true performance was actually 58.8%. These data suggest that in the class of embryos deemed non-viable due to a negative pregnancy outcome, there are many embryos that are viable and just wrongly classified. The UDC is a unique technique that is effective at identifying these mis-classified cases, which when removed from the AI training datasets results in improved AI performance and enables the true reporting of AI performance. This also calls into question whether it is even possible to achieve the high accuracy (above 90%) reported by others in the literature when embryo viability data is inherently poor quality.
STUDY QUESTION: Can an artificial intelligence (AI)-based model predict human embryo viability using images captured by optical light microscopy? SUMMARY ANSWER: We have combined computer vision image processing methods and deep learning techniques to create the non-invasive Life Whisperer AI model for robust prediction of embryo viability, as measured by clinical pregnancy outcome, using single static images of Day 5 blastocysts obtained from standard optical light microscope systems. WHAT IS KNOWN ALREADY: Embryo selection following IVF is a critical factor in determining the success of ensuing pregnancy. Traditional morphokinetic grading by trained embryologists can be subjective and variable, and other complementary techniques, such as time-lapse imaging, require costly equipment and have not reliably demonstrated predictive ability for the endpoint of clinical pregnancy. AI methods are being investigated as a promising means for improving embryo selection and predicting implantation and pregnancy outcomes. STUDY DESIGN, SIZE, DURATION: These studies involved analysis of retrospectively collected data including standard optical light microscope images and clinical outcomes of 8886 embryos from 11 different IVF clinics, across three different countries, between 2011 and 2018. PARTICIPANTS/MATERIALS, SETTING, METHODS: The AI-based model was trained using static two-dimensional optical light microscope images with known clinical pregnancy outcome as measured by fetal heartbeat to provide a confidence score for prediction of pregnancy. Predictive accuracy was determined by evaluating sensitivity, specificity and overall weighted accuracy, and was visualized using histograms of the distributions of predictions. Comparison to embryologists' predictive accuracy was performed using a binary classification approach and a 5-band ranking comparison. MAIN RESULTS AND THE ROLE OF CHANCE: The Life Whisperer AI model showed a sensitivity of 70.1% for viable embryos while maintaining a specificity of 60.5% for non-viable embryos across three independent blind test sets from different clinics. The weighted overall accuracy in each blind test set was >63%, with a combined accuracy of 64.3% across both viable and non-viable embryos, demonstrating model robustness and generalizability beyond the result expected from chance. Distributions of predictions showed clear separation of correctly and incorrectly classified embryos. Binary comparison of viable/non-viable embryo classification demonstrated an improvement of 24.7% over embryologists' accuracy (P = 0.047, n = 2, Student's t test), and 5-band ranking comparison demonstrated an improvement of 42.0% over embryologists (P = 0.028, n = 2, Student's t test). LIMITATIONS, REASONS FOR CAUTION: The AI model developed here is limited to analysis of Day 5 embryos; therefore, further evaluation or modification of the model is needed to incorporate information from different time points. The endpoint described is clinical pregnancy as measured by fetal heartbeat, and this does not indicate the probability of live birth. The current investigation was performed with retrospectively collected data, and hence it will be of importance to collect data prospectively to assess real-world use of the AI model. WIDER IMPLICATIONS OF THE FINDINGS: These studies demonstrated an improved predictive ability for evaluation of embryo viability when compared with embryologists' traditional morphokinetic grading methods. The superior accuracy of the Life Whisperer AI model could lead to improved pregnancy success rates in IVF when used in a clinical setting. It could also potentially assist in standardization of embryo selection methods across multiple clinical environments, while eliminating the need for complex time-lapse imaging equipment. Finally, the cloud-based software application used to apply the Life Whisperer AI model in clinical practice makes it broadly applicable and globally scalable to IVF clinics worldwide.