Is single-cell lineage tree reconstruction feasible in clinical human embryos? Single-cell lineage trees can be generated from cleavage stage human embryos with the help of artificial intelligence (AI). Single-cell lineage trees (SCLTs) are important for understanding cell differentiation and tissue formation in early embryo development. They also provide insights into embryo genetics and developmental mechanisms. However, SCLT reconstruction in the preimplantation human embryo has largely been confined to research settings, employing techniques such as dye injection. Despite a limited number of studies and small sample sizes, significant discoveries have been made. One of the largest studies to date manually reconstructed lineage trees in 54 clinical cleavage stage embryos using time-lapse incubators, revealing that clones from the first mitotic division contribute unequally to the ICM and TE. Retrospective analysis of embryoscope time-lapse data captured from three clinics. The dataset comprised 498 embryo focal stacks used for 3D reconstruction pipeline training and evaluation; 15 timelapses each of 20 frames captured between 35-70 hours post-insemination (hpi) for cell tracking system training and evaluation; and 3 full timelapse videos captured from 0-60 hpi to evaluate SCLT reconstruction. All data was manually annotated with segmentation masks, cleavage timings and single-cell lineage trees as appropriate. We take the approach of first detecting cells at each frame before linking them together into tracks. For the first stage, we trained the 3D reconstruction pipeline introduced by He et al. to detect the positions of individual cells. These positions were then fed into a variant of the Crocker-Grier algorithm to track cells over time. The algorithm was adapted to detect cleavage events. From tracks and histories of cleavage events, SCLTs were automatically reconstructed. On the whole, our results suggest that reconstructing SCLTs is feasible in a clinical context. We evaluated the performance of our system at cell tracking using established cell tracking metrics including Multi-Object Tracking Accuracy (MOTA), Multi-Object Tracking Precision (MOTP) and IDF1. We evaluated the detection of cleavage events with the precision, recall and F1 metrics. Following 5-fold cross-validation, our system achieved a mean MOTA of 0.38, IDF1 of 0.74 and MOTP of 0.77. Our system also detected cleavage events with a precision of 0.89, recall of 0.74 and F1 of 0.81. Cleavage event detection timings were, on average, 0.24 frames off from ground truth. Further in-depth manual evaluation was employed to identify the specific failure modes of the system. The system made several errors, including missing reverse cleavages (N = 2), incorrectly propagating the identities of parent cells (N = 2), assigning child cells to incorrect parents (N = 1), and swapping cell identities (N = 4). As such, while the system is not yet suitable for fully autonomous operation, it may have potential to significantly accelerate annotation, supporting a human-in-the-loop approach to SCLT reconstruction - a task previously infeasible in most clinical settings due to time constraints. While this study serves as a proof-of-concept for SCLT reconstruction in clinical embryos, larger datasets are needed. This may be tricky, as lineage ground-truth is very time-consuming to annotate. Moreover, future studies should assess SCLTs’ clinical relevance by correlating reconstructed lineage trees with single-cell sequencing and outcome data. This study demonstrates the feasibility of reconstructing SCLTs in clinical embryos. From a biological standpoint, the system could aid in detecting mitotic aneuploidy and serve as a marker for mosaicism. Additionally, it highlights a more explainable, biologically-grounded application of AI, providing additional insights over conventional morphokinetics. No
Abstract Study question 3-dimensional quantitative assessment of the inner cell mass (ICM) is more effective at selecting viable embryos than the widely used Gardner grading system? Summary answer 3-dimensional quantitative assessment of the ICM is significantly more predictive of implantation and live birth than the widely used Gardner grading of the ICM What is known already ICM assessment forms a key part of most widely used blastocyst grading systems. Grading under these systems, however, is often qualitative, subjective, and subject to inter/intra-operator variability. To reduce subjectivity in ICM grading, it has been suggested that quantitative criteria are used. Previous studies have focused on the assessment of such parameters (e.g. ICM area) in 2 dimensions. However, such approaches fail to account for the various embryo orientations which could lead to different measurements. Therefore, quantitative 3-dimensional analysis of the ICM is needed to improve the reliability and consistency of grading. Study design, size, duration This was a retrospective analysis of 153 randomly selected tEB embryos from two clinics. All data was captured on Embryoscope between 2018 and 2020. Data for maternal age, blastocyst grade, blastocyst ploidy, implantation and live birth rates were also recorded. Participants/materials, setting, methods Embryos were imaged at 7+ focal planes and in-focus sections of each ICM were annotated using ImageJ. These sections were combined into a 3D model using Blender, and sphericity (ψ) computed. Embryos were categorised as spheroid (ψ ≥ 0.8), ellipsoid (ψ < 0.8), crescent and irregular. Chi-squared tests were performed, comparing classification and clinical outcome. The predictive power of the system was compared to the Gardner scale (spheroid/ellipsoid classifications and A grades considered predictive of a positive outcome respectively). Main results and the role of chance Most ICM were either ellipsoid (48%) and spheroid (37%) rather than crescent (6%) or irregular (8%, p < 0.001). ICM 3D classification was associated with implantation, X2(3, N = 119)= 11.10, p = 0.01 and live birth, X2(3, N = 119)= 10.78, p = 0.01. Blastocysts with ellipsoid (70%, 28/40) and spheroid (58%, 18/31) shaped ICM had the greatest chances of implantation, whilst those with crescent (40%, 6/15) and irregularly shaped ICM (33%, 11/33) had lower chances of implantation (p < 0.011). Similarly, blastocysts with ellipsoid (58%, 23/40) and spheroid (48%, 15/31) shaped ICM had the greatest chances of live birth, whilst those with crescent (33%, 5/15) and irregularly shaped ICM (21%, 7/33) had lower chances of live birth (p < 0.013). 3D assessment of ICM outperformed Gardner ICM grading with regards to prediction of implantation (accuracy 64.7% vs 49.6%, area-under-the-curve 0.642 vs 0.486, precision 64.8% vs 51.9%, sensitivity 73% vs 65.1%, specificity 55.4% vs 32.1%, F1 score 68.7% vs 57.8%, 3D vs Gardner) and live birth (accuracy 62.2% vs 50.4%, area-under-the-curve 0.641 vs 0.531, precision 53.5% vs 44.3%, sensitivity 76% vs 70%, specificity 52% vs 36.2%, F1 score 62.2% vs 54.3%, 3D vs Gardner). Limitations, reasons for caution Manually rendering the 3D images for each ICM was time consuming, making it impractical for clinical application. Further work is now focusing on automating this process using image recognition and augmented reality. Wider implications of the findings This is a novel ICM assessment that provides an improved perspective of the 3-dimensionality of the ICM which is more consistent and effective than traditional methods of embryo assessment, potentially contributing towards improving embryo selection efficacy and time to pregnancy. Trial registration number none
RESEARCH QUESTION:What can three-dimensional cell contact networks tell us about the developmental potential of cleavage-stage human embryos? DESIGN:This pilot study was a retrospective analysis of two Embryoscope imaging datasets from two clinics. An artificial intelligence system was used to reconstruct the three-dimensional structure of embryos from 11-plane focal stacks. Networks of cell contacts were extracted from the resulting embryo three-dimensional models and each embryo's mean contacts per cell was computed. Unpaired t-tests and receiver operating characteristic curve analysis were used to statistically analyse mean cell contact outcomes. Cell contact networks from different embryos were compared with identical embryos with similar cell arrangements. RESULTS:At t4, a higher mean number of contacts per cell was associated with greater rates of blastulation and blastocyst quality. No associations were found with biochemical pregnancy, live birth, miscarriage or ploidy. At t8, a higher mean number of contacts was associated with increased blastocyst quality, biochemical pregnancy and live birth. No associations were found with miscarriage or aneuploidy. Mean contacts at t4 weakly correlated with those at t8. Four-cell embryos fell into nine distinct cell arrangements; the five most common accounted for 97% of embryos. Eight-cell embryos, however, displayed a greater degree of variation with 59 distinct cell arrangements. CONCLUSIONS:Evidence is provided for the clinical relevance of cleavage-stage cell arrangement in the human preimplantation embryo beyond the four-cell stage, which may improve selection techniques for day-3 transfers. This pilot study provides a strong case for further investigation into spatial biomarkers and three-dimensional morphokinetics.
BACKGROUND AND AIM Inner cell mass (ICM) assessment forms a key part of most widely used blastocyst grading systems. Grading under these systems, however, is often qualitative, subjective, and subject to inter/intra-operator variability. To reduce subjectivity in ICM grading, it has been suggested that quantitative criteria are used. Previous studies have focused on the assessment of such parameters (e.g. ICM area) in 2-dimensions. However, such approaches fail to account for the various embryo orientations which lead to different measurements. Therefore, we aimed to design a novel quantitative 3-dimensional grading system to improve the reliability and consistency of grading, focussing on ICM reconstruction. METHODS This study retrospectively analysed embryoscope images of 153 randomly selected blastocysts from two clinics. Embryos were imaged at 7+ focal planes and in-focus sections of each ICM were annotated using ImageJ. Sections were combined into a 3D model using Blender, and sphericity (ψ) computed. Embryos were categorised as spheroid, ellipsoid, crescent and irregular based on their sphericity and number of planes of symmetry. Chi-squared tests were performed, comparing classification and clinical outcome. The predictive power (AUC, accuracy, precision) of the 3D assessment was compared to that of the Gardner scale. RESULTS Our study has found that blastocysts with ellipsoid (70%) and spheroid (58%) shaped ICM had the greatest chances of implantation, whilst those with crescent (40%) and irregularly shaped ICM (33%) had lower chances of implantation (p<0.01). Similarly, blastocysts with ellipsoid (58%) and spheroid (48%) shaped ICM had the greatest chances of live birth, whilst those with crescent (33%) and irregularly shaped ICM (21%) had lower chances of live birth (p<0.01). 3D assessment of ICM outperformed Gardner ICM grading with regards to prediction of implantation (area-under-the-curve 0.647 vs 0.496, accuracy 64.7% vs 49.6% and precision 0.648 vs 0.519, 3D vs Gardner) and live birth (accuracy 62.1% vs 50.4%, area-under-the-curve 0.641 vs 0.531, precision 0.535 vs 0.443). CONCLUSIONS In conclusion, our study shows that 3D ICM shape predicts IVF outcomes better than the Gardner system. Our findings also emphasize the importance of 3D structure and quantitative assessment in the embryo selection process, and offers the potential for integration into existing embryo grading systems.
Embryo selection is a critical step in the process of in-vitro fertilisation in which embryologists choose the most viable embryos for transfer into the uterus. In recent years, numerous works have used computer vision to perform embryo selection. However, many of these works have neglected the fact that the embryo is a 3D structure, instead opting to analyse embryo images captured at a single focal plane. In this paper we present a method for the 3D reconstruction of cleavage-stage human embryos. Through a user study, we validate that our reconstructions align with expert assessments. Furthermore, we demonstrate the utility of our approach by generating graph representations that capture biologically relevant features of the embryos. In pilot experiments, we train a graph neural network on these representations and show that it outperforms existing methods in predicting live birth from euploid embryo transfers. Our findings suggest that incorporating 3D reconstruction and graph-based analysis can improve automated embryo selection.
Abstract Study question Can a generalizable computer vision foundation model be trained using self-supervised learning on time-lapse embryo images to perform multiple clinically relevant downstream tasks? Summary answer We developed FEMI, a foundation model trained on eight million time-lapse images, to perform multiple clinical tasks, including blastocyst quality scoring, ploidy prediction, and segmentation. What is known already In vitro fertilization success critically depends on choosing viable embryos, a process hampered by current limited diagnostic tools, high costs, and ethical concerns. In recent years, various medical fields have increasingly explored foundation models using vision transformer architectures. These models, trained in a self-supervised manner on vast unlabeled image datasets, perform various clinically relevant tasks. Once trained on a massive unlabeled dataset, the encoder portion of the foundation model can be extracted and subsequently fine-tuned with a labeled dataset to perform various clinically relevant tasks like classification, regression, and segmentation. Study design, size, duration The foundation model was trained on 8 million Embryoscope® (E-SD) and Embryoscope+® (E+) images spanning clinics in the United States, Canada, and Europe. Datasets contained various additional information, such as ploidy status, blastocyst scores (numerical values from 3 to 14 based on Zhan et al. 2020), and segmentation masks (for trophectoderm [TE], inner cell mass [ICM], and zona pellucida [ZP]) that were used for training downstream tasks. Participants/materials, setting, methods A masked autoencoder architecture is trained on time-lapse images to create a foundation model (FEMI). The encoder from the autoencoder is extracted and fine-tuned on three image-based downstream tasks: ploidy prediction, blastocyst scoring, and embryo component segmentation. The performance of the fine-tuned foundation model is compared to VGG16 architectures, specifically trained for each downstream task. We evaluated performances using the area under the receiver-operating-characteristic (AUROC), mean absolute error (MAE), and mean intersection over union (mIoU). Main results and the role of chance The training dataset for the downstream tasks consisted of image data and labels, including PGT-A ploidy results (euploid or aneuploid), blastocyst scores (3-14), and segmentation masks. No clinical data like maternal age was used in order to assess FEMI’s learning capabilities from only time-lapse images against VGG16 model baselines. Both FEMI and baseline models used the same data splits for consistent comparisons. For ploidy prediction, models trained on a Weill Cornell Medicine (WCM) dataset were validated using WCM, Florida, and Spain data. FEMI achieved a 0.610 ± 0.004 AUROC, outperforming the baseline’s 0.590 ± 0.005. In blastocyst score prediction, FEMI attained a superior MAE of 0.099 ± 0.002 compared to the baseline’s 0.118 ± 0.001. For embryo segmentation (TE, ICM, ZP), a publicly available dataset from Simon Fraser University was split for training and validation. FEMI exceeded the baseline in mIoU for TE segmentation (0.738 ± 0.002 vs. 0.726 ± 0.002) and was comparable in ICM (0.832 ± 0.008 vs. 0.821 ± 0.011) and ZP segmentation (0.779 ± 0.002 vs. 0.778 ± 0.010). Limitations, reasons for caution The current version of FEMI is trained on a subset of our available datasets and can be improved through further training. Moreover, the downstream ploidy and blastocyst score models may be biased by errors in PGT results and the subjectivity of embryologists, respectively. Wider implications of the findings A generalizable foundational model for IVF time-lapse imaging aids clinicians in embryo selection by providing various clinical insights for better decision-making. FEMI, once fully trained, will be publicly accessible to the scientific community allowing for researchers to fine-tune FEMI on their own clinic-specific applications with their own image datasets. Trial registration number not applicable
(Abstracted from Nat Commun 2024;15(1):7756 Assisted reproductive technology, including in vitro fertilization, has become an important treatment for individuals who desire familial expansion but are unable to conceive. A crucial part of this process is the determination of embryo viability and the selection of the highest quality embryos for transfer into the uterus.
Abstract Study question Does MAGENTA, an AI assessment of oocyte quality, indicate ploidy potential of the developed blastocysts in the presence of confounding variables? Summary answer MAGENTA correlates to oocytes that develop into euploid embryos across diverse demographics, highlighting its robust ability to assess oocyte quality and genetic potential. What is known already MAGENTA is a non-invasive AI-image analysis tool that assesses images of mature oocytes and provides a score from 0-10, where higher scores indicate higher chances of developing into a blastocyst. MAGENTA has also displayed a correlation to blastocyst morphological quality. The ploidy status of blastocysts plays a significant role in cycle success and is often utilized for the selection of embryos to transfer. Despite the selection of embryos based on their genetic normality, many transfers are unsuccessful. With most chromosomal errors arising in the oocyte, assessment at the gamete level is critical to improving outcomes. Study design, size, duration A retrospective study using CRM Weill Cornell Medicine data included 3,342 mature oocytes that developed into blastocysts and underwent PGT-A in 2021. Blastocysts tested as mosaic (n = 753) or inconclusive (n = 13) were excluded. Images of the oocytes were obtained immediately post-ICSI using EmbryoScope time-lapse incubators (Vitrolife, Sweden). Blastocyst morphological grades were categorized by quality: highest (ICM+TE grade A), high (ICM+TE grade A/B), medium (ICM+TE grade B), or low quality (ICM or TE grade C/D). Participants/materials, setting, methods Data was obtained from 629 patients (females 26-45 years old) with an average of 5.3 blastocysts (range 1-20) with PGT-A results per patient. MAGENTA assessed each oocyte image and provided a score (0-10), with higher scores representing a greater potential to develop into a blastocyst embryo. In this dataset, 1827 (55%) blastocysts were tested as euploid and 1515 (45%) were tested as aneuploid. PGT-A was conducted for various specific and non-specific indications in these patients. Main results and the role of chance MAGENTA was positively associated with oocytes that developed into blastocysts of increasing quality, with significance between highest and high-quality (7.4 vs 7.0,p<0.05) and medium and low-quality (6.9 vs 6.1,p<0.01). Of the blastocysts, the MAGENTA score further displayed significant differences between oocytes that developed into euploid (7.0) compared to aneuploid blastocysts (6.6) (p < 0.001;Welch’s t-test). Although MAGENTA does not predict ploidy, a threshold of 8.8 MAGENTA score was determined to best distinguish between ploidy outcomes with 61% euploidy rate above this threshold compared to 51% below, a significantly different proportion (p < 0.001;Two-Sample Proportions z-test). Therefore, oocytes scored ≥8.8 display a 19.6% relative increase in chance of developing into a euploid blastocyst. To account for female age, subgroup analysis was conducted and found MAGENTA scores to be significantly different between oocytes that developed into euploid and aneuploid blastocysts for patients <35 years old (7.2 vs 6.6, p < 0.01;Welch’s t-test) and ≥35 years old (6.8 vs. 6.5, p < 0.05;Welch’s t-test). Blastocyst morphological grade was also investigated as a confounding factor. No significant differences existed between MAGENTA scores of euploid and aneuploid blastocysts among the low, high, or highest-quality. Significantly, differences were only found within medium quality blastocysts (7.0 vs 6.6,p<0.05), which had the greatest sample size (n = 1854). Limitations, reasons for caution MAGENTA was not trained to predict euploid embryo development from images of mature oocytes. Further data is required to discern the correlation of MAGENTA to chromosomal ploidy status, apart from blastocyst quality. Additionally, investigation into sperm DNA fragmentation as a potential contributor to ploidy determination is warranted. Wider implications of the findings MAGENTA can identify oocytes of better quality that not only have higher chances of blastocyst development and higher morphological grades, but also indicates oocytes that are more likely to become euploid embryos across patients of various PGT-A indications and a broad age demographic. Trial registration number not applicable
Background: In recent times, various algorithms have been developed to assist in the selection of embryos fortransfer based on artificial intelligence (AI). Nevertheless, the majority of AI models employed in this context werecharacterized by a lack of transparency. To address these concerns, we aim to design an interpretable tool to automatehuman embryo evaluation by combining artificial neural networks (ANNs) and genetic algorithms (GA).Materials and Methods: This retrospective cohort study included 223 human blastocyst time-lapse (TL) imagestaken at 110 hours post-injection. All the images were evaluated by five embryologists from different clinics in termsof blastocyst expansion (BE), quality of the inner cell mass (ICM), and trophectoderm (TE). The embryo databasewas used to develop an AI system (70% training, 15% validation, and 15% test) for automate blastocyst assessment.The entire set of images underwent a standardization process, followed by processing and segmentation using Matlabsoftware. The resulting quantified variables were utilized in AI techniques (ANN and GA). Finally, the accuracy andperformance of the automation tool was assessed with the area under the receiver operating characteristic (ROC)curve (AUC). Then, the level of agreement among embryologists and between embryologists and the AI system wascompared with Kappa Index.Results: The overall agreement among embryologists was low (Kappa: 0.4 for BE; and 0.3 for TE and ICM). The AItool achieved higher consistency (Kappa 0.7 for BE and ICM; and 0.4 for TE). The AI exhibited high accuracy in classifyingBE (test 81.5%), ICM (test 78.8%), and TE (test 78.3%) and better performance for BE (AUC 0.888-0.956)than for ICM (AUC 0.605-0.854) and TE (AUC 0.726-0.769) assessment.Conclusion: Our AI tool highlighted the superior consistency of AI compared to human operators in grading blastocystmorphology. This research represents an important step towards fully automating objective embryo evaluation.
Abstract Study question What can we learn from 3D reconstructions of cleavage-stage embryos derived from Hoffman modulation contrast (HMC) time-lapses? Summary answer A simple spatial biomarker extracted from 3D embryo reconstructions at the t4 and t8 stages is associated with blastulation, blastocyst quality, pregnancy and live birth. What is known already Several works have demonstrated significant associations between t4 cell arrangement and blastulation potential. However, no studies have investigated the impacts of cell arrangement beyond the t4 stage in a clinical setting owing to difficulties visualising the 3D structure of embryos in a safe, cost-effective manner. In the previous ESHRE meeting, He et al. presented a deep learning system for the 3D reconstruction of cleavage-stage embryos from HMC focal stacks recorded in standard timelapse incubators. In this work, we use the aforementioned system to understand the spatial networks present in cleavage-stage embryos and investigate their associations with clinical outcomes. Study design, size, duration The study was a retrospective analysis of two imaging datasets from two different clinics. The first dataset (DS1) consisted of 162 t4 embryos with information on blastulation and Gardner grade. The second dataset (DS2) consisted of 202 embryos at t4 and t8 with information on blastocyst grade, pregnancy, live birth and PGT-A. All data was captured at 11 focal planes on Embryoscope incubators between 2018 and 2020. Participants/materials, setting, methods The system proposed by He et al. was used to reconstruct the 3D structure of embryos from their focal stacks. Networks of cell contacts were extracted from the resulting embryo 3D models and each embryo’s mean contacts per cell was computed. Statistical analysis of average cell contacts with respect to outcomes was carried out using unpaired t-tests. Moreover, cell contact networks from different embryos were compared to identify embryos with similar cell arrangements. Main results and the role of chance At t4, a higher average number of contacts per cell was associated with greater rates of blastulation in DS1 (2.59 vs 2.36, blastulated vs non-blastulated, p = 0.029) and blastocyst quality in both DS1 (2.59 vs 2.37, good vs poor, p = 0.010) and DS2 (2.51 vs 2.35, good vs poor, p = 0.017) where a ‘good’ embryo is defined as having an embryologist-provided Gardner grade with EXP>2, ICM>C and TE>C. At t8, a higher average number of contacts was associated with increased blastocyst quality (3.36 vs 3.09, good vs poor, p = 0.017), pregnancy (3.32 vs 2.87, pregnant vs not pregnant, p = 0.003) and live birth (3.40 vs 2.90, live birth vs no live birth, p = 0.0003). No associations were found with miscarriage or aneuploidy. Moreover, average contacts at t4 were not correlated with those at t8 (r = 0.15, 95% CI [-0.043, 0.34]). While 4-cell embryos fell neatly into 9 distinct cell arrangements with the 5 most common (tetrahedral, pseudotetrahedral, planar, closed-Y and linear) accounting for 95% of embryos, 8-cell embryos displayed a great degree of variation with 59 distinct cell arrangements, the largest such group representing only 8 embryos. Limitations, reasons for caution The datasets used in this study were small. Moreover, DS2 was subject to survivorship bias as it only contained embryos with PGT-A results which necessarily entailed successful blastulation. Furthermore, the 3D reconstruction system was not applicable to all cleavage-stage embryos, especially those obscured by the well or having undergone compaction. Wider implications of the findings This work provides evidence for the clinical relevance of cleavage-stage cell arrangement in the human preimplantation embryo beyond the 4-cell stage, which may improve selection techniques for D3 transfers. Moreover, our work provides a strong case for further investigation into spatial biomarkers derived from 3D embryo reconstruction and 3D morphokinetics. Trial registration number N/A
Assessing fertilized human embryos is crucial for in vitro-fertilization (IVF), a task being revolutionized by artificial intelligence and deep learning. Existing models used for embryo quality assessment and chromosomal abnormality (ploidy) detection could be significantly improved by effectively utilizing time-lapse imaging to identify critical developmental time points for maximizing prediction accuracy. Addressing this, we developed and compared various embryo ploidy status prediction models across distinct embryo development stages. We present BELA (Blastocyst Evaluation Learning Algorithm), a state-of-the-art ploidy prediction model surpassing previous image- and video-based models, without necessitating subjective input from embryologists. BELA uses multitask learning to predict quality scores that are used downstream to predict ploidy status. By achieving an AUC of 0.76 for discriminating between euploidy and aneuploidy embryos on the Weill Cornell dataset, BELA matches the performance of models trained on embryologists' manual scores. While not a replacement for preimplantation genetic testing for aneuploidy (PGT-A), BELA exemplifies how such models can streamline the embryo evaluation process, reducing time and effort required by embryologists.