OBJECTIVE:To examine intrapatient variability in retrieved oocyte numbers across consecutive in vitro fertilization (IVF) ovarian stimulation (OS) cycles with an identical OS protocol. DESIGN:Cross-continental, multicenter retrospective cohort study. SUBJECTS:Patients undergoing OS for IVF (2014-2024) with ≥2 OS cycles within 6 months using the same OS protocol, gonadotropin type, and initial and daily dose; all underwent freeze-all-cycles and had ≥1 oocyte retrieved in each of the two consecutive cycles. For each patient, the earliest consecutive pair meeting criteria was analyzed. EXPOSURE:Oocyte yield in the first vs. consecutive OS cycle. MAIN OUTCOME MEASURES:Primary outcomes included: (i) average percentage change in oocyte yield between cycles (higher divided by lower oocyte yield); and (ii) 25th, 50th (median), and 75th percentiles of oocyte-yield percentage change, overall and by age groups (≤30, 31-35, 36-39, ≥40 years). Secondary outcomes included: (i) coefficient of determination (R2) between each cycle's oocyte count and the patient's average oocyte count across both cycles, representing the extent of variation explained by the patient's baseline profile; (ii) shifts between ovarian response categories, poor (1-3 oocytes), suboptimal (4-9), normal (10-14), and hyper-response (≥15); (iii) average and median percentage change in mature-oocyte yield. RESULTS:Overall, 801 cycle pairs met the inclusion criteria. Mean daily gonadotropin dosage was 361.5 ± 112.6 IU; with comparable demographic and cycle characteristics between cycles. Overall, the average percentage change in oocyte yield was 62.7%; the 25th, 50th (median), and 75th percentiles were 16.7%, 40%, and 80%, respectively. Fifty-percent of patients showed >33% difference in retrieved oocytes, and 381/801 (47.57%) shifted ovarian response categories, with 29/381 (7.61%) shifting across two categories. Median oocyte yield percentage change was 44.4% in women ≥40 vs. 33.3% in those ≤30 years. The coefficient of determination between each cycle and the average of the two cycles was 0.834, representing the optimal performance any prediction model could achieve when predicting oocyte yield given baseline characteristics alone. The average percentage change in mature oocyte yield was 74.5%, with a median of 50%. CONCLUSION:Cycle-to-cycle variations in retrieved oocyte yield exist despite the same cycle conditions, across all age groups, reflecting fluctuations in ovarian follicular readiness and response, challenging ovarian response categorization based on oocyte yield and stressing the importance of key performance indicators in IVF OS cycles.
RESEARCH QUESTION:How can outcomes in programmed frozen embryo transfers (FET) be improved, and what are the key areas where more research is needed? DESIGN:Using a Delphi-consensus framework, a scientific committee comprising five experts and a Scientific Coordinator formulated 17 literature-supported and expert opinion-supported statements, which were presented to an Extended Panel of international experts (five scientific committee members, excluding the Scientific Coordinator, and an additional 22 experts) who voted on their level of agreement or disagreement with each statement using a 5-point Likert-type scale (1 = Absolutely agree; 2 = More than agree; 3 = Agree; 4 = Disagree; 5 = Absolutely disagree). Consensus was reached if over 66% of participants agreed or disagreed. RESULTS:All statements exceeded the threshold for agreement after one round of voting. The statements covered aspects such as the decision-making process for choosing a programmed FET cycle; the route, dose and duration of oestrogen supplementation, and the benefit of monitoring serum oestrogen, LH and progesterone concentrations; the route and duration of progesterone treatments; the use of combination progesterone therapy; the duration and cessation of luteal phase support; and pregnancy outcomes from different FET protocols. CONCLUSIONS:The decision to undergo a programmed FET cycle should be made jointly between the patient and the physician and should be individualized according to the documented advantages (predictability and flexibility) and disadvantages (adverse obstetric and neonatal outcomes).
IMPORTANCE:While prospective registration is a key measure to prevent selective reporting, excluding unregistered or retrospectively registered trials would greatly reduce the statistical power of systematic reviews. OBJECTIVE:We aimed to explore the association between trial registration status and the reporting of large treatment effects on pregnancy or live birth. EVIDENCE REVIEW:This is a meta-epidemiologic study of infertility randomized controlled trials (RCTs) published between 2012 and 2023, identified through systematic searches of Embase, Medline, and CENTRAL. Eligible RCTs involved infertile couples and reported outcomes on biochemical, clinical, ongoing pregnancy, or live birth. Conference abstracts, secondary analyses of RCTs, or trials with unclear registration timing were excluded. The primary outcome was reporting of large treatment effects, defined as a statistically significant relative risk for pregnancy or live birth below 0.80 or above 1.25. The association between trial registration status and reporting of large treatment effects was analyzed using a Poisson regression with a log link, adjusted for potential confounders, with prospective registration as the reference group. FINDINGS:Among 1,369 infertility RCTs, 19.8% were prospectively registered, 29.6% were retrospectively registered, and 50.6% were unregistered. Compared with prospectively registered trials, there was no statistically significant difference in the likelihood of reporting large treatment effects among retrospectively registered trials (adjusted relative risk [aRR] = 0.77; 95% confidence interval [CI]: 0.49-1.20) or unregistered trials (aRR = 1.07; 95% CI, 0.71-1.66). These findings remained consistent in analyses restricted to trials published after 2018. In trials published before 2018, retrospective registration was associated with a significantly lower likelihood of reporting large treatment effects (aRR = 0.40; 95% CI, 0.21-0.75). CONCLUSION AND RELEVANCE:The reporting of large treatment effects has not been shown to be greater in unregistered or retrospectively registered infertility trials. Our findings do not support the blanket exclusion of all unregistered or retrospectively registered RCTs from systematic reviews.
Delayed parenthood is a defining feature of contemporary reproductive medicine, fundamentally reshaping reproductive planning and increasing reliance on medically assisted reproduction (MAR). For women of advanced reproductive age (ARA), this shift exposes a potential disconnect between conventional treatment pathways and the biological constraints of ovarian ageing and a rapidly narrowing reproductive window. This risks preventable delays, cumulative treatment burden, and suboptimal outcomes in this population. Addressing these challenges requires deliberate adaptation of current clinical strategies to minimize avoidable delays in assessment and treatment, while ensuring timely, optimized, and outcome-focused care. In this Perspectives article, which is based on expert consensus opinion, we propose a hypothesis-generating framework — the Faster-Fewer-and-Finished (3F) approach — in which expedited evaluation, tailored cycle sequencing, and a patient-tailored ‘finish’ criterion are applied to maximize time-to-first-live birth or to arrive at a shared decision on how or if to continue, with treatment decisions anchored by individual age and ovarian reserve. The 3F approach prioritizes achieving a live birth or ideal family size in the shortest possible time by advocating the ‘fewest’ number of ovarian stimulation cycles necessary. While more than one cycle may be unavoidable, the emphasis is on early strategic tailoring of treatment according to age category, efficiency of intervention, and minimizing biological and emotional attrition. By aligning treatment strategies with both the realities of ovarian ageing and the wishes and life priorities of ARA patients, the 3F approach integrates established principles of MAR into a time-critical, outcome-driven theoretical framework for infertility care.
What is the performance of a machine-learning) ML( model in predicting the number of oocytes retrieved during ovarian stimulation, compared to prediction by fertility specialists? The ML model consistently outperformed fertility specialists in predicting the number of oocytes in real-world IVF patients treated with r-hFSH-α originator. Fertility specialists often encounter challenges in accurately predicting oocyte retrieval outcomes before commencing the ovarian stimulation cycle, with frequent under- or overestimations. This variability stems from the subjective interpretation of patient-specific factors, such as age and ovarian reserve markers, during clinical decision-making. While existing research highlights the potential of AI tools to enhance clinical practice, further studies are needed to systematically compare the predictive performance of AI models with the expertise of fertility specialists. Twelve real-world IVF patient cases were provided in a web-based survey to 29 experienced fertility specialists from 29 clinics worldwide (Oct–Dec 2024). Patient profile data (age, height, weight, indication for ART treatment, AMH, AFC, freeze-all/fresh transfer) were provided. The r-hFSH-α starting dose was shared with the specialists (assuming no dose adjustment), who were requested to estimate the number of oocytes retrieved. Specialists completed a tutorial and quiz before the first case evaluation. A previously developed XGBoost (advanced tree-based that excels in speed and performance, particularly for structured data) machine learning model prediction of the number of oocytes retrieved according to patient profile and r-hFSH-α starting dose was compared to fertility specialists’ predictions. The predictions were evaluated against the ground truth number of oocytes retrieved. The accuracy of the machine learning model was then compared to specialist predictions. Performance was assessed using mean absolute error (MAE) and mean error (ME) to compare the predictions of the model and fertility specialists against the ground truth number of retrieved oocytes for each case. Physicians achieved an MAE of 3.77 oocytes, compared to an MAE of 1.94 oocytes achieved by the ML model, highlighting the model’s greater accuracy. Physicians achieved an ME of 1.97±4.43 (mean±standard deviation), showing a consistent tendency to underestimate the number of retrieved oocytes. In contrast, the ML model displayed a far more balanced ME of 0.16±2.64, demonstrating significantly lower bias and variability in predictions. To further compare performance, we ranked the physicians and the ML model for each case based on the accuracy of their predictions. A rank of 1 was assigned to the most accurate prediction (closest estimation to the ground-truth per case), while a rank of 30 was assigned to the least accurate. Across all cases, the ML model achieved the best average rank of 8.08, outperforming even the most accurate physicians. The three top-performing physicians had average ranks of 8.17, 10.79, and 10.96. The median physician rank was 15.62, highlighting that the ML model consistently delivered more accurate predictions than the majority of physicians. This survey involved a small cohort of reproductive specialists, assessing only 12 clinical cases. Additionally, the machine learning model’s predictions were based exclusively on r-hFSH-α originator-treated cycles. AI models may enhance prediction accuracy and consistency, enabling personalized ovarian stimulation protocols and supporting informed decision-making during patient consultation. Such tools can also provide AI-based second opinion in challenging cases and potentially improve consistency in r-hFSH-α dosing strategy among reproductive specialists. No
The paper aims to investigate the biological role of microRNAs secreted by preimplantation embryo into the blastocoel fluid and to detect a distinctive molecular signature for identifying embryos with the highest implantation potential. We carried on a multicenter retrospective study involving five European IVF centers. We collected 112 blastocoel fluid samples from embryos on day 5 post-fertilization, cultured individually, along with data on blastocyst grade and embryo transfer outcomes. Using a custom TLDA Array, we compared the expression levels of 89 miRNAs between 33 fluids from high-quality implanted embryos and 30 fluids from high-quality not-implanted embryos. Expression differences were assessed using SAM and t-test. Additionally, correlation and function enrichment analysis and network construction were conducted to identify the biological roles of deregulated microRNAs. We identified six up-regulated microRNAs in the blastocoel fluid from implanted embryos, significantly and positively correlated across all samples (r ≥ 0.7; P ≤ 0.05). They could take part in pluripotency circuits, regulating and being regulated by transcription factors associated with stemness, cell growth, and embryo development. The ROC curve analysis confirmed the potential of these miRNAs as implantation classifiers. The six miRNAs up-regulated in blastocoel fluid from implanted embryos may represent a functional molecular signature for evaluating blastocyst quality and identifying the most competent embryos. Their evaluation associated with non-invasive preimplantation genetic testing, integrating epigenetic and genomic analyses, could enhance implantation grade and allow for identification of the euploid embryo not able to implant.
STUDY QUESTION:What are the trial characteristics, geographic distribution, and selected methodological issues of randomized controlled trials (RCTs) in infertility published from 2012 to 2023? SUMMARY ANSWER:Of the 1425 infertility RCTs, over two-thirds focused on IVF, nearly two-fifths did not use pregnancy or live birth as the primary outcome, a third lacked a primary outcome, a half were unregistered, and just over half were conducted in China (22%), Iran (20%), or Egypt (10%). WHAT IS KNOWN ALREADY:RCTs are the main source of evidence on the effectiveness of interventions. Knowledge about RCTs in infertility from the recent past will help to pinpoint research gaps and prioritize the future research agenda. Here, we aim to present a descriptive analysis of trial characteristics, geographic distribution, and selected methodological issues in infertility trials published in the last decade. STUDY DESIGN SIZE DURATION:This is a systematic review. We systematically searched Embase, Medline, and Cochrane Central for RCTs in infertility from January 2012 to August 2023. RCTs involving subfertile women and women who reported pregnancy endpoints were eligible, while conference abstracts or secondary analyses were not. We did not limit our search based on the language of the articles. PARTICIPANTS/MATERIALS SETTING METHODS:The full articles were text-mined and manually extracted for the description of trials' characteristics (e.g. sample size, blinding method, types of intervention), the country where the patients were recruited, and methodological issues (trial registrations and specification of primary outcomes). We extracted funding statements from Dimensions, a literature database chosen for its comprehensive and robust metadata. Gross domestic product (GDP) data were obtained from the United Nations' official website. The accuracy of extracted data was validated in a random sample of 50 articles, and false positivity and false negativity were all at or below 8%. We used descriptive statistics, including frequencies and percentages to illustrate the overall and temporal trends. MAIN RESULTS AND THE ROLE OF CHANCE:Among 8757 records, we found 1425 eligible RCTs, with a median sample size of 140, and 33.3% had a sample size <100. Most (69.6%) of the trials focused on IVF, with the rest focusing on ovulation induction (12.4%), intrauterine insemination (10.6%), surgeries (4.8%), or other interventions (2.6%). Regarding the geographic distribution, China (n = 310), Iran (n = 284), and Egypt (n = 138) contributed to 51% of the RCTs, followed by Turkey (n = 82), India (n = 71), and the USA (n = 69); mainland Europe produced 343 trials. Ranked by publications of trials per trillion GDP, Greece had the most papers with 4.6, followed by Iraq at 3.9, and Iran at 2.5. Regarding trial registration, 47.8% of trials were unregistered, the proportion of studies that were unregistered halved from 70.0% in 2012 to 34.6% in 2022. Of all RCTs, 37.6% had primary outcomes unspecified; the proportion of trials specifying primary outcomes increased from 49.5% in 2012 to 61.4% in 2022. The proportion of trials which declared receiving no funding was 76.9%. LIMITATIONS REASONS FOR CAUTION:We primarily used text mining for data extraction. Despite optimizing the algorithm to identify all outcome definitions and manually curating the extracted data, there were inaccuracies in data extraction; however, the false positivity and false negativity of data extraction were all at or below 8%. Also, we focused on trials reporting pregnancy outcomes, as these are of primary interest to patients and carry significant implications on clinical practice. However, we acknowledge that early-stage trials with only upstream endpoints also play an important role and should be considered when evaluating the full spectrum of infertility trials. Finally, we only included published RCTs and hence, our results cannot be extrapolated to unpublished RCTs. WIDER IMPLICATIONS OF THE FINDINGS:The domination of RCTs on IVF calls for a reconsideration of other topics to be studied and a realignment of research priorities. The imbalanced geographic distribution of infertility trials raises questions about the generalizability of study results and equity in the distribution of healthcare resources. The prevalence of trials without registration or primary outcomes specified highlights the imperative to improve trial design and reporting quality. Encouragingly, the improving trial registrations suggest the enforcement of trial registrations from the journals is effective. STUDY FUNDING/COMPETING INTERESTS:B.W.M. is supported by an NHMRC Investigator grant (GNT1176437). W.T.L. is supported by an NHMRC Investigator grant (GTN2016729). W.L.L. reports receiving a PhD scholarship from the China Scholarship Council. Q.F. reports receiving a PhD scholarship from Merck. B.W.M. reports receiving consultancy fees, travel support, and research funding from Merck; consultancy fees from Organon and Norgine; and stock ownership in ObsEva. T.D.H and S.L. are employees of Merck. W.T.L., W.L.L., and J.C. report no conflicts of interest. REGISTRATION NUMBER:PROSPERO CRD42024498624.
OBJECTIVE:To provide a framework for conducting rigorous nonrandomized studies of interventions in fertility treatment research, addressing their role as complements to randomized controlled trials (RCTs) in evaluating treatment outcomes. DESIGN:Multidisciplinary expert consensus on best practices for nonrandomized studies of interventions, informed by advancements in novel methodologies, including causal inference. SUBJECTS:Patients undergoing assisted reproductive technologies (ARTs) procedures, such as ovarian stimulation, laboratory techniques, and embryo transfer. INTERVENTION:None. MAIN OUTCOME MEASURES:Guidance on methodological rigor, transparency, and relevance in nonrandomized studies of interventions study design and analysis. RESULTS:Randomized controlled trials are the gold standard for determining the efficacy and safety of fertility treatment/ART interventions but can face logistical, practical, and sometimes ethical challenges. Nonrandomized studies of interventions, when conducted with high methodological rigor, complement RCTs by offering insights into real-world clinical practices and diverse patient populations. Key limitations of nonrandomized studies of interventions include susceptibility to confounding and selection bias, which require meticulous study design and advanced analytical techniques to address. Recent innovations, such as target trial emulation studies, have enhanced the validity of causal inferences based on nonrandomized studies of interventions. This article outlines 7 recommendations to improve the credibility of nonrandomized studies of interventions in ART research: clearly define research questions with precise estimands; design nonrandomized studies of interventions as emulated trials; use directed acyclic graphs to clarify causal assumptions; preregister study protocols; separate data analysis from study planning; incorporate negative controls to detect biases; and use appropriate analytical methods to account for confounding and selection bias. CONCLUSION:Integrating evidence from RCTs and well-conducted nonrandomized studies of interventions enhances clinical decision making in fertility treatment research. By adhering to these recommendations, researchers can improve the quality, transparency, and impact of nonrandomized studies of interventions, ultimately fostering robust, evidence-based clinical practices in fertility treatment/ART.
OBJECTIVE:To address the reporting of minimal clinically important differences (MCIDs) in the published literature and investigate the utility of absolute and relative differences when defining this parameter. DESIGN:Expert opinion. SUBJECTS:Not applicable. EXPOSURE:Not applicable. MAIN OUTCOME MEASURES:It is essential to consider both statistical significance and clinical significance when interpreting study findings. Key to this is establishing the MCID that is beneficial for patients. In the context of assisted reproductive technology, several benchmarks are used to evaluate differences between treatments, including the mean number of retrieved oocytes, pregnancy rates, clinical pregnancies, live births, and cumulative live births. RESULTS:Determining the MCID for assisted reproductive technology procedures is a subjective process that can be influenced by several factors, depending on whether an absolute or relative difference is selected. Furthermore, various and often overlapping MCIDs have been used in different studies, meaning that the same difference between treatment and comparator can be interpreted as evidence of superiority in some studies and as evidence of noninferiority in others. To address these inconsistencies, we recommend that comparative studies should, by design, include a clearly defined and justified MCID threshold, with differences expressed as relative measures. CONCLUSION:We recommend that the MCID should be defined in advance and expressed as a relative measure, which can be interpreted across various groups of women with diverse prognoses. However, absolute measures should also be reported for completeness.
Are unregistered infertility randomized controlled trials (RCTs) more likely to report inflated findings than unregistered trials? Of 1,369 infertility RCTs, half were unregistered. Effect sizes did not differ between registered and non-registered RCTs, nor did registration timing play a role. There have been calls for a complete exclusion of retrospectively or not registered trials from systematic reviews, citing their heightened risks of reporting larger treatment effects. We aim to assess the prevalence of unregistered, retrospectively, and prospectively registered RCTs in infertility research, and examine the association between trial registration status and reported effect sizes on pregnancy or live birth outcomes. This systematic review included RCTs in infertility published from 1 January 2012 to 30 August 2023, identified through systematic searches of Embase, Medline, and CENTRAL. RCTs involving infertile women that reported either biochemical, clinical, ongoing pregnancy or live birth were eligible, while conference abstracts or secondary analyses of RCTs were not. Two authors independently screened articles. Trial registration status and contingency tables on pregnancy or live birth were manually extracted from the trial publications. For each article, we registered trial registration which is further stratified into prospective and retrospective categories. The outcome is the reporting of large effect sizes, defined as relative risks for pregnancy or live birth below 0.75 or above 1.25. The association between trial registration and large effect sizes was analyzed using a Poisson regression model adjusted for variables significant in univariate analysis. Among 1,369 infertility RCTs, 19.8% (n = 271) were prospectively registered, 29.6% (n = 405) were retrospectively registered and 50.6% (n = 693) were unregistered. In Quantile 1 journals, 48.7% (n = 132) of trials were prospectively registered, 29.1% (n = 121) were retrospectively registered, and 15.9% (n = 110) were unregistered. Between 2012 and 2023, prospectively registered trials surged from 3.1% (4/128) to 38.0% (46/121). The proportion of retrospectively registered RCTs remained stable during this timeframe, at around 20.0%, while the proportion of unregistered trials declined from 69.5% (89/128) to 39.7% (48/121). China contributed the most trials (n = 294), with 24.9% prospectively registered, 14.6% retrospectively registered and 60.5% unregistered, followed by Iran (n = 262), with 9.5% prospectively registered, 47.3% retrospectively registered and 43.2% unregistered. The European Union ranked third (n = 228), with 23.2% prospectively registered, 38.6% retrospectively registered and 38.2% unregistered. Adjusted multivariate analysis found no association between trial registration and large effect sizes (adjusted risk ratio: 1.02, 95% confidence interval: 0.82 -1.28), even when stratified by registration type (prospective: 1.05, 0.78 -1.40; retrospective: 1.00, 0.78 -1.29). We relied on trial publications to determine registration timing, which may not accurately reflect actual practices during trial execution. Due to the large volume of trials included, conducting trustworthiness assessments for all trials was infeasible. Our findings highlight pervasive noncompliance with trial registration in infertility. Trial registration alone should not be considered a guarantee of accurate reporting of treatment effects. Our results challenge the blanket exclusion of unregistered or retrospectively registered RCTs from systematic reviews and call for a more nuanced assessment of RCTs. No
BACKGROUND:The ovarian response to gonadotropin stimulation varies widely among women, and could impact the probability of live birth as well as treatment risks. Many studies have evaluated the impact of different gonadotropin starting doses, mainly based on predictive variables like ovarian reserve tests (ORT) including anti-Müllerian hormone (AMH), antral follicle count (AFC), and basal follicle-stimulating hormone (bFSH). A Cochrane systematic review revealed that individualizing the gonadotropin starting dose does not affect efficacy in terms of ongoing pregnancy/live birth rates, but may reduce treatment risks such as the development of ovarian hyperstimulation syndrome (OHSS). An individual patient data meta-analysis (IPD-MA) offers a unique opportunity to develop and validate a universal prediction model to help choose the optimal gonadotropin starting dose to minimize treatment risks without affecting efficacy. OBJECTIVE AND RATIONALE:The objective of this IPD-MA is to develop and validate a gonadotropin dose-selection model to guide the choice of a gonadotropin starting dose in IVF/ICSI, with the purpose of minimizing treatment risks without compromising live birth rates. SEARCH METHODS:Electronic databases including MEDLINE, EMBASE, and CRSO were searched to identify eligible studies. The last search was performed on 13 July 2022. Randomized controlled trials (RCTs) were included if they compared different doses of gonadotropins in women undergoing IVF/ICSI, presented at least one type of ORT, and reported on live birth or ongoing pregnancy. Authors of eligible studies were contacted to share their individual participant data (IPD). IPD and information within publications were used to determine the risk of bias. Generalized linear mixed multilevel models were applied for predictor selection and model development. OUTCOMES:A total of 14 RCTs with data of 3455 participants were included. After extensive modeling, women aged 39 years and over were excluded, which resulted in the definitive inclusion of 2907 women. The optimal prediction model for live birth included six predictors: age, gonadotropin starting dose, body mass index, AFC, IVF/ICSI, and AMH. This model had an area under the curve (AUC) of 0.557 (95% confidence interval (CI) from 0.536 to 0.577). The clinically feasible live birth model included age, starting dose, and AMH and had an AUC of 0.554 (95% CI from 0.530 to 0.578). Two models were selected as the optimal model for combined treatment risk, as their performance was equal. One included age, starting dose, AMH, and bFSH; the other also included gonadotropin-releasing hormone (GnRH) analog. The AUCs for both models were 0.769 (95% CI from 0.729 to 0.809). The clinically feasible model for combined treatment risk included age, starting dose, AMH, and GnRH analog, and had an AUC of 0.748 (95% CI from 0.709 to 0.787). WIDER IMPLICATIONS:The aim of this study was to create a model including patient characteristics whereby gonadotropin starting dose was predictive of both live birth and treatment risks. The model performed poorly on predicting live birth by modifying the FSH starting dose. On the contrary, predicting treatment risks in terms of OHSS occurrence and management by modifying the gonadotropin starting dose was adequate. This dose-selection model, consisting of easily obtainable patient characteristics, aids in the choice of the optimal gonadotropin starting dose for each individual patient to lower treatment risks and potentially reduce treatment costs.
STUDY QUESTION:How frequently do infertility trials report live birth and pregnancy, and how consistently were their definitions reported? SUMMARY ANSWER:One-third of 1425 infertility trials published in the last decade reported live birth, with one in eight reporting clinical pregnancy, ongoing pregnancy, and live birth concurrently; absent, ambiguous, or heterogeneous definitions were common. WHAT IS KNOWN ALREADY:Absent or inconsistent outcome definitions in randomized controlled trials (RCTs) limit their interpretation and complicate subsequent evidence synthesis. While reporting live birth in infertility trials has been a long-running recommendation, the extent to which this is adhered to, and the temporal trend of adherence, is unclear. Furthermore, it is unknown if outcome reporting in infertility trials is clear and consistent. STUDY DESIGN, SIZE, DURATION:We studied all RCTs in infertility published between 2012 and 2023. We aimed to assess (i) whether biochemical pregnancy, clinical pregnancy, ongoing pregnancy, and live birth were reported; the temporal trends in reporting these pregnancy outcomes, and compare the characteristics of trials reporting each type of outcome; (ii) whether and how these pregnancy outcomes were defined. PARTICIPANTS/MATERIALS, SETTING, METHODS:We systematically searched Embase, Medline, and CENTRAL for RCTs in infertility from January 2012 to August 2023. RCTs involving infertile women that reported either biochemical pregnancy, clinical pregnancy, ongoing pregnancy, or live birth were eligible. Secondary analyses, interim analyses, or conference abstracts were not eligible. Two authors independently screened articles. We extracted pregnancy definitions and trial characteristics primarily using text mining in R, a programming environment for data analysis, and supplemented by manual checking. The accuracy of extracted data was validated in a random sample of 50 articles, with sensitivity and specificity all at or above 90%. MAIN RESULTS AND THE ROLE OF CHANCE:We included 1425 infertility RCTs. Among these, 419 (29.4%) reported biochemical pregnancy. While 1359 (95.4%) RCTs reported clinical pregnancy, 404 (28.4%) reported ongoing pregnancy, and 484 (34.0%) reported live birth, only 174 (12.2%) reported all three outcomes. The proportion of trials reporting live birth increased from 23.1% in 2012 to 33.7% in 2023. Trials reporting up to biochemical pregnancy or clinical pregnancy were more likely to be unregistered, smaller, single-centered, and published in non-first quarter journals. Definitions for biochemical, clinical, ongoing pregnancy, and live birth were provided in 68.5% (287/419), 64.5% (876/1359), 70.5% (285/404), and 41.1% (199/484) of articles reporting on these outcomes. Among 876 clinical pregnancy definitions, 63.4% (n = 555) specified the pregnancy confirmation timing. Of the 220 definitions that reported gestational weeks (ranging from 4 to 16 weeks), the most common cut-off was 6 weeks, used in 48.2% (n = 106) of cases. For ongoing pregnancy definitions, 96.1% (n = 274) of the 285 definitions included gestational age in weeks (ranging from 6 to 32 weeks), with 12 weeks being the most common cut-off used in 49.1% (n = 140) of definitions. Among 199 live birth definitions, 62.3% (n = 124) used a gestational age threshold (ranging from 20 to 37 weeks), with 24 weeks being the most common cut-off, used in 28.6% (n = 57) of trials. LIMITATIONS, REASONS FOR CAUTION:Due to the vast data we needed to extract, we used text-mining supplemented by manual data extraction. While we optimized the text-mining algorithm attempting to identify all types of outcome definitions and manually curated all extracted definitions, definitions were missed in less than 10% of randomly checked studies, which is a limitation of this study. We only described definition patterns in published RCTs, and our results cannot be extrapolated to unpublished RCTs. WIDER IMPLICATIONS OF THE FINDINGS:Despite long-standing recommendations to report live birth in infertility trials, in the last decade only a third of RCTs did so. This highlights a disconnection between the advocated outcome and what researchers are reporting. We observed an encouraging trend that there has been a consistent rise in the proportion of trials reporting live birth. Furthermore, the significant lack and variability of pregnancy definitions underscore the imperative to increase the dissemination and uptake of standardized pregnancy outcomes. STUDY FUNDING/COMPETING INTEREST(S):No funding was received for the study. Q.F. reports receiving a PhD scholarship from Merck. B.W.M. is supported by an NHMRC Investigator grant (GNT1176437). B.W.M. reports consultancy, travel support, and research funding from Merck and consultancy for Organon and Norgine. B.W.M. holds stock from ObsEva. W.T.L. is supported by an NHMRC Investigator grant (GTN2016729). W.L.L. reports receiving a PhD scholarship from the China Scholarship Council. T.D.H and S.L. are employees of Merck Healthcare KGaA, Darmstadt, Germany. R.W. is supported by an NHMRC Investigator grant (GTN2009767). The other author has no conflict of interest to declare. REGISTRATION NUMBER:CRD42024498624.
Objective: To assess the adequate ovarian follicular development and oocyte recovery between ovarian potential (antral follicle count [AFC]) before the start of ovarian stimulation (OS) and oocyte quantity and quality at oocyte retrieval. A holistic overview of the current key performance indicators (KPIs) was applied to identify the complementary strengths and identify where the current repertoire can be expanded. Design: Expert opinion. Intervention: None. Main Outcome Measures: To formulate a proposal for a refined and expanded repertoire of KPIs for individualized OS for assisted reproductive technology. Results: The performance and outcomes of OS on ovarian follicular development can be evaluated through the application of defined KPIs. Current KPIs for OS are the ovarian sensitivity index, follicular output rate (FORT), oocyte retrieval rate, and follicle-to-oocyte index (FOI). Notably, there are no specific KPIs dedicated to the assessment of follicular development (i.e., recruitment, selection, growth, and dominance). In light of this, we recommend expanding the current KPIs for OS to include "early FORT"(accounting for the number of follicles measuring >= 10 to 11 mm on day 5/6 of OS relative to AFC) and "modified FORT"(the ratio between the number of follicles measuring >= 12 mm at the time of oocyte maturation triggering and AFC); the extension of oocyte retrieval rate to include two discrete categories at oocyte retrieval-follicles measuring >= 12 mm and >= 16 mm-to ensure that all responsive follicles are accounted for; and FOI to be measured at oocyte maturation triggering and oocyte retrieval ("advanced FOI"). Conclusion: Once validated and adopted in clinical practice, we envisage that the proposed expanded KPIs measuring the effect of OS on follicular development (recruitment, selection, growth, and dominance) will increase the understanding of the relationship between ovarian reserve, measured by AFC, and oocyte quantity and quality at oocyte retrieval. This understanding will enable physicians to better evaluate the direct effect of different gonadotropins and doses on ovarian response, leading to a more personalized approach to OS in the context of assisted reproductive technology treatment. (Fertil Steril (R) 2025;123:653-64. (c) 2024 by American Society for Reproductive Medicine.)
Does enhanced follicular growth and health during dynamic culture of human ovarian cortical tissue (HOCT) in perfusion bioreactor (PB) correlate with improved stromal tissue quality? Dynamic PB culture preserves stromal cell viability and enhances extracellular matrix (ECM) rearrangement improving follicle growth compared with static culture Despite the unsatisfactory results achieved in large mammals over the past 25 years, in vitro folliculogenesis remains an ultimate goal for preserving female fertility. Recently, dynamic PB culture has yielded a higher number of high-quality secondary follicles compared to static culture. However, the health of ovarian stroma and its involvement in early folliculogenesis has been largely overlooked. The ovarian stromal cells and the ECM environment play a crucial role in follicle activation and growth, as well as in the differentiation of theca cells. Recent research demonstrated reciprocal paracrine crosstalk between follicles and the surrounding stroma, highlighting their mutual interaction HOCT strips (1x1x0.5mm) were cultured 7 days in dynamic PB and conventional dishes (CD). The viability of the ovarian stromal cells and follicles was assessed under confocal microscopy through live-dead far-red and propidium iodide labeling; collagen thickness and packing density degree were assessed on PicroSirius red (PSR) stained sections under polarized light. Follicle stages were evaluated through histology Biopsies from 6 patients (mean age of 34.4) undergoing surgery for benign diseases. Stromal cells viability in cultured tissue was determined at different depths (100, 150, 200μm) from the cortex. Neosynthesis/remodeling of collagen was analyzed in cortical areas surrounding the follicles. In PSR stained sections, the degree of collagen maturation was assessed as follows: green, thin fibers (neosynthesized); yellow, medium-sized fibers (low assembly); orange, medium-sized fibers (medium assembly); and red, mature thick fibers (high assembly) Analysis of stromal cell viability of cultured tissues showed a significant beneficial effect of dynamic culture at all different depths analyzed compared to static culture (100μm: D0 95.0, PB 89.3, CD 67.5%; 150μm: D0 95.2, PB 90.8, CD 63.4%; 200μm: D0 89.6, PB 90.4, CD 68.5%; P < 0.001). Analysis of PSR-stained samples revealed a significantly higher production of green fibers (PB 2.50 vs CD 1.77, P < 0.05) and lower production of orange and red fibers (orange, PB 13.65 vs CD 17.10; red, PB 1.61 vs CD 2.66; P < 0.01) indicating greater collagen neosynthesis and remodeling in PB culture compared to CD. Follicular analyses showed an improved follicular viability (PB 80.9 vs CD 41.7%, P < 0.01) and a greater number of secondary follicles (PB 40.6 vs CD 21.6%, P < 0.01), highlighting the benefits of dynamic culture HOCT was donated by patients with benign diseases and had a reduced number of follicles. Therefore, the findings may not fully reflect a physiological condition The maintenance of physiological conditions through dynamic culture in PB enhances follicle growth and health as well as preserves the ovarian stromal component, which it’s well-known to play a fundamental role in folliculogenesis No
What is the extent of intra-patient variability in the number of oocytes retrieved across consecutive ovarian stimulation cycles, and how do the responder groups change? 50% of patients experienced over 33% difference in the number of oocytes retrieved in two consecutive cycles. 49.3% of patients shifted in ovarian response categories. Inherent biological fluctuations leading to inter-cycle variability in the ovarian response is well-known, but its extent is not well defined. The reason for this variation includes both biological factors, mainly variations in the cohort of antral follicles available for recruitment by ovarian stimulation (OS) with gonadotropins and variability of ovarian follicular sensitivity to ovarian stimulation. Additionally, procedural factors such as OS length, oocyte maturation trigger timing and type (human chorionic gonadotropin [hCG] or GnRH agonist or dual/double trigger), and inter-cycle differences in the oocyte retrieval technique can also contribute to differences in the number of retrieved oocytes. A retrospective analysis included 802 patients who underwent two or more ART OS cycles within a six-month period. For each patient, the first two consecutive cycles meeting the following criteria were selected: the same OS protocol, same gonadotropin type, same gonadotropin starting and daily dosage (treatment duration was unrestricted due to no observed dependence with retrieved oocyte), and freeze-all cycles only. Data were collected from the United States, Europe, and Asia. Percentage change was calculated by dividing the higher number by the lower number of retrieved oocytes in the patient’s two consecutive cycles. Variability was assessed based on the average percentage change and the 25th, 50th (median), 75th percentiles, for three age groups: ≤35, 36-39, ≥40 years. Further analysis was conducted on the changes in patient ovarian response – categorized as follows: Retrieved OocytePatientsPoor1-392Suboptimal4-9364Normal10-15248High≥1698 The results of the variability assessment for each age group are presented below: Oocytes Average Change[25th, 50th, 75th percentiles]Age ≤ 35:54.5%[14.3%, 33.3%, 66.7%]Age 36-39:57.2%[14.8%, 33.3%, 75.0%]Age ≥ 40:69.4%[18.8%, 44.4%, 80.0%] This shows that 50% of patients up to age 39 will experience over 33% difference in the number of oocytes retrieved in consecutive cycles. For patients aged 40 and older, 50% will have over 44% difference. The percentage change in the patient ovarian response category between the first and second OS cycle is as follows (the number of cycles is presented in parentheses): 1st ->PoorSuboptimalNormalHighPoor ->44.6% (41)50.0% (46)5.4% (5)0.0% (0)Suboptimal ->14.6% (53)55.8% (203)25.8% (94)3.8% (14)Normal ->2.4% (6)30.6% (76)43.5% (108)23.4% (58)High ->0.0% (0)10.2% (10)33.7% (33)56.1% (55) In total, 395 (49.3%) patients changed their ovarian response category between consecutive cycles from which 25 (6.3%) shifted two categories (from poor to normal, suboptimal to high, or vice versa) away. None shifted from poor to high, or vice versa. The main limitation of this study is its retrospective design, which may introduce biases due to unaccounted factors. Additionally, patient-specific variables, such as lifestyle factors or ovarian reserve markers, as well as cycle-specific factors including OS length and related total gonadotropin dose or oocyte maturation trigger timing, were not considered. Cycle-to-cycle variations in retrieved oocytes persist despite the same cycle conditions, reflecting fluctuations in ovarian follicular availability and response to OS, challenging ovarian response categorization based on oocyte yield. With the emergence of AI tools designed to optimize outcomes, this variability can serve benchmark for assessing the best achievable performance. No
IMPORTANCE:Live births are the "gold standard" for assessing infertility treatments but are less frequently reported than clinical pregnancy. Identifying approaches to interpret randomized controlled trials (RCTs) without live birth data is essential. OBJECTIVE:To evaluate the correlation between treatment effects and the consistency of statistical significance in conclusions drawn from live birth versus early-stage pregnancy endpoints (an umbrella term defined as biochemical, clinical, or ongoing pregnancy) in RCTs reporting both outcomes. DATA SOURCES:We systematically searched EMBASE, MEDLINE, and CENTRAL for RCTs in infertility from January 1, 2012 to August 30, 2023. STUDY SELECTION AND SYNTHESIS:Randomized controlled trials involving subfertile women reporting contingency tables for live birth and at least one early-stage pregnancy endpoint were eligible. Contingency tables on pregnancy or live births were manually extracted from the trial publications. We calculated Spearman's rho for treatment effects between early-stage pregnancy endpoints and live birth and compared their statistical significance using Chi-square tests. The above analyses were then conducted in prespecified subgroups. MAIN OUTCOMES:The correlation between treatment effects and the consistency of statistical significance in conclusions drawn from live birth versus early-stage pregnancy endpoints. RESULTS:Among 8,757 records, 1,281 infertility RCTs were eligible. The Spearman's rho for treatment effects was 0.78 (95% confidence interval [CI]: 0.69-0.86, P<.001; 169 RCTs) between biochemical pregnancy and live birth, 0.87 (95% CI: 0.83-0.90, P<.01; 429 RCTs) between clinical pregnancy and live birth, and 0.96 (95% CI: 0.92-0.98, P<.01; 138 RCTs) between ongoing pregnancy and live birth. Statistical significance between early-stage pregnancy endpoints and live birth was consistent in above 88% of trials. The correlation of treatment effects between clinical or ongoing pregnancy and live birth remained robust in subgroups of infertility treatments, including ovarian stimulation for medically assisted reproduction treatment. CONCLUSION AND RELEVANCE:A strong correlation was observed between clinical or ongoing pregnancy and live birth, supporting the rationale for using clinical or ongoing pregnancy data to assess treatment effectiveness when live birth data are limited or unavailable. However, the risk of intervention-related pregnancy loss must be considered during the assessment.
BackgroundParenthood is a key life goal for many, but infertility affects about 1 in 6 globally. While fertility treatments offer solutions, their high costs limit access. Many health systems provide public funding, yet budget constraints prevent fully funded access, often leaving patients with significant out-of-pocket costs. Policy makers face the challenge of prioritizing individuals for publicly funded treatments, but how to do this remains unclear and underresearched. Worldwide, funding policies vary widely, often adopting controversial access criteria.MethodsWe investigated Belgian population preferences for prioritizing in vitro fertilization (IVF) funding through a discrete-choice experiment with a representative sample of 3,000 Belgians. Attributes included maternal and partner age, infertility cause, civil status, prior biological children, and treatment cost. Using a Bayesian D-optimal design and panel mixed logit model, we assessed criteria relevance. The resulting multiattribute utility function created a priority ranking of couples, which we compared to the ranking under the current Belgian policy, which focuses only on maternal age (<43 y).ResultsAnalysis of 29,670 prioritization choices identified maternal age, infertility cause, and prior biological children as key criteria. Maternal age of 35 y was prioritized highest, age 25 y as high as 40 y, followed by declining priority until 55 y. Biomedical malfunctions were prioritized over same-sex relationships or unhealthy lifestyles, with the latter prioritized lowest. Having no prior biological children was prioritized categorically higher than having 1, 2, or 3 children, all prioritized equally. Preferences were homogeneous across sociodemographic groups.ConclusionsHow to set IVF funding priorities remains a matter of debate. Our study shows that the Belgian population considers multiple criteria beyond maternal age to prioritize couples, calling for further discussion on ethical justifiability and access implications.HighlightsParenthood is a key life goal to many, but about 1 in 6 are affected by infertility. However, in most countries, public funding for fertility treatment is not provided to everyone who could benefit, and hard choices are inevitable.This study used a discrete-choice experiment in a representative sample of the Belgian population to investigate which criteria should be used for prioritization.Results indicated that maternal age, cause of infertility, and the number of prior biological children were the most significant factors in determining public support for IVF funding. Partner age, civil status of the couple, and cost of IVF treatment were not important.People use multiple criteria to set IVF funding priorities, beyond maternal age (the only criterion used in the current Belgian funding policy). Future research should explore the ethical justifiability and practical implications of using cause of infertility and number of prior children as additional criteria.
Abstract Background Currently, there is no consensus on the optimal management of women with low prognosis in ART. In this Delphi consensus, a panel of international experts provided real-world clinical perspectives on a series of literature-supported consensus statements regarding the overall relevance of the POSEIDON criteria for women with low prognosis in ART. Methods Using a Delphi-consensus framework, twelve experts plus two Scientific Coordinators discussed and amended statements and supporting references proposed by the Scientific Coordinators (Round 1). Statements were distributed via an online survey to an extended panel of 53 experts, of whom 36 who voted anonymously on their level of agreement or disagreement with each statement using a six-point Likert-type scale (1 = Absolutely agree; 2 = More than agree; 3 = Agree; 4 = Disagree; 5 = More than disagree; 6 = Absolutely disagree) (Round 2). Consensus was reached if > 66% of participants agreed or disagreed. Results The extended panel voted on seventeen statements and subcategorized them according to relevance. All but one statement reached consensus during the first round; the remaining statement reached consensus after rewording. Statements were categorized according to impact, low-prognosis validation, outcomes and patient management. The POSEIDON criteria are timely and clinically sound. The preferred success measure is cumulative live birth and key management strategies include the use of recombinant FSH preparations, supplementation with r-hLH, dose increases and oocyte/embryo accumulation through vitrification. Tools such as the ART Calculator and Follicle-to-Oocyte Index may be considered. Validation data from large, prospective studies in each POSEIDON group are now needed to corroborate existing retrospective data. Conclusions This Delphi consensus provides an overview of expert opinion on the clinical implications of the POSEIDON criteria for women with low prognosis to ovarian stimulation.
Female fertility depends on the ovarian reserve of follicles, which is determined at birth. Primordial follicle development and oocyte maturation are regulated by multiple factors and pathways and classified into gonadotropin-independent and gonadotropin-dependent phases, according to the response to gonadotropins. Folliculogenesis has always been considered to be gonadotropin-dependent only from the antral stage, but evidence from the literature highlights the role of follicle-stimulating hormone (FSH) and luteinizing hormone (LH) during early folliculogenesis with a potential role in the progression of the pool of primordial follicles. Hormonal and molecular pathway alterations during the very earliest stages of folliculogenesis may be the root cause of anovulation in polycystic ovary syndrome (PCOS) and in PCOS-like phenotypes related to antiepileptic treatment. Excessive induction of primordial follicle activation can also lead to premature ovarian insufficiency (POI), a condition characterized by menopause in women before 40 years of age. Future treatments aiming to suppress initial recruitment or prevent the growth of resting follicles could help in prolonging female fertility, especially in women with PCOS or POI. This review will briefly introduce the impact of gonadotropins on early folliculogenesis. We will discuss the influence of LH on ovarian reserve and its potential role in PCOS and POI infertility.