Background:Young children with medulloblastoma (MB) are especially susceptible for tumor and treatment-related brain damage. Craniospinal irradiation (CSI) has been identified as a main risk factor for cognitive and motor dysfunction 2-5 years after diagnosis. Longitudinal follow-up data on these children beyond 10 years after diagnosis are scarce. Methods:To characterize their long-term neurocognitive outcome, we analyzed a cohort of 24 MB patients below the age of 3 years at diagnosis treated within the HIT-SKK'87 (systemic chemotherapy and deferred CSI) and HIT-SKK'92 (systemic chemotherapy and intraventricular methotrexate (MTXi.vt.), no radiotherapy) trials after 5 and 15 years. A comprehensive test battery was used to measure cognitive operations, executive functions with selective attention, and psychomotor abilities. Results:Overall, MB survivors showed subnormal test results independent of treatment, especially in the domain of motor function. CSI had an additional detrimental effect on general and fluid intelligence and simultaneous processing, whereas MTXi.vt.-recipients never scored worse than irradiated patients. Upon longitudinal follow-up, tests assessing sequential processing and general IQ showed stable or even improved results. In contrast, fluid intelligence, visual-motor integration, and most pronounced tapping speed further declined at 15 years in both the treatment groups and reached the disability zone. Conclusion:Impairment of motor and cognitive skills represents a major handicap for MB survivors. Long-term development of fluid intelligence and visuomotor integration is hindered in all children. CSI further aggravated these dysfunctions. Visual or auditory domains and memory performance are less affected. Long-term neurocognitive support is warranted in young pediatric survivors of MB.
In oncological clinical trials, overall survival (OS) is the gold-standard endpoint, but long follow-up and treatment switching can delay or dilute detectable effects. Progression-free survival (PFS) often provides earlier evidence and is therefore frequently used together with OS as multiple primary endpoints. Since in certain scenarios trial success may be defined if one of the two hypotheses involved can be rejected, a correction for multiple testing may be deemed necessary. Because PFS and OS are generally highly dependent, their test statistics are typically correlated. Ignoring this dependency (e.g. via a simple Bonferroni correction) is not power optimal. We develop a group-sequential testing procedure for the multiple primary endpoints PFS and OS that fully exhausts the family-wise error rate (FWER) by exploiting their dependence. Specifically, we characterize the joint asymptotic distribution of log-rank statistics across endpoints and multiple event-driven analysis cutoffs. Furthermore, we show that we can consistently estimate the covariance structure. Embedding these results in a closed testing procedure, we can recalculate critical values of the test statistics in order to spend the available type I error optimally. An important extension to the current literature is that we allow for both interim and final analysis to be event-driven. Simulations based on illness-death multi-state models empirically confirm FWER control for moderate to large sample sizes. Compared with a simple Bonferroni correction, the proposed methods recover roughly two thirds of the power loss for OS, increase disjunctive and conjunctive power, and enable meaningful early stopping. In planning, these gains translate into about 5
Classic adaptive designs for time-to-event trials are based on the log-rank statistic and its increments. Thereby, only information from the time-to-event endpoint on which the selected log-rank statistic is based may be used for data-dependent design modifications in interim analyses. Further information (e.g. surrogate parameters) may not be used. As pointed out in a letter by P. Bauer and M. Posch in 2004, adaptive tests on overall survival (OS) based on the log-rank statistic do in general not control the significance level if interim information on progression-free survival (PFS) is used for sample size adjustments, because progression is associated with increased risk of death. In contrast, in adaptive designs for time-to-event trials, which are constructed according to the principle of patient-wise separation, all trial data observed in interim analyses may be used for design modifications without compromizing type one error rate control. But by design, this comes at the price of incomplete use of the primary endpoint data in the final test decision or worst-case considerations which lead to a loss of power. Thus, the patient-wise separation approach cannot be regarded as a general solution to the problem described by Bauer and Posch. We address this problem within the framework of a comprehensive independent increments approach. We develop adaptive tests on OS in which sample size adjustments may be based on the observed interim data of both OS and PFS, while avoiding the problems of the patient-wise separation approach. We provide this methodology for both single-arm trials, in which a new therapy is compared with a pre-specified deterministic reference, and randomized trials, in which a new therapy is compared with a concurrent control group. The underlying assumption is that the joint distribution of OS and PFS is induced by a Markovian illness-death model.
Single-arm studies in the early development phases of new treatments are not uncommon in the context of rare diseases or in paediatrics. If an assessment of efficacy is to be made at the end of such a study, the observed endpoints can be compared with reference values that can be derived from historical data. For a time-to-event endpoint, a statistical comparison with a reference curve can be made using the one-sample log-rank test. In order to ensure the interpretability of the results of this test, the role of the reference curve is crucial. This quantity is often estimated from a historical control group using a parametric procedure. Hence, it should be noted that it is subject to estimation uncertainty. However, this aspect is not taken into account in the one-sample log-rank test statistic. We analyse this estimation uncertainty for the common situation that the reference curve is estimated parametrically using the maximum likelihood method, and indicate how the variance estimation of the one-sample log-rank test can be adapted in order to take this variability into account. The resulting test procedures are illustrated using a data example and analysed in more detail using simulations.
AbstractIntroductionNeurosurgery is considered the mainstay of treatment for pediatric low‐grade glioma (LGG); the extent of resection determines subsequent stratification in current treatment protocols. Yet, surgical radicality must be balanced against the risks of complications that may affect long‐term quality of life. We investigated whether this consideration impacted surgical resection patterns over time for patients of the German LGG studies.Patients and MethodsFour thousand two hundred and seventy pediatric patients from three successive LGG studies (median age at diagnosis 7.6 years, neurofibromatosis (NF1) 14.7%) were grouped into 5 consecutive time intervals (TI1‐5) for date of diagnosis and analyzed for timing and extent of first surgery with respect to tumor site, histology, NF1‐status, sex, and age.ResultsThe fraction of radiological LGG diagnoses increased over time (TI1 12.6%; TI5 21.7%), while the extent of the first neurosurgical intervention (3440/4270) showed a reduced fraction of complete/subtotal and an increase of partial resections from TI1 to TI5. Binary logistic regression analysis for the first intervention within the first year following diagnosis confirmed the temporal trends (p < 0.001) and the link with tumor site for each extent of resection (p < 0.001). Higher age is related to more complete resections in the cerebellum and cerebral hemispheres.ConclusionsThe declining extent of surgical resections over time was unrelated to patient characteristics. It paralleled the evolution of comprehensive treatment algorithms; thus, it may reflect alignment of surgical practice to recommendations in respect to age, tumor site, and NF1‐status integrated as such into current treatment guidelines. Further investigations are needed to understand how planning, performance, or tumor characteristics impact achieving surgical goals.
The analysis of multiple time-to-event outcomes in a randomised controlled clinical trial can be accomplished with exisiting methods. However, depending on the characteristics of the disease under investigation and the circumstances in which the study is planned, it may be of interest to conduct interim analyses and adapt the study design if necessary. Due to the expected dependency of the endpoints, the full available information on the involved endpoints may not be used for this purpose. We suggest a solution to this problem by embedding the endpoints in a multi-state model. If this model is Markovian, it is possible to take the disease history of the patients into account and allow for data-dependent design adaptiations. To this end, we introduce a flexible test procedure for a variety of applications, but are particularly concerned with the simultaneous consideration of progression-free survival (PFS) and overall survival (OS). This setting is of key interest in oncological trials. We conduct simulation studies to determine the properties for small sample sizes and demonstrate an application based on data from the NB2004-HR study.
TPS10067 Background: Pediatric low-grade gliomas (pLGGs) are the most common CNS tumors of childhood. Genomic alterations of BRAF are oncogenic drivers in almost all pLGGs. Approximately 50%‒60% of pLGGs harbor a KIAA1549-BRAF fusion and 5%‒15% a BRAF V600E mutation. Tovorafenib is an investigational, oral, selective, CNS-penetrant, small molecule, type II pan-RAF inhibitor. The registrational, phase 2 FIREFLY-1 (NCT04775485) study of tovorafenib in pediatric patients with recurrent/progressive LGG is ongoing and interim analysis has shown encouraging anticancer activity. Methods: LOGGIC/FIREFLY-2 (NCT05566795) is a registrational, 2-arm, randomized, multicenter, global (~100 sites across Australia, Canada, Europe, New Zealand, Singapore, South Korea, and USA), phase 3 trial being conducted in collaboration with the SIOPe Brain Tumor Group LOGGIC Consortium. The study is evaluating the efficacy, safety, and tolerability of tovorafenib vs. standard of care (SoC) chemotherapy in patients < 25 years old with pLGG harboring an activating RAF-alteration and requiring first-line systemic therapy. Approximately 400 patients will be randomized 1:1 to receive oral tovorafenib, 420 mg/m 2 (≤600 mg) once weekly (tablet or liquid suspension), or an investigator’s choice of SoC chemotherapy: COG-V/C regimen (60 weeks), SIOPe-LGG-V/C regimen (81 weeks), or single-agent vinblastine (70 weeks). Tovorafenib will be continued until the occurrence of radiographic progression (based on Response Assessment in Neuro-Oncology [RANO] criteria as determined by the investigator and confirmed by independent review) or unacceptable toxicity; patients with radiographic progression may be allowed to continue tovorafenib if, in the opinion of the treating investigator, they are deriving clinical benefit from study treatment. Patients who progress in the SoC arm during or after completion of chemotherapy are eligible to cross-over to receive tovorafenib. The primary endpoint is the objective response rate based on RANO criteria, as determined by independent review. Key secondary endpoints are progression-free survival and duration of response per RANO criteria by independent review and overall survival. Other secondary endpoints include efficacy assessments per Response Assessment in Pediatric Neuro-Oncology (RAPNO) criteria, changes in neurological and visual function, and safety and tolerability. Exploratory endpoints include efficacy assessments per investigator, tumor volume, adaptive behavior and quality of life. Prognostic and predictive molecular biomarkers, including senescence profiles for treatment outcome, response prediction, and treatment resistance, will be explored in parallel studies. Clinical trial information: NCT05566795 .
The one-sample log-rank test is the preferred method for analysing the outcome of single-arm survival trials. It compares the survival distribution of patients with a prefixed reference survival curve that usually represents the expected outcome under standard of care. However, classical one-sample log-rank tests assume that the reference curve is known, ignoring that it is frequently estimated from historical data and therefore susceptible to sampling error. Neglecting the variability of the reference curve can lead to an inflated type I error rate, as shown in a previous paper. Here, we propose a new survival test that allows to account for the sampling error of the reference curve without knowledge of the full underlying historical survival time data. Our new test allows to perform a valid historical comparison of patient survival times when only a historical survival curve rather than the full historic data is available. It thus applies in settings where the two-sample log-rank test is not applicable as method of choice due to non-availability of historic individual patient survival time data. We develop sample size calculation formulas, give an example application and study the performance of the new test in a simulation study.
Supplementary Data 1. This supplementary data comprises a step-by-step description on how our gene-expression based classification models were generated and validated
Supplementary Data 2. This supplementary data comprises the r-algorithm scripts required to perform the both model selection and validation as described in the step-by-step protocol
Legends to Supplementaries. This documents comprises the legends to the supplementary figures and tables
Time-to-event endpoints show an increasing popularity in phase II cancer trials. The standard statistical tool for such one-armed survival trials is the one-sample log-rank test. Its distributional properties are commonly derived in the large sample limit. It is however known from the literature, that the asymptotical approximations suffer when sample size is small. There have already been several attempts to address this problem. While some approaches do not allow easy power and sample size calculations, others lack a clear theoretical motivation and require further considerations. The problem itself can partly be attributed to the dependence of the compensated counting process and its variance estimator. For this purpose, we suggest a variance estimator which is uncorrelated to the compensated counting process. Moreover, this and other present approaches to variance estimation are covered as special cases by our general framework. For practical application, we provide sample size and power calculations for any approach fitting into this framework. Finally, we use simulations and real world data to study the empirical type I error and power performance of our methodology as compared to standard approaches.
Supplementary Table 3. This table summarizes the results of the Kaplan-Meier estimates for both EFS and OS of clinically relevant subgroups of neuroblastoma patients according to classification by the four remaining classifiers SVM_th22, SVM_th24, SVM_th26 and SVM_th44.
Supplementary Table 1. Clinical co-variates for the 709 patients who participated in the study. This supplementary table comprises detailed information on clinical co-variates for all 709 neuroblastoma patients who participated in the study.
Supplementary Figure 1. Highlights the EFS and OS for the cohort of neuroblastoma patients with MYCN-amplified disease (n=114, Fig 1a) and for th subcohort of patients >18 months of age with stage 4, MYCN non-amplified disease (n=102, fig. 1b). F, favorable; UF, unfavorable.
Supplementary Table 2. External performance validation of the top five classifiers. This supplementary table comprises the external performance metrics of the top five classifiers. Indicated are the values for classification accuracy, sensitivity , specificity and Matthew's correlation coefficient (MCC) in the prediction of 325 patients of the test set who fulfilled the criteria for classifier training (Favorable, n=187; Unfavorable, n=138).
Supplementary Table 4. Transcripts that contribute to the SVM_th10 classifier. This supplementary table highlights the transcripts that constitute the SVM_th10 classifier. Indicated for each feature are the Oligo-ID, the Gene symbols, the Entrez gene IDs, the RefSeq IDs, the Gene IDs and the Transcript IDs.
Abstract Confirmatory adaptive designs comprise a range of statistical methods that allow to modify the sample size of an ongoing trial in a data-dependent way without compromising control of the type I error rate. For short-term endpoints (e.g., 3-month response rate), comprehensive methodology of adaptive designs exists. However, clinical trials in oncology often have a special focus on long-term outcome and therefore often choose a time-to-event endpoint as the primary endpoint. Typical examples are progression-free survival (PFS) or overall survival (OS). But subtle statistical problems arise when adaptively analysing survival trials. Classical designs for survival trials are therefore commonly limited to a single primary endpoint, which combines the occurrence of progression, toxicities, deaths, and other events of potential interest into a single statistical measure (composite endpoint). However, the complexity of oncological diseases can be mapped more accurately using multi-stage models, where the occurrence of progressions, toxicities and deaths is modelled jointly instead of combining them into a single composite endpoint. We present and discuss adaptive design methodology for single-arm phase II survival trials for testing hypotheses on the joint distribution of several time-to-event endpoints in the context of multi-state models. We illustrate the methodology using the example of adaptive hypothesis tests for the joint distribution of progression-free survival (PFS) and overall survival (OS) in the context of an illness-death model. The methodology is motivated from application in pediatric oncology.
The one-sample log-rank test is the method of choice for single-arm Phase II trials with time-to-event endpoint. It allows to compare the survival of patients to a reference survival curve that typically represents the expected survival under standard of care. The one-sample log-rank test, however, assumes that the reference survival curve is known. This ignores that the reference curve is commonly estimated from historic data and thus prone to sampling error. Ignoring sampling variability of the reference curve results in type I error rate inflation. We study this inflation in type I error rate analytically and by simulation. Moreover we derive the actual distribution of the one-sample log-rank test statistic, when the sampling variability of the reference curve is taken into account. In particular, we provide a consistent estimate of the factor by which the true variance of the one-sample log-rank statistic is underestimated when reference curve sampling variability is ignored. Our results are further substantiated by a case study using a real world data example in which we demonstrate how to estimate the error rate inflation in the planning stage of a trial.
Abstract Objectives The aim of this study was to compare the second trimester thymus-thorax-ratio (TTR) between fetuses born preterm (study group) and those born after 37 weeks of gestation were completed (control group). Methods This study was conducted as a retrospective evaluation of the ultrasound images of 492 fetuses in the three vessel view. The TTR was defined as the quotient of a.p. thymus diameter and a.p. thoracic diameter. Results Fetuses that were preterm showed larger TTR (p<0.001) the second trimester than those born after 37 weeks of gestation were completed. The sensitivity of a binary classifier based on TTR for predicting preterm birth (PTB) was 0.792 and the specificity 0.552. Conclusions In our study, fetuses affected by PTB showed enlarged thymus size. These findings led us to hypothesize, that inflammation and immunomodulatory processes are altered early in pregnancies affected by PTB. However, TTR alone is not able to predict PTB.