Skin segmentation from clinical photography is a crucial step in dermatological image analysis. However, the variability in skin tones, lighting conditions, anatomical regions, and the presence of additional objects introduces significant challenges. Due to these complexities, the segmentation process is often performed manually, as developing an algorithm capable of handling such diverse conditions is particularly difficult. Recently, openworld foundation models have emerged, offering the potential to generalize across diverse and unseen conditions. These models present a promising opportunity for dermatology. In this work, we adopt two such models-Grounding DINO and SAM 2-to construct a pipeline for zero-shot skin segmentation in dermatology. We evaluated our approach on two clinical skin photography datasets comprising 27,378 images. Based on a manual rating protocol, 77.1% of the segmentations were deemed acceptable, demonstrating robustness in handling realworld clinical photographs. Our results highlight the potential of open-world foundation models to address a challenging problem in dermatology with minimal human involvement.
BACKGROUND:Response-adaptive randomization is controversial even in the best circumstances when based on a quickly determined primary outcome. In disease settings in which the primary outcome requires long follow-up, an intermediate endpoint may be chosen to update randomization allocations. The aim of our study is to evaluate the impact of response-adaptive randomization applied to an imperfect intermediate endpoint. We use tuberculosis trials as the motivating example. METHODS:We simulated a response-adaptive randomization design, adapting randomization allocations using an imperfect intermediate endpoint, in a superiority trial of two experimental regimens and one control arm. The primary study outcome was treatment success after 73 weeks from randomization; the intermediate endpoint was culture conversion at 8 weeks. We compared different sensitivity (Se) and specificity (Spe) scenarios for the intermediate endpoint, while varying the true treatment efficacy. We evaluated the performance of response-adaptive randomization to achieve its primary goal of allocating more participants to the better arm and the impact of time-trends on type I error rate. RESULTS:Even in an ideal state of perfect accuracy (i.e. intermediate endpoint with Se = 100% and Spe = 100%), response-adaptive randomization did not always live up to its main purpose of allocating more patients to the better arm. Lower accuracy of the intermediate endpoint leads to greater divergence from the goal of more allocations to the better arm. The larger the difference in treatment efficacy between the arms, the more striking the impact of an intermediate endpoint with poor diagnostic accuracy. Time-trends inflate the type I error rate, and while stratified tests can correct this, they do so at the cost of a power loss. Allocating more patients to the worst arm increases power for comparisons with this arm but reduces power for comparisons of the best arm to control. CONCLUSION:Given the objective of evaluating several new therapeutic regimens in a timely manner, response-adaptive randomization is tempting. However, it requires at least reliance on highly accurate intermediate endpoints, which are still no guarantee of response-adaptive randomization's trustworthiness.
Abstract The MVA-BN vaccine is considered safe and effective and has been widely deployed to prevent monkeypox. However, safety and tolerability data from endemic areas are limited. We performed a single-arm clinical trial of the safety of the MVA-BN vaccine among adults in the Democratic Republic of the Congo (ClinicalTrials.gov: NCT05734508; Registered 21 February 2023). Participants were personnel working on a monkeypox therapeutics trial (PALM-007) conducted in a high-risk, endemic setting for monkeypox. Participants received two doses of vaccine 28 days apart and were actively followed up to 28 days after each dose to assess the safety profile of MVA-BN. From March 2023 to June 2024, 500 participants were enrolled and received their first vaccine dose; 494 (99%) received the second dose and associated follow-up. There were no adverse events (AEs) of grade 3 or higher observed during the study. 175 participants (35%) reported at least one vaccination site AE, the most common of which was mild or moderate vaccination site pain (120/500 first dose recipients (24%); 48/494 second dose recipients (10%)). There was one recorded case of breakthrough monkeypox occurring approximately 3 months after the second dose. This observational study provides further evidence that the MVA-BN monkeypox vaccine is generally safe and well-tolerated. As previously reported, the most common AEs were mild or moderate injection site reactions.
Six months of drug treatment is standard of care for drug-sensitive pulmonary tuberculosis (TB). Understanding the factors determining the length of treatment required for durable cure would allow individualization of treatment durations. We conducted a prospective, randomized, controlled noninferiority trial (PredictTB) of 4 versus 6 months of chemotherapy in patients with pulmonary TB in South Africa and China. Seven hundred and four participants with newly diagnosed, drug-sensitive TB were enrolled and stratified on the basis of radiographic disease characteristics assessed by FDG PET/CT imaging. Participants with less extensive disease (n = 273) were randomly assigned at week 16 to complete therapy after 4 months or continue receiving treatment for 6 months. This study was stopped early after an interim analysis revealed that patients assigned to the 4-month treatment arm had a higher risk of relapse. Among participants who received 4 months of chemotherapy, 17 of 141 (12.1%) experienced TB-specific unfavorable outcomes compared with only 2 of 132 (1.5%) who completed 6 months of treatment. In the nonrandomized arm that included participants with more extensive disease, only 8 of 248 (3.2%) experienced unfavorable outcomes. Total lung cavity volume and lesion glycolysis at week 16 were associated with the risk of unfavorable outcomes. PET/CT imaging at TB recurrence showed that bacteriological relapses predominantly occurred in active cavities originally present at baseline. Subsequent post hoc automated segmentation of serial PET/CT scans combined with machine learning enabled the classification of participants according to their likelihood of relapse.
We designed and successfully implemented photography protocols for the PALM007 (Pamoja Tulinde Maisha - Together Save Lives in Swahili) randomised clinical trial evaluating tecovirimat for the treatment of mpox in the Democratic Republic of the Congo (DRC). We addressed the unique challenges and limitations of conducting standardised photography procedures on patients with mpox in high-risk 'red zones' in remote study treatment centre sites. Key considerations included standardised photography protocols, training, quality assurance, participant privacy and comfort, and interpersonal communication. Our developed procedures enabled acquisition of 61 926 standardised, clinical-quality photographs of 597 patients with mpox. These considerations may guide future studies collecting standardised photographs in a low-resource setting.
For emerging and re-emerging viral diseases, the timing and location of outbreaks are typically unpredictable and may require rapid implementation of randomized controlled trials (RCTs) to evaluate the safety and efficacy of candidate vaccines or therapeutics. In austere and hard-to-reach settings, establishing effective laboratory operations for RCTs requires considerable effort to address both anticipated and unforeseen challenges. In this opinion article, we describe key principles and practical challenges associated with the effective operationalization of research laboratories in the context of a double-blind, placebo-controlled RCT for the evaluation of a candidate therapeutic for mpox in the Democratic Republic of the Congo. The solutions implemented and lessons learned may inform the planning and operationalization of clinical trial laboratories in resource-limited outbreak settings and provide a framework to guide future efforts.
This Perspective discusses the use of bayesian methods in clinical trials and the importance of having US Food and Drug Administration (FDA) guidance that provides substantive insights about the methods’ proper role.
BACKGROUND Tecovirimat is available for the treatment of mpox (formerly known as monkeypox) in Europe and the United States, on the basis of findings from efficacy studies in animals and safety evaluations in healthy humans. Evidence from randomized, controlled trials of safety and efficacy in patients with mpox is lacking. METHODS We conducted a double-blind, randomized, placebo-controlled trial of tecovirimat in patients with mpox in the Democratic Republic of Congo (DRC). Patients with at least one mpox skin lesion and positive polymerase-chain-reaction results for clade I MPXV were assigned in a 1:1 ratio to receive tecovirimat or placebo. All patients received supportive care. The primary end point was resolution of mpox lesions, measured in number of days after randomization. Safety was also assessed. RESULTS From October 7, 2022, through July 9, 2024, a total of 597 patients underwent randomization - 295 to receive tecovirimat and 302 to receive placebo. The median time from randomization to lesion resolution was 7 days with tecovirimat and 8 days with placebo; the competing-risks hazard ratio for lesion resolution was 1.13 (95% confidence interval [CI], 0.97 to 1.31; P=0.14). Results were similar whether patients began the trial regimen within 7 days after the reported onset of symptoms (competing-risks hazard ratio, 1.16; 95% CI, 0.98 to 1.37) or more than 7 days after onset (competing-risks hazard ratio, 1.00; 95% CI, 0.71 to 1.40). Overall mortality was 1.7%, which was lower than the case fatality rate of 4.6% reported in the DRC in 2023. At 14 days, the percentages of patients who had blood, lesion, and oropharyngeal samples negative for MPXV by PCR were similar in the two groups. Adverse events occurred in 72.9% of the patients in the tecovirimat group and 70.5% of those in the placebo group, and serious adverse events were reported in 5.1% and 5.0%, respectively. CONCLUSIONS Tecovirimat did not reduce the number of days to lesion resolution in patients with mpox caused by clade I MPXV. No safety concerns were identified.
The US Food and Drug Administration (FDA) and National Institutes of Health (NIH) share a mutual interest in facilitating efficient, well-designed clinical studies of drugs, devices, and biological products. Recent advances in science and technology, as well as innovative approaches to research design and methodology, provide opportunities to enhance efficiency in medical product development and improve participant engagement in clinical trials. Recent initiatives across the FDA and NIH focus on evidence modernization approaches. Fostering appropriate use of novel designs and sources of evidence, such as real-world data (RWD) to support marketing authorizations and satisfy postapproval study requirements, may be enhanced by using consensus terminology for innovative study designs. To facilitate effective communication within the scientific community, FDA and NIH formed an interagency collaborative initiative to define clinical research terms related to innovative study designs, with a focus on studies using RWD, for FDA-regulated medical products or broader research and foster a shared understanding of terms across the clinical research ecosystem. The FDA-NIH Modernizing Research and Evidence (MoRE) Glossary Working Group (MGWG) was initiated in April 2023 to evaluate terms inadequately defined within the clinical research community that would benefit from development of a consensus definition. The MGWG conducted a landscape evaluation of common innovative design terminology that may lack clarity or concordance. Subsequently, the MGWG reviewed whether and how existing regulations, guidance, and policies use or define such terms. Following the landscape evaluation, the MGWG engaged in rigorous review to seek consensus definitions. In addition, federal agencies sought public input via a request for information before publishing the included terms and definitions. The MGWG developed the MoRE Consensus Definitions, comprising 40 clinical research terms and definitions related to innovative clinical study designs that support scientific, patient, clinical, and regulatory decision-making. The MoRE Consensus Definitions are intended to facilitate effective communication about clinical research and enable transparency around innovative clinical study designs. This publication makes available the glossary developed through this collaboration and serves as an accessible resource for the clinical research enterprise. Furthermore, as clinical research is continuously evolving, additional efforts may focus on emerging new vocabulary and evolving use of current terms to benefit medical product development.
BACKGROUND:Swift regulatory approval of therapeutic interventions is crucial during emerging infectious disease outbreaks. However, variability in treatment effects based on disease severity or subgroups complicates trial design and endpoint selection. Prioritized composite endpoints can capture treatment effects across diverse clinical courses; however, their performance under heterogeneous treatment effects remains uncertain. This study uses simulation to evaluate trial design strategies in such contexts. METHODS:This study examines eight combinations of population and endpoint strategies to optimize trial design in emerging infectious diseases: evaluating treatment in the overall population and subgroups, with various endpoint choices including single, multiple, and prioritized composite endpoints. Simulated data was generated using multistate models based on the ACTT-1 study. Eight treatment effect scenarios, some exhibiting heterogeneity, were considered to evaluate ability to demonstrate efficacy. RESULTS:In scenarios without heterogeneous treatment effects, analyses in the overall population generally showed higher power than subgroup analyses. Time to recovery had relatively high power, while prioritized composite and multiple endpoints were comparable. In scenarios with treatment effect heterogeneity by baseline disease severity, power was higher in effective subgroups than in the overall population. Prioritized composite endpoints showed high power in scenarios where the treatment was effective on distinct endpoints in each subgroup. CONCLUSIONS:For drug development in emerging infectious diseases with limited information, it is preferable to focus on evaluating prioritized composite or multiple endpoints in the overall population. Stratified analysis can be more powerful than unstratified analysis and should be considered for the primary analysis in the overall population.
BACKGROUND:Monkeypox virus (MPXV) has been linked to vertical transmission, but systematic data are scarce. We aimed to describe the sociodemographic, clinical, and virological characteristics and assess the frequency and determinants of adverse outcomes in pregnant women with MPXV clade I infection. METHODS:In this prospective cohort study, we pooled data from three cohort studies (MBOTE-SK, PREGMPOX, and Uvira mpox) and one randomised controlled trial (PALM007) conducted in the South Kivu, Maniema, and Sankuru provinces of DR Congo between Dec 29, 2022, and June 20, 2025. Pregnant women and adolescent girls with a PCR-confirmed diagnosis of mpox were followed up throughout hospitalisation for mpox, delivery, and until discharge during the postpartum period. We extracted data on sociodemographic characteristics, MPXV exposure, clinical and obstetric presentation, and laboratory results. In a univariable analysis, we examined factors associated with the following adverse outcomes: spontaneous or missed abortion (<20 weeks of gestation), stillbirth (≥20 weeks of gestation), preterm birth (<37 weeks of gestation), live birth of a neonate with macroscopic mpox-like lesions, early (first 7 days) neonatal death, congenital anomaly, or maternal death (during pregnancy or discharge postpartum). FINDINGS:We collected data from 89 pregnant women in the first (25 [28%]), second (31 [35%]), and third (33 [37%]) trimesters across all four studies: MBOTE-SK (36 [40%]), PREGMPOX (24 [27%]), PALM007 (25 [28%]), and Uvira mpox (four [4%]). All participants recovered from mpox; no maternal deaths were reported. During hospitalisation for mpox, fetal loss was reported in 17 (19%) women. Final pregnancy outcomes were known for 69 (78%) participants; adverse outcomes were reported in 35 (51%) women (95% CI 38-63), including fetal loss in 31 (45%; 95% CI 33-57; 16 [52%] spontaneous abortions, four [13%] missed abortions, and 11 [35%] stillbirths). Of the 38 live births, four neonates had congenital mpox-like lesions; one infant died a few hours after birth. No preterm births or congenital abnormalities were recorded. MPXV infection during the first trimester was associated with a higher risk of adverse pregnancy outcomes than during the second (risk ratio [RR] 0·6 [95% CI 0·4-0·9]) and third (0·2 [0·1-0·4]) trimesters (p=0·0008). Adverse outcomes were also associated with high viral load in skin lesions (PCR cycle threshold ≤30; RR 3·5 [95% CI 1·0-12·3]; p=0·045), direct sexual contact with the index case (1·6 [1·1-2·4]; p=0·026), positive HIV status (2·0 [1·4-2·9]; p=0·0002), and the presence of genital lesions (1·9 [1·1-3·2]; p=0·025). INTERPRETATION:MPXV clade I infection in pregnancy is associated with a high risk of fetal loss and congenital infection, particularly during the first trimester. Targeted preventive and clinical strategies are urgently needed to protect pregnant women and their infants in settings that are endemic and epidemic for mpox. FUNDING:The European and Developing Countries Clinical Trials Partnership, the Belgian Directorate-General Development Cooperation and Humanitarian Aid, the Swiss National Science Foundation, the Research Foundation-Flanders, the Gates Foundation, the Intramural Research Program of the National Institutes of Health, and the National Cancer Institute.
Background: Platform trials typically feature a shared control arm and multiple experimental treatment arms. Staggered entry and exit of arms splits the control group into two cohorts: those randomized during the same period in which the experimental arm was open (concurrent controls) and those randomized outside that period (nonconcurrent controls). Combining these control groups may offer increased statistical power but can lead to bias if analyses do not account for time trends in the response variable. Proposed methods of adjustment for time may increase type I error rates when time trends impact arms unequally or when large, sudden changes to the response rate occur. However, there has been limited exploration of the degree of type I error inflation one can plausibly expect in real-world scenarios.Methods: We use data from the Adaptive COVID-19 Treatment Trial (ACTT) to mimic a realistic platform trial with a remdesivir control arm. We compare four strategies for estimating the effect of interferon beta-1a (the ACTT-3 experimental arm) relative to remdesivir (data from ACTT-1, ACTT-2, and ACTT-3) on recovery and death by day 29: utilizing concurrent controls only (the prespecified analysis), pooling all remdesivir arm data without adjustment (the "unadjusted-pooled" analysis), adjusting for time as a categorical variable, and a Bayesian hierarchical model implementation which adjusts for time trends using smoothing techniques (the "Bayesian time machine"). We compare type I error rates and relative efficiency of each method in simulation settings based on observed ACTT remdesivir arm data.Results: The unadjusted-pooled approach provided substantially different estimates of the effect of interferon beta-1a relative to remdesivir compared with the concurrent-only and model-based approaches, indicating that changes in recovery and death rates over time were not ignorable across different stages of ACTT. The model-based approaches rely on an assumption of constant treatment effects for each arm in the platform relative to control; error rates more than doubled in settings where this was not satisfied. Relative efficiency of the model-based approaches compared with the concurrent-only analysis was moderate.Conclusions: In simulation settings where key model assumptions were not met, potential efficiency gains from incorporation of nonconcurrent controls were outweighed by the risk of substantial type I error rate inflation. This leads us to advise against these strategies for primary analyses in confirmatory clinical trials, aligning with current FDA guidance advising against comparisons to nonconcurrent controls in COVID-19 settings. The model-based adjustment methods may be useful in other settings, but we recommend performing the concurrent-only analysis as a reference for assessing the degree to which nonconcurrent controls drive results.
Clinical trials conducted during the COVID-19 pandemic demonstrated the value of adaptive design methods in emerging disease settings, when there can be considerable uncertainty around disease natural history, anticipated endpoint effect sizes and population size. In such settings, there may also be uncertainty regarding the most appropriate primary endpoint. This might lead to an externally-driven decision to change the primary endpoint during the course of an adaptive trial. If information on the new primary endpoint is already being collected, initially as a secondary endpoint, the trial could continue with a new primary endpoint. In this case it is unclear how statistical inference on the final primary endpoint should be adjusted for interim analyses monitoring the initial primary endpoint so as to control the overall type I error rate as adjusting for monitoring as if this was based on the new endpoint could be conservative whereas failing to make any adjustment could lead to type I error rate inflation if the new and original endpoint are correlated. This paper shows how group-sequential methods can be modified to control the type I error rate for the analysis of the new primary endpoint irrespective of the true treatment effect on the initial primary endpoint. The method is illustrated using a simulated data example based on a clinical trial of remdesivir in COVID-19. Construction of critical values for the test of the new primary endpoint require a value for the correlation between this and the initial primary endpoint. We present simulation studies to demonstrate that the type I error rate is controlled when this value is estimated from the data on the two endpoints obtained from the trial.
Purpose Mpox is a viral illness with symptoms similar to smallpox. A key clinical metric to monitor disease progression is the number of skin lesions. Manually counting mpox skin lesions is labor-intensive and susceptible to human error. Approach We previously developed an mpox lesion counting method based on the UNet segmentation model using 66 photographs from 18 patients. We have compared four additional methods: the instance segmentation methods Mask R-CNN, YOLOv8, and E2EC, in addition to a UNet++ model. We designed a patient-level leave-one-out experiment, assessing their performance using F1 score and lesion count metrics. Finally, we tested whether an ensemble of the networks outperformed any single model. Results Mask R-CNN model achieved an F1 score of 0.75, YOLOv8 a score of 0.75, E2EC a score of 0.70, UNet++ a score of 0.81, and baseline UNet a score of 0.79. Bland-Altman analysis of lesion count performance showed a limit of agreement (LoA) width of 62.2 for Mask R-CNN, 91.3 for YOLOv8, 94.2 for E2EC, and 62.1 for UNet++, with the baseline UNet model achieving 69.1. The ensemble showed an F1 score performance of 0.78 and LoA width of 67.4. Conclusions Instance segmentation methods and UNet-based semantic segmentation methods performed equally well in lesion counting. Furthermore, the ensemble of the trained models showed no performance increase over the best-performing model UNet, likely because errors are frequently shared across models. Performance is likely limited by the availability of high-quality photographs for this complex problem, rather than the methodologies used.
BACKGROUND:Although antivirals remain important for the treatment COVID-19, methods to assess treatment efficacy are lacking. Here, we investigated the impact of remdesivir on viral dynamics and their contribution to understanding antiviral efficacy in the multicenter Adaptive COVID-19 Treatment Trial 1, which randomized patients to remdesivir or placebo. METHODS:Longitudinal specimens collected during hospitalization from a substudy of 642 patients with COVID-19 were measured for viral RNA (upper respiratory tract and plasma), viral nucleocapsid antigen (serum), and host immunologic markers. Associations with clinical outcomes and response to therapy were assessed. RESULTS:Higher baseline plasma viral loads were associated with poorer clinical outcomes, and decreases in viral RNA and antigen in blood but not the upper respiratory tract correlated with enhanced benefit from remdesivir. The treatment effect of remdesivir was most pronounced in patients with elevated baseline nucleocapsid antigen levels: the recovery rate ratio was 1.95 (95% CI, 1.40-2.71) for levels >245 pg/mL vs 1.04 (95% CI, .76-1.42) for levels <245 pg/mL. Remdesivir also accelerated the rate of viral RNA and antigen clearance in blood, and patients whose blood levels decreased were more likely to recover and survive. CONCLUSIONS:Reductions in SARS-CoV-2 RNA and antigen levels in blood correlated with clinical benefit from antiviral therapy. CLINICAL TRIAL REGISTRATION:NCT04280705 (ClinicalTrials.gov).
In this issue of NEJM Evidence, de Boer et al.1 investigate a linkage between acute myocardial infarction and influenza infection with the application of a self-controlled risk interval study, a design commonly applied to evaluate increased risk of rare side effects from vaccines.2,3 The authors studied whether acute myocardial infarction was more likely to occur within 1 week after influenza diagnosis through a comparison of the relative likelihood of two types of events: acute myocardial infarction within 7 days after a positive influenza test result and acute myocardial infarction with influenza infection occurring outside a 7-day window.1 Basically, the authors evaluated whether the sequential but nearly coincident occurrence of influenza and acute myocardial infarction was more likely than the two events occurring far apart in time.
In 1970, the first case of mpox (formerly known as monkeypox) was documented in an infant in Equateur Province, Democratic Republic of Congo (DRC).1 Infections with clade I monkeypox virus (MPXV) are endemic in the rainforest regions of central Africa and result from both zoonotic and human-to-human transmission. The cessation of smallpox vaccination in 1980 because of the eradication of smallpox has led to an increase in the number of individuals who are orthopox immune naïve and is felt to be responsible for a recent increase in mpox cases in the DRC. Comparisons of active surveillance in Sankuru Province from 2005 through 2007 revealed a 20-fold increase in the incidence of mpox compared with the 1980s, with a 5-fold-lower incidence among those with a smallpox vaccination scar.2.
Motivated by the experience of COVID-19 trials, we consider clinical trials in the setting of an emerging disease in which the uncertainty of natural disease course and potential treatment effects makes advance specification of a sample size challenging. One approach to such a challenge is to use a group sequential design to allow the trial to stop on the basis of interim analysis results as soon as a conclusion regarding the effectiveness of the treatment under investigation can be reached. As such a trial may be halted before a formal stopping boundary is reached, we consider the final analysis under such a scenario, proposing alternative methods for when the decision to halt the trial is made with or without knowledge of interim analysis results. We address the problems of ensuring that the type I error rate neither exceeds nor falls unnecessarily far below the nominal level. We also propose methods in which there is no maximum sample size, the trial continuing either until the stopping boundary is reached or it is decided to halt the trial.