Objectives Transvaginal ultrasound is typically the initial diagnostic approach in patients with postmenopausal bleeding for detecting endometrial atypical hyperplasia/cancer. Although transvaginal ultrasound demonstrates notable sensitivity, its specificity remains limited. The objective of this study was to enhance the diagnostic accuracy of transvaginal ultrasound through the integration of artificial intelligence. By using transvaginal ultrasound images, we aimed to develop an artificial intelligence based automated segmentation model and an artificial intelligence based classifier model. Methods Patients with postmenopausal bleeding undergoing transvaginal ultrasound and endometrial sampling at Mayo Clinic between 2016 and 2021 were retrospectively included. Manual segmentation of images was performed by four physicians (readers). Patients were classified into cohort A (atypical hyperplasia/cancer) and cohort B (benign) based on the pathologic report of endometrial sampling. A fully automated segmentation model was developed, and the performance of the model in correctly identifying the endometrium was compared with physician made segmentation using similarity metrics. To develop the classifier model, radiomic features were calculated from the manually segmented regions-of-interest. These features were used to train a wide range of machine learning based classifiers. The top performing machine learning classifier was evaluated using a threefold approach, and diagnostic accuracy was assessed through the F1 score and area under the receiver operating characteristic curve (AUC-ROC). Results 302 patients were included. Automated segmentation-reader agreement was 0.790.21 using the Dice coefficient. For the classification task, 92 radiomic features related to pixel texture/shape/intensity were found to be significantly different between cohort A and B. The threefold evaluation of the top performing classifier model showed an AUC-ROC of 0.90 (range 0.88-0.92) on the validation set and 0.88 (range 0.86-0.91) on the hold-out test set. Sensitivity and specificity were 0.87 (range 0.77-0.94) and 0.86 (range 0.81-0.94), respectively. Conclusions We trained an artificial intelligence based algorithm to differentiate endometrial atypical hyperplasia/cancer from benign conditions on transvaginal ultrasound images in a population of patients with postmenopausal bleeding.
Automatic abnormality identification of brachial plexus (BP) from normal magnetic resonance imaging to localize and identify a neurologic injury in clinical practice (MRI) is still a novel topic in brachial plexopathy. This study developed and evaluated an approach to differentiate abnormal BP with artificial intelligence (AI) over three commonly used MRI sequences, i.e. T1, FLUID sensitive and post-gadolinium sequences. A BP dataset was collected by radiological experts and a semi-supervised artificial intelligence method was used to segment the BP (based on nnU-net). Hereafter, a radiomics method was utilized to extract 107 shape and texture features from these ROIs. From various machine learning methods, we selected six widely recognized classifiers for training our Brachial plexus (BP) models and assessing their efficacy. To optimize these models, we introduced a dynamic feature selection approach aimed at discarding redundant and less informative features. Our experimental findings demonstrated that, in the context of identifying abnormal BP cases, shape features displayed heightened sensitivity compared to texture features. Notably, both the Logistic classifier and Bagging classifier outperformed other methods in our study. These evaluations illuminated the exceptional performance of our model trained on FLUID-sensitive sequences, which notably exceeded the results of both T1 and post-gadolinium sequences. Crucially, our analysis highlighted that both its classification accuracies and AUC score (area under the curve of receiver operating characteristics) over FLUID-sensitive sequence exceeded 90%. This outcome served as a robust experimental validation, affirming the substantial potential and strong feasibility of integrating AI into clinical practice.
Automatic speech recognition (ASR) systems have been shown to have large quality disparities between the language varieties they are intended or expected to recognize. One way to mitigate this is to train or fine-tune models with more representative datasets. But this approach can be hindered by limited in-domain data for training and evaluation. We propose a new way to improve the robustness of a US English short-form speech recognizer using a small amount of out-of-domain (long-form) African American English (AAE) data. We use CORAAL, YouTube and Mozilla Common Voice to train an audio classifier to approximately output whether an utterance is AAE or some other variety including Mainstream American English (MAE). By combining the classifier output with coarse geographic information, we can select a subset of utterances from a large corpus of untranscribed short-form queries for semi-supervised learning at scale. Fine-tuning on this data results in a 38.5% relative word error rate disparity reduction between AAE and MAE without reducing MAE quality.
Transiting exoplanets orbiting young nearby stars are ideal laboratories for testing theories of planet formation and evolution. However, to date only a handful of stars with age <1 Gyr have been found to host transiting exoplanets. Here we present the discovery and validation of a sub-Neptune around HD 18599, a young (300 Myr), nearby (d=40 pc) K star. We validate the transiting planet candidate as a bona fide planet using data from the TESS, Spitzer, and Gaia missions, ground-based photometry from IRSF, LCO, PEST, and NGTS, speckle imaging from Gemini, and spectroscopy from CHIRON, NRES, FEROS, and Minerva-Australis. The planet has an orbital period of 4.13 d, and a radius of 2.7Rearth. The RV data yields a 3-sigma mass upper limit of 30.5Mearth which is explained by either a massive companion or the large observed jitter typical for a young star. The brightness of the host star (V 9 mag) makes it conducive to detailed characterization via Doppler mass measurement which will provide a rare view into the interior structure of young planets.
Introduction/Background Postmenopausal vaginal bleeding (PMB) is usually the first manifestation of endometrial cancer (EC) and endometrial atypical hyperplasia (EAH). Transvaginal ultrasound (TVUS) is often the first diagnostic step for PMB. Although TVUS has a high sensitivity, specificity is low and a high rate of invasive biopsy procedures are performed, the majority of which are found negative on pathologic evaluation. This study developed an Artificial Intelligence (AI) model based on TVUS images to improve the accuracy of TVUS in EAH/EC early recognition in patients with PMB. Methodology 300 patients with PMB were enrolled. All patients underwent TVUS and endometrial sampling within three months from TVUS. Manual segmentation of the endometrium on two static images for each patient was performed independently by two radiologists. Patients were classified into cohort A (EAH/EC) and cohort B (benign) based on the endometrial sampling report. A fully automated segmentation model (ASE) was developed. For the second phase, radiomic features were calculated from the regions-of-interest and individual feature analysis was evaluated. These features were also used to train a wide range of machine learning-based classifiers. Results ASE-reader agreement shows similar performance to inter-reader agreement (ASE-Reader agreement: Dice similarity of 0.79±0.21). For the classification task, the deep learning model identified 92 features related to image texture and pixel intensity that were significantly different between cohort A and B. The top performing classifier model was a Support Vector Classifier using Minimum Redundancy Maximum Relevance feature selection. For the 3-fold evaluation, the AUC was 0.90 [0.88–0.92] for validation, and 0.88 [0.86–0.91] on the hold-out test set. Conclusion We have trained an AI-based algorithm to differentiate EC/EAH from benign conditions based on TVUS images in a PMB population. Based on our preliminary results, we plan to expand this work in larger cohorts and evaluate the AI model in external datasets.
Uterine leiomyosarcoma (LMS) is a rare but aggressive malignancy. On imaging, it is difficult to differentiate LMS from, for example, degenerated leiomyoma (LM), a prevalent but benign condition. We curated a data set of 115 axial T2-weighted MRI images from 110 patients (mean [range] age=45 [17-81] years) with UTs that included five different tumor types. These data were randomly split stratifying on tumor volume into training (n=85) and test sets (n=30). An independent second reader (reader 2) provided manual segmentations for all test set images. To automate segmentation, we applied nnU-Net and explored the effect of training set size on performance by randomly generating subsets with 25, 45, 65 and 85 training set images. We evaluated the ability of radiomic features to distinguish between types of UT individually and when combined through feature selection and machine learning. Using the entire training set the mean [95% CI] fibroid DSC was measured as 0.87 [0.59-1.00] and the agreement between the two readers was 0.89 [0.77-1.0] on the test set. When classifying degenerated LM from LMS we achieve a test set F1-score of 0.80. Classifying UTs based on radiomic features we identify classifiers achieving F1-scores of 0.53 [0.45, 0.61] and 0.80 [0.80, 0.80] on the test set for the benign versus malignant, and degenerated LM versus LMS tasks. We show that it is possible to develop an automated method for 3D segmentation of the uterus and UT that is close to human-level performance with fewer than 150 annotated images. For distinguishing UT types, while we train models that merit further investigation with additional data, reliable automatic differentiation of UTs remains a challenge.
BACKGROUND & AIMS: Our purpose was to detect pancreatic ductal adenocarcinoma (PDAC) at the prediagnostic stage (336 months before clinical diagnosis) using radiomics-based machine-learning (ML) models, and to compare performance against radiologists in a case-control study. METHODS: Volumetric pancreas segmentation was performed on prediagnostic computed tomography scans (CTs) (median interval between CT and PDAC diagnosis: 398 days) of 155 patients and an age-matched cohort of 265 subjects with normal pancreas. A total of 88 first-order and gray-level radiomic features were extracted and 34 features were selected through the least absolute shrinkage and selection operator-based feature selection method. The dataset was randomly divided into training (292 CTs: 110 prediagnostic and 182 controls) and test subsets (128 CTs: 45 prediagnostic and 83 controls). Four ML classifiers, k-nearest neighbor (KNN), support vector machine (SVM), random forest (RM), and extreme gradient boosting (XGBoost), were evaluated. Specificity of model with highest accuracy was further validated on an independent internal dataset (n = 176) and the public National Institutes of Health dataset (n = 80). Two radiologists (R4 and R5) independently evaluated the pancreas on a 5-point diagnostic scale. RESULTS: Median ( range) time between prediagnostic CTs of the test subset and PDAC diagnosis was 386 (97-1092) days. SVM had the highest sensitivity (mean; 95% confidence interval) (95.5; 85.5-100.0), specificity ( 90.3; 84.3-91.5), F1-score (89.5; 82.3-91.7), area under the curve (AUC) (0.98; 0.94-0.98), and accuracy (92.2%; 86.7-93.7) for classification of CTs into prediagnostic versus normal. All 3 other ML models, KNN, RF, and XGBoost, had comparable AUCs (0.95, 0.95, and 0.96, respectively). The high specificity of SVM was generalizable to both the independent internal ( 92.6%) and the National Institutes of Health dataset (96.2%). In contrast, interreader radiologist agreement was only fair (Cohen's kappa 0.3) and their mean AUC (0.66; 0.46-0.86) was lower than each of the 4 ML models (AUCs: 0.95-0.98) (P < .001). Radiologists also recorded false positive indirect findings of PDAC in control subjects (n = 83) (7% R4, 18% R5). CONCLUSIONS: Radiomics-based ML models can detect PDAC from normal pancreas when it is beyond human interrogation capability at a substantial lead time before clinical diagnosis. Prospective validation and integration of such models with complementary fluid- based biomarkers has the potential for PDAC detection at a stage when surgical cure is a possibility.
In spite of the advent of extremely large telescopes in the UV/optical/NIR range, the current generation of 8-10m facilities is likely to remain competitive at ground-UV wavelengths for the foreseeable future. The Cassegrain U-Band Efficient Spectrograph (CUBES) has been designed to provide high-efficiency (>40 300-420 nm goal) at a spectral resolving power of R>20,000, although a lower-resolution, sky-limited mode of R 7,000 is also planned. CUBES will offer new possibilities in many fields of astrophysics, providing access to key lines of stellar spectra: a tremendous diversity of iron-peak and heavy elements, lighter elements (in particular Beryllium) and light-element molecules (CO, CN, OH), as well as Balmer lines and the Balmer jump (particularly important for young stellar objects). The UV range is also critical in extragalactic studies: the circumgalactic medium of distant galaxies, the contribution of different types of sources to the cosmic UV background, the measurement of H2 and primordial Deuterium in a regime of relatively transparent intergalactic medium, and follow-up of explosive transients. The CUBES project completed a Phase A conceptual design in June 2021 and has now entered the Phase B dedicated to detailed design and construction. First science operations are planned for 2028. In this paper, we briefly describe the CUBES project development and goals, the main science cases, the instrument design and the project organization and management.
To determine if pancreas radiomics-based AI model can detect the CT imaging signature of type 2 diabetes (T2D). Total 107 radiomic features were extracted from volumetrically segmented normal pancreas in 422 T2D patients and 456 age-matched controls. Dataset was randomly split into training (300 T2D, 300 control CTs) and test subsets (122 T2D, 156 control CTs). An XGBoost model trained on 10 features selected through top-K-based selection method and optimized through threefold cross-validation on training subset was evaluated on test subset. Model correctly classified 73 (60
Objective Machine learning, deep learning, and artificial intelligence (AI) are terms that have made their way into nearly all areas of medicine. In the case of medical imaging, these methods have become the state of the art in nearly all areas from image reconstruction to image processing and automated analysis. In contrast to other areas, such as brain and breast imaging, the impacts of AI have not been as strongly felt in gynecologic imaging. In this review article, we: (i) provide a background of clinically relevant AI concepts, (ii) describe methods and approaches in computer vision, and (iii) highlight prior work related to image classification tasks utilizing AI approaches in gynecologic imaging. Data sources A comprehensive search of several databases from each database's inception to March 18th, 2021, English language, was conducted. The databases included Ovid MEDLINE(R) and Epub Ahead of Print, In-Process & Other Non-Indexed Citations, and Daily, Ovid EMBASE, Ovid Cochrane Central Register of Controlled Trials, and Ovid Cochrane Database of Systematic Reviews and ClinicalTrials.gov. Methods of study selection We performed an extensive literature review with 61 articles curated by three reviewers and subsequent sorting by specialists using specific inclusion and exclusion criteria. Tabulation, integration, and results We summarize the literature grouped by each of the three most common gynecologic malignancies: endometrial, cervical, and ovarian. For each, a brief introduction encapsulating the AI methods, imaging modalities, and clinical parameters in the selected articles is presented. We conclude with a discussion of current developments, trends and limitations, and suggest directions for future study. Conclusion This review article should prove useful for collaborative teams performing research studies targeted at the incorporation of radiological imaging and AI methods into gynecological clinical practice.
Muons from extensive air showers appear as rings in images taken with imaging atmospheric Cherenkov telescopes, such as VERITAS. These muon-ring images are used for the calibration of the VERITAS telescopes, however the calibration accuracy can be improved with a more efficient muon-identification algorithm. Convolutional neural networks (CNNs) are used in many state-of-the-art image-recognition systems and are ideal for muon image identification, once trained on a suitable dataset with labels for muon images. However, by training a CNN on a dataset labelled by existing algorithms, the performance of the CNN would be limited by the suboptimal muon-identification efficiency of the original algorithms. Muon Hunters 2 is a citizen science project that asks users to label grids of VERITAS telescope images, stating which images contain muon rings. Each image is labelled 10 times by independent volunteers, and the votes are aggregated and used to assign a `muon' or `non-muon' label to the corresponding image. An analysis was performed using an expert-labelled dataset in order to determine the optimal vote percentage cut-offs for assigning labels to each image for CNN training. This was optimised so as to identify as many muon images as possible while avoiding false positives. The performance of this model greatly improves on existing muon identification algorithms, identifying approximately 30 times the number of muon images identified by the current algorithm implemented in VEGAS (VERITAS Gamma-ray Analysis Suite), and roughly 2.5 times the number identified by the Hough transform method, along with significantly outperforming a CNN trained on VEGAS-labelled data.
In the era of Extremely Large Telescopes, the current generation of 8-10m facilities are likely to remain competitive at ground-UV wavelengths for the foreseeable future. The Cassegrain U-Band Efficient Spectrograph (CUBES) has been designed to provide high-efficiency (> 40%) observations in the near UV (305-400 nm requirement, 300-420 nm goal) at a spectral resolving power of R >20, 000 (with a lower-resolution, sky-limited mode of R ~7, 000). With the design focusing on maximizing the instrument throughput (ensuring a Signal to Noise Ratio (SNR) ~20 per high-resolution element at 313 nm for U ~18.5 mag objects in 1h of observations), it will offer new possibilities in many fields of astrophysics, providing access to key lines of stellar spectra: a tremendous diversity of iron-peak and heavy elements, lighter elements (in particular Beryllium) and light-element molecules (CO, CN, OH), as well as Balmer lines and the Balmer jump (particularly important for young stellar objects). The UV range is also critical in extragalactic studies: the circumgalactic medium of distant galaxies, the contribution of different types of sources to the cosmic UV background, the measurement of H2 and primordial Deuterium in a regime of relatively transparent intergalactic medium, and follow-up of explosive transients. The CUBES project completed a Phase A conceptual design in June 2021 and has now entered the detailed design and construction phase. First science operations are planned for 2028.
The de facto standard of dynamic histogram binning for radiomic feature extraction leads to an elevated sensitivity to fluctuations in annotated regions. This may impact the majority of radiomic studies published recently and contribute to issues regarding poor reproducibility of radiomic-based machine learning that has led to significant efforts for data harmonization; however, we believe the issues highlighted here are comparatively neglected, but often remedied by choosing static binning. The field of radiomics has improved through the development of community standards and open-source libraries such as PyRadiomics. But differences in image acquisition, systematic differences between observers' annotations, and preprocessing steps still pose challenges. These can change the distribution of voxels altering extracted features and can be exacerbated with dynamic binning.
Significance Statement Volumetric measurements are needed to characterize kidney structural findings on CT images to evaluate and test their potential utility in clinical decision making. Deep learning can enable this task in a scalable and reliable manner. Although automated kidney segmentation has been previously explored, methods for distinguishing cortex from medulla have never been done before. In addition, automated methods are typically evaluated at a single institution, without testing generalizability and robustness across different institutions. The tool developed in this study performs at the level of human readers and could enable large diverse population studies to evaluate how kidney, cortex, and medulla volumes can be used in various clinical settings, and establish normative values at large scale. Background In kidney transplantation, a contrast CT scan is obtained in the donor candidate to detect subclinical pathology in the kidney. Recent work from the Aging Kidney Anatomy study has characterized kidney, cortex, and medulla volumes using a manual image-processing tool. However, this technique is time consuming and impractical for clinical care, and thus, these measurements are not obtained during donor evaluations. This study proposes a fully automated segmentation approach for measuring kidney, cortex, and medulla volumes. Methods A total of 1930 contrast-enhanced CT exams with reference standard manual segmentations from one institution were used to develop the algorithm. A convolutional neural network model was trained ( n =1238) and validated ( n =306), and then evaluated in a hold-out test set of reference standard segmentations ( n =386). After the initial evaluation, the algorithm was further tested on datasets originating from two external sites ( n =1226). Results The automated model was found to perform on par with manual segmentation, with errors similar to interobserver variability with manual segmentation. Compared with the reference standard, the automated approach achieved a Dice similarity metric of 0.94 (right cortex), 0.90 (right medulla), 0.94 (left cortex), and 0.90 (left medulla) in the test set. Similar performance was observed when the algorithm was applied on the two external datasets. Conclusions A fully automated approach for measuring cortex and medullary volumes in CT images of the kidneys has been established. This method may prove useful for a wide range of clinical applications.
Total kidney volume (TKV) is the most important imaging biomarker for quantifying the severity of autosomal-dominant polycystic kidney disease (ADPKD). 3D ultrasound (US) can accurately measure kidney volume compared to 2D US; however, manual segmentation is tedious and requires expert annotators. We investigated a deep learning-based approach for automated segmentation of TKV from 3D US in ADPKD patients. We used axially acquired 3D US-kidney images in 22 ADPKD patients where each patient and each kidney were scanned three times, resulting in 132 scans that were manually segmented. We trained a convolutional neural network to segment the whole kidney and measure TKV. All patients were subsequently imaged with MRI for measurement comparison. Our method automatically segmented polycystic kidneys in 3D US images obtaining an average Dice coefficient of 0.80 on the test dataset. The kidney volume measurement compared with linear regression coefficient and bias from human tracing were R2 = 0.81, and − 4.42%, and between AI and reference standard were R2 = 0.93, and − 4.12%, respectively. MRI and US measured kidney volumes had R2 = 0.84 and a bias of 7.47%. This is the first study applying deep learning to 3D US in ADPKD. Our method shows promising performance for auto-segmentation of kidneys using 3D US to measure TKV, close to human tracing and MRI measurement. This imaging and analysis method may be useful in a number of settings, including pediatric imaging, clinical studies, and longitudinal tracking of patient disease progression.
ABSTRACT We present an analysis of spectropolarimetric observations of the low-mass weak-line T Tauri stars TWA 25 and TWA 7. The large-scale surface magnetic fields have been reconstructed for both stars using the technique of Zeeman Doppler imaging. Our surface maps reveal predominantly toroidal and non-axisymmetric fields for both stars. These maps reinforce the wide range of surface magnetic fields that have been recovered, particularly in pre-main sequence stars that have stopped accreting from the (now depleted) central regions of their discs. We reconstruct the large scale surface brightness distributions for both stars, and use these reconstructions to filter out the activity-induced radial velocity jitter, reducing the RMS of the radial velocity variations from 495 to 32 m s −1 for TWA 25, and from 127 to 36 m s −1 for TWA 7, ruling out the presence of close-in giant planets for both stars. The TWA 7 radial velocities provide an example of a case where the activity-induced radial velocity variations mimic a Keplerian signal that is uncorrelated with the spectral activity indices. This shows the usefulness of longitudinal magnetic field measurements in identifying activity-induced radial velocity variations.
Short-orbit gas giant planet formation/evolution mechanisms are still not well understood. One promising pathway to discriminate between mechanisms is to constrain the occurrence rate of these peculiar exoplanets at the earliest stage of the system's life. However, a major limitation when studying newly born stars is stellar activity. This cocktail of phenomena triggered by fast rotation, strong magnetic fields and complex internal dynamics, especially present in very young stars, compromises our ability to detect exoplanets. In this paper, we investigated the limitations of such detections in the context of already acquired data solely using radial velocity data acquired with a non-stabilised spectrograph. We employed two strategies: Doppler Imaging and Gaussian Processes and could confidently detect Hot Jupiters with semi-amplitude of 100 m.s^-1 buried in the stellar activity. We also showed the advantages of the Gaussian Process approach in this case. This study serves as a proof of concept to identify potential candidates for follow-up observations or even discover such planets in legacy datasets available to the community.
In 2017, the Muon Hunter project on the Zooniverse.org citizen science platform successfully gathered more than two million classification labels for nearly 140,000 camera images from VERITAS. The aim was to select and parameterize muon events for use in training convolutional neural networks. The success of this project proved that crowdsourcing labels for IACT image analysis is a viable avenue for further development of advanced machine-learning algorithms. These algorithms could potentially lend themselves to improving class separation between gamma-ray and hadronic event types. Nonetheless, it took two months to gather these labels from volunteers, which could be a bottleneck for future applications of this method. Here we present Muon Hunters 2.0: the follow-on project that demonstrates the development of unsupervised clustering techniques to gather muon labels more efficiently from volunteer classifiers.