Supplementary Materials give a detailed introduction to various types of textural features, the anatomy of the breast, and how mammography works.
Table S2 shows the discrepancies in the image conditions of the mammograms that were used for performance analysis. The image conditions include the types of mammograms (FFDM/SFM), view of the mammogram (CC/MLO), digitizing machine (SFM only) or digital machine (FFDM only), regions of interest, and grey level and size of the pixels.
Table S1 shows the discrepancies in sample characteristics of the 18 studies used to compare the discrimination of textural features with PMD. The sample characteristics include the country where the study was initiated, sample size, validation methods, mean age at mammography and baseline, mean years that mammograms were taken before diagnosis, menopausal status, and variables match and adjustment.
Figure S1 shows that 18 publications were included in the analysis after applying our search strategy. We searched the PubMed database for original publications on mammographic texture analysis published up to November 2023. The searching strategy included (“texture” OR “texture analysis” OR “feature” OR “parenchymal pattern” OR “parenchymal analysis”) AND (“mammogram” OR “mammograms” OR “mammography”) AND (“risk” OR “risk assessment” OR “risk prediction”). The literature inclusion criteria were: 1) the publication is published in English, 2) full text is available, 3) the discriminatory performance of both texture-based features and Cumulus or mammographic density measures similar to Cumulus using the same dataset were reported.
Supplementary Figure S2 shows the estimation of genetic ancestry of AGOG, GliomaScan and GICC samples by way of scatter plot of PC1 versus PC2.
The non-breast non-ovarian cancers associated with BRCA1 and BRCA2 pathogenic variants (PVs) are controversial. We aimed to examine this using a prospective cohort design. This study included 1260 BRCA1 and 1058 BRCA2 PV carriers (91
Supplementary Table S6 shows the DEPTH posterior log odds in favour of association (PLO) score and the logistic regression p-value of the lowest p-value variant for each of the 34 known glioma risk regions using GliomaScan (S6a), AGOG (S6b), and GICC (S6c).
BACKGROUND:Artificial intelligence (AI)-based algorithms are being implemented in breast screening to detect breast cancers on mammographic images. We aimed to apply an epidemiological approach to demonstrate how a cancer detection algorithm can be leveraged as an intermediate-term predictor of breast cancer (current and 4-year risk) to deliver greater risk-based personalisation in screening mammography. METHODS:In this population cohort study, we used detection scores from an AI cancer detection algorithm (BRAIx AI Reader), which was calibrated using a training dataset of 397 648 women aged 40 years to 97 years from women who screened at BreastScreen Victoria, Australia between Jan 1, 2016, and Dec 31, 2017, to create a woman-specific mammography-based score for breast cancer risk, the BRAIx risk score. Subsequently, the BRAIx risk score was evaluated on an independent test dataset of women from BreastScreen Victoria, Australia, comprising a random population cohort of 96 348 women who screened from Jan 1, 2016, to Dec 31, 2017, aged 40 years to 74 years, and an independent, external dataset from woman screened at Karolinska University Hospital, Stockholm, Sweden. We applied logistic regression, using the BRAIx risk score to estimate risks of invasive breast cancers on the test dataset: (1) detected at cohort entry (n=525); and (2) for women given an all clear, diagnosed during the next 4 years either at future screens (n=790) or during intervals between screens (n=308). We also trained full multivariate risk models (logistic regression and elastic net) using the training dataset and evaluated their predictive performance on the test and external validation data, with assessment of familial aspects of the BRAIx risk score achieved with inference about causation from examining changes in regression coefficients in an innovative statistical analysis framework. FINDINGS:In both Australian and Swedish test datasets, the BRAIx risk score predicted cancer detection at cohort entry and future cancer risk (all p<0·0001). The BRAIx risk score was the strongest tested explanatory factor for cancer detection at cohort entry (odds ratio 13·80 [95% CI 9·54-20·80] in Australian data; 8·89 [3·19-37·49] in Swedish data) and for intermediate-term cancer risk (2·29 [2·13-2.47] in Australian data; 2·15 [1·85-2·50] in Swedish data). We found that adding a thresholded binary version of the BRAIx risk score significantly improved model fit (p<2·2 × 10-16, Australian and Swedish data) and women with BRAIx risk scores of more than 2 were significantly at many-fold increased risk of intermediate-term cancer than women below that threshold (12·34 [7·33-20·91], Australia; 44·7 [11·9-184·9], Sweden; p<0·0001). For the top 2% of women given an all clear with the highest BRAIx risk score, the probability of a cancer diagnosis within 4 years was 9·7%. The BRAIx risk score explained 23% of why family history predicts 4-year risk (p<0·0001). After fitting the BRAIx risk score in a multivariate model, mammographic density was no longer significantly associated with breast cancer risk in the Australian test data (p>0·05) and became associated with lower risk for intermediate-term cancer in the external Swedish test dataset (0·83 [0·73-0·95]). INTERPRETATION:The BRAIx risk score is a strong intermediate-term predictor of breast cancer (current to 4-year risk). Calibrating the score on a training dataset produces population-specific probabilities for calculating individual-specific risk scores for screening clients based on their mammogram images. These risk scores enable future development of personalised screening pathways to transform population breast cancer screening and save lives. Identification of women given an all clear but at very high risk, similar to those carrying BRCA1 and BRCA2 mutations, could reveal insights into both familial and non-familial causes of breast cancer. FUNDING:Australian Government Medical Research Future Fund, the Ramaciotti Foundation, the National Breast Cancer Foundation, Cancer Australia, and the National Health and Medical Research Council.
Supplementary Table S9 shows the generalised Berk-Jones statistic (GBJ) p-values for the potential novel glioma susceptibility regions by study, glioma type and sex.
Floods and tropical cyclones (TCs), two of the most frequent and costliest climate-related disasters worldwide, have been linked to sustained health risks extending beyond acute hazards. However, evidence on the underlying epigenetic mechanisms remains scarce. We aimed to characterize DNA methylation patterns associated with exposure to floods and TCs of varying intensities. We collected peripheral blood samples from 479 women (132 twin pairs and 215 of their sisters) across Australia. Blood-derived DNA methylation profiles were assessed using the Illumina HumanMethylation450 BeadChip array. Daily flood and TC exposure data for the 6 years preceding each blood draw were obtained from the Dartmouth Flood Observatory and the International Best Track Archive for Climate Stewardship, respectively, and linked to participants based on residential addresses. Using a within-sibship analytical framework that accounted for shared familial factors and other relevant covariates, we examined associations between flood and TC exposures of varying intensities and site-specific methylation at each cytosine-guanine dinucleotide (CpG). Differentially methylated regions (DMRs) were identified using a combination of the comb-p and DMRcate algorithms. There were 164 CpGs and 219 DMRs associated with flood and TC exposures (Bonferroni-adjusted p value < 0.05), mapping to 242 genes enriched in pathways related to inflammation and immune regulation. These genes have been implicated in a wide range of human diseases or phenotypes. The number of differentially methylated CpGs increased with more recent and higher-intensity exposures. Intensity-dependent gene regulation was observed, with genes such as AMT and C22orf45 consistently implicated across various exposure levels, whereas RNF39 and ACY3 emerged only at higher intensities. Exposures to floods and TCs were associated with differentially DNA methylated signals across the human genome, exhibiting intensity-dependent patterns. The identified signals and related gene pathways may shed light on the biological mechanism underlying the profound health effects of climate-related disasters.
Mammographic (or breast) density is an established risk factor for breast cancer, previously measured using a variety of quantitative, semi-automated and automated approaches. We present a new automated measure, AutoCumulus, learned from applying deep learning to semi-automated measures. We studied the mammograms of 9,057 population-screened women in the BRAIx program for which semi-automated measurements of mammographic density had been made by experienced readers using the CUMULUS software. The dataset was split into training, testing, and validation sets (80
The Hierarchical Taxonomy of Psychopathology (HiTOP) model is an empirically based dimensional alternative to current consensus based diagnostic classification systems. It was developed through the synthesis of decades of empirical studies, but its construct and structural validity have never been comprehensively evaluated within a single sample using a confirmatory modeling approach. We assessed 108 psychopathology phenotypes, incorporating 37 novel components, in 542 monozygotic and 155 dizygotic twin pairs with diverse mental health histories to validate the hierarchical dimensional structure of psychopathology and quantify its genetic and environmental influences. Analyses largely replicated the theorized dimensional structure at lower, but not upper, tiers of the model. Heritability was lower and more variable at lower (median h2 = 43.5) than upper (median h2 = 57.1) tiers. Our findings provide partial construct and structural validation for the HiTOP model while identifying novel psychopathological phenotypes and heritability estimates to guide future biological studies.
Studies have identified genetic and epidemiologic factors associated with mammographic density (MD) phenotypes. However, MD-associated genetic variants only account for a small proportion of the total estimated heritability. Interrogating interactions between genetic and epidemiologic factors could potentially identify additional MD-associated loci, expand our understanding of the genetic basis of MD phenotypes, and clarify how epidemiologic factors modulate relationships between genetic variants and MD. We conducted six separate genome-wide, gene-environment (GxE) interaction analyses, applying 2 degrees of freedom (df) and 1df interaction tests, for each of three MD phenotypes (percent density, dense area (DA), and nondense area (NDA)). The six epidemiologic factors considered were height, ever parous, parity, ever menopausal hormone therapy, ever breastfeeding, and months of breastfeeding. We included European ancestry participants from multiple studies within the Markers of Density consortium and the Breast Cancer Association Consortium (n = 4895-16 218 depending on specific analyses). We identified 11 loci with genome-wide significant (P < 5 × 10-8) interaction tests including two novel common genetic signals interacting with parity (8p21.2) and ever breastfeeding (19p13.2) for NDA. Our results suggest that epidemiologic risk factors might influence relationships between common genetic variants and MD phenotypes at particular genomic loci.
Supplementary Figure S1 showing the flow chart of the analyses for identifying novel, unique, and common risk regions from DEPTH and logistic regression analyses.
BACKGROUND:The adverse gut microbiome may underlie the variability in risks of colorectal cancer (CRC) and metachronous CRC in people with Lynch syndrome (LS). The role of pks+/-Escherichia coli (pks+/-E. coli), Enterotoxigenic Bacteroides fragilis (ETBF), and Fusobacterium nucleatum (Fn) in CRCs and adenomas in people with LS is unknown. METHODS:A total of 358 LS cases, including 386 CRCs, 90 adenomas, 195 normal colonic mucosa DNA from the Australasian Colon Cancer Family Registry were tested using multiplex TaqMan qPCR. Logistic regression was used to compare the intratumoural prevalence of each bacteria in Lynch CRCs with 1336 sporadic CRCs. Cox proportional-hazards regression estimated the associations of each bacteria with the risk of metachronous CRC and neoplasia. FINDINGS:Pks+ E. coli (odds ratio [95% confidence interval] = 1.60 [1.08-2.35], P = 0.017), pks-E. coli (3.87 [2.58-5.80], P < 0.001) and Fn (19.47 [13.32-28.87], P < 0.001) were significantly enriched in LS CRCs when compared with sporadic CRCs. Pks+ E. coli in the initial CRC was associated with an increased risk of metachronous CRC (hazard ratio [95% confidence interval] = 2.32 [1.29-4.17], P = 0.005) and metachronous colorectal neoplasia (1.51 [1.02-2.23], P = 0.040) when compared with CRCs without pks+ E. coli. INTERPRETATION:Pks+ E. coli, pks-E. coli, and Fn are enriched within LS CRCs, suggesting possible roles in CRC development in LS. Having intratumoural pks+ E. coli is associated with increased risk of metachronous CRC, suggesting that, if validated, people with LS might benefit from pks+ E. coli screening and eradication. FUNDING:This work was funded by an NHMRC Investigator grant (GNT1194896) and a Cancer Australia/Cancer Council NSW co-funded grant (GNT2012914).
Supplementary Table S6 shows the number of risk regions identified by first stage DEPTH analysis.
Supplementary Table S7 shows the number of risk regions that are common to both CCFR and GECCO datasets in stage 1 DEPTH analysis, as well as whether they are detected by logistic regression analysis or were in previous CRC GWAS studies.
Supplementary Table S4 shows the number of samples excluded from the GICC data and reasons for exclusion.