Predicting isocitrate dehydrogenase (IDH) mutations in gliomas using magnetic resonance imaging (MRI) is clinically important for treatment planning. This study compared two artificial intelligence (AI) models, GliomaDepth-IDH (ResNet34-based) and GliomaVista-IDH (Vision Transformer-based), with 18 physicians (eight neuroradiologists, five neurosurgeons, and five neurosurgery residents) in predicting IDH mutation status. On the Brain Tumor Segmentation Challenge dataset, the GliomaVista-IDH AI model achieved an area under the curve (AUC) value of 0.97, significantly outperforming all physician groups. However, external validation on a Japanese cohort revealed performance degradation: GliomaDepth-IDH declined to an AUC of 0.75 and GliomaVista-IDH to 0.82, with GliomaVista-IDH showing significant calibration issues (Brier score = 0.32). High-performing physicians achieved comparable results (AUC = 0.88) with superior calibration (Brier score = 0.19). Inter-rater reliability analysis revealed substantial variability across physician groups. These findings suggest that AI models can assist many physicians, while experienced practitioners remain competitive with better-calibrated predictions in challenging domains.
Abstract Polyploidization is a key evolutionary force in plants, but the reasons behind its prevalence remain unclear. While the potential ecological benefits of established polyploids are well studied, little is known about the short-term genomic and epigenetic responses immediately after polyploidization, which are crucial for successful speciation. In this study, we assemble the genomes of the two progenitors of Arabidopsis kamchatica ( A. halleri and A. lyrata ) and examine the epigenome of synthetic and natural tetraploids of A. kamchatica to investigate the combined effect of allopolyploidization and environment on DNA methylation changes. We find the most significant methylation changes at allopolyploidization, followed by smaller changes in subsequent generations. Offspring grown under different conditions show divergent patterns, suggesting environmental effects, while their methylation patterns converge toward those of natural tetraploids over generations. Our findings highlight two key epigenetic changes post-polyploidization: convergence toward established polyploids and divergence driven by environmental factors.
Abstract Time-series RNA sequencing provides a powerful framework for studying dynamic gene regulation, yet conventional analyses usually represent gene expression profiles as real-valued vectors in Euclidean space and quantify similarity using correlation or distance. Inspired by quantum information theory, we present a framework for encoding time-series gene expression profiles as complex-valued vectors comprising amplitude and phase components in Hilbert space. We designed multiple encoding models to represent gene expression in the amplitude of complex-valued vectors, encode temporal differences in the phase, and extend the phase representation to incorporate the direction of local expression changes. Gene-gene similarity was then quantified using fidelity, which measures the overlap between two encoded vectors. Evaluation using time-series RNA-seq datasets across diverse species and biological contexts showed that different encoding models produced distinct fidelity distributions that were related to, but distinct from, conventional correlation measures. We then constructed gene-gene networks using pairwise fidelity values and detected communities containing genes with similar temporal profiles. Although fidelity distributions differed across encoding models, the resulting communities captured major temporal expression programs, and functional annotations based on gene ontology and Kyoto encyclopedia of genes and genomes pathway analyses provided exploratory biological context. The detected communities were comparable to those obtained using conventional methods, including weighted correlation network analysis and fuzzy c-means clustering. Furthermore, as a proof-of-concept, we performed SWAP-test circuit simulations to mimic fidelity computation on a quantum computer; under noise-aware conditions, these simulations produced less accurate fidelity estimates with higher computational cost than classical computation. As a proof-of-concept, this study provides a complementary view of temporal transcriptome organization, rather than a uniformly superior alternative to conventional methods.
Allopolyploids arise through hybridization between related species, carrying multiple sets of chromosomes from distinct progenitors, referred to as subgenomes. Within allopolyploids, duplicated genes across subgenomes, called homeologs, are thought to enhance environmental robustness by shifting their expression ratios depending on environmental and developmental changes. However, existing methods for detecting such ratio shifts, including HomeoRoq and Fisher's exact test, are limited to allopolyploids with two subgenome sets and thus cannot handle more complex cases. We present the HOmeolog Bias Identification Test (HOBIT), a statistical method for detecting shifts in homeolog expression ratios across different conditions using RNA-seq count data. HOBIT performs a likelihood ratio test for each homeolog, comparing a full model allowing homeolog expression ratios to vary across conditions with a reduced model assuming constant ratios. Simulation benchmarks for allotetraploids and allohexaploids demonstrated that HOBIT outperforms existing methods in both area under the receiver operating characteristic curve and F1 score. Application to real RNA-seq datasets from allotetraploid Cardamine flexuosa, allotriploid Cardamine insueta, and allohexaploid Triticum aestivum (wheat) produced biologically consistent results reflecting experimental settings. HOBIT provides a promising framework for uncovering homeolog regulation and adaptive responses in allopolyploids, without restrictions of ploidy complexities.
Allopolyploids arise through hybridization between related species, carrying multiple sets of chromosomes from distinct progenitor, referred to as subgenomes. Within allopolyploids, duplicated genes across subgenomes, called homeologs, are thought to enhance environmental robustness by shifting their expression ratios depending on environmental and developmental changes. However, existing methods for detecting such ratio shifts, including HomeoRoq and Fisher’s exact test, are limited to allopolyploids with two subgenome sets inherited from two progenitors, and thus cannot handle more complex cases (e.g., allohexaploid wheat). Here, we present the HOmeolog Bias Identification Test (HOBIT), a statistical method for detecting shifts in homeolog expression ratios across different conditions using RNA-Seq count data. HOBIT performs a likelihood ratio test for each homeolog, comparing a full model allowing homeolog expression ratios to vary across conditions with a reduced model assuming constant ratios. Simulation benchmarks for allotetraploids and allohexaploids demonstrated that HOBIT outperforms existing methods in both area under the receiver operating characteristic curve and F1 score. Application to real RNA-Seq datasets from allotetraploid Cardamine flexuosa , allotriploid Cardamine insueta , and Triticum aestivum (wheat) produced biologically consistent results reflecting experimental settings. HOBIT provides a promising framework for uncovering homeolog regulation and adaptive responses in allopolyploids, without restrictions of ploidy complexities. ### Competing Interest Statement The authors have declared no competing interest. Japan Society for the Promotion of Science, JP25K09719, JP23K23582, JP22H05179 Japan Science and Technology Agency, JPMJCR16O3 Swiss National Science Foundation, 31003A_212551
Background:Sublobar resection for small peripheral non-small cell lung cancer (NSCLC) (≤2 cm) became one of the standard procedures. Retrospective studies demonstrated that pathological pleural invasion (pPL) is associated with a higher risk of local recurrence during sublobar resection. If pPL can be properly assessed intraoperatively, converting to lobectomy may reduce the risk of local recurrence associated with sublobar resection. The study objective was to develop a deep learning algorithm predicting pPL from thoracoscopic images. Methods:Among consecutive patients who underwent radical thoracoscopic surgery for cT1N0M0 NSCLC (TNM 8th) from 5/2020 to 3/2022, 80 patients with pleural surface changes due to tumor (excluding cTis/1mi or peritumoral adhesions) were included. A tumor recognition deep learning model using the ResNet50 architecture was constructed from images and the focus was visualized using gradient-weighted class activation mapping (Grad-CAM). Among images in which a tumor is visible, the presence of pPL was predicted (trained on 64, validated on 16). Predictive ability was compared with the surgeons' intraoperative evaluation using McNemar's test. Results:Among 80 patients (age 69±10 years, 42.5% female, tumor diameter 20±7 mm), pPL was found in 22 patients. Compared to the pPL- group, the pPL+ group was significantly older, with larger solid diameter, more pure solid nodules, and higher SUV max. Among the 422,873 images extracted from all 80 videos, 2,074 images showed tumors, of which 608 images were pPL+. The tumor recognition algorithm had an image-level accuracy of 0.78 and F1 score of 0.60. The pPL model had a patient-level accuracy of 0.69, while the accuracy of thoracic surgeons was 0.75 (P=0.32). Conclusions:Deep learning analysis of thoracoscopic images of lung cancer surgery showed the possibility of prediction of pPL to a comparable degree to surgeons.
BackgroundVarious skin diseases exist; some require specialist treatment. Consequently, patients with these conditions are often referred from primary dermatological clinics to other medical institutions for further secondary or tertiary care. However, it has yet to be fully evaluated at primary dermatological clinics which skin conditions necessitate such referrals.ObjectivesThis study aims to identify skin conditions that often require further care when treated at a primary dermatological clinic.MethodsThis study enroled 14,306 patients who had visited the Ueo Dermatology Clinic (a primary dermatological clinic in Saiki City, Oita prefecture, Japan) from 1 January 2020 to 31 December 2022. The following clinical information was examined: the primary disease, age, sex, whether referrals (general or emergency) were made after the clinic visit or not, and the reasons for the referrals.Results'Eczema and dermatitis' was the most frequent category in the clinic, although the referral rates for this category were not exceptionally high compared to other categories. A large number of emergency referrals was observed for 'Viral infections (herpes zoster),' 'Drug-induced skin reactions,' and 'Bacterial infections (cellulitis)'; a large number of general referrals was observed for 'Malignant skin tumours and melanomas' and 'Benign skin tumours.' Although low numbers, the rates for general and emergency referrals were high in the categories of 'Blistering Diseases,' 'Connective Tissue Diseases,' and 'Vasculitis, Purpura, and Other Vascular Diseases.' In an analysis of the reasons for the referrals, the following reasons ranked highest: requirement for high-level medical treatment, surgical operation, or hospitalisation, and consultation with doctors in departments other than dermatology.ConclusionsThis study identified several skin conditions often requiring additional care when treated at a primary dermatological clinic. Based on these findings, seamless cooperation between primary clinics and specialised medical institutions is anticipated.
Genome-wide association studies have enabled the identification of important genetic factors in many trait studies. However, only a fraction of the heritability can be explained by known genetic factors, even in the most common diseases. Genetic loci combinations, or epistatic contributions expressed by combinations of single nucleotide polymorphisms (SNPs), have been argued to be one of the critical factors explaining some of the missing heritability, especially in oligogenic/polygenic diseases. Rheumatoid arthritis (RA) is a complex disease with more than 100 reported SNP associations, as well as various HLA haplotypes and amino acids; however, many associations between RA and inter-chromosomal SNP combinations are unknown. To discover novel associations of epistatic interactions with high odds ratios in RA, we applied the LAMPLINK method, a systematic enumerative procedure for identifying high-order SNP combinations, to a Japanese RA cohort (discovery cohort; 4024 patients with RA and 7731 controls). We validated the identified associations in a different Japanese cohort (validation cohort; 810 RA patients and 6303 controls). In this study, we identified 90 significant genetic associations in the discovery cohort. Among these, 74 (82.2%) associations were replicated in the validation cohort, and eight combinations were inter-chromosomal, all of which comprised rs7765379 or rs35265698 located in the HLA region. These two SNPs exhibited strong correlations with valine at amino acid position 11 in HLA-DRB1 (HLA-DRB1-11-Val). Finally, we discovered that rs9624 showed an association with RA through an epistatic interaction with HLA-DRB1-11-Val. Overall, LAMPLINK showed high reliability for identifying epistatic genetic contributions hidden in complex traits.
BACKGROUND:Digital health technologies using mobile apps and wearable devices are a promising approach to the investigation of substance use in the real world and for the analysis of predictive factors or harms from substance use. Moreover, consecutive repeated data collection enables the development of predictive algorithms for substance use by machine learning methods. OBJECTIVE:We developed a new self-monitoring mobile app to record daily substance use, triggers, and cravings. Additionally, a wearable activity tracker (Fitbit) was used to collect objective biological and behavioral data before, during, and after substance use. This study aims to describe a model using machine learning methods to determine substance use. METHODS:This study is an ongoing observational study using a Fitbit and a self-monitoring app. Participants of this study were people with health risks due to alcohol or methamphetamine use. They were required to record their daily substance use and related factors on the self-monitoring app and to always wear a Fitbit for 8 weeks, which collected the following data: (1) heart rate per minute, (2) sleep duration per day, (3) sleep stages per day, (4) the number of steps per day, and (5) the amount of physical activity per day. Fitbit data will first be visualized for data analysis to confirm typical Fitbit data patterns for individual users. Next, machine learning and statistical analysis methods will be performed to create a detection model for substance use based on the combined Fitbit and self-monitoring data. The model will be tested based on 5-fold cross-validation, and further preprocessing and machine learning methods will be conducted based on the preliminary results. The usability and feasibility of this approach will also be evaluated. RESULTS:Enrollment for the trial began in September 2020, and the data collection finished in April 2021. In total, 13 people with methamphetamine use disorder and 36 with alcohol problems participated in this study. The severity of methamphetamine or alcohol use disorder assessed by the Drug Abuse Screening Test-10 or the Alcohol Use Disorders Identification Test-10 was moderate to severe. The anticipated results of this study include understanding the physiological and behavioral data before, during, and after alcohol or methamphetamine use and identifying individual patterns of behavior. CONCLUSIONS:Real-time data on daily life among people with substance use problems were collected in this study. This new approach to data collection might be helpful because of its high confidentiality and convenience. The findings of this study will provide data to support the development of interventions to reduce alcohol and methamphetamine use and associated negative consequences. INTERNATIONAL REGISTERED REPORT IDENTIFIER (IRRID):DERR1-10.2196/44275.
Long-term field monitoring of leaf pigment content is informative for understanding plant responses to environments distinct from regulated chambers but is impractical by conventional destructive measurements. We developed PlantServation, a method incorporating robust image-acquisition hardware and deep learning-based software that extracts leaf color by detecting plant individuals automatically. As a case study, we applied PlantServation to examine environmental and genotypic effects on the pigment anthocyanin content estimated from leaf color. We processed >4 million images of small individuals of four Arabidopsis species in the field, where the plant shape, color, and background vary over months. Past radiation, coldness, and precipitation significantly affected the anthocyanin content. The synthetic allopolyploid A. kamchatica recapitulated the fluctuations of natural polyploids by integrating diploid responses. The data support a long-standing hypothesis stating that allopolyploids can inherit and combine the traits of progenitors. PlantServation facilitates the study of plant responses to complex environments termed "in natura".
Abstract Although allopolyploid species are common among natural and crop species, it is not easy to distinguish duplicated genes, known as homeologs, during their genomic analysis. Yet, cost-efficient RNA sequencing (RNA-seq) is to be developed for large-scale transcriptomic studies such as time-series analysis and genome-wide association studies in allopolyploids. In this study, we employed a 3′ RNA-seq utilizing 3′ untranslated regions (UTRs) containing frequent mutations among homeologous genes, compared to coding sequence. Among the 3′ RNA-seq protocols, we examined a low-cost method Lasy-Seq using an allohexaploid bread wheat, Triticum aestivum. HISAT2 showed the best performance for 3′ RNA-seq with the least mapping errors and quick computational time. The number of detected homeologs was further improved by extending 1 kb of the 3′ UTR annotation. Differentially expressed genes in response to mild cold treatment detected by the 3′ RNA-seq were verified with high-coverage conventional RNA-seq, although the latter detected more differentially expressed genes. Finally, downsampling showed that even a 2 million sequencing depth can still detect more than half of expressed homeologs identifiable by the conventional 32 million reads. These data demonstrate that this low-cost 3′ RNA-seq facilitates large-scale transcriptomic studies of allohexaploid wheat and indicate the potential application to other allopolyploid species.
Abstract PURPOSE To compare the performance and explainability of the visual transformer (ViT) and convolutional neural network (CNN) architectures in predicting genomic mutations from brain MRI. METHODS The performances of the ViT and CNN classification models in predicting the IDH mutation status of gliomas were compared. The two models were fine-tuned on the TCIA dataset. The fine-tuned models were evaluated on the TCIA dataset and an external independent dataset, namely the Japanese Cohort (JC) dataset. To evaluate their explanatory power, the gradient-weighted class activation mapping (Grad-CAM) visualization of the CNNs model and attention map visualization of the ViT model were compared. RESULTS The visual transformer model consistently outperforms the convolutional neural network on both the TCIA and JC datasets (p-value = 0.021, p-value < 0.001, statistically different). The attention map of the ViT model accurately highlighted the tumor, and the Grad-CAM of the CNN model sometimes highlighted non-tumor areas. CONCLUSION The ViT model was more robust against differences in the image domain. The ViT model's attention map had superior explainability.
We propose an estimation method of subjects' physical/mental health condition from their heart rate (HR) and evaluate it on the newly collected data including 25 million points over 97 participants. The accurate health condition estimation is important for an employee's mental health care and an objective understanding of our condition. For the estimation, the heart rate variability (HRV) has been widely used, but there are some technical difficulties with measuring the HRV, such as maintaining a good quality of data for a long period of time. Here, we predict the subjects' physical/mental health only from the HR measured by Fitbit instead of the HRV. We first measured more than 25 million points of HR and steps data from 97 participants over 3 months using the Fitbit Inspire HRTM. We also conducted questionnaires to check their physical conditions each day. We then predict their condition by focusing on the inactive period of HR and applying the support vector machine to the preprocessed data. The best balanced accuracy of our method achieved 0.582, which was higher than the state-of-the-art method with HRV whose accuracy is 0.565.
The aim of the present study was to determine which individual or combined CpG sites among O6-methylguanine DNA methyltransferase CpG 74–89 in glioblastoma mainly affects the response to temozolomide resulting from CpG methylation using statistical analyses focused on the tumor volume ratio (TVR). We retrospectively examined 44 patients who had postoperative volumetrically measurable residual tumor tissue and received adjuvant temozolomide therapy for at least 6 months after initial chemoradiotherapy. TVR was defined as the tumor volume 6 months after the initial chemoradiotherapy divided by that before the start of chemoradiotherapy. Predictive values for TVR as a response to adjuvant therapy were compared among the averaged methylation percentages of individual or combined CpGs using the receiver operating characteristic curve. Our data revealed that combined CpG 78 and 79 showed a high area under the curve (AUC) and a positive likelihood ratio and that combined CpG 76–79 showed the highest AUC among all combinations. AUCs of consecutive CpG combinations tended to be higher for CpG 74–82 in exon 1 than for CpG 83–89 in intron 1. In conclusion, the methylation status at CpG sites in exon 1 was strongly associated with TVR reduction in glioblastoma.
Purpose We are attempting to develop a navigation system for safe and effective peripancreatic lymphadenectomy in gastric cancer surgery. As a preliminary study, we examined whether or not the peripancreatic dissection line could be learned by a machine learning model (MLM). Methods Among the 41 patients with gastric cancer who underwent radical gastrectomy between April 2019 and January 2020, we selected 6 in whom the pancreatic contour was relatively easy to trace. The pancreatic contour was annotated by a trainer surgeon in 1242 images captured from the video recordings. The MLM was trained using the annotated images from five of the six patients. The pancreatic contour was then segmented by the trained MLM using images from the remaining patient. The same procedure was repeated for all six combinations. Results The median maximum intersection over union of each image was 0.708, which was higher than the threshold (0.5). However, the pancreatic contour was misidentified in parts where fatty tissue or thin vessels overlaid the pancreas in some cases. Conclusion The contour of the pancreas could be traced relatively well using the trained MLM. Further investigations and training of the system are needed to develop a practical navigation system.