Discovery of sensitive and biologically grounded biomarkers is essential for early detection and monitoring of Alzheimer's disease (AD). Structural MRI is widely available but typically relies on hand-crafted features such as cortical thickness or volume. We ask whether self-supervised learning (SSL) can uncover more powerful biomarkers from the same data. Existing SSL methods underperform FreeSurfer-derived features in disease classification, conversion prediction, and amyloid status prediction. We introduce Residual Noise Contrastive Estimation (R-NCE), a new SSL framework that integrates auxiliary FreeSurfer features while maximizing additional augmentation-invariant information. R-NCE outperforms traditional features and existing SSL methods across multiple benchmarks, including AD conversion prediction. To assess biological relevance, we derive Brain Age Gap (BAG) measures and perform genome-wide association studies. R-NCE-BAG shows high heritability and associations with MAPT and IRAG1, with enrichment in astrocytes and oligodendrocytes, indicating sensitivity to neurodegenerative and cerebrovascular processes.
BACKGROUND:Transposable elements (TEs) occupy nearly half of the human genome and play diverse biological roles. Despite their abundance, the extent to which TEs contribute to three-dimensional (3D) genome structure remains unclear. RESULTS:To investigate this, we generate a modified Hi-C analysis pipeline to probe TE-associated chromatin interactions. Our analysis reveals that TE sequences are responsible for 3D genome structure in interphase nuclei. This phenomenon is mediated by the recruitment of specific epigenetic/transcription factors to TEs, which both promote and impair chromatin contacts. We computationally identified known factors positively associated with chromatin contacts (CTCF, RAD21, SMC3) and chromatin contact impairing proteins (RNF2). Additionally, we identiy potential novel factors (SMARCA4, MAFK), which, when knocked down, lead to decreased chromatin contacts and loops at and between TEs. Notably, SMARCA4 knockdown selectively reduce short-range contacts, highlighting its role in maintaining 3D genome structure through TE binding. CONCLUSIONS:Overall, our findings demonstrate that TEs are crucial determinants of 3D genome organization in mammalian cells.
Recent advances in large language models (LLMs) have demonstrated significant promise in document understanding and question-answering. Despite the progress, existing approaches can only process short documents due to limited context length or fail to fully leverage multi-modal information. In this work, we introduce DocAgent, a multi-agent framework for long-context document understanding that imitates human reading practice. Specifically, we first extract a structured, tree-formatted outline from documents to help agents identify relevant sections efficiently. Further, we develop an interactive reading interface that enables agents to query and retrieve various types of content dynamically. To ensure answer reliability, we introduce a reviewer agent that cross-checks responses using complementary sources and maintains a task-agnostic memory bank to facilitate knowledge sharing across tasks. We evaluate our method on two long-context document understanding benchmarks, where it bridges the gap to human-level performance by surpassing competitive baselines, while maintaining a short context length. Our code is available at https://github.com/lisun-ai/DocAgent.
Large language models (LLMs), pre-trained on vast amounts of text, have shown remarkable abilities in understanding general knowledge and commonsense. There-fore, it's desirable to leverage pre-trained LLM to help solve computer vision tasks. Previous works on multi-modal LLM mainly focus on the generation capability. In this work, we propose LLM-augmented visual representation learning (LMVR). Our approach involves initially using a vision encoder to extract features, which are then projected into the word embedding space of the LLM. The LLM then generates responses based on the visual representation and a text prompt. Finally, we aggregate sequence-level features from the hidden layers of the LLM to obtain image-level representations. We conduct extensive experiments on multiple datasets, and have the following findings: (a) LMVR outperforms traditional vision encoder on various down-stream tasks, and effectively learns the correspondence between words and image regions; (b) LMVR improves the generalizability compared to using a vision encoder alone, as evidenced by its superior resistance to domain shift; (c) LMVR improves the robustness of models to corrupted and perturbed visual data. Our findings demonstrate LLM-augmented visual representation learning is effective as it learns object-level concepts and commonsense knowledge.
The nuclear matrix, a proteinaceous gel composed of proteins and RNA, is an important nuclear structure that supports chromatin architecture, but its role in human pluripotent stem cells (hPSCs) has not been described. Here we show that by disrupting heterogeneous nuclear ribonucleoprotein U (HNRNPU) or the nuclear matrix protein, Matrin-3, primed hPSCs adopted features of the naive pluripotent state, including morphology and upregulation of naive-specific marker genes. We demonstrate that HNRNPU depletion leads to increased chromatin accessibility, reduced DNA contacts and increased nuclear size. Mechanistically, HNRNPU acts as a transcriptional co-factor that anchors promoters of primed-specific genes to the nuclear matrix with POLII to promote their expression and their RNA stability. Overall, HNRNPU promotes cell-type stability and when reduced promotes conversion to earlier embryonic states. Ma et al. show that heterogeneous nuclear ribonucleoprotein U promotes the primed state in human pluripotent stem cells by interacting with nuclear matrix protein, Matrin-3, and regulating primed-specific genes.
Vision-language models pre-trained on large scale of unlabeled biomedical images and associated reports learn generalizable semantic representations. These multi-modal representations can benefit various downstream tasks in the biomedical domain. Contrastive learning is widely used to pre-train vision-language models for general natural images and associated captions. Despite its popularity, we found biomedical texts have complex and domain-specific semantics that are often neglected by common contrastive methods. To address this issue, we propose a novel method, perturbed report discrimination, for pre-train biomedical vision-language models. First, we curate a set of text perturbation methods that keep the same words, but disrupt the semantic structure of the sentence. Next, we apply different types of perturbation to reports, and use the model to distinguish the original report from the perturbed ones given the associated image. Parallel to this, we enhance the sensitivity of our method to higher level of granularity for both modalities by contrasting attention-weighted image sub-regions and sub-words in the image-text pairs. We conduct extensive experiments on multiple downstream tasks, and our method outperforms strong baseline methods. The results demonstrate that our approach learns more semantic meaningful and robust multi-modal representations.
Epigenetic control of cell fates is a critical determinant to maintain cell type stability and permit differentiation during embryonic development. However, the epigenetic control mechanisms are not well understood. Here, it is shown that the histone acetyltransferase reader protein BRD8 impairs the conversion of primed mouse EpiSCs (epiblast stem cells) to naive mouse ESCs (embryonic stem cells). BRD8 works by maintaining histone acetylation on promoters and transcribed gene bodies. BRD8 is responsible for maintaining open chromatin at somatic genes, and histone acetylation at naive-specific genes. When Brd8 expression is reduced, chromatin accessibility is unchanged at primed-specific genes, but histone acetylation is reduced. Conversely, naive-specific genes has reduced repressive chromatin marks and acquired accessible chromatin more rapidly during the cell type conversion. It is shown that this process requires active histone deacetylation to promote the conversion of primed to naive. This data supports a model for BRD8 reading histone acetylation to accurately localize the genome-wide binding of the histone acetyltransferase KAT5. Overall, this study shows how the reading of the histone acetylation state by BRD8 maintains cell type stability and both enables and impairs stem cell differentiation.
This paper introduces an innovative methodology for producing high-quality 3D lung CT images guided by textual information. While diffusion-based generative models are increasingly used in medical imaging, current state-of-the-art approaches are limited to low-resolution outputs and underutilize radiology reports’ abundant information. The radiology reports can enhance the generation process by providing additional guidance and offering fine-grained control over the synthesis of images. Nevertheless, expanding text-guided generation to high-resolution 3D images poses significant memory and anatomical detail-preserving challenges. Addressing the memory issue, we introduce a hierarchical scheme that uses a modified UNet architecture. We start by synthesizing low-resolution images conditioned on the text, serving as a foundation for subsequent generators for complete volumetric data. To ensure the anatomical plausibility of the generated samples, we provide further guidance by generating vascular, airway, and lobular segmentation masks in conjunction with the CT images. The model demonstrates the capability to use textual input and segmentation tasks to generate synthesized images. Algorithmic comparative assessments and blind evaluations conducted by 10 board-certified radiologists indicate that our approach exhibits superior performance compared to the most advanced models based on GAN and diffusion techniques, especially in accurately retaining crucial anatomical features such as fissure lines and airways. This innovation introduces novel possibilities. This study focuses on two main objectives: (1) the development of a method for creating images based on textual prompts and anatomical components, and (2) the capability to generate new images conditioning on anatomical elements. The advancements in image generation can be applied to enhance numerous downstream tasks.
Current state-of-the-art models for natural language understanding require a preprocessing step to convert raw text into discrete tokens. This process known as tokenization relies on a pre-built vocabulary of words or sub-word morphemes. This fixed vocabulary limits the model's robustness to spelling errors and its capacity to adapt to new domains. In this work, we introduce a novel open-vocabulary language model that adopts a hierarchical two-level approach: one at the word level and another at the sequence level. Concretely, we design an intra-word module that uses a shallow Transformer architecture to learn word representations from their characters, and a deep inter-word Transformer module that contextualizes each word representation by attending to the entire word sequence. Our model thus directly operates on character sequences with explicit awareness of word boundaries, but without biased sub-word or word-level vocabulary. Experiments on various downstream tasks show that our method outperforms strong baselines. We also demonstrate that our hierarchical model is robust to textual corruption and domain shift.
Supervised learning based object detectors suffer from the high cost and difficulty of labeling datasets. Self-supervised learning methods require no manual annotations. However, the misalignment between the pretext task designed for image classification and the downstream task affects the detection performance. Therefore, this paper proposes a self-supervised dense contrastive learning method to improve performance of object detection in remote sensing images. Specifically, first, Swin Transformer substitutes popular CNN to extract features of augmented multiple views. Second, global and local features are extracted using parallel global and dense projector heads, respectively. Third, a predictor head is added to increase the nonlinear transformations in the network. Extensive experiments on the NWPU VHR-10 dataset show that the proposed method outperforms two representative strong baseline methods, including MoCoV2 and DenseCL.
This paper introduces an innovative methodology for producing high-quality 3D lung CT images guided by textual information. While diffusion-based generative models are increasingly used in medical imaging, current state-of-the-art approaches are limited to low-resolution outputs and underutilize radiology reports' abundant information. The radiology reports can enhance the generation process by providing additional guidance and offering fine-grained control over the synthesis of images. Nevertheless, expanding text-guided generation to high-resolution 3D images poses significant memory and anatomical detail-preserving challenges. Addressing the memory issue, we introduce a hierarchical scheme that uses a modified UNet architecture. We start by synthesizing low-resolution images conditioned on the text, serving as a foundation for subsequent generators for complete volumetric data. To ensure the anatomical plausibility of the generated samples, we provide further guidance by generating vascular, airway, and lobular segmentation masks in conjunction with the CT images. The model demonstrates the capability to use textual input and segmentation tasks to generate synthesized images. Algorithmic comparative assessments and blind evaluations conducted by 10 board-certified radiologists indicate that our approach exhibits superior performance compared to the most advanced models based on GAN and diffusion techniques, especially in accurately retaining crucial anatomical features such as fissure lines and airways. This innovation introduces novel possibilities. This study focuses on two main objectives: (1) the development of a method for creating images based on textual prompts and anatomical components, and (2) the capability to generate new images conditioning on anatomical elements. The advancements in image generation can be applied to enhance numerous downstream tasks.
Rationale: Chronic obstructive pulmonary disease (COPD) is characterized by pathologic changes in the airways, lung parenchyma, and persistent inflammation, but the links between lung structural changes and blood transcriptome patterns have not been fully described. Objectives: The objective of this study was to identify novel relationships between lung structural changes measured by chest computed tomography (CT) and blood transcriptome patterns measured by blood RNA sequencing (RNA-seq). Methods: CT scan images and blood RNA-seq gene expression from 1223 participants in the COPD Genetic Epidemiology (COPDGene (R)) study were jointly analyzed using deep learning to identify shared aspects of inflammation and lung structural changes that we labeled image-expression axes (IEAs). We related IEAs to COPD-related measurements and prospective health outcomes through regression and Cox proportional hazards models and tested them for biological pathway enrichment. Results: We identified 2 distinct IEAs: IEAemph which captures an emphysema-predominant process with a strong positive correlation to CT emphysema and a negative correlation to forced expiratory volume in 1 second and body mass index (BMI); and IEAairway which captures an airway-predominant process with a positive correlation to BMI and airway wall thickness and a negative correlation to emphysema. Pathway enrichment analysis identified 29 and 13 pathways significantly associated with IEAemph and IEAairway, respectively (adjusted p<0.001). Conclusions: Integration of CT scans and blood RNA-seq data identified 2 IEAs that capture distinct inflammatory processes associated with emphysema and airway-predominant COPD.
Somatic cell reprogramming and oncogenic transformation share surprisingly similar features, yet transformed cells are resistant to reprogramming. Epigenetic barriers must block transformed cells from reprogramming, but the nature of those barriers is unclear. In this study, we generated a systematic panel of transformed mouse embryonic fibroblasts (MEFs) using oncogenic transgenes and discovered transformed cell lines compatible with reprogramming when transfected with Oct4/Sox2/Klf4/Myc. By comparing the reprogramming-capable and incapable transformed lines we identified multiple stages of failure in the reprogramming process. Some transformed lines failed at an early stage, whilst other lines seemed to progress through a conventional reprogramming process. Finally, we show that MEK inhibition overcomes one critical reprogramming barrier by indirectly suppressing a hyperacetylated active epigenetic state. This study reveals that diverse epigenetic barriers underly resistance to reprogramming of transformed cells.
Large-scale volumetric medical images with annotation are rare, costly, and time prohibitive to acquire. Self-supervised learning (SSL) offers a promising pre-training and feature extraction solution for many downstream tasks, as it only uses unlabeled data. Recently, SSL methods based on instance discrimination have gained popularity in the medical imaging domain. However, SSL pre-trained encoders may use many clues in the image to discriminate an instance that are not necessarily disease-related. Moreover, pathological patterns are often subtle and heterogeneous, requiring the ability of the desired method to represent anatomy-specific features that are sensitive to abnormal changes in different body parts. In this work, we present a novel SSL framework, named DrasCLR, for 3D lung CT images to overcome these challenges. We propose two domain-specific contrastive learning strategies: one aims to capture subtle disease patterns inside a local anatomical region, and the other aims to represent severe disease patterns that span larger regions. We formulate the encoder using conditional hyper-parameterized network, in which the parameters are dependant on the anatomical location, to extract anatomically sensitive features. Extensive experiments on large-scale datasets of lung CT scans show that our method improves the performance of many downstream prediction and segmentation tasks. The patient-level representation improves the performance of the patient survival prediction task. We show how our method can detect emphysema subtypes via dense prediction. We demonstrate that fine-tuning the pre-trained model can significantly reduce annotation efforts without sacrificing emphysema detection accuracy. Our ablation study highlights the importance of incorporating anatomical context into the SSL framework. Our codes are available at https://github.com/batmanlab/DrasCLR.
Somatic cell reprogramming and oncogenic transformation share surprisingly similar features, yet transformed cells are highly resistant to reprogramming. There must be barriers that block transformed cells from reprogramming, but the nature of those barriers is unclear. In this study, we generated a systematic panel of transformed mouse embryonic fibroblasts (MEFs) using a variety of oncogenic transgenes, and discovered transformed cell lines that remain compatible with reprogramming when transfected with Oct4/Sox2/Klf4/Myc. By comparing the reprogramming-capable and incapable transformed lines we identified multiple stages of failure in the reprogramming process. Some transformed lines failed very early, whilst other lines seemed to progress through a normal-looking reprogramming process. Finally, we show that MEK inhibition overcomes one critical reprogramming barrier by indirectly suppressing a hyperactive epigenetic state in some of the transformed cells. This study reveals that the barriers underlying resistance to reprogramming vary between the different transformation methods. Key findings Somatic cell reprogramming of transformed cells is context-specific Inhibition of MEK converts some cell lines to reprogramming-capable Transformed cell lines are characterized by a hyperactive chromatin state MEK inhibition indirectly affects chromatin to enable reprogramming
Around 60% of in vitro fertilized (IVF) human embryos irreversibly arrest before compaction between the 3- to 8-cell stage, posing a significant clinical problem. The mechanisms behind this arrest are unclear. Here, we show that the arrested embryos enter a senescent-like state, marked by cell cycle arrest, the down-regulation of ribosomes and histones and down-regulation of MYC and p53 activity. The arrested embryos can be divided into 3 types. Type I embryos fail to complete the maternal-zygotic transition, and Type II/III embryos have low levels of glycolysis and either high (Type II) or low (Type III) levels of oxidative phosphorylation. Treatment with the SIRT agonist resveratrol or nicotinamide riboside (NR) can partially rescue the arrested phenotype, which is accompanied by changes in metabolic activity. Overall, our data suggests metabolic and epigenetic dysfunctions underlie the arrest of human embryos.
Generative Adversarial Networks (GAN) have many potential medical imaging applications, including data augmentation, domain adaptation, and model explanation.Due to the limited memory of Graphical Processing Units (GPUs), most current 3D GAN models are trained on low-resolution medical images, these models either cannot scale to high-resolution or are prone to patchy artifacts.In this work, we propose a novel end-to-end GAN architecture that can generate high-resolution 3D images.We achieve this goal by using different configurations between training and inference.During training, we adopt a hierarchical structure that simultaneously generates a low-resolution version of the image and a randomly selected sub-volume of the high-resolution image.The hierarchical design has two advantages: First, the memory demand for training on high-resolution images is amortized among sub-volumes.Furthermore, anchoring the high-resolution sub-volumes to a single low-resolution image ensures anatomical consistency between sub-volumes.During inference, our model can directly generate full high-resolution images.We also incorporate an encoder with a similar hierarchical structure into the model to extract features from the images.Experiments on 3D thorax CT and brain MRI demonstrate that our approach outperforms state of the art in image generation.We also demonstrate clinical applications of the proposed model in data augmentation and clinical-relevant feature extraction.
Rationale Chronic obstructive pulmonary disease (COPD) is characterized by pathologic changes in the airways, lung parenchyma, and persistent inflammation, but the links between lung structural changes and patterns of systemic inflammation have not been fully described. Objectives To identify novel relationships between lung structural changes measured by chest computed tomography (CT) and systemic inflammation measured by blood RNA sequencing. Methods CT scan images and blood RNA-seq gene expression from 1,223 subjects in the COPDGene study were jointly analyzed using deep learning to identify shared aspects of inflammation and lung structural changes that we refer to as Image-Expression Axes (IEAs). We related IEAs to COPD-related measurements and prospective health outcomes through regression and Cox proportional hazards models and tested them for biological pathway enrichment. Measurements and Main Results We identified two distinct IEAs: IEAemph captures an emphysema-predominant process with a strong positive correlation to CT emphysema and a negative correlation to FEV1 and Body Mass Index (BMI); IEAairway captures an airway-predominant process with a positive correlation to BMI and airway wall thickness and a negative correlation to emphysema. Pathway enrichment analysis identified 29 and 13 pathways significantly associated with IEAemph and IEAairway, respectively (adjusted p<0.001). Conclusions Integration of CT scans and gene expression data identified two IEAs that capture distinct inflammatory processes associated with emphysema and airway-predominant COPD. Scientific Knowledge on the Subject Chronic obstructive pulmonary disease (COPD) is characterized by lung structural changes and has a prominent systemic inflammatory component, but the links between lung structural changes and patterns of systemic inflammation in COPD have not been fully described. What This Study Adds to the Field We identified novel relationships between lung structural changes and systemic inflammation by simultaneously analyzing CT scans and blood RNA-sequencing gene expression using deep learning models. We identified two distinct Image-Expression Axes (IEAs) that characterize different inflammatory processes associated with emphysema and airway predominant COPD. This article has an online data supplement, which is accessible from this issue’s table of content online at [www.atsjournals.org][1]. ### Competing Interest Statement Peter J. Castaldi has received grant support from Bayer and consulting fees from Novartis and GSK. Craig P. Hersh reports grant support from Bayer, Boehringer-Ingelheim, and Vertex, and consulting fees from AstraZeneca and Takeda. Edwin K. Silverman has received grant support from Bayer and GSK. ### Funding Statement This work was supported by NHLBI K08 HL141601, R01 HL124233, R01 HL126596, R01 HL147326, U01 HL089897, and U01 HL089856. The COPDGene study ([NCT00608764][2]) is also supported by the COPD Foundation through contributions made to an Industry Advisory Committee comprised of AstraZeneca, Bayer Pharmaceuticals, Boehringer-Ingelheim, Genentech, GlaxoSmithKline, Novartis, Pfizer and Sunovion. ### Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes The details of the IRB/oversight body that provided approval or exemption for the research described are given below: Institutional review board (IRB) approval was obtained. IRB Protocol Title: Genetic Epidemiology of COPD. IRB Protocol Number: Brigham and Women's Hospital / 2007P000554. I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines and uploaded the relevant EQUATOR Network research reporting checklist(s) and other pertinent material as supplementary files, if applicable. Yes All data produced are available online at dbGaP [1]: http://www.atsjournals.org [2]: /lookup/external-ref?link_type=CLINTRIALGOV&access_num=NCT00608764&atom=%2Fmedrxiv%2Fearly%2F2022%2F10%2F14%2F2022.09.26.22280242.atom
Although self-supervised learning enables us to bootstrap the training by exploiting unlabeled data, the generic self-supervised methods for natural images do not sufficiently incorporate the context. For medical images, a desirable method should be sensitive enough to detect deviation from normal-appearing tissue of each anatomical region; here, anatomy is the context. We introduce a novel approach with two levels of self-supervised representation learning objectives: one on the regional anatomical level and another on the patient-level. We use graph neural networks to incorporate the relationship between different anatomical regions. The structure of the graph is informed by anatomical correspondences between each patient and an anatomical atlas. In addition, the graph representation has the advantage of handling any arbitrarily sized image in full resolution. Experiments on large-scale Computer Tomography (CT) datasets of lung images show that our approach compares favorably to baseline methods that do not account for the context. We use the learned embedding for staging lung tissue abnormalities related to COVID-19.
Knowledge distillation has been used to capture the knowledge of a teacher model and distill it into a student model with some desirable characteristics such as being smaller, more efficient, or more generalizable. In this paper, we propose a framework for distilling the knowledge of a powerful discriminative model such as a neural network into commonly used graphical models known to be more interpretable (e.g., topic models, autoregressive Hidden Markov Models). Posterior of latent variables in these graphical models (e.g., topic proportions in topic models) is often used as feature representation for predictive tasks. However, these posterior-derived features are known to have poor predictive performance compared to the features learned via purely discriminative approaches. Our framework constrains variational inference for posterior variables in graphical models with a similarity preserving constraint. This constraint distills the knowledge of the discriminative model into the graphical model by ensuring that input pairs with (dis)similar representation in the teacher model also have (dis)similar representation in the student model. By adding this constraint to the variational inference scheme, we guide the graphical model to be a reasonable density model for the data while having predictive features which are as close as possible to those of a discriminative model. To make our framework applicable to a wide range of graphical models, we build upon the Automatic Differentiation Variational Inference (ADVI), a black-box inference framework for graphical models. We demonstrate the effectiveness of our framework on two real-world tasks of disease subtyping and disease trajectory modeling.