The continued expansion in size and resolution of volumetric datasets generated by electron microscopy (3DEM) is making high-performance computing (HPC) an essential component. HPC provides the scalable, parallelized infrastructure required to process and analyze these increasingly large datasets. The Texas Advanced Computing Center’s (TACC) Core Experience Portal (CEP) is a science gateway specialized for leveraging HPC. The CEP can be adapted into customized versions based on a research community’s needs. Here we have constructed 3dem.org as a gateway for the community exploring volumetric data generated by electron microscopy (3DEM). This gateway provides the 3DEM community with user-friendly browser-based access to raw and processed datasets including both private and shared data, image processing tools for alignment and segmentation, and simulation and analysis environments. All this is linked to the underlying HPC environment at TACC with the ability to connect to other data storage and compute systems. 3dem.org bridges advanced electron microscopy with HPC, providing the research community with scalable, accessible infrastructure for discovery.
Abstract Background. ClinicalTrials.gov lists >500,000 studies, raising the question of how often “new” trials are truly novel versus incremental design variants. Lung cancer remains a leading cause of cancer-related death worldwide, underscoring the need for innovative over incremental trials. Prior work has used clinical trial knowledge graphs (KGs) to support design recommendations. Here, we apply a small cell lung cancer (SCLC) patient-anchored KG to classify lung cancer targeted therapy trials as “novel additions” versus “aggregate similarity” and to identify gaps where new, need-aligned trial designs are warranted. Methods. We curated 286 adult lung cancer targeted therapy trials from ClinicalTrials.gov that were retrieved as matches to a SCLC case. We built a trial-to-trial KG by linking each trial to key entities (e.g., tumor type, genomic alterations, targets/pathways, drug classes) and learned graph-based embeddings to obtain trial-level representations. Cosine similarity was used to quantify trial similarity between embeddings; community detection was used to define trial clusters. Clusters with ≥5 trials and median within-cluster similarity ≥0.80 were labeled “aggregate similar.” A novelty score (1 − maximum similarity to any other trial) and betweenness centrality were combined to label “novel additions”. A predefined set of SCLC molecular report-matched trials was cross-referenced to characterize their distribution across identified clusters. Results. Altogether, we defined 4 trial clusters (median size 67), three of which met our criteria for aggregate similarity, comprising 74% (n=212) of large community trials of highly similar designs. Across all 286 trials, median pairwise (0.74) and nearest-neighbor similarity (0.98) were consistent with dense replication of existing biomarker- and line-of-therapy-defined templates. Only 38 trials (13%) were classified as novel additions based on combined novelty (median=0.070) and betweenness centrality (median 0.004) thresholds, and showed significantly (p < 0.01) higher median scores in both as compared to non-novel trials. Novel additions were enriched in a single aggregate similarity cluster (24/74 trials) that contained most case molecular report-matched trials (65/68). The remaining trial clusters had 1-10% novel additions, indicating that case-level matching occurs within dense targeted therapy communities, whereas a KG analysis can surface structurally distinctive, novel addition trials. Conclusions. KG analysis of lung cancer trials provides a principled way to operationalize “novel addition” versus “aggregate similarity” at the portfolio level. This patient-anchored framework can support sponsors, investigators, and regulators in prioritizing innovative trials, reducing redundancy in an already crowded lung cancer trial landscape, and ultimately aligning trial development with unmet patient needs. Citation Format: Mahitha Simhambhatla, Maya Ylagan, Praneeth Sajja, Daruka Mahadevan, Erik S. Ferlanti, James Carson, Boone Goodgame, Ehsan Irajizad, Samir M. Hanash, Jeanne Kowalski. A knowledge graph screen of innovative versus incremental lung cancer trials for need-aligned designs [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2026; Part 1 (Regular Abstracts); 2026 Apr 17-22; San Diego, CA. Philadelphia (PA): AACR; Cancer Res 2026;86(7 Suppl):Abstract nr 6865.
The lung is a vital organ that undergoes extensive morphological and functional changes during postnatal development. To disambiguate how different cell populations contribute to organ development, we performed proteomic and transcriptomic analyses of four sorted cell populations from the lung of human subjects 0-8 years of age with a focus on early life. The cell populations analyzed included epithelial, endothelial, mesenchymal, and immune cells. Our results revealed distinct molecular signatures for each of the sorted cell populations that enable the description of molecular shifts occurring in these populations during postnatal development. We confirmed that the proteome of the different cell populations was distinct regardless of age and identified functions specific to each population. We identified a series of cell population protein markers, including those located at the cell surface, that show differential expression and distribution on RNA in situ hybridization and immunofluorescence imaging. We validated the spatial distribution of alveolar type 1 and endothelial cell surface markers. Temporal analyses of the proteomes of the four populations revealed processes modulated during postnatal development and clarified the findings obtained from whole-tissue proteome studies. Finally, the proteome was compared with a transcriptomics survey performed on the same lung samples to evaluate processes under post-transcriptional control.
Bronchopulmonary dysplasia (BPD) is a neonatal lung disease characterized by inflammation and scarring leading to long-term tissue damage. Previous whole tissue proteomics identified BPD-specific proteome changes and cell type shifts. Little is known about the proteome-level changes within specific cell populations in disease. Here, we sorted epithelial (EPI) and endothelial (ENDO) cell populations based on their differential surface markers from normal and BPD human lungs. Using a low-input compatible sample preparation method (MicroPOT), proteins were extracted and digested into peptides and subjected to liquid chromatography-tandem mass spectrometry (LC-MS/MS) proteome analysis. Of the 4,970 proteins detected, 293 were modulated in abundance or detection in the EPI population and 422 were modulated in ENDO cells. Modulation of proteins associated with actin-cytoskeletal function, such as SCEL, LMO7, and TBA1B was observed in the BPD EPIs. Using confocal imaging and analysis, we validated the presence of aberrant multilayer-like structures comprising SCEL and LMO7, known to be associated with epidermal cornification, in the human BPD lung. This is the first report of the accumulation of cornification-associated proteins in BPD. Their localization in the alveolar parenchyma, primarily associated with alveolar type 1 (AT1) cells, suggests a role in the BPD postinjury response. In the ENDOs, redox balance and mitochondrial function pathways were modulated. Alternative mRNA splicing and cell proliferative functions were elevated in both populations, suggesting potential dysregulation of cell progenitor fate. This study characterized the proteome of epithelial and endothelial cells from the BPD lung for the first time, identifying population-specific changes in BPD pathogenesis.NEW & NOTEWORTHY The study is the first to perform proteomics on sorted pulmonary epithelial and endothelial populations from bronchopulmonary dysplasia (BPD) and age-matched control human donors. We identified an increase in cornification-associated proteins in BPD (e.g., SCEL and LMO7), and evidenced the presence of multilayered structures unique to BPD alveolar regions, associated with alveolar type 1 (AT1) cells. By changing the nature and/or biomechanical properties of the epithelium, these structures may alter the behavior of other alveolar cell types potentially contributing to the arrested alveolarization observed in BPD. Finally, our data suggest the modulation of cell proliferation and redox homeostasis in BPD providing potential mechanisms for the reduced vascular growth associated with BPD.
An improved understanding of the human lung necessitates advanced systems models informed by an ever-increasing repertoire of molecular omics, cellular imaging, and pathological datasets. To centralize and standardize information across broad lung research efforts, we expanded the LungMAP.net website into a new gateway portal. This portal connects a broad spectrum of research networks, bulk and single-cell multiomics data, and a diverse collection of image data that span mammalian lung development and disease. The data are standardized across species and technologies using harmonized data and metadata models that leverage recent advances, including those from the Human Cell Atlas, diverse ontologies, and the LungMAP CellCards initiative. To cultivate future discoveries, we have aggregated a diverse collection of single-cell atlases for multiple species (human, rhesus, and mouse) to enable consistent queries across technologies, cohorts, age, disease, and drug treatment. These atlases are provided as independent and integrated queryable datasets, with an emphasis on dynamic visualization, figure generation, reanalysis, cell-type curation, and automated reference-based classification of user-provided single-cell genomics datasets (Azimuth). As this resource grows, we intend to increase the breadth of available interactive interfaces, supported data types, data portals and datasets from LungMAP, and external research efforts.
The Human BioMolecular Atlas Program (HuBMAP) aims to create a multi-scale spatial atlas of the healthy human body at single-cell resolution by applying advanced technologies and disseminating resources to the community. As the HuBMAP moves past its first phase, creating ontologies, protocols and pipelines, this Perspective introduces the production phase: the generation of reference spatial maps of functional tissue units across many organs from diverse populations and the creation of mapping tools and infrastructure to advance biomedical research. The Human BioMolecular Atlas Program (HuBMAP) presents its production phase: the generation of spatial maps of functional tissue units across organs from diverse populations and the creation of tools and infrastructure to advance biomedical research.
Neuroscience research has expanded dramatically over the past 30 years by advancing standardization and tool development to support rigor and transparency. Consequently, the complexity of the data pipeline has also increased, hindering access to FAIR data analysis to portions of the worldwide research community. brainlife.io was developed to reduce these burdens and democratize modern neuroscience research across institutions and career levels. Using community software and hardware infrastructure, the platform provides open-source data standardization, management, visualization, and processing and simplifies the data pipeline. brainlife.io automatically tracks the provenance history of thousands of data objects, supporting simplicity, efficiency, and transparency in neuroscience research. Here brainlife.io's technology and data services are described and evaluated for validity, reliability, reproducibility, replicability, and scientific utility. Using data from 4 modalities and 3,200 participants, we demonstrate that brainlife.io's services produce outputs that adhere to best practices in modern neuroscience research.
Virtual screening is a key step of the drug discovery process which utilizes computational resources to simulate the behavior of small molecules in the binding site of a target protein. [13] Researchers often test millions of molecules when searching for an early hit compound, requiring significant CPU hours. An accessible, convenient, fast, and computationally efficient means for virtual screening is desirable in order for researchers to conserve resources in the early phase of drug discovery. We developed an application programming interface (API) integrated workflow that allows researchers to submit virtual screening batch jobs to the Lonestar6 supercomputer through a web portal. The containerized [7] workflow employs parallelized Python scripting using mpi4py [2] to efficiently distribute molecular docking tasks performed by AutoDock Vina. [3] The Texas Advanced Computing Center (TACC) API (TAPIS) framework [16], a REST API framework for research computing, was used to integrate the workflow into the University of Texas System Research Cyberinfrastructure (UTRC) web portal. [12] Five large libraries representing commercially-available small molecules or fragments were prepared and are available for screening. Here, we discuss our experience developing this service, as well as the results of extensive internal benchmarks to determine the most efficient parallelization scheme to employ for each molecule library when submitting batch jobs. Regardless of the chosen ligand library, the core, node, and parallel task specifications allow the user to run a virtual drug screening and receive their resulting top docking scores in 24 hours. The service is available to registered academic users, and more information can be found at the Drug Discovery at TACC website. [18]
Human disease states are biomolecularly multifaceted and can span across phenotypic states, therefore it is important to understand diseases on all levels, across cell types, and within and across microanatomical tissue compartments. To obtain an accurate and representative view of the molecular landscape within human lungs, this fragile tissue must be inflated and embedded to maintain spatial fidelity of the location of molecules and minimize molecular degradation for molecular imaging experiments. Here, we evaluated agarose inflation and carboxymethyl cellulose embedding media and determined effective tissue preparation protocols for performing bulk and spatial mass spectrometry-based omics measurements. Mass spectrometry imaging methods were optimized to boost the number of annotatable molecules in agarose inflated lung samples. This optimized protocol permitted the observation of unique lipid distributions within several airway regions in the lung tissue block. Laser capture microdissection of these airway regions followed by high-resolution proteomic analysis allowed us to begin linking the lipidome with the proteome in a spatially resolved manner, where we observed proteins with high abundance specifically localized to the airway regions. We also compared our mass spectrometry results to lung tissue samples preserved using two other inflation/embedding media, but we identified several pitfalls with the sample preparation steps using this preservation method. Overall, we demonstrated the versatility of the inflation method, and we can start to reveal how the metabolome, lipidome, and proteome are connected spatially in human lungs and across disease states through a variety of different experiments.
Rationale: The current understanding of human lung development derives mostly from animal studies. Although transcript-level studies have analyzed human donor tissue to identify genes expressed during normal human lung development, protein-level analysis that would enable the generation of new hypotheses on the processes involved in pulmonary development are lacking. Objectives: To define the temporal dynamic of protein expression during human lung development. Methods: We performed proteomics analysis of human lungs at 10 distinct times from birth to 8 years to identify the molecular networks mediating postnatal lung maturation. Measurements and Main Results: We identified 8,938 proteins providing a comprehensive view of the developing human lung proteome. The analysis of the data supports the existence of distinct molecular substages of alveolar development and predicted the age of independent human lung samples, and extensive remodeling of the lung proteome occurred during postnatal development. Evidence of post-transcriptional control was identified in early postnatal development. An extensive extracellular matrix remodeling was supported by changes in the proteome during alveologenesis. The concept of maturation of the immune system as an inherent part of normal lung development was substantiated by flow cytometry and transcriptomics. Conclusions: This study provides the first in-depth characterization of the human lung proteome during development, providing a unique proteomic resource freely accessible at Lungmap.net. The data support the extensive remodeling of the lung proteome during development, the existence of molecular substages of alveologenesis, and evidence of post-transcriptional control in early postnatal development.
The Identifier Services (IDS) project conducted research into and built a prototype to manage distributed genomics datasets remotely and over time. Inspired by archival concepts, IDS allows researchers to track dataset evolution through multiple copies, modifications, and derivatives, independent of where data are located – both symbolically, in the research lifecycle, and physically, in a repository or storage facility. The prototype implementation is based on a three-step data modeling process involving: a) understanding and recording of different researcher workflows, b) mapping the workflows and data to a generic data model and identifying functions, and c) integrating the data model as architecture and interactive functions into cyberinfrastructure (CI). Identity functions are operationalized as continuous tracking of authenticity attributes including data location, differences between seemingly identical datasets, metadata, data integrity, and the roles of different types of local and global identifiers used during the research lifecycle. CI resources were used to conduct identity functions at scale, including scheduling content comparison tasks on high-performance computing resources. The prototype was developed and evaluated considering six data test cases, and feedback was received through a focus-group activity. While there are some technical roadblocks to overcome, our project demonstrates that identity functions are innovative solutions to manage large distributed genomic datasets.
Lipids are a naturally occurring group of molecules that not only contribute to the structural integrity of the lung preventing alveolar collapse but also play important roles in the anti-inflammatory responses and antiviral protection. Alteration in the type and spatial localization of lipids in the lung plays a crucial role in various diseases, such as respiratory distress syndrome (RDS) in preterm infants and oxidative stress-influenced diseases, such as pneumonia, emphysema, and lung cancer following exposure to environmental stressors. The ability to accurately measure spatial distributions of lipids and metabolites in lung tissues provides important molecular insights related to lung function, development, and disease states. Nanospray desorption electrospray ionization (nano-DESI) and other ambient ionization mass spectrometry techniques enable label-free imaging of complex samples in their native state with minimal to absolutely no sample preparation. However, lipid coverage obtained in nano-DESI mass spectrometry imaging (MSI) experiments has not been previously characterized. In this work, the depth of lipid coverage in nano-DESI MSI of mouse lung tissues was compared to liquid chromatography tandem mass spectrometry (LC-MS/MS) lipidomics analysis of tissue extracts prepared using two different procedures: standard Folch extraction method of the whole lung samples and extraction into a 90% methanol/10% water mixture used in nano-DESI MSI experiments. A combination of positive and negative ionization mode nano-DESI MSI identified 265 unique lipids across 20 lipids subclasses and 19 metabolites (284 in total) in mouse lung tissues. Except for triacylglycerols (TG) species, nano-DESI MSI provided comparable coverage to LC-MS/MS experiments performed using methanol/water tissue extracts and up to 50% coverage in comparison with the Folch extraction-based whole lung lipidomics analysis. These results demonstrate the utility of nano-DESI MSI for comprehensive spatially resolved analysis of lipids in tissue sections. A combination of nano-DESI MSI and LC-MS/MS lipidomics is particularly useful for exploring changes in lipid distributions during lung development, as well as resulting from disease or exposure to environmental toxicants.
Cell type-resolved proteome analyses of the brain, heart and liver have been reported, however a similar effort on the lipidome is currently lacking. Here we applied liquid chromatography-tandem mass spectrometry to characterize the lipidome of major lung cell types isolated from human donors, representing the first lipidome map of any organ. We coupled this with cell type-resolved proteomics of the same samples (available at Lungmap.net). Complementary proteomics analyses substantiated the functional identity of the isolated cells. Lipidomics analyses showed significant variations in the lipidome across major human lung cell types, with differences most evident at the subclass and intra-subclass (i.e. total carbon length of the fatty acid chains) level. Further, lipidomic signatures revealed an overarching posture of high cellular cooperation within the human lung to support critical functions. Our complementary cell type-resolved lipid and protein datasets serve as a rich resource for analyses of human lung function.
Lung diseases and disorders are a leading cause of death among infants. Many of these diseases and disorders are caused by premature birth and underdeveloped lungs. In addition to developmentally related disorders, the lungs are exposed to a variety of environmental contaminants and xenobiotics upon birth that can cause breathing issues and are progenitors of cancer. In order to gain a deeper understanding of the developing lung, we applied an activity-based chemoproteomics approach for the functional characterization of the xenometabolizing cytochrome P450 enzymes, active ATP and nucleotide binding enzymes, and serine hydrolases using a suite of activity-based probes (ABPs). We detected P450 activity primarily in the postnatal lung; using our ATP-ABP, we characterized a wide range of ATPases and other active nucleotide- and nucleic acid-binding enzymes involved in multiple facets of cellular metabolism throughout development. ATP-ABP targets include kinases, phosphatases, NAD- and FAD-dependent enzymes, RNA/DNA helicases, and others. The serine hydrolase-targeting probe detected changes in the activities of several proteases during the course of lung development, yielding insights into protein turnover at different stages of development. Select activity-based probe targets were then correlated with RNA in situ hybridization analyses of lung tissue sections.
This data is a curated collection of visual images of gene expression patterns from the pre- and post-natal mouse lung, accompanied by associated mRNA probe sequences and RNA-Seq expression profiles. Mammalian lungs undergo significant growth and cellular differentiation before and after the transition to breathing air. Documenting normal lung development is an important step in understanding abnormal lung development, as well as the challenges faced during a preterm birth. Images in this dataset indicate the spatial distribution of mRNA transcripts for over 500 different genes that are active during lung development, as initially determined via RNA-Seq. Images were systematically acquired using high-throughput in situ hybridization with non-radioactive digoxigenin-labeled mRNA probes across mouse lungs from developmental time points E16.5, E18.5, P7, and P28. The dataset was produced as part of The Molecular Atlas of Lung Development Program (LungMAP) and is hosted at https://lungmap.net. This manuscript describes the nature of the data and the protocols for generating the dataset.
Lung immaturity is a major cause of morbidity and mortality in premature infants. Understanding the molecular mechanisms driving normal lung development could provide insights on how to ameliorate disrupted development. While transcriptomic and proteomic analyses of normal lung development have been previously reported, characterization of changes in the lipidome is lacking. Lipids play significant roles in the lung, such as dipalmitoylphosphatidylcholine in pulmonary surfactant; however, many of the roles of specific lipid species in normal lung development, as well as in disease states, are not well defined. In this study, we used liquid chromatography-mass spectrometry (LC-MS/MS) to investigate the murine lipidome during normal postnatal lung development. Lipidomics analysis of lungs from post-natal day 7, day 14 and 6-8 week mice (adult) identified 924 unique lipids across 21 lipid subclasses, with dramatic alterations in the lipidome across developmental stages. Our data confirmed previously recognized aspects of post- natal lung development and revealed several insights, including in sphingolipid-mediated apoptosis, inflammation and energy storage/usage. Complementary proteomics, metabolomics and chemical imaging corroborated these observations. This multi-omic view provides a unique resource and deeper insight into normal pulmonary development.
The National Heart, Lung, and Blood Institute is funding an effort to create a molecular atlas of the developing lung (LungMAP) to serve as a research resource and public education tool. The lung is a complex organ with lengthy development time driven by interactive gene networks and dynamic cross talk among multiple cell types to control and coordinate lineage specification, cell proliferation, differentiation, migration, morphogenesis, and injury repair. A better understanding of the processes that regulate lung development, particularly alveologenesis, will have a significant impact on survival rates for premature infants born with incomplete lung development and will facilitate lung injury repair and regeneration in adults. A consortium of four research centers, a data coordinating center, and a human tissue repository provides high-quality molecular data of developing human and mouse lungs. LungMAP includes mouse and human data for cross correlation of developmental processes across species. LungMAP is generating foundational data and analysis, creating a web portal for presentation of results and public sharing of data sets, establishing a repository of young human lung tissues obtained through organ donor organizations, and developing a comprehensive lung ontology that incorporates the latest findings of the consortium. The LungMAP website (www.lungmap.net) currently contains more than 6,000 high-resolution lung images and transcriptomic, proteomic, and lipidomic human and mouse data and provides scientific information to stimulate interest in research careers for young audiences. This paper presents a brief description of research conducted by the consortium, database, and portal development and upcoming features that will enhance the LungMAP experience for a community of users.