In recent years, high dimensional omics analyses have become more commonplace for investigating complex biological systems. Typically, these studies attempt to identify key biomolecules associated with a particular biological process. Often, machine learning (ML) is used to identify these biomolecules, typically by learning which biomolecules are highly predictive of a treatment, biological outcome, or phenotype. A major challenge of applying ML to high throughput omics is overcoming noise when sample size is limited and unbalanced with respect to tens of thousands of biomolecules measured. Thus, feature selection (the process of reducing the number of predictors) is both a critical and common step in the ML analysis pipeline. While much attention has been given to embedding and wrapping techniques for feature selection in the omics space, filter-based methods for model-free feature selection have appealing theoretical properties. This manuscript evaluates sure screening, a class of filter-based feature selection methods which provide analytical guarantees for true feature set retention. Here, we cover existing feature screening methods based on the sure screening principal, available software, methods to improve feature screening, and contextualize feature screening in the larger discussion of feature selection for omics data analysis. Additionally, a suite of model-free sure screening approaches is applied and compared for several omics biomedical applications in a ML classification context. We identified BcorSIS as the most effective and computationally efficient screening method across various omics datasets, consistently outperforming others like CSIS and DCSIS in runtime.
Type 1 diabetes (T1D) results from dysfunction and loss of the insulin-producing pancreatic β cells. The body’s lipid metabolism is strongly regulated during this process but there is a need to understand how this regulation contributes to the β-cell death. Here, we investigated the role of free fatty acids (FFAs) in T1D development. Lipidomics data from the case-control study of The Determinants of Diabetes in the Young (TEDDY) consortium were re-analyzed to determine temporal changes in lipid profiles during T1D development. Fatty acid distribution across the pancreas was measured by mass spectrometry imaging. To model islet inflammation, FFAs were measured in human islets treated with the pro-inflammatory cytokines IL-1β + IFNγ by gas chromatography-mass spectrometry. Further, the effects of FFAs on MIN6 insulin-producing cells were measured by proteomics analysis along with biochemical and cell biology assays. We investigated if similar effects occur in vivo during T1D on the islets’ single-cell RNA sequencing data from Human Pancreas Analysis Program (HPAP). During islet autoimmunity, prior to T1D onset, plasma phosphatidylcholine and triacylglycerols are reduced with a simultaneous increase of FFAs. Similarly in human islets, IL-1β + IFNγ induce the increase of palmitate. Moreover, FFAs are abundantly detected across islets and surrounding exocrine tissue. FFAs synergistically enhanced cytokine-mediated apoptosis in MIN6 cells by downregulating the production of nicotinamide adenosine dinucleotide (NAD) via downregulating nicotinamide phosphoribosyltransferase (NAMPT), a rate-limiting enzyme of the NAD salvage pathway. The enhancement of cytokine-mediated apoptosis was reverted by supplementing cells with nicotinic acid and nicotinamide mononucleotide, metabolites that bypass NAMPT in NAD biosynthesis. NAMPT downregulation was further observed during T1D development, supporting that NAD production might be compromised in vivo. Our findings show that fatty acids are released during islet autoimmunity. These fatty acids enhance pro-inflammatory cytokine-mediated apoptosis through impaired NAD metabolism.
Plasma extracellular vesicles (EVs) are considered excellent sources for biomarker discovery since they carry signatures of their cellular origin and disease processes. In this paper, we evaluate the potential of plasma EV proteomics analysis for identifying predictive biomarkers of developing type 1 diabetes (T1D), which results from autoimmune destruction of insulin-producing β cells in the islet. We used strong anion exchange beads (Mag-Net) to capture plasma EVs from 19 donors with islet autoimmunity (diagnosed by circulating autoantibodies against islet proteins-AAB+) versus 17 control individuals and analyzed their protein cargo by mass spectrometry. The analysis identified and quantified 5,480 proteins, a 3.2-fold increase in proteome coverage compared to our previous T1D biomarker proteomics study that used whole plasma depleted of the 14 most abundant proteins. The Mag-Net approach also detected 1306 out of the 1717 proteins (76%) that we previously verified as EV proteins. Statistical tests revealed 448 proteins to be differentially abundant in AAB+ versus control volunteers, including 69 previously verified EV proteins. A functional-enrichment analysis resulted in overrepresentation of 25 pathways among the differentially abundant proteins, including pathways related to autoimmune response and lipid metabolism. The capacity of this data to predict AAB+ was tested with a machine learning analysis using a random forest model, resulting in a receiver operating characteristic-area under the curve of 0.81. Overall, our study indicates that plasma EV proteomics analysis can be an exciting approach for studying biomarkers for developing T1D. SIGNIFICANCE OF THE STUDY: Type 1 diabetes (T1D) is a disease characterized by the body's inability to produce insulin and consequently, to control blood glucose levels. Despite the initial trigger being unclear, the disease development process involves an autoimmune response to the islets of Langerhans, resulting in the death of insulin-producing β cells. There is no cure for the disease, and treatment relies on exogenous administration of insulin. Therefore, preventive therapies that block the autoimmune process are attractive for treating T1D. In fact, anti-CD3 antibody (Teplizumab) delays the onset of T1D by 2 years by targeting T cells. Predictive biomarkers for developing T1D are needed to aid the development and implementation of new therapies and to identify the initial trigger and mechanisms of the islet autoimmune process. In this paper, we assess the potential of plasma extracellular vesicle (EV) proteomics analysis for identifying predictive biomarkers of T1D. Our results show excellent potential of the approach, opening opportunities to perform broader studies to identify biomarkers for developing T1D.
Graphical networks are useful, widely used modeling approaches to represent complex biological processes with biological measurements generated by platforms such as mass spectrometry. Bayesian analyses of graphical networks for omics data have several advantages over their frequentist counterparts such as the inclusion of prior knowledge in the estimation of models. However, Bayesian approaches to date have only been feasible for data with a couple of hundred biomolecules due to prohibitive computational time, but omics data often contain tens of thousands of biomolecules. Here, we present and illustrate a more computationally efficient approach named BPlane (Bayesian PseudoLikelihood-based Algorithm for Network Estimation) to extend Bayesian modeling capabilities for larger-sized data sets, such as most untargeted proteomics data. Via simulation, we demonstrate that BPlane produces substantial computational savings over a current state-of-the-art Bayesian algorithm while maintaining competitive edge detection accuracy. We further demonstrate the computational benefit of BPlane on a SARS-CoV-2 proteomics data set with 7000 proteins.
Ribosome inactivating proteins (RIPs) such as ricin and abrin depurinate an adenine base in the sarcin/ricin loop in the large ribosomal subunit, leading to inhibtion of protein synthesis and cell death. Here, we demonstrate that RIP toxin activity can be detected via nanopore-based DNA sequencing using synthetic oligonucleotide substrates. This is achieved by monitoring the mismatch proportion at the canonical target sequences incorporated into the synthetic substrate and determining the sequence length distribution throughout the entire substrate sequence. The mismatch proportion increases and sequence length distribution decreases with increasing toxin concentration for both ricin and abrin in buffer as well as in more complex backgrounds such as saliva and nasal secretions.
Reporting back of research results (RBRR) is becoming a recognized component of ethical community-engaged and human subjects research. Scholars emphasize the importance of involving participants in developing reports and methods of report-back, arguing that RBRR enhances comprehension, trust, and engagement with findings. Yet, despite growing recognition, standardized guidelines for ethical RBRR remain limited. To address this gap, we conducted a systematic literature review of peer-reviewed studies- primarily in genomics, environmental health, and biomonitoring- to identify RBRR development strategies. Secondarily, we assessed how these strategies may align with the bioethical principles of autonomy, beneficence, nonmaleficence, and justice to inform future RBRR design. A systematic search of Web of Science, PubMed, and Google Scholar yielded 2,164 records; after removing 748 duplicates, 1,416 unique studies were screened, and 32 met the inclusion criteria (peer-reviewed, published 2016–2024, written in English). Studies were required to include primary report-back (e.g., direct return of results to participants) or detailed descriptions of RBRR methods or recommendations. The review followed a PE/IO (Population, Exposure, Intervention, Outcome) framework and adhered to PRISMA guidelines. Risk of bias was assessed using the STROBE cohort study checklist and GRADE criteria. Across studies, RBRR was framed as an ethical obligation, and an opportunity to improve understanding of environmental influences on health. Most emphasized plain-language communication (n = 8), multimodal dissemination (n = 10), and culturally responsive design (n = 4). However, only three studies applied formal communication or evaluation frameworks, and only one-third described how materials were developed. Common evaluation methods included post-report surveys (n = 22), interviews (n = 18), and focus groups (n = 14), however, we noted a lack of consistency in evaluation methods. Collectively, these 32 studies underscored the importance of tailoring materials to population characteristics, providing multiple formats, and experimenting with visual and digital tools to enhance comprehension. Although none cited a bioethical framework, the core principles were reflected in practice: respect for autonomy through participants’ right to know their results; beneficence through the development of accessible, actionable materials; nonmaleficence through anticipating and mitigating anxiety or confusion; and justice through culturally and linguistically appropriate design. Yet, gaps remain, with inconsistent characterization of RBRR methods and limited evaluation of RBRR. This study was limited by the risk of bias in participant selection, as many studies included participants with prior interest in the field of environmental health or emotional investment in the studies. Ethical RBRR supports both individual and collective knowledge and action when paired with ongoing community engagement. Developing consistent, evidence-based best practices that balance feasibility with contextual relevance could strengthen trust, comprehension, and the translation of scientific findings into meaningful public-health action.
Immunosenescence, the age-associated decline in immune function, is a key feature of human aging. In human lymphoid organs, however, the specific immune cell populations that acquire senescence-associated phenotypes during aging and how they influence the surrounding tissue microenvironment remain poorly understood. A spatially resolved map of these senescence-associated immune states in human lymphoid tissues could help clarify their relationship with aging and their potential contributions to the progressive decline of immune function. Here, we integrated single-cell and spatial multi-omics to systematically characterize age-related senescence in human lymph nodes (LNs). Single-cell transcriptomics of lymphoid tissues from donors aged 18 to 100 years old identified 34 immune and stromal cell types and revealed age-associated upregulation of senescence signatures in specific populations. Spatial proteomic profiling of 99 LN sections from 51 donors (18-86 years) using high-plex immunofluorescence (∼20 million cells) mapped senescence markers (p16, p21, HMGB1, 𝛾 -H2AX) at single-cell resolution, revealing diverse senescent-like cell types ("senotypes") and a stepwise shift from extrafollicular to germinal center (GC) localization with age. Notably, we observed focal clonal-like senescence in GC B cells in older donor LNs. Spatial transcriptomics, epigenomics, and metabolic imaging of selected samples further elucidate the multi-omics signatures and underlying mechanisms of functional impairment, metabolic remodeling, and distinct regulatory programs in senescent-like GC B cells. This study presents a comprehensive spatial atlas of senescence-associated immune states in human lymph nodes, revealing cell-type-specific and spatial heterogeneity that may contribute to immunosenescence and the decline of immune function during aging.
Bottom-up proteomic workflows rely on sequential preprocessing steps, commonly including peptide-to-protein aggregation ("roll-up"), to enhance data reliability and interpretability. While roll-up is effective for protein-centered analyses, it may be suboptimal for applications focused on post-translational modifications (PTMs) or protein structural changes, such as limited proteolysis-mass spectrometry (LiP-MS). Here, we investigate how different roll-up strategies influence site-level quantification in PTM differential analysis. Moreover, we introduce a novel site-centric roll-up approach tailored for LiP-MS, which quantifies proteolytic fragments rather than solely tryptic peptides. We benchmark these methods through simulation studies, comparing their sensitivity and specificity in detecting structural and PTM-driven changes. We found that the median and mean roll-up methods outperform the sum method in both PTM and LiP proteomics, and site-level quantification in LiP outperforms peptide-level quantification. Our findings offer the first systematic, data-driven guidance for selecting roll-up techniques in site-level proteomic analyses, with implications for both PTM-focused and structural proteomics studies.
Immunosenescence is a hallmark of human aging and contributes to age-related immune decline, yet the development of senescence-associated phenotypes in the human lymphoid organs remains poorly understood. Here, we integrate single-cell and spatial multi-omics to systematically characterize age-related senescence in human lymph nodes (LNs) across the lifespan. Spatial proteomic profiling of 99 LN sections from 51 donors (18-86 years) using high-plex immunofluorescence (~20 million cells) mapped senescence markers (p16, p21, HMGB1, and γ-H2AX) at single-cell resolution, revealing diverse senescent-like cell types ("senotypes") and a stepwise shift from extrafollicular to germinal-center localization with age. In aged LNs, germinal-center B cells exhibit focal accumulation of senescence-associated programs, accompanied by impaired functional signatures, metabolic remodeling, and altered regulatory networks. These findings define a spatially organized landscape of immunosenescence in human lymphoid tissue and highlight germinal-center B cells as a key locus of age-associated immune dysfunction.
Expression-based omics technologies (e.g., proteomics, metabolomics, transcriptomics, etc.) increasingly rely on supervised and unsupervised machine learning (ML) models to find key biomolecules distinguishing conditions, identify natural groupings in biological data, or generate predictions for outcomes of interest. Fitting ML models to omics data presents several challenges, including handling missing data, selecting a normalization method, choosing a valid model, and optimizing hyperparameters, all requiring statistical programming skills to address these challenges. Thus, the open-source web application OmicsMLMentor was designed to lower the barrier to ML modeling for omics data. OmicsMLMentor supports the fitting of 15 ML models (10 supervised and 5 unsupervised) tailored to omics data sets, such as proteomics, metabolomics, lipidomics, and transcriptomics. OmicsMLMentor offers several omics-specific features, including methods for handling missingness (imputation, conversion, removal), normalization tests, ranking of models based on the structure of a user’s data and user input, and optimal hyperparameter selections using cross-validation splits. By streamlining ML workflows for omics analysis, OmicsMLMentor addresses critical gaps in existing online web tools, facilitating a broader adoption of these models for omics research. Here, OmicsMLMentor is applied to data from a lignin exposure study to highlight example workflows for fitting both supervised and unsupervised models to data.
Understanding condition-specific molecular interactions in complex biological systems requires scalable and accurate network inference from high-dimensional omics data, often generated by mass spectrometry platforms. Accurately estimating condition-specific interaction networks is critical for highlighting biological mechanisms, yet it remains challenging due to high dimensionality, noise, and the difficulty of accurately comparing networks across experimental groups. We propose a scalable and flexible Bayesian framework, clustering-focused iterative (CFI) estimation, for joint inference of Gaussian graphical models in multicondition omics data. CFI leverages hierarchical clustering of pooled data to identify consistent subnetwork structures, parallel Bayesian estimation within clusters for each condition, and an iterative merging step to recover intercluster dependencies without relying on restrictive assumptions or bridge variables. Through extensive simulations, we demonstrate that CFI achieves substantial computational gains with up to 64% average reduction in runtime relative to running the same network methods without CFI, while maintaining or improving accuracy compared to traditional approaches. The approach is scalable, with demonstrated improvement for large networks with thousands of nodes. We applied CFI to a mass spectrometry-based proteomics data set comparing host responses for samples subjected to mock and SARS-CoV-2 infection. Out of 6721 proteins, CFI identified 576 edges in the mock condition and 589 in the SARS-CoV-2 condition, 159 altered connections involving 63 proteins, reflecting substantial differences in networks between conditions. Enrichment analysis of these differential subnetworks revealed key biological pathways implicated in viral response and host regulation. These results illustrate CFI's ability to scale to realistic omics data, uncover condition-specific network rewiring, and provide interpretable biological insights.
Spatial omics technologies, such as mass spectrometry imaging (MSI), can capture biomolecular distributions and their spatial locations directly from a tissue, but these distributions are not easily associated with tissue morphology without additional microscopy performed on the sample. To identify tissue regions of interest (ROIs), a reference image (such as a stained tissue, microscopy images, etc.) may be used and paired with mass spectrometry data. Due to the intense time requirements of manually labeling ROIs, segmentation models have been demonstrated as time-saving alternatives, as they automatically annotate ROIs by grouping like pixel intensities together. Supervised approaches trained on specific image types are preferred, but when those cases are not available, generalizable unsupervised segmentation models may then be used. Here, eight unsupervised semantic segmentation algorithms in R and Python, representing both statistical and machine learning algorithms, were compared for their ability to match manual annotations of 30 tiled PAS-stained kidney images and 25 tiled plant root images. Noise reduction techniques such as dimension reduction were tested to see whether they improved segmentations, and all models were applied to the full stitched images. Performance metrics were calculated to provide recommendations on the highest performing models to demonstrate their potential for automated annotation of tissue ROIs. At small cluster sizes, k-means and pytorch-tip tended to have the best performance in terms of balanced accuracy and time, though all algorithms had decreased performance at higher cluster numbers. Lastly, the segmentation model choice was demonstrated to have an impact on downstream statistics, highlighting the importance of testing and selecting the best segmentation model on a case-by-case basis, as no one model had the best performance in every comparison.
Background: Amniotic fluid (AF) plays a key role in fetal development, yet the evolving composition of AF and its effects on hemostasis and thrombosis are poorly understood. Objectives: To characterize the procoagulant properties of AF as a function of gestation in humans and nonhuman primates. Methods: We analyzed the proteomes, lipidomes, and procoagulant properties of AF obtained by amniocentesis from rhesus macaque and human pregnancies at gestational age-matched time points. Results: When added to human plasma, both rhesus and human AF accelerated clotting time and fibrin generation. We identified proteomic modules associated with clotting time and enriched for coagulation-related pathways. Proteins known to be involved in hemostasis were highly correlated with each other, and their intensity of expression varied across gestation in both rhesus and humans. Inhibition of the contact pathway did not affect the procoagulant effect of AF. Blocking tissue factor pathway inhibitor reversed the ability of AF to block the generation of activated factor X. The prothrombinase activity of AF was inhibited by phospholipid inhibitors. The levels of phosphatidylserine in AF were inversely correlated with clotting time. AF promoted platelet activation and secretion in plasma. Conclusion: Overall, our findings reveal that the addition of AF to plasma enhances coagulation in a manner dependent on phospholipids as well as the presence of proteases and other proteins that directly regulate coagulation. We describe a correlation between clotting time and expression of coagulation proteins and phosphatidylserine in both rhesus and human AF, supporting the use of rhesus models for future studies of AF biology.
Americans spend approximately 90% of their time indoors, with more than 66% of that time spent in residential buildings. Factors pertaining to household behavior or environmental factors may influence types of semi-volatile organic compounds (SVOC) found indoors. Paired indoor and outdoor passive samplers were deployed at twenty-four locations across the United States. Samples were analyzed for >1500 SVOCs to identify common patterns in exposure profiles and investigate influences of household behavior and environmental factors. Unique differences between indoor and outdoor profiles were identified, with indoor air typically having greater frequency and concentration of SVOCs relative to outdoor air. A significant relationship between fragrance chemicals and scented consumer products was identified. When considering a multifactorial approach, chemical exposures were most influenced by environmental and demographic factors. Our data highlights specific groups of chemicals identified at higher concentrations indoors and their potential influences, as well as the complexity of identifying specific sources of chemical exposures.
The rapid evolution of SARS-CoV-2 has led to the emergence of numerous variants with enhanced transmissibility and immune evasion. Despite widespread vaccination, infections persist, and the mechanisms by which SARS-CoV-2 reprograms host metabolism remain incompletely understood. Here, we investigated whether virus-induced lipid remodeling is conserved across variants and whether changes in lipid abundance correlate with alterations in lipid biosynthetic enzymes. Using global untargeted lipidomics and quantitative proteomics, we analyzed A549-ACE2 cells infected with the Delta (B.1.617.2) or Omicron (B.1.1.529) variants and compared them to cells infected with the ancestral WA1 strain. In parallel, we conducted quantitative proteomics to assess virus-induced changes in the host proteome. Our results reveal that SARS-CoV-2 drives a remarkably consistent pattern of metabolic rewiring at both the lipidomic and proteomic levels across all three variants. We mapped changes in the expression of host metabolic enzymes and compared these to corresponding shifts in lipid abundance. This integrative analysis identified key host proteins involved in virus-mediated lipid remodeling, including fatty acid synthase (FASN), lysosomal acid lipase (LIPA), and ORM1-like protein 2 (ORMDL2). Together, these findings highlight conserved metabolic dependencies of SARS-CoV-2 variants and underscore host lipid metabolism as a potential target for broad-spectrum antiviral strategies.
Though chemical exposures are known to potentially have negative impacts on health, including contributing to chronic diseases such as cancer, the quantitative contribution of risk is not fully understood for every chemical. A commonly used approach to quantify levels of risk is to measure the proportion of organisms (such as a total number of zebrafish on a plate or mice in a cage) with abnormal behavioral responses or morphology at increasing concentrations of chemical exposure. A particular challenge with processing the proportional data from these assays is the appropriate estimation of chemical concentration levels that result in malformations or acute toxicity, as these values typically vary between experimental measurements. The recommended approach by the Environmental Protection Agency (EPA) is to fit benchmark dose curves with specific filters and model fitting steps, which are crucial to properly processing the proportional data. Several tools exist for the fitting of benchmark dose response curves, but none are standalone Python libraries built to process both morphological and behavioral data as proportions with all the EPA recommended filters, filter parameters, models, and model parameters. Thus, here we present the benchmark dose response curve (bmdrc) Python library, which was built to closely follow these EPA guidelines with helpful visualizations of filters and fitted model curves, and reports for reproducibility purposes. bmdrc is open-source and has demonstrated utility as a support package to an existing web portal for information on chemicals (https://srp.pnnl.gov). Our package will support any toxicology analysis where the response is a proportional value at increasing levels of a concentration of a chemical or chemical mixture.
Quantifying the similarity between two mass spectra─a known reference mass spectrum and an unidentified sample mass spectrum─is at the heart of compound identification workflows in gas chromatography-mass spectrometry (GC-MS). The reference spectrum most like the sample is assigned as its identification (provided some quantitative similarity threshold is met, e.g., 80%) and thus accurately measuring similarity is essential. Significant research has gone toward developing metrics for this purpose, each of which has attempted to improve upon existing methods by incorporating GC-MS-specific information (e.g., peak ratios or retention times) or adopting various statistical and algorithmic frameworks. While this active development has led to a plethora of similarity metrics with demonstrated value across different contexts, the unfortunate consequence has been confusion surrounding which metric should be used as a global standard. No such metric is currently accepted as the standard method because different metrics have demonstrated optimal performance in different contexts. In this work, we propose an ensemble approach to spectral similarity scoring that combines the collective information from across existing similarity metrics to form an improved, globally representative similarity metric as a step toward establishing a global standard method. The resulting ensemble metrics are evaluated on over 88,000 spectra of varying complexity and demonstrate improved abilities to accurately rank the correct reference spectrum as the top-matching candidate for a sample relative to the rankings generated by individual similarity scores.
Though data acquisition and initial signal preprocessing of nuclear magnetic resonance (NMR) spectra have achieved high degrees of automation, downstream processing─specifically the profiling of spectra─has bottlenecked the overall NMR analysis workflow. Several efforts have been made to mitigate this bottleneck, but these solutions often trade an increase in automation for limitations elsewhere. In this work, we introduce nmRanalysis, a user-friendly web application that integrates the strengths of existing profiling tools for a more automated profiling workflow. nmRanalysis additionally incorporates novel features, including a machine-learning-driven recommender system for metabolite identification, further increasing the utility of nmRanalysis over the individual tools that it incorporates.
Introduction The placenta uses lipids and other nutrients to support its own metabolism hence impacting the type and amount of these substrates available to the growing fetus. Maternal obesity and gestational diabetes (GDM) can disrupt placental lipid metabolism and thus lead to altered fetal growth contributing to adverse pregnancy outcomes and developmentally programing the offspring for disease in later life. Understanding obesity and GDM driven changes in placental lipid metabolism is thus important. Methods We collected maternal plasma and placental villous tissue following elective cesarean section at term from women who were lean (pre-pregnancy BMI 18.5-24.9), obese (BMI>30) or obese with type A2 GDM n=8 each group (4 male and 4 female placentas). Fatty acid composition of different lipid classes was analyzed by LC-MS/MS analysis. Significant changes in GDM vs obese, GDM vs lean, and obese vs lean were determined in both a fetal sex-dependent and independent manner. Results In placenta 436 lipids were identified, among which 85 showed significant changes. We report significant changes in placental triglyceride, phosphatidylcholine, and phosphatidylinositol lipids containing essential fatty acids- DHA and AA in GDM, with male placentas driving these changes. In maternal plasma, 284 lipids were identified with 14 showing significant changes, but we observed no changes based on fetal sex. Discussion Maternal obesity and GDM impact placental lipid composition in a sexually dimorphic manner. The alteration in specific lipid classes can impact cellular energetics and placental function.
This study integrates quantitative data on personal exposure to polycyclic aromatic hydrocarbons (PAHs) in 162 silicone wristbands with demographics, behavioral information, and housing characteristics to explore contributions to residential exposure in a community influenced by historic and current industrial activities over the course of a year. Forty-six residents completed questionnaires and wore silicone wristbands as personal passive samplers for seven consecutive days on up to four separate occasions between November 2022 and June 2023; the repeated measures in this study lend insight into exposure sources with variability. It was hypothesized that individual behaviors and housing characteristics would be sources of dependence and correlation between personal PAH exposures. Fifty PAHs were detected, seventeen of which were alkylated PAHs. Exposure to PAHs of similar molecular weight was often correlated, notably between naphthalenes (2-rings) and PAHs of 3 or more rings. Generalized linear mixed models identified flooring type, participant age, and sampling month as predictors of increased PAH exposure. Flooring type and use of wood stoves or heavy machinery were identified as predictors of increased naphthalene exposure relative to larger PAHs. Personal behaviors and housing characteristics were sources of dependence and correlation between personal PAH exposures across repeated measures. We demonstrate the value of collection and integration of questionnaire data with exposure data, and repeated measures that consider intra-individual variability. Identification of influential exposure factors through repeated measures of chemical exposure and characterization of variability in personal exposure as performed in this study is important in the development of exposure mitigation strategies.