Abstract Background Differential expression analysis is a central tool for studying the biological processes altered in human diseases via transcriptomic signatures. However, transcriptomic datasets are systematically confounded by latent variables from two distinct sources: unmeasured technical and biological heterogeneity within the expression data, and expression differences driven by population stratification. Correction using expression-based surrogate variables (SVs) and genotype-based principal components (PCs) addresses these sources independently, yet no study has directly evaluated their combined use against either method alone within a differential expression framework. In this study we hypothesised that simultaneously including both correction layers would produce more biologically valid and reproducible results than either approach alone, and tested this in two independent RNA-seq datasets of amyotrophic lateral sclerosis (ALS) cases and controls with matching genotype data. Results Four nested differential expression models (corrected for PC-only, SV-only, both SV and PC, and neither PCs nor SVs) were evaluated across the KCLBB (96 cases and 52 controls) and ALS Consortium (272 cases and 35 controls) datasets. Models were evaluated on: cross-dataset effect size concordance, cross-dataset replicability quantified by the Jaccard Similarity Index, and biological recall against a curated reference set of 66 known ALS genes. The combined SV+PC framework consistently outperformed simpler models across all metrics. Replicability improved nearly ten-fold compared to the non-corrected model, (Jaccard index: 2.28% to 19.5%), and the combined framework exhibited a statistically significant 2.1% gain over the SV-only model. The biological recall ALS genes recovered doubled comparing to the SV correction alone. Crucially, effect size stability was preserved, with the combined model expanding the shared transcriptomic signal without sacrificing consistency. These findings remained generally robust to PC number in sensitivity analyses. Conclusions This study found that SVs and genotype PCs address non-redundant sources of confounding, and we recommend their combined use as standard practice in differential expression analysis where matched genotype data are available. Notably PCs capturing population structure can also be derived directly from RNA-seq data, extending the applicability of this framework to studies lacking matched genotype data. Although this analysis was restricted to ALS datasets, we expect these findings to generalise to other traits.
Amyotrophic lateral sclerosis (ALS) is a heritable disorder where rare variants with low-to-moderate penetrance are thought to dominate genetic risk. To identify such rare variants, we harmonized and analyzed exome data from 22 cohorts, totaling 17,919 individuals with ALS and 200,703 controls across discovery and replication phases. Rare variant analyses identified several new risk genes, with replication confirming association of YKT6 and supporting HTR3C, GBGT1 and KNTC1. We also provide strong, independent validation for genes with limited previous evidence: ARPP21, DNAJC7 and CFAP410. Notably, in ARPP21, we identified a new high-effect variant (p.P747L) and confirmed that p.P563L is an ALS-associated variant leading to an aggressive disease course. Beyond new discoveries, our analyses largely recapitulated the known genetic architecture of ALS, identifying risk variants in over 20% of cases and supporting a cumulative oligogenic risk model. These findings highlight new translational targets and show that rare variant analyses capture substantially more genetic risk than common variant genome-wide association studies.
Background The pathogenic G4C2 repeat expansion in the C9ORF72 gene is the most common genetic cause of amyotrophic lateral sclerosis (ALS) and frontotemporal dementia (FTD). Studies focused on delineating the underlying perturbed mechanisms resulting from this genetic mutation are often confounded by the heterogeneity present in current disease models, such as patient-derived iPSC lines, with estimations of up to 50% of the variation in iPSC cell phenotypes resulting from inter-individual differences. Isogenic models, in which the pathogenic mutation is introduced into a defined genetic background, offer a powerful approach to isolating mutation-specific effects and enable high-resolution comparison across distinct ALS/FTD-associated mutations. Such models are essential for uncovering convergent disease mechanisms and improving reproducibility in ALS/FTD research. Methods A two-step scarless CRISPR/Cas9 genome editing strategy was used to generate isogenic human iPSC lines carrying a de novo knock-in of a disease-length G4C2 repeat expansion in the C9ORF72 locus. The resulting lines underwent thorough quality control and were differentiated into lower motor neurons and assessed for the presence of key ALS/FTD pathologies, including changes to C9ORF72 mRNA and protein expression, RNA foci and dipeptide repeat proteins. Results Two C9ORF72 knock-in iPSC lines were generated with 631 and 600 G4C2 repeats, alongside an isogenic genome editing control line. The C9ORF72 G4C2 repeat expansion knock-in iPSC lines exhibit both loss-of-function and gain-of-function pathological features characteristic of ALS/FTD. Compared to the parental wild-type KOLF2.1J line and isogenic (wild-type) CRISPR control line, these exhibit a significant reduction in C9ORF72 mRNA and protein levels, the presence of RNA foci accumulation, and a marked increase in poly(GA) and poly(GP) dipeptide repeat protein levels in iPSCs and motor neurons. Conclusions This is one of the first reports of a successful knock-in of the pathogenic C9ORF72 G4C2 repeat expansion into a human iPSC line, establishing a genetically defined and physiologically relevant model of ALS/FTD. These isogenic lines recapitulate both key loss- and gain-of-function disease pathologies, providing a crucial complement to existing patient-derived iPSC banks. By eliminating confounding genetic background variability, these cell lines will enable more precise interrogation of C9ORF72 -linked pathomechanisms and offer a robust platform for comparative studies across the ALS and FTD spectrum, mechanistic investigations, and future therapeutic targeting with enhanced translational relevance. ### Competing Interest Statement The authors have declared no competing interest. * ALS : amyotrophic lateral sclerosis ASO : antisense oligonucleotide DPR : dipeptide repeat FTD : frontotemporal dementia HRE : hexanucleotide repeat iPSC : induced pluripotent stem cell liMNs : lower induced motor neurons RAN : repeat associated non-ATG translation rpPCR : repeat primed PCR ssODN : single-stranded oligodeoxyribonucleotide Motor Neurone Disease Association, https://ror.org/02gq0fg61, Ruepp/Apr19/872-791 UK Dementia Research Institute, https://ror.org/02wedp412, UK DRI-6204, UK DRI-6203, UK DRI 1203
Accurately forecasting streamflow is essential for effectively managing water resources. High-quality operational forecasts allow us to prepare for extreme weather events, optimize hydropower generation, and minimize the impact of human development on the natural environment. However, streamflow forecasts are inherently limited by the quality and availability of upstream weather sources. The weather forecasts that drive hydrological modeling vary in their temporal resolutions and are prone to outages, such as the ECMWF data outage in November of 2023. Here, we present HydroForecast Short Term 3 (ST-3), a state-of-the-art probabilistic deep learning model for medium-term (10-day) streamflow forecasts. ST-3 combines long short-term memory architecture with Boolean tensors representing data availability and dense embeddings for processing of the information in these tensors. This architecture allows for a training routine that implements data augmentation to synthesize varying amounts of availability of weather inputs. The result is a model that 1) makes accurate forecasts even in the case of an upstream data outage, 2) achieves higher accuracy by leveraging data of varying temporal resolutions including regional weather inputs with shorter lead times than the most common medium term weather inputs, and 3) generates individual forecast traces for each individual weather source, facilitating inference across regions where weather data availability is limited. Initial results across CAMELS sites in North America indicate that the incorporation of near-term high resolution weather data increases early horizon forecast KGE by nearly 0.25 with meaningful improvements in metrics seen across our customers’ operational sites. Validation metrics across individual weather sources, as well as model interrogation through integrated gradients highlights a high level of fidelity in the model’s learned physical relationships across forecast scenarios.
Background Mobile element insertions, particularly transposable elements (TEs) such as Alu, LINE-1 (L1), SVA, and endogenous retroviruses (ERVs), represent a major source of human genetic variation and have been implicated in evolution, genomic instability, and disease. Although long-read sequencing generally outperforms short-read sequencing for the characterisation of such elements, their accurate detection with long-reads remains challenging, with different computational tools adopting varying approaches and producing divergent call sets. As gold standards currently do not exist for TE detection, benchmarking these methods is essential to understand their strengths, limitations, and biases. Here, we systematically evaluate the performance of available state-of-the-art TE detection tools on both simulated and real human genome data using highly characterised samples from the Genome in a Bottle consortium, population level reference databases and an in-house collection for which matching short-read sequencing data are available. Results Our results show significant differences in calling strategies, leading to substantial variation in precision, recall, and the spectrum of TE families detected across tools. Our benchmark also displays the differences between short-read and long-read calls, highlighting the importance of appropriate method selection. Conclusions The benchmarking results presented here will aid TE researchers make better informed decisions on which tool to use in their long-read TE analyses. Strengths and limitations of different tools have been highlighted in depth as well as their computational requirements, which will result in less time spent finding the best tool for the job and promote faster TE research. ### Competing Interest Statement The authors have declared no competing interest.
Sex is an important covariate in all genetic and epigenetic research due to its role in the incidence, progression and outcome of many phenotypic characteristics and human diseases. Amyotrophic lateral sclerosis (ALS) is a motor neuron disease with a sex bias towards higher incidence in males. Here, we report for the first time a blood-based epigenome-wide association study meta-analysis in 9274 individuals after stringent quality control (5529 males and 3975 females). We identified a total of 226 ALS saDMPs (sex-associated DMPs) annotated to a total of 159 unique genes. These ALS saDMPs were depleted at transposable elements yet significantly enriched at enhancers and slightly enriched at 3'UTRs. These ALS saDMPs were enriched for transcription factor motifs such as ESR1 and REST. Moreover, we identified an additional 10 genes associated with ALS saDMPs through chromatin loop interactions, suggesting a potential regulatory role for these saDMPs on distant genes. Furthermore, we investigated the relationship between DNA methylation at specific CpG sites and overall survival in ALS using Cox proportional hazards models. We identified two ALS saDMPs, cg14380013 and cg06729676, that showed significant associations with survival. Overall, our study reports a reliable catalogue of sex-associated ALS saDMPs in ALS and elucidates several characteristics of these sites using a large-scale dataset. This resource will benefit future studies aiming to investigate the role of sex in the incidence, progression and risk for ALS.
Lake trophic state is a key ecosystem property that integrates a lake’s physical, chemical, and biological processes. Despite the importance of trophic state as a gauge of lake water quality, standardized and machine-readable observations are uncommon. Remote sensing presents an opportunity to detect and analyze lake trophic state with reproducible, robust methods across time and space. We used Landsat surface reflectance data to create the first compendium of annual lake trophic state for 55,662 lakes of at least 10 ha in area throughout the contiguous United States from 1984 through 2020. The dataset was constructed with FAIR data principles (Findable, Accessible, Interoperable, and Reproducible) in mind, where data are publicly available, relational keys from parent datasets are retained, and all data wrangling and modeling routines are scripted for future reuse. Together, this resource offers critical data to address basic and applied research questions about lake water quality at a suite of spatial and temporal scales.
Salinity dynamics in the Delaware Bay estuary are a critical water quality concern as elevated salinity can damage infrastructure and threaten drinking water supplies. Current state-of-the-art modeling approaches use hydrodynamic models, which can produce accurate results but are limited by significant computational costs. We developed a machine learning (ML) model to predict the 250 mg L-1 Cl- isochlor, also known as the "salt front," using daily river discharge, meteorological drivers, and tidal water level data. We use the ML model to predict the location of the salt front, measured in river miles (RM) along the Delaware River, during the period 2001-2020, and we compare predictions of the ML model to the hydrodynamic Coupled Ocean-Atmosphere-Wave-Sediment Transport (COAWST) model. The ML model predicts the location of the salt front with greater accuracy (root mean squared error [RMSE] = 2.52 RM) than the COAWST model does (RMSE = 5.36); however, the ML model struggles to predict extreme events. Furthermore, we use functional performance and expected gradients, tools from information theory and explainable artificial intelligence, to show that the ML model learns physically realistic relationships between the salt front location and drivers (particularly discharge and tidal water level). These results demonstrate how an ML modeling approach can provide predictive and functional accuracy at a significantly reduced computational cost compared to process-based models. In addition, these results provide support for using ML models in operational forecasting, scenario testing, management decisions, hindcasting, and resulting opportunities to understand past behavior and develop hypotheses.
This paper proposes a meta-transfer-learning method for predicting daily maximum water temperature in stream networks with explicit modeling of extreme events. Accurate prediction of these extreme events is challenging because of their sparsity in the training data and their distinct responses to external drivers when compared to non-extreme observations. To overcome these challenges, we propose a sample reweighting strategy to escalate the importance of extreme events in the training process while preserving the predictive performance in normal time periods. The sample weight for each training data point is estimated as the similarity with the target test data point using contextual information and physical simulation. The obtained sample weight values are then used to fine-tune the initial model to transfer it to the test data. This method is further enhanced by an extreme value theory-based loss function to enforce the distribution of extreme data points and accelerated by a clustering algorithm based on the estimated similarities. Additionally, we introduce an online learning strategy to further refine the predictive model using newly collected observed data. The experimental results using real stream data from the Delaware River Basin over the past 36 years demonstrate that our meta-transfer-learning method produces more accurate predictions in both normal and extreme time periods when compared to baselines without the sample re-weighting scheme. The similarity learning method can reveal meaningful relationships amongst data points. We also show that the clustering algorithm can be used to accelerate the prediction while not compromising the predictive performance. The online learning strategy is shown to further improve predictive performance using recently observed data.
Stream temperature is a fundamental control on ecosystem health. Recent efforts incorporating process guidance into deep learning models for predicting stream temperature have been shown to outperform existing statistical and physical models. This performance is in part because deep learning architectures can actively learn spatiotemporal relationships that govern how water and energy propagate through a river network. However, exploration of how spatiotemporal awareness and process guidance influence a model's generalizability under shifting environmental conditions such as climate change is limited. Here, we use Explainable Artificial Intelligence (XAI) to interrogate how differing deep learning architectures affect a model's learned spatial and temporal dependencies, and how those learned dependencies affect a model's ability to maintain high accuracy when applied to unseen environmental conditions. Using the Delaware River Basin in the northeastern United States as a test case, we compare two spatiotemporally aware process‐guided deep learning models for predicting stream temperature (a recurrent graph convolution network—RGCN, and a temporal convolution graph model—Graph WaveNet). Both models achieve equally high predictive performance when testing data are well represented in the training data (test root mean squared errors of 1.64°C and 1.65°C); however, Graph WaveNet significantly outperforms RGCN in 4 out of 5 experiments where test partitions represent different types of unseen environmental conditions. XAI results show that the architecture of Graph WaveNet leads to learned spatial relationships with greater fidelity to physical processes, and that this fidelity improves the generalizability of the model when applied to shifting and/or unseen environmental conditions.
Trophic state (TS) characterizes a waterbody’s biological productivity and depends on its morphometry, physics, chemistry, biology, climate, and history. However, multiple TS operational definitions have emerged to meet use-specific classification needs. These differing operational definitions can create inconsistent understanding, can lead to miscommunication, and can result in siloed management strategies for TS. For example, some regulatory agencies use TS to signify ecological integrity as opposed to biological productivity, where TS classification may trigger intervention efforts. These inconsistencies may be compounded when interdisciplinary projects employ varied TS frameworks. To emphasize the consequences of using multiple TS classification schemes, we present three scenarios for which an improved understanding of the TS concept could advance limnological research, management efforts, and interdisciplinary collaboration. As the field of limnology continues to expand, we highlight the importance of re-evaluating even the most fundamental limnological concepts, such as TS, to ensure congruence with evolving, cutting-edge science.
Objective: Variants in the superoxide dismutase (SOD1) gene are among the most common genetic causes of amyotrophic lateral sclerosis. Reflecting the wide spectrum of putatively deleterious variants that have been reported to date, it has become clear that SOD1-linked ALS presents a highly variable age at symptom onset and disease duration.Methods: Here we describe an open access web tool for comparative phenotype analysis in ALS: https://sod1-als-browser.rosalind.kcl.ac.uk/. The tool contains a built-in dataset of clinical information from 1383 people with ALS harboring a SOD1 variant resulting in one of 162 unique amino acid sequence alterations and from a non-SOD1 comparator ALS cohort of 13,469 individuals. We present two examples of analyses possible with this tool, testing how the ALS phenotype relates to SOD1 variants that alter amino acid residue hydrophobicity and to distinct variants at the 94th residue of SOD1, where six are sampled.Results and conclusions: The tool provides immediate access to the datasets and enables bespoke analysis of phenotypic trends associated with different protein variants, including the option for users to upload their own datasets for integration with the server data. The tool can be used to study SOD1-ALS and provides an analytical framework to study the differences between other user-uploaded ALS groups and our large reference database of SOD1 and non-SOD1 ALS. The tool is designed to be useful for clinicians and researchers, including those without programming expertise, and is highly flexible in the analyses that can be conducted.
Elevated impulsivity is a key component of attention-deficit hyperactivity disorder (ADHD), bipolar disorder and juvenile myoclonic epilepsy (JME). We performed a genome-wide association, colocalization, polygenic risk score, and pathway analysis of impulsivity in JME ( n = 381). Results were followed up with functional characterisation using a drosophila model. We identified genome-wide associated SNPs at 8q13.3 ( P = 7.5 × 10 −9 ) and 10p11.21 ( P = 3.6 × 10 −8 ). The 8q13.3 locus colocalizes with SLCO5A1 expression quantitative trait loci in cerebral cortex ( P = 9.5 × 10 −3 ). SLCO5A1 codes for an organic anion transporter and upregulates synapse assembly/organisation genes. Pathway analysis demonstrates 12.7-fold enrichment for presynaptic membrane assembly genes ( P = 0.0005) and 14.3-fold enrichment for presynaptic organisation genes ( P = 0.0005) including NLGN1 and PTPRD . RNAi knockdown of Oatp30B , the Drosophila polypeptide with the highest homology to SLCO5A1 , causes over-reactive startling behaviour ( P = 8.7 × 10 −3 ) and increased seizure-like events ( P = 6.8 × 10 −7 ). Polygenic risk score for ADHD genetically correlates with impulsivity scores in JME ( P = 1.60 × 10 −3 ). SLCO5A1 loss-of-function represents an impulsivity and seizure mechanism. Synaptic assembly genes may inform the aetiology of impulsivity in health and disease.
Humans have drastically disrupted the global sediment cycle. Suspended sediment flux and concentration are key controls over both river morphology and river ecosystems. Our ability to understand sediment dynamics within river corridors is limited by observations. Here, we present RivSed, a database of satellite observations of suspended sediment concentration (SSC) from 1984 to 2018 across 460 large (>60 m wide) US rivers that provides a new, spatially explicit view of river sediment. We found that 32% of US rivers have a declining temporal trend in sediment concentration, with a mean reduction of 40% since 1984, whereas only 2% have an increasing trend. Most rivers (52%) show decreasing sediment concentration longitudinally moving downstream, typically due to a few large dams rather than the accumulated effect of many small dams. Comparing our observations with modeled 'pre-dam' longitudinal SSC, most rivers (53%) show different patterns. However, contemporary longitudinal patterns in concentration are remarkably stable from year to year since 1984, with more stability in large, highly managed rivers with less cropland. RivSed has broad applications for river geomorphology and ecology and highlights anthropogenic effects on river corridors across the US.
For over a century, ecologists have used the concept of trophic state (TS) to characterize an aquatic ecosystem’s biological productivity. Because measuring productivity can be challenging within an ecosystem and across landscapes, multiple TS classification schemes, each relying on a variety of proxies for productivity, have emerged to meet use-specific needs. Most commonly, chlorophyll a, phosphorus, and Secchi depth are used to discriminate TS based on autotrophic production, whereas phosphorus, dissolved organic carbon, and true color are used to discriminate TS based on autotrophic and heterotrophic production. Both classification schemes aim to characterize an ecosystem’s function broadly, but the relative emphasis on heterotrophic and autotrophic processes masks nuances in how an ecosystem’s function is understood. Moreover, differing classification schemes can create inconsistent understanding and can lead to narrowed interpretation of ecosystem integrity. For example, the U.S. Clean Water Act focuses exclusively on threats to autotrophic water quality, framed in terms of eutrophication in response to nutrient loading. This usage lacks information about non-algal threats to water quality, such as dystrophication in response to dissolved organic carbon loading. Consequently, the TS classification schemes used to identify eutrophication and dystrophication may refer to ecosystems similarly (e.g., oligotrophic and eutrophic), yet these categories are derived from different proxies. These inconsistencies in TS classification schemes may be compounded when interdisciplinary projects employ varied TS frameworks. Even with these shortcomings, TS can still be used to distill information on complex aquatic ecosystem function into a set of generalizable expectations, which can then be used to contextualize, compare, and project ecosystems across scales. However, to emphasize the consequences of using multiple TS classification schemes, we present three scenarios for which an improved understanding of the TS concept advances freshwater research, management efforts, and interdisciplinary collaboration. To increase clarity in TS, the aquatic sciences could benefit from including information about the proxy variables as well as the spatiotemporal domains used to classify TS. As the field of aquatic sciences expands and climatic irregularity increases, we highlight the importance of re-evaluating fundamental concepts, such as TS, to ensure their compatibility with evolving science.
Abstract Although water color is a fundamental property of freshwater ecosystems, the global distribution of lake color remains unknown. Here, we used 5.14 million color records derived from satellite images during 2013–2020 to determine the modal water color of 85,360 representative lakes worldwide. We find that globally, lake modal colors span the visible spectrum with a bimodal distribution of blue versus non‐blue (green‐brown), with 31% of lakes being blue and 69% being non‐blue. Both climate and lake morphology influence lake modal color. Blue lakes are associated with cooler summer air temperature, winter ice cover, and higher precipitation. Under a 3°C increase in summer air temperature, 14% of blue lakes could shift to a non‐blue regime, representing a substantial change to their underlying ecology. As lake ecosystems continue to face a range of stressors, this study provides a critical baseline for understanding lake responses to global environmental change.
Global change may contribute to ecological changes in high-elevation lakes and reservoirs, but a lack of data makes it difficult to evaluate spatiotemporal patterns. Remote sensing imagery can provide more complete records to evaluate whether consistent changes across a broad geographic region are occurring. We used Landsat surface reflectance data to evaluate spatial patterns of contemporary lake color (2010-2020) in 940 lakes in the U.S. Rocky Mountains, a historically understudied area for lake water quality. Intuitively, we found that most of the lakes in the region are blue (66%) and were found in steep-sided watersheds (>22.5º) or alternatively were relatively deep (>4.5m) with mean annual air temperature (MAAT) <4.5ºC. Most green/brown lakes were found in relatively shallow sloped watersheds with MAAT ≥4.5ºC. We extended the analysis of contemporary lake color to evaluate changes in color from 1984-2020 for a subset of lakes with the most complete time series (n=527). We found limited evidence of lakes shifting from blue to green states, but rather, 55% of the lakes had no trend in lake color. Surprisingly, where lake color was changing, 32% of lakes were trending toward bluer wavelengths, and only 13% shifted toward greener wavelengths. Lakes and reservoirs with the most substantial shifts toward blue wavelengths tended to be in urbanized, human population centers at relatively lower elevations. In contrast, lakes that shifted to greener wavelengths did not relate clearly to any lake or landscape features that we evaluated, though declining winter precipitation and warming summer and fall temperatures may play a role in some systems. Collectively, these results suggest that the interactions between local landscape factors and broader climatic changes can result in heterogeneous, context-dependent changes in lake color.
Superoxide dismutase (SOD1) gene variants may cause amyotrophic lateral sclerosis, some of which are associated with a distinct phenotype. Most studies assess limited variants or sample sizes. In this international, retrospective observational study, we compare phenotypic and demographic characteristics between people with SOD1-ALS and people with ALS and no recorded SOD1 variant. We investigate which variants are associated with age at symptom onset and time from onset to death or censoring using Cox proportional-hazards regression. The SOD1-ALS dataset reports age of onset for 1122 and disease duration for 883 people; the comparator population includes 10,214 and 9010 people respectively. Eight variants are associated with younger age of onset and distinct survival trajectories; a further eight associated with younger onset only and one with distinct survival only. Here we show that onset and survival are decoupled in SOD1-ALS. Future research should characterise rarer variants and molecular mechanisms causing the observed variability.
Artisanal and small-scale gold mining (ASGM) is the primary global source of anthropogenic mercury (Hg) emissions and a large source of landscape change. ASGM occurs throughout the world, including in the Peruvian Amazon. This data set contains measurements of surface water, precipitation, throughfall, leaves, sediment, soil, and air samples from across the Madre de Dios region of Peru, in locations near and remote from ASGM. These data were collected to determine the fate and transport of Hg across the landscape. Samples were collected in 2018 and 2019. Data predominantly included total Hg and methyl Hg concentrations in surface water, precipitation, throughfall, leaves, sediment, soil, and air. Additional water and soil parameters were also measured to better characterize their chemistry. There are no copyright restrictions; please cite this data paper when the data are used in publication.