We present the scaffold-level genome assemblies of four Danio species, released in 2020: Danio albolineatus, Danio choprai, Danio jaintianensis and Danio tinwini (Chordata; Actinopterygii; Cypriniformes; Cyprinidae). Genome sizes range from 1,102.16 Mb for Danio choprai to 1,497.36 Mb for Danio tinwini.
We present the genome assemblies of three females of the Danio rerio strains, AB, Nadia and Cooch Behar (zebrafish; Chordata; Actinopteri; Cypriniformes; Cyprinidae). These assemblies were released in 2020 as part of the Danioninae Sequencing Project. The genome sequence of the strain AB is 1,405.10 megabases, the Nadia strain 1,465.10, and the Cooch Behar strain 1,421.80 megabases in length. Most of the assembly is scaffolded into 25 chromosomal pseudomolecules in each case. For each strain, the mitochondrial genome was also assembled and is 16.6 kilobases in length.
We present the scaffold-level genome assemblies of four Danio species, released in 2020: Danio albolineatus, Danio choprai, Danio jaintianensis and Danio tinwini (Chordata; Actinopterygii; Cypriniformes; Cyprinidae). Genome sizes range from 1,102.16 Mb for Danio choprai to 1,497.36 Mb for Danio tinwini.
We present a genome assembly from a female specimen of Danio aesculapii (the Panther Danio; Chordata; Actinopteri; Cypriniformes; Cyprinidae). The genome sequence is 1,381.5 megabases in span. Most of the assembly is scaffolded into 25 chromosomal pseudomolecules. Gene annotation of this assembly on Ensembl identified 23,884 protein coding genes.
We present a genome assembly from a female specimen of Danio aesculapii (the Panther Danio; Chordata; Actinopteri; Cypriniformes; Cyprinidae). The genome sequence is 1,381.5 megabases in span. Most of the assembly is scaffolded into 25 chromosomal pseudomolecules. Gene annotation of this assembly on Ensembl identified 23,884 protein coding genes.
We present the trio-binned, haplotype-resolved genome assemblies (released in 2020) of both haplotypes of an individual of the Danio rerio SAT strain, a cross between Tuebingen (maternal) and AB (paternal) strains (zebrafish; Chordata; Actinopteri; Cypriniformes; Cyprinidae). The genome sequence of the paternal haplotype (fDreABH1) is 1,354.1 megabases long, while the genome sequence of the maternal haplotype (fDreTuH) is 1,360.5 megabases long. Most of the assembly is scaffolded into 25 chromosomal pseudomolecules. The mitochondrial genome has also been assembled and is 16.6 kilobases in length.
OBJECTIVE:Although previous research shows that generalized and focal epilepsies have at least some distinct genetic influences, it remains uncertain why some families manifest both types of epilepsy. We tested two hypotheses: (1) families with both generalized and focal epilepsy carry separate risk alleles for both types; and (2) within mixed families, the type of epilepsy each individual manifests is influenced by the relative burden of separate risk alleles for generalized epilepsies and focal epilepsies. METHODS:The Epi4K cohort included 711 individuals with epilepsy from 257 families (113 generalized families, 66 focal families, 78 mixed families). We calculated polygenic risk scores (PRSs) for genetic generalized epilepsy (GGE_PRS) and for focal epilepsy (Focal_PRS). We used mixed-effects models to compare these PRSs between and within families, accounting for relatedness. RESULTS:Compared to population controls, individuals in generalized families had elevated GGE_PRS (p < .001) but not elevated Focal_PRS (p = .50); focal family individuals had elevated Focal_PRS (p = .008) but not elevated GGE_PRS (p = .22); and individuals in mixed families had both elevated GGE_PRS and elevated Focal_PRS (both p < .001). Within mixed families, GGE_PRS was higher in individuals with generalized epilepsy than in individuals with focal epilepsy (p < .001), whereas we did not detect a difference in Focal_PRS between individuals with generalized and focal epilepsy (p = .46). The GGE_PRS value explained 10% of the variance in phenotype within mixed families. SIGNIFICANCE:The occurrence of families with both generalized and focal epilepsy in separate individuals is explained at least partly by the chance co-occurrence of distinct genetic risk alleles for generalized and focal epilepsies. Within mixed families, an individual's epilepsy type can be explained at least in part by the relative burden of risk alleles for genetic generalized epilepsy.
We present the trio-binned, haplotype-resolved genome assemblies (released in 2020) of both haplotypes of an individual of the Danio rerio SAT strain, a cross between Tuebingen (maternal) and AB (paternal) strains (zebrafish; Chordata; Actinopteri; Cypriniformes; Cyprinidae). The genome sequence of the paternal haplotype (fDreABH1) is 1,354.1 megabases long, while the genome sequence of the maternal haplotype (fDreTuH) is 1,360.5 megabases long. Most of the assembly is scaffolded into 25 chromosomal pseudomolecules. The mitochondrial genome has also been assembled and is 16.6 kilobases in length.
We present the genome assemblies of three females of the Danio rerio strains, AB, Nadia and Cooch Behar (zebrafish; Chordata; Actinopteri; Cypriniformes; Cyprinidae). These assemblies were released in 2020 as part of the Danioninae Sequencing Project. The genome sequence of the strain AB is 1,405.10 megabases, the Nadia strain 1,465.10, and the Cooch Behar strain 1,421.80 megabases in length. Most of the assembly is scaffolded into 25 chromosomal pseudomolecules in each case. For each strain, the mitochondrial genome was also assembled and is 16.6 kilobases in length.
OBJECTIVE:To assess the possible effects of genetics on seizure outcome by estimating the familial aggregation of three outcome measures: seizure remission, history of ≥4 tonic-clonic seizures, and seizure control for individuals taking antiseizure medication. METHODS:We analyzed families containing multiple persons with epilepsy in four previously collected retrospective cohorts. Seizure remission was defined as being 5 and 10 years seizure-free at last observation. Total number of tonic-clonic seizures was dichotomized at <4 and ≥4 seizures. Seizure control in patients taking antiseizure medication was defined as no seizures for 1, 2, and 3 years. We used Bayesian generalized linear mixed-effects model (GLMM) to estimate the intraclass correlation coefficient (ICC) of the family-specific random effect, controlling for epilepsy type, age at epilepsy onset, and age at last data collection as fixed effects. We analyzed each cohort separately and performed meta-analysis using GLMMs. RESULTS:The combined cohorts included 3644 individuals with epilepsy from 1463 families. A history of ≥4 tonic-clonic seizures showed strong familial aggregation in three separate cohorts and meta-analysis (ICC .28, 95% confidence interval [CI] .21-.35, Bayes factor 8 × 1016). Meta-analyses did not reveal significant familial aggregation of seizure remission (ICC .08, 95% CI .01-.17, Bayes factor 1.46) or seizure control for individuals taking antiseizure medication (ICC .13, 95% CI 0-.35, Bayes factor 0.94), with heterogeneity among cohorts. SIGNIFICANCE:A history of ≥4 tonic-clonic seizures aggregated strongly in families, suggesting a genetic influence, whereas seizure remission and seizure control for individuals taking antiseizure medications did not aggregate consistently in families. Different seizure outcomes may have different underlying biology and risk factors. These findings should inform the future molecular genetic studies of seizure outcomes.
OBJECTIVE:To understand the etiological landscape and phenotypic differences between 2 developmental and epileptic encephalopathy (DEE) syndromes: DEE with spike-wave activation in sleep (DEE-SWAS) and epileptic encephalopathy with spike-wave activation in sleep (EE-SWAS). METHODS:All patients fulfilled International League Against Epilepsy (ILAE) DEE-SWAS or EE-SWAS criteria with a Core cohort (n = 91) drawn from our Epilepsy Genetics research program, together with 10 etiologically solved patients referred by collaborators in the Expanded cohort (n = 101). Detailed phenotyping and analysis of molecular genetic results were performed. We compared the phenotypic features of individuals with DEE-SWAS and EE-SWAS. Brain-specific gene co-expression analysis was performed for D/EE-SWAS genes. RESULTS:We identified the etiology in 42/91 (46%) patients in our Core cohort, including 29/44 (66%) with DEE-SWAS and 13/47 (28%) with EE-SWAS. A genetic etiology was identified in 31/91 (34%). D/EE-SWAS genes were highly co-expressed in brain, highlighting the importance of channelopathies and transcriptional regulators. Structural etiologies were found in 12/91 (13%) individuals. We identified 10 novel D/EE-SWAS genes with a range of functions: ATP1A2, CACNA1A, FOXP1, GRIN1, KCNMA1, KCNQ3, PPFIA3, PUF60, SETD1B, and ZBTB18, and 2 novel copy number variants, 17p11.2 duplication and 5q22 deletion. Although developmental regression patterns were similar in both syndromes, DEE-SWAS was associated with a longer duration of epilepsy and poorer intellectual outcome than EE-SWAS. INTERPRETATION:DEE-SWAS and EE-SWAS have highly heterogeneous genetic and structural etiologies. Phenotypic analysis highlights valuable clinical differences between DEE-SWAS and EE-SWAS which inform clinical care and prognostic counseling. Our etiological findings pave the way for the development of precision therapies. ANN NEUROL 2024;96:932-943.
BACKGROUND:Phenotypic variability within families with epilepsy is often observed, even when relatives share the same monogenic cause. We aimed to investigate whether common polygenic risk for epilepsy could explain the penetrance and phenotypic expression of rare pathogenic variants in familial epilepsies. METHODS:We studied 58 clinically heterogeneous families with genetic epilepsy with febrile seizures plus (GEFS+). Relatives were coded as either unaffected or affected with epilepsy, and graded according to phenotype severity: no seizures, febrile seizures (FS) only, febrile seizures plus (FS+), generalised/focal epilepsy, or developmental and epileptic encephalopathy (DEE). Epilepsy polygenic risk scores (PRSs) were tested for association with epilepsy phenotype. Within families, the mean PRS difference was compared between pairs concordant versus discordant for phenotype severity. Statistical analyses were performed using mixed-effect regression models. FINDINGS:304 individuals segregating a known, or presumed, rare variant of large effect, were studied. Within families, higher epilepsy polygenic risk was associated with an epilepsy diagnosis (OR = 1.39, 95% CI 1.08, 1.80, padj = 0.040). Relatives with a more severe phenotype had a mean pairwise PRS difference of +0.19 higher than relatives with a milder phenotype (padj = 0.010). The difference increased with greater phenotype discordance between relatives. As the cohort included two rare variants with >30 relatives each, variant-specific genotype-phenotype associations could also be analysed. Whilst the epilepsy PRS effect was strong for relatives segregating the GABRG2 p.Arg82Gln pathogenic variant (padj = 0.0010), the effect was not significant for SCN1B p.Cys121Trp. INTERPRETATION:We provide support for genetic background modifying the penetrance and phenotypic expression of rare variants associated with 'monogenic' epilepsies. In GEFS+ families, relatives with higher epilepsy PRSs were more likely to show penetrance (epilepsy diagnosis) and a more severe phenotype. Variant-specific analyses suggest that some rare variants may be more susceptible to PRS modification, carrying important genetic counselling and disease prognostication implications for patients. FUNDING:National Health and Medical Research Council of Australia, Medical Research Future Fund of Australia.
Polygenic risk scores (PRSs) are an important tool for understanding the role of common genetic variants in human disease. Standard best practices recommend that PRSs be analyzed in cohorts that are independent of the genome-wide association study (GWAS) used to derive the scores without sample overlap or relatedness between the two cohorts. However, identifying sample overlap and relatedness can be challenging in an era of GWASs performed by large biobanks and international research consortia. Although most genomics researchers are aware of best practices and theoretical concerns about sample overlap and relatedness between GWAS and PRS cohorts, the prevailing assumption is that the risk of bias is small for very large GWASs. Here, we present two real-world examples demonstrating that sample overlap and relatedness is not a minor or theoretical concern but an important potential source of bias in PRS studies. Using a recently developed statistical adjustment tool, we found that excluding overlapping and related samples was equal to or more powerful than adjusting for overlap bias. Our goal is to make genomics researchers aware of the magnitude of risk of bias from sample overlap and relatedness and to highlight the need for mitigation tools, including independent validation cohorts in PRS studies, continued development of statistical adjustment methods, and tools for researchers to test their cohorts for overlap and relatedness with GWAS cohorts without sharing individual-level data.
We present a genome assembly from an individual Danionella dracula (the Dracula fish; Chordata; Actinopterygii; Cypriniformes; Danionidae; Danioninae). The genome sequence is 665.21 megabases in span. This is a scaffold-level assembly, with a scaffold N50 of 10.29 Mb.
Cartilaginous fishes (chondrichthyans: chimeras and elasmobranchs -sharks, skates, and rays) hold a key phylogenetic position to explore the origin and diversifications of jawed vertebrates. Here, we report and integrate reference genomic, transcriptomic, and morphological data in the small-spotted catshark Scyliorhinus canicula to shed light on the evolution of sensory organs. We first characterize general aspects of the catshark genome, confirming the high conservation of genome organization across cartilaginous fishes, and investigate population genomic signatures. Taking advantage of a dense sampling of transcriptomic data, we also identify gene signatures for all major organs, including chondrichthyan specializations, and evaluate expression diversifications between paralogs within major gene families involved in sensory functions. Finally, we combine these data with 3D synchrotron imaging and in situ gene expression analyses to explore chondrichthyan-specific traits and more general evolutionary trends of sensory systems. This approach brings to light, among others, novel markers of the ampullae of Lorenzini electrosensory cells, a duplication hotspot for crystallin genes conserved in jawed vertebrates, and a new metazoan clade of the transient-receptor potential (TRP) family. These resources and results, obtained in an experimentally tractable chondrichthyan model, open new avenues to integrate multiomics analyses for the study of elasmobranchs and jawed vertebrates.
We present a genome assembly from an individual Danionella dracula (the Dracula fish; Chordata; Actinopterygii; Cypriniformes; Danionidae; Danioninae). The genome sequence is 665.21 megabases in span. This is a scaffold-level assembly, with a scaffold N50 of 10.29 Mb.
Long-range sequencing grants insight into additional genetic information beyond what can be accessed by both short reads and modern long-read technology. Several new sequencing technologies, such as "Hi-C" and "Linked Reads", produce long-range datasets for high-throughput and high-resolution genome analyses, which are rapidly advancing the field of genome assembly, genome scaffolding, and more comprehensive variant identification. In this review, we focused on five major long-range sequencing technologies: high-throughput chromosome conformation capture (Hi-C), 10X Genomics Linked Reads, haplotagging, transposase enzyme linked long-read sequencing (TELL-seq), and single- tube long fragment read (stLFR). We detailed the mechanisms and data products of the five platforms and their important applications, evaluated the quality of sequencing data from different platforms, and discussed the currently available bioinformatics tools. This work will benefit the selection of appropriate long-range technology for specific biological studies.
Numerous novel adaptations characterise the radiation of notothenioids, the dominant fish group in the freezing seas of the Southern Ocean. To improve understanding of the evolution of this iconic fish group, here we generate and analyse new genome assemblies for 24 species covering all major subgroups of the radiation, including five long-read assemblies. We present a new estimate for the onset of the radiation at 10.7 million years ago, based on a time-calibrated phylogeny derived from genome-wide sequence data. We identify a two-fold variation in genome size, driven by expansion of multiple transposable element families, and use the long-read data to reconstruct two evolutionarily important, highly repetitive gene family loci. First, we present the most complete reconstruction to date of the antifreeze glycoprotein gene family, whose emergence enabled survival in sub-zero temperatures, showing the expansion of the antifreeze gene locus from the ancestral to the derived state. Second, we trace the loss of haemoglobin genes in icefishes, the only vertebrates lacking functional haemoglobins, through complete reconstruction of the two haemoglobin gene clusters across notothenioid families. Both the haemoglobin and antifreeze genomic loci are characterised by multiple transposon expansions that may have driven the evolutionary history of these genes.
ABSTRACT We describe FoundHaplo, a novel identity-by-descent algorithm designed to identify individuals with known, untyped, disease-causing variants using only SNP array data. FoundHaplo leverages knowledge of shared disease haplotypes for inherited disease-causing variants to identify individuals who share the disease haplotype and are, therefore, likely to carry the rare (MAF<0.01) variant. We performed a simulation study to evaluate the performance of FoundHaplo across 33 known disease-harbouring loci. We demonstrated the ability of FoundHaplo to infer the presence of two rare (MAF<0.01) pathogenic variants, SCN1B c.363C>G (p.Cys121Trp) and WWOX c.49G>A (p.E17K), which can cause mild dominant and severe recessive epilepsy respectively, in two large cohorts including 1,573 individuals with epilepsy from the Epi25 cohort and 468,481 individuals from the UK Biobank. We demonstrate that FoundHaplo performs substantially better at inferring the presence of these variants than existing genome-wide imputation approaches. FoundHaplo is a valuable, low-cost screening tool that can be applied to search SNP genotyping array data for disease-causing variants with known founder effects based on shared disease haplotypes. FoundHaplo is available at https://github.com/bahlolab/FoundHaplo .
The National Collection of Type Cultures (NCTC) was founded on 1 January 1920 in order to fulfil a recognized need for a centralized repository for bacterial and fungal strains within the UK. It is among the longest-established collections of its kind anywhere in the world and today holds approximately 6000 type and reference bacterial strains – many of medical, scientific and veterinary importance – available to academic, health, food and veterinary institutions worldwide. Recently, a collaboration between NCTC, Pacific Biosciences and the Wellcome Sanger Institute established the NCTC3000 project to long-read sequence and assemble the genomes of up to 3000 NCTC strains. Here, at the beginning of the collection’s second century, we introduce the resulting NCTC3000 sequence read datasets, genome assemblies and annotations as a unique, historically and scientifically relevant resource for the benefit of the international bacterial research community.