Agricultural research has long faced challenges with data sharing, often relying on informal networks and requiring significant effort to clean and harmonize data. This hampers collaboration and limits data reuse. While FAIR (Findable, Accessible, Interoperable, and Reusable) principles are widely adopted in biomedical research, their uptake in agricultural genomics has lagged. The AgBioData Standards for Genetic Variation Working Group aims to close this gap by promoting FAIR data practices. We surveyed current standards for managing agricultural genetic variation and recommend adopting reference SNP identifiers (rsIDs) as a key step. We present examples from crop research communities with varying data maturity, including those without reference assemblies. Milestones include introducing nearly 220 million rsIDs to Gramene and pangenome databases, projecting rsIDs from reference to pangenome varieties in sorghum and maize, and developing an agricultural FAIR guide for rsID adoption. Better coordination among data producers, repositories, and breeding platforms is essential to improve interoperability, consistency, and accelerate genetic variant discovery for crop trait improvement.
Agriculture is a key component of the One Health Joint Plan of Action (OH JPA), a global initiative focused on developing sustainable solutions to address health threats. Our work supports the OH JPA pathway of data, evidence, information systems, and knowledge exchange. FAIR data standards are crucial for integrating genetic and phenotypic datasets, enabling comparisons, and providing broader insights. While these standards are well-established in biomedical research, they have only recently been adopted in agriculture. Historically, agricultural data sharing relied on personal networks, requiring significant effort to clean and harmonize data, limiting collaboration. To improve this, we worked with the AgBioData Standards for Genetic Variation Working Group to evaluate practices and review data formats. We introduced rsIDs for genetic markers in the SorghumBase and Gramene databases, ensuring consistency across breeding lines and landraces, and partnered with germplasm repositories to implement standard biosample identifiers, thus enhancing interoperability across Ag data resources. Our efforts show that data standards not only improve data sharing but also support re-analysis, and boost research visibility, collaboration, and impact.
Reference genomes are foundational to genomics but suffer from widespread ambiguity and incompatibility due to inconsistent naming, undocumented differences, and lack of formal mechanisms for comparison. To address this, we introduce the GA4GH refget Sequence Collections (seqcol) standard. Refget seqcol is a framework for unambiguous representation, retrieval, and comparison of sequence collections such as reference genomes and transcriptomes. The seqcol standard comprises four components: a structured data schema, a canonical encoding algorithm that produces content-based, globally unique identifiers, a retrieval API, and a comparison protocol. This standard enables precise identification of sequence collections, even across decentralized or private systems, and allows compatibility assessments beyond exact identity, such as order-relaxed matches or shared coordinate systems. We applied the refget seqcol standard to 60 human and 36 mouse reference genomes sourced from major providers. Using digest-based comparisons, we quantified levels of similarity across attributes including sequence names, lengths, coordinate systems, and actual sequence content. Our analysis revealed some consistent subsets of sequences or coordinate systems, as well as substantial incompatibility among references and duplicate references under different names. To support adoption of refget seqcol, we provide a Python package implementing the full standard, a web API, and a comparison interface allowing users to assess local references against a curated database. This work offers a scalable, reproducible solution to the reference genome compatibility crisis, enabling improved transparency, reuse, and integration in genomic analyses. Refget seqcol enhances interoperability across tools and datasets, making genomic research more robust and reproducible.
The COVID-19 pandemic has seen large-scale pathogen genomic sequencing efforts, becoming part of the toolbox for surveillance and epidemic research. This resulted in an unprecedented level of data sharing to open repositories, which has actively supported the identification of SARS-CoV-2 structure, molecular interactions, mutations and variants, and facilitated vaccine development and drug reuse studies and design. The European COVID-19 Data Platform was launched to support this data sharing, and has resulted in the deposition of several million SARS-CoV-2 raw reads. In this paper we describe (1) open data sharing, (2) tools for submission, analysis, visualisation and data claiming (e.g. ORCiD), (3) the systematic analysis of these datasets, at scale via the SARS-CoV-2 Data Hubs as well as (4) lessons learnt. This paper describes a component of the Platform, the SARS-CoV-2 Data Hubs, which enable the extension and set up of infrastructure that we intend to use more widely in the future for pathogen surveillance and pandemic preparedness.
Previous studies indicated that in some species phylogeographic patterns obtained in the analysis of nuclear and mitochondrial DNA (mtDNA) markers can be different. Such mitonuclear discordance can have important evolutionary and ecological consequences. In the present study, we aimed to check whether there was any discordance between mtDNA and nuclear DNA in the bank vole population in the contact zone of its two mtDNA lineages. We analysed the population genetic structure of bank voles using genome-wide genetic data (SNPs) and diversity of sequenced heart transcriptomes obtained from selected individuals from three populations inhabiting areas outside the contact zone. The SNP genetic structure of the populations confirmed the presence of at least two genetic clusters, and such division was concordant with the patterns obtained in the analysis of other genetic markers and functional genes. However, genome-wide SNP analyses revealed the more detailed structure of the studied population, consistent with more than two bank vole recolonisation waves, as recognised previously in the study area. We did not find any significant differences between individuals representing two separate mtDNA lineages of the species in functional genes coding for protein-forming complexes, which are involved in the process of cell respiration in mitochondria. We concluded that the contemporary genetic structure of the populations and the width of the contact zone were shaped by climatic and environmental factors rather than by genetic barriers. The studied populations were likely isolated in separate Last Glacial Maximum refugia for insufficient amount of time to develop significant genetic differentiation.
The European Variation Archive (EVA) is a primary open repository for archiving, accessioning, and distributing genome variation including single nucleotide variants, short insertions and deletions (indels), and larger structural variants (SVs) in any species. The EVA has archived more than 3.8 billion variants across 2,374 studies and 227 species since its launch in 2014. A key function of the EVA as a long-term data archive is to provide globally unique persistent identifiers for all discovered variant loci (RS identifiers) so that they can be referenced in publications, cross-linked between databases, and integrated with successive reference genome builds. The EVA periodically releases stable versions of all the RS identifiers it creates. The third release made available in February 2022 featured 1.2 billion RS loci across its supported species. An important new feature of this release is that almost half RS identifiers loci have been updated to a newer assembly to synchronise with the one supported by Ensembl making it much easier to use the two resources in conjunction. The release data is available primarily via FTP1 but all the underlying RS identifiers are queryable on the EVA website[2] and the EVA REST API[3]. Other services for researchers include: standard variant annotation, calculation of population statistics, and an intuitive browser to view and download queried variants in either Variant Call Format (VCF) or Comma-Separated Value (CSV) files backed up by a comprehensive REST API to query and export variant and genotypes programmatically. The EVA also contributes to the maintenance of the Variant Call Format (VCF) specification and has implemented a validation suite to ensure correctness of all the submissions made to the archive. This suite, as well as the rest of our software is freely available on GitHub[4] and has amassed ~7300 downloads worldwide. 1 ftp.ebi.ac.uk/pub/databases/eva/rs_releases/release_3/ 2 https://www.ebi.ac.uk/eva 3 https://www.ebi.ac.uk/eva/webservices/identifiers/swagger-ui.html 4 https://github.com/ebivariation
In this opinion article, we discuss the formatting of files from (plant) genotyping studies, in particular the formatting of metadata in Variant Call Format (VCF) files. The flexibility of the VCF format specification facilitates its use as a generic interchange format across domains but can lead to inconsistency between files in the presentation of metadata. To enable fully autonomous machine actionable data flow, generic elements need to be further specified. We strongly support the merits of the FAIR principles and see the need to facilitate them also through technical implementation specifications. They form a basis for the proposed VCF extensions here. We have learned from the existing application of VCF that the definition of relevant metadata using controlled standards, vocabulary and the consistent use of cross-references via resolvable identifiers (machine-readable) are particularly necessary and propose their encoding. VCF is an established standard for the exchange and publication of genotyping data. Other data formats are also used to capture variant data (for example, the HapMap and the gVCF formats), but none currently have the reach of VCF. For the sake of simplicity, we will only discuss VCF and our recommendations for its use, but these recommendations could also be applied to gVCF. However, the part of the VCF standard relating to metadata (as opposed to the actual variant calls) defines a syntactic format but no vocabulary, unique identifier or recommended content. In practice, often only sparse descriptive metadata is included. When descriptive metadata is provided, proprietary metadata fields are frequently added that have not been agreed upon within the community which may limit long-term and comprehensive interoperability. To address this, we propose recommendations for supplying and encoding metadata, focusing on use cases from plant sciences. We expect there to be overlap, but also divergence, with the needs of other domains.
Abstract The European Variation Archive (EVA; https://www.ebi.ac.uk/eva/) is a resource for sharing all types of genetic variation data (SNPs, indels, and structural variants) for all species. The EVA was created in 2014 to provide FAIR access to genetic variation data and has since grown to be a primary resource for genomic variants hosting >3 billion records. The EVA and dbSNP have established a compatible global system to assign unique identifiers to all submitted genetic variants. The EVA is active within the Global Alliance of Genomics and Health (GA4GH), maintaining, contributing and implementing standards such as VCF, Refget and Variant Representation Specification (VRS). In this article, we describe the submission and permanent accessioning services along with the different ways the data can be retrieved by the scientific community.
This study reveals the genomic architecture of a rapidly evolving mutation which segregates as a single-locus, X-linked trait -- flatwing -- in wild Hawaiian field crickets (Teleogryllus oceanicus). Flatwingsilences males by eliminating sound-producing structures on their forewings. Silence protects them from an acoustically-orienting parasitoid fly (Ormia ochracea), but interferes with their ability to attract and court females for mating. Silent crickets spread rapidly on several Hawaiian islands under pressure from the flies, representing one of the fastest rates of evoutionary change documented in the wild. Here we present an annotated genome sequence of T. oceanicus along with a linkage map and QTL analysis of the trait derived from RAD-sequencing of a backcrossed mapping population. RNA-seq was used to probe the functional pathways affected by the mutation during early development, and pleiotropic effects on another signaling trait, cuticular hydrocarbons, were assessed and genetically mapped.
Oestrogenic wastewater treatment works (WwTW) effluents discharged into UK rivers have been shown to affect sexual development, including inducing intersex, in wild roach (Rutilus rutilus). This can result in a reduced breeding capability with potential population level impacts. In the absence of a sex probe for roach it has not been possible to confirm whether intersex fish in the wild arise from genetic males or females, or whether sex reversal occurs in the wild, as this condition can be induced experimentally in controlled exposures to WwTW effluents and a steroidal oestrogen. Using restriction site-associated DNA sequencing (RAD-seq), we identified a candidate for a genetic sex marker and validated this marker as a sex probe through PCR analyses of samples from wild roach populations from nonpolluted rivers. We also applied the sex marker to samples from roach exposed experimentally to oestrogen and oestrogenic effluents to confirm suspected phenotypic sex reversal from males to females in some treatments, and also that sex-reversed males are able to breed as females. We then show, unequivocally, that intersex in wild roach populations results from feminisation of males, but find no strong evidence for complete sex reversal in wild roach at river sites contaminated with oestrogens. The discovered marker has utility for studies in roach on chemical effects, wild stock assessments, and reducing the number of fish used where only one sex is required for experimentation. Furthermore, we show that the marker can be applied nondestructively using a fin clip or skin swab, with animal welfare benefits.
RNA-seq data collected from Teleogryllus oceanicus of silent and singing morphs at embryonic stages
Evolutionary adaptation is generally thought to occur through incremental mutational steps, but large mutational leaps can occur during its early stages. These are challenging to study in nature due to the difficulty of observing new genetic variants as they arise and spread, but characterizing their genomic dynamics is important for understanding factors favoring rapid adaptation. Here, we report genomic consequences of recent, adaptive song loss in a Hawaiian population of field crickets (Teleogryllus oceanicus). A discrete genetic variant, flatwing, appeared and spread approximately 15 years ago. Flatwing erases sound-producing veins on male wings. These silent flatwing males are protected from a lethal, eavesdropping parasitoid fly. We sequenced, assembled and annotated the cricket genome, produced a linkage map, and identified a flatwing quantitative trait locus covering a large region of the X chromosome. Gene expression profiling showed that flatwing is associated with extensive genome-wide effects on embryonic gene expression. We found that flatwing male crickets express feminized chemical pheromones. This male feminizing effect, on a different sexual signaling modality, is genetically associated with the flatwing genotype. Our findings suggest that the early stages of evolutionary adaptation to extreme pressures can be accompanied by greater genomic and phenotypic disruption than previously appreciated, and highlight how abrupt adaptation might involve suites of traits that arise through pleiotropy or genomic hitchhiking.
Secondary trait loss is widespread and has profound consequences, from generating diversity to driving adaptation. Sexual trait loss is particularly common. Its genomic impact is challenging to reconstruct because most reversals occurred in the distant evolutionary past and must be inferred indirectly, and questions remain about the extent of disruption caused by pleiotropy, altered gene expression and loss of homeostasis. We tested the genomic signature of recent sexual signal loss in Hawaiian field crickets, Teleogryllus oceanicus . Song loss is controlled by a sex-linked Mendelian locus, flatwing , which feminises male wings by erasing sound-producing veins. This variant spread rapidly under pressure from an eavesdropping parasitoid fly. We sequenced, assembled and annotated the T. oceanicus genome, produced a high-density linkage map, and localised flatwing on the X chromosome. We characterised pleiotropic effects of flatwing , including changes in embryonic gene expression and alteration of another sexual signal, chemical pheromones. Song loss is associated with pleiotropy, hitchhiking and genome-wide regulatory disruption which feminises flatwing male pheromones. The footprint of recent adaptive trait loss illustrates R. A. Fisher's influential prediction that variants with large mutational effect sizes can invade genomes during the earliest stages of adaptation to extreme pressures, despite having severely disruptive genomic consequences.
The heritability of classical Hodgkin lymphoma (cHL) has yet to be fully deciphered. We report a family with five members diagnosed with nodular sclerosis cHL. Genetic analysis of the family provided evidence of linkage at chromosomes 2q35-37, 3p14-22 and 21q22, with logarithm of odds score >2. We excluded the possibility of common genetic variation influencing cHL risk at regions of linkage, by analysing GWAS data from 2,201 cHL cases and 12,460 controls. Whole exome sequencing of affected family members identified the shared missense mutations p.(Arg76Gln) in FAM107A and p.(Thr220Ala) in SLC26A6 at 3p21 as being predicted to impact on protein function. FAM107A expression was shown to be low or absent in lymphoblastoid cell lines and SLC26A6 expression lower in lymphoblastoid cell lines derived from p.(Thr220Ala) mutation carriers. Expression of FAM107A and SLC26A6 was low or absent in Hodgkin Reed-Sternberg (HRS) cell lines and in HRS cells in Hodgkin lymphoma tissue. No sequence variants were detected in KLHDC8B, a gene previously suggested as a cause of familial cHL linked to 3p21. Our findings provide evidence for candidate gene susceptibility to familial cHL.
Linking intraspecific and interspecific divergence is an important challenge in speciation research. X chromosomes are expected to evolve faster than autosomes and disproportionately contribute to reproductive barriers, and comparing genetic variation on X and autosomal markers within and between species can elucidate evolutionary processes that shape genome variation. We performed RADseq on a 16 population transect of two closely related Australian cricket species, Teleogryllus commodus and T. oceanicus, covering allopatry and sympatry. This classic study system for sexual selection provides a rare exception to Haldane's rule, as hybrid females are sterile. We found no evidence of recent introgression, despite the fact that the species coexist in overlapping habitats in the wild and interbreed in the laboratory. Putative X‐linked loci showed greater differentiation between species compared with autosomal loci. However, population differentiation within species was unexpectedly lower on X‐linked markers than autosomal markers, and relative X‐to‐autosomal genetic diversity was inflated above neutral expectations. Populations of both species showed genomic signatures of recent population expansions, but these were not strong enough to account for the inflated X/A diversity. Instead, most of the excess polymorphism on the X could better be explained by sex‐biased processes that increase the relative effective population size of the X, such as interspecific variation in the strength of sexual selection among males. Taken together, the opposing patterns of diversity and differentiation at X versus autosomal loci implicate a greater role for sex‐linked genes in maintaining species boundaries in this system.