The rapid identification of allele variants in target genes is crucial for accelerating marker-assisted selection (MAS) in plant breeding. Although current high-throughput genotyping methods are efficient in detecting known polymorphisms, they are limited when multiple variant sites are scattered along the gene. This study presents a target amplicon sequencing approach using Oxford Nanopore Technologies (ONT-TAS) to rapidly sequence full-length genes and identify allele variants in sunflower and wheat collections. This procedure combines multiplex PCR and a rapid sequencing kit, significantly reducing the time and cost compared to previous methods. The efficiency of the approach was demonstrated by sequencing four genes (Ahasl1, Ahasl2, Ahasl3, and FAD2) in 40 sunflower genotypes and three genes (Ppo, Wx, and Lox) in 30 wheat genotypes. The ONT-TAS revealed a complete picture of SNPs and InDels distributed over the individual alleles, enabling rapid (4.5 h for PCR and sequencing) characterization of the genetic diversity of the target genes in the germplasm collections. The results showed a significant diversity of the Ahasl1/Ahasl3 and Wx-A/Lox-B genes in the sunflower and wheat collections, respectively. This method offers a high-throughput, cost-effective (USD 3.4 per gene) solution for genotyping and identifying novel allele variants in plant breeding programs.
Transposable element insertions (TEIs) are an important source of genomic innovation by contributing to plant adaptation, speciation, and the production of new varieties. The often large, complex plant genomes make identifying TEIs from short reads difficult and expensive. Moreover, rare somatic insertions that reflect mobilome dynamics are difficult to track using short reads. To address these challenges, we combined Cas9-targeted Nanopore sequencing (CANS) with the novel pipeline NanoCasTE to trace both genetically inherited and somatic TEIs in plants. We performed CANS of the EVADÉ ( EVD ) retrotransposon in wild-type Arabidopsis thaliana and rapidly obtained up to 40x sequence coverage. Analysis of hemizygous T-DNA insertion sites and genetically inherited insertions of the EVD transposon in the ddm1 genome uncovered the crucial role of DNA methylation in shaping EVD insertion preference. We also investigated somatic transposition events of the ONSEN transposon family, finding that genes that are downregulated during heat stress are preferentially targeted by ONSEN s. Finally, we detected hypomethylation of novel somatic insertions for two ONSEN s. CANS and NanoCasTE are effective tools for detecting TEIs and exploring mobilome organization in plants in response to stress and in different genetic backgrounds, as well as screening T-DNA insertion mutants and transgenic plants.
Developing seed is a unique stage of plant development with highly dynamic changes in transcriptome. Here, we aimed to detect the novel previously unannotated, genes of the triticale (x Triticosecale Wittmack, AABBRR genome constitution) genome that are expressed during different stages and at different parts of the developing seed. For this, we carried out the Oxford Nanopore sequencing of cDNA obtained for middle (15 days post-anthesis, dpa) and late (20 dpa) stages of seed development. The obtained data together with our previous direct RNA sequencing of early stage (10 dpa) of seed development revealed 39,914 expressed genes including 7128 (17.6%) genes that were not previously annotated in A, B, and R genomes. The bioinformatic analysis showed that the identified genes belonged to long non-coding RNAs (lncRNAs), protein-coding RNAs, and TE-derived RNAs. The gene set analysis revealed the transcriptome dynamics during seed development with distinct patterns of over-represented gene functions in early and middle/late stages. We performed analysis of the lncRNA genes polymorphism and showed that the genes of some of the tested lncRNAs are indeed polymorphic in the triticale collection. Altogether, our results provide information on thousands of novel loci expressed during seed development that can be used as new targets for GWAS analysis, the marker-assisted breeding of triticale, and functional elucidation.
Тритикале (×Triticosecale Wittm.) в основном является кормовой культурой, однако она обладает не только высокими потенциальными возможностями урожайности, но и повышенной устойчивостью к биотическим и абиотическим факторам окружающей среды [1]. Поэтому изучение принципов и закономерностей регуляции биологических процессов тритикале имеет высокий потенциал для производства и селекции растений. Тритикале имеет огромный геном (24 Gb), большая часть которого состоит из неизвестных генов, таких как, например, некодирующие РНК (нРНК). Известно, что у растений длинные некодирующие РНК (днРНК) регулируют важные процессы в развитии и реакции растений на стресс [2-3]. К настоящему моменту днРНК малоизучены изучены в отношении тритикале. Triticale (×Triticosecale Wittm.) is mainly a fodder crop, but it has not only high yield potential, but also increased resistance to biotic and abiotic environmental factors [1]. Therefore, the study of the principles and regularities of the regulation of biological processes of triticale has a high potential for the production and breeding of plants. Triticale has a huge genome (24 Gb), most of which consists of unknown genes, such as, for example, non-coding RNA (nRNA). In plants, long noncoding RNAs (lncRNAs) are known to regulate important processes in the development and response of plants to stress [2–3]. To date, lncRNAs have been little studied in relation to triticale.
За счет активного роста посевных площадей, тритикале все чаще подвергается воздействию биотических стрессов [1,2]. Поиск среди генотипов тритикале генов устойчивости и их последующее применение в селекционных программах является благоприятным способом защиты зерновых культур от возбудителя стеблевой ржавчины за счёт гибридного происхождения данной зерновой культуры [3]. Коллекция сортов и линий тритикале, среди которых присутствовало 37 яровых и 81 озимых генотипов, из лаборатории маркерной геномной селекции растений и коллекции НЦЗ им. П.П. Лукьяненко, была оценена на наличие продуктов амплификации характерных для широко используемых в селекционных процессах генов устойчивости. Контролем служила коллекция сортов дифференциаторов ФГБНУ ФНЦБЗР. Due to the active growth of sown areas, triticale is increasingly exposed to biotic stresses [1,2]. The search for resistance genes among triticale genotypes and their subsequent use in breeding programs is a favorable way to protect crops from the stem rust pathogen due to the hybrid origin of this crop [3]. A collection of varieties and lines of triticale, among which there were 37 spring and 81 winter genotypes, from the Laboratory of Marker Genomic Plant Breeding and the collection of the N.N. P.P. Lukyanenko, was evaluated for the presence of amplification products characteristic of resistance genes widely used in breeding processes. The collection of varieties of differentiators of the FGBNU FNTsBZR served as the control.
Transposable elements (TEs) contribute not only to genome diversity but also to transcriptome diversity in plants. To unravel the sources of LTR retrotransposon (RTE) transcripts in sunflower, we exploited a recently developed transposon activation method (‘TEgenesis’) along with long-read cDNA Nanopore sequencing. This approach allows for the identification of 56 RTE transcripts from different genomic loci including full-length and non-autonomous RTEs. Using the mobilome analysis, we provided a new set of expressed and transpositional active sunflower RTEs for future studies. Among them, a Ty3/Gypsy RTE called SUNTY3 exhibited ongoing transposition activity, as detected by eccDNA analysis. We showed that the sunflower genome contains a diverse set of non-autonomous RTEs encoding a single RTE protein, including the previously described TR-GAG (terminal repeat with the GAG domain) as well as new categories, TR-RT-RH, TR-RH, and TR-INT-RT. Our results demonstrate that 40% of the loci for RTE-related transcripts (nonLTR-RTEs) lack their LTR sequences and resemble conventional eucaryotic genes encoding RTE-related proteins with unknown functions. It was evident based on phylogenetic analysis that three nonLTR-RTEs encode GAG (HadGAG1-3) fused to a host protein. These HadGAG proteins have homologs found in other plant species, potentially indicating GAG domestication. Ultimately, we found that the sunflower retrotranscriptome originated from the transcription of active RTEs, non-autonomous RTEs, and gene-like RTE transcripts, including those encoding domesticated proteins.
High-copy tandemly organized repeats (TRs), or satellite DNA, is an important but still enigmatic component of eukaryotic genomes. TRs comprise arrays of multi-copy and highly similar tandem repeats, which makes the elucidation of TRs a very challenging task. Oxford Nanopore sequencing data provide a valuable source of information on TR organization at the single molecule level. However, bioinformatics tools for de novo identification of TRs in raw Nanopore data have not been reported so far. We developed NanoTRF, a new python pipeline for TR repeat identification, characterization and consensus monomer sequence assembly. This new pipeline requires only a raw Nanopore read file from low-depth (<1×) genome sequencing. The program generates an informative html report and figures on TR genome abundance, monomer sequence and monomer length. In addition, NanoTRF performs annotation of transposable elements (TEs) sequences within or near satDNA arrays, and the information can be used to elucidate how TR–TE co-evolve in the genome. Moreover, we validated by FISH that the NanoTRF report is useful for the evaluation of TR chromosome organization—clustered or dispersed. Our findings showed that NanoTRF is a robust method for the de novo identification of satellite repeats in raw Nanopore data without prior read assembly. The obtained sequences can be used in many downstream analyses including genome assembly assistance and gap estimation, chromosome mapping and cytogenetic marker development.
Россия является одной из лидирующих стран по производству подсолнечника, получая порядка 15,3 млн. тон в год с площади 10033 тыс. га (1). Признаки, по которым ведется селекция подсолнечника и улучшение его агрокультурных качеств, сильно разграничены в зависимости от производственных задач. Одними из основных направлений селекции подсолнечника являются устойчивость к гербицидам и болезням, повышение качества масла, селекция на декоративность культуры и др. Russia is one of the leading countries in the production of sunflower, receiving about 15.3 million tons per year from an area of 10,033 thousand hectares (1). The traits by which sunflower breeding and improvement of its agricultural qualities are carried out are strongly delimited depending on production tasks. One of the main areas of sunflower breeding is resistance to herbicides and diseases, improving the quality of oil, breeding for ornamental culture, etc.
Long-read data is a great tool to discover new active transposable elements (TEs). However, no ready-to-use tools were available to gather this information from low coverage ONT datasets. Here, we developed a novel pipeline, nanotei, that allows detection of TE-contained structural variants, including individual TE transpositions. We exploited this pipeline to identify TE insertion in the Arabidopsis thaliana genome. Using nanotei, we identified tens of TE copies, including ones for the well-characterized ONSEN retrotransposon family that were hidden in genome assembly gaps. The results demonstrate that some TEs are inaccessible for analysis with the current A. thaliana (TAIR10.1) genome assembly. We further explored the mobilome of the ddm1 mutant with elevated TE activity. Nanotei captured all TEs previously known to be active in ddm1 and also identified transposition of non-autonomous TEs. Of them, one non-autonomous TE derived from (AT5TE33540) belongs to TR-GAG retrotransposons with a single open reading frame (ORF) encoding the GAG protein. These results provide the first direct evidence that TR-GAGs and other non-autonomous LTR retrotransposons can transpose in the plant genome, albeit in the absence of most of the encoded proteins. In summary, nanotei is a useful tool to detect active TEs and their insertions in plant genomes using low-coverage data from Nanopore genome sequencing.
Зерно злаковых культур содержит несколько фракций белков: альбумины, глобулины, проламины, глютелины. Последние две фракции формируют клейковину – нерастворимый в воде комплекс, определяющий основные технологические свойства хлеба. Тритикале является новым и перспективным гибридом пшеницы и ржи, который требует подробный анализ аллельного состава субъединиц запасных белков, что имеет большую практическую значимость, и, в совокупности с молекулярными методами, используется в селекции при создании сортов зерновых культур с высокими хлебопекарными качествами [1]. Гидрофобные белки глютенины вносят наибольший вклад в эластичность и вязкость теста, из которого производятся мучные изделия. Глютенины кодируются полиморфными генами Glu-1, находящиеся на длинном плече хромосом первой группы [2]. В каждом локусе (Glu-A1, Glu-B1 и Glu-R1) расположены два близко сцепленных гена: х– и у–типа. Ген х–типа определяет структуру высокомолекулярных субъединиц с большей молекулярной массой, а ген у-типа – с меньшей [3]. The grain of cereal crops contains several fractions of proteins: albumins, globulins, prolamins, glutelins. The last two fractions form gluten - a water-insoluble complex that determines the main technological properties of bread. Triticale is a new and promising hybrid of wheat and rye, which requires a detailed analysis of the allelic composition of storage protein subunits, which is of great practical importance, and, in combination with molecular methods, is used in breeding to create varieties of cereals with high baking qualities [1]. The hydrophobic proteins glutenin make the greatest contribution to the elasticity and viscosity of the dough from which flour products are made. Glutenins are encoded by polymorphic Glu-1 genes located on the long arm of chromosomes of the first group [2]. Each locus (Glu-A1, Glu-B1 and Glu-R1) contains two closely linked genes: x- and y-type. The x-type gene determines the structure of high-molecular-weight subunits with a higher molecular weight, while the y-type gene determines the structure of smaller ones [3].
Sequencing and epigenetic profiling of target genes in plants are important tasks with various applications ranging from marker design for plant breeding to the study of gene expression regulation. This is particularly interesting for plants with big genome size for which whole-genome sequencing can be time-consuming and costly. In this study, we asked whether recently proposed Cas9-targeted nanopore sequencing (nCATS) is efficient for target gene sequencing for plant species with big genome size. We applied nCATS to sequence the full-length glutenin genes (Glu-1Ax, Glu-1Bx and Glu-1By) and their promoters in hexaploid triticale (X Triticosecale, AABBRR, genome size is 24 Gb). We showed that while the target gene enrichment per se was quite high for the three glutenin genes (up to 645×), the sequencing depth that was achieved from two MinION flowcells was relatively low (5–17×). However, this sequencing depth was sufficient for various tasks including detection of InDels and single-nucleotide variations (SNPs), read phasing and methylation profiling. Using nCATS, we uncovered SNP and InDel variation of full-length glutenin genes providing useful information for marker design and deciphering of variation of individual Glu-1By alleles. Moreover, we demonstrated that glutenin genes possess a ‘gene-body’ methylation epigenetic profile with hypermethylated CDS part and hypomethylated promoter region. The obtained information raised an interesting question on the role of gene-body methylation in glutenin gene expression regulation. Taken together, our work disclosures the potential of the nCATS approach for sequencing of target genes in plants with big genome size.
The intergenic space of plant genomes encodes many functionally important yet unexplored RNAs. The genomic loci encoding these RNAs are often considered “junk”, DNA as they are frequently associated with repeat-rich regions of the genome. The latter makes the annotations of these loci and the assembly of the corresponding transcripts using short RNAseq reads particularly challenging. Here, using long-read Nanopore direct RNA sequencing, we aimed to identify these “junk” RNA molecules, including long non-coding RNAs (lncRNAs) and transposon-derived transcripts expressed during early stages (10 days post anthesis) of seed development of triticale (AABBRR, 2n = 6x = 42), an interspecific hybrid between wheat and rye. Altogether, we found 796 lncRNAs and 20 LTR retrotransposon-related transcripts (RTE-RNAs) expressed at this stage, with most of them being previously unannotated and located in the intergenic as well as intronic regions. Sequence analysis of the lncRNAs provide evidence for the frequent exonization of Class I (retrotransposons) and class II (DNA transposons) transposon sequences and suggest direct influence of “junk” DNA on the structure and origin of lncRNAs. We show that the expression patterns of lncRNAs and RTE-related transcripts have high stage specificity. In turn, almost half of the lncRNAs located in Genomes A and B have the highest expression levels at 10–30 days post anthesis in wheat. Detailed analysis of the protein-coding potential of the RTE-RNAs showed that 75% of them carry open reading frames (ORFs) for a diverse set of GAG proteins, the main component of virus-like particles of LTR retrotransposons. We further experimentally demonstrated that some RTE-RNAs originate from autonomous LTR retrotransposons with ongoing transposition activity during early stages of triticale seed development. Overall, our results provide a framework for further exploration of the newly discovered lncRNAs and RTE-RNAs in functional and genome-wide association studies in triticale and wheat. Our study also demonstrates that Nanopore direct RNA sequencing is an indispensable tool for the elucidation of lncRNA and retrotransposon transcripts.
The centromere is a unique part of the chromosome combining a conserved function with an extreme variability in its DNA sequence. Most of our knowledge about the functional centromere organization is obtained from species with small and medium genome/chromosome sizes while the progress in plants with big genomes and large chromosomes is lagging behind. Here, we studied the genomic organization of the functional centromere in Allium fistulosum and A. cepa, both species with a large genome (13 Gb and 16 Gb/1C, 2n = 2x = 16) and large-sized chromosomes. Using low-depth DNA sequencing for these two species and previously obtained CENH3 immunoprecipitation data we identified two long (1.2 Kb) and high-copy repeats, AfCen1K and AcCen1K. FISH experiments showed that AfCen1K is located in all centromeres of A. fistulosum chromosomes while no AcCen1K FISH signals were identified on A. cepa chromosomes. Our molecular cytogenetic and bioinformatics survey demonstrated that these repeats are partially similar but differ in chromosomal location, sequence structure and genomic organization. In addition, we could conclude that the repeats are transcribed and their RNAs are not polyadenylated. We also observed that these repeats are associated with insertions of retrotransposons and plastidic DNA and the landscape of A. cepa and A. fistulosum centromeric regions possess insertions of plastidic DNA. Finally, we carried out detailed comparative satellitome analysis of A. cepa and A. fistulosum genomes and identified a new chromosome- and A. cepa-specific tandem repeat, TR2CL137, located in the centromeric region. Our results shed light on the Allium centromere organization and provide unique data for future application in Allium genome annotation.
LTR retrotransposons (RTEs) play a crucial role in plant genome evolution and adaptation. Although RTEs are generally silenced in somatic plant tissues under non-stressed conditions, some expressed RTEs (exRTEs) escape genome defense mechanisms. As our understanding of exRTE organization in plants is rudimentary, we systematically surveyed the genomic and transcriptomic organization and mobilome (transposition) activity of sunflower (Helianthus annuus L.) exRTEs. We identified 44 transcribed RTEs in the sunflower genome and demonstrated their distinct genomic features: more recent insertion time, longer open reading frame (ORF) length, and smaller distance to neighboring genes. We showed that GAG-encoding ORFs are present at significantly higher frequencies in exRTEs, compared with non-expressed RTEs. Most exRTEs exhibit variation in copy number among sunflower cultivars and one exRTE Gagarin produces extrachromosomal circular DNA in seedling, demonstrating recent and ongoing transposition activity. Nanopore direct RNA sequencing of full-length RTE RNA revealed complex patterns of alternative splicing in RTE RNAs, resulting in isoforms that carry ORFs for distinct RTE proteins. Together, our study demonstrates that tens of expressed sunflower RTEs with specific genomic organization shape the hidden layer of the transcriptome, pointing to the evolution of specific strategies that circumvent existing genome defense mechanisms.
The results of PCR analysis of the collection of spring triticale accessions for the presence of genes Lr9, Lr12, Lr19, Lr24, Lr25, Lr28, Lr29, and Lr47 (conferring resistance to wheat leaf (brown) rust caused by Puccinia triticina Erikss.) with the use of molecular markers and isogenic lines carrying target genes (as a positive control) are presented in this article. The absence of positive PCR amplification of the DNA markers for the Lr9, Lr24, Lr28, Lr29, and Lr47 genes is observed in all the studied accessions of the spring triticale collection. The accessions Lena 1270, 25AD20, k-1763, k-3256, and Arta 59 are found to carry the Xgwm251 marker allele of the same size as that of isogenic Thatcher line with Lr25. PCR analysis using the LrAg marker shows that such triticale accessions as Pamyati Merezhko, Ulyana, V20-140, S17, PRAG 554/1, C95, 08871, RIL-130 R22-2, 172-1-16, C250, 08857, 09228, 131/17, A2-16-11, POPW9, PRAG 500, C260, Arta116/2, PRAG 554, AVS19883, k-1220, PRAG 553/1, C254, PRAG 518, PRAG 418, R-7-5 RIL202, L2413, and L8-6 carry a fragment close in size to that of the isogenic Thatcher line with Lr19 (used as positive control). Thus, we have shown that the gene pool of spring triticale is extremely depleted in leaf rust resistance genes. Active work is required on the introgression of new resistance genes both from the known donor lines of triticale and from bread wheat.
Bread-making quality is a crucial trait for wheat and triticale breeding. Several genes significantly influence these characteristics, including glutenin genes and the wheat bread-making (wbm) gene. World wheat collection screening showed that only a few percent of cultivars carry the valuable wbm variant, providing a useful source for wheat breeding. In contrast, no such analysis has been performed for triticale (wheat (AABB genome) × rye (RR) amphidiploid) collections. Despite the importance of the wbm gene, information about its origin and genomic organization is lacking. Here, using modern genomic resources available for wheat and its relatives, as well as PCR screening, we aimed to examine the evolution of the wbm gene and its appearance in the triticale genotype collection. Bioinformatics analysis revealed that the wheat Chinese Spring genome does not have the wbm gene but instead possesses the orthologous gene, called wbm-like located on chromosome 7A. The analysis of upstream and downstream regions revealed the insertion of LINE1 (Long Interspersed Nuclear Elements) retrotransposons and Mutator DNA transposon in close vicinity to wbm-like. Comparative analysis of the wbm-like region in wheat genotypes and closely related species showed low similarity between the wbm locus and other sequences, suggesting that wbm originated via introgression from unknown species. PCR markers were developed to distinguish wbm and wbm-like sequences, and triticale collection was screened resulting in the detection of three genotypes carrying wbm-specific introgression, providing a useful source for triticale breeding programs.