Ultra-multiplex PCR has become essential for targeted amplicon sequencing in molecular biology and clinical genetics. We implemented PCRpanel, a command-line Java application with a companion web interface that jointly optimises primer thermodynamics, linguistic sequence complexity, primer-dimer interactions, and multiplex-aware, reference-guided specificity and repeat masking for automated ultra-multiplex panel design. We benchmarked PCRpanel in silico against NGS-PrimerPlex and Olivar and validated it experimentally by designing tiled primer pools targeting all exonic regions of COL4A3, COL4A4, COL4A5 and COL4A6 and amplifying DNA from four clinical Alport syndrome samples. As this was a pilot run aimed at initial validation of the design, sequencing is reported for one representative patient sample, in which each primer pool was amplified separately and sequenced at two template dilutions, giving 14 libraries on the Illumina MiSeq platform; pool 1.1 was run as a 50-primer sub-pool during protocol optimisation. Reads were aligned to GRCh38 with the DRAGEN Bio-IT Platform v4.4.4 and coverage was quantified over an exon-restricted target interval of 45,372 bp derived from GENCODE v49. Duplicate marking was disabled, because amplicon molecules generated by a common primer pair are positionally indistinguishable from PCR duplicates. Here we present PCRpanel, a command-line Java application for the automated design of thermodynamically optimised primer panels for short- and long-read amplicon sequencing. PCRpanel supports workflows from simple two-primer assays to ultra-high-plex designs containing hundreds of amplicons and accommodates targets ranging from complete viral genomes and eukaryotic gene families to structural-variant breakpoints and environmental metagenomes. The software jointly optimises sequence complexity, thermodynamic stability, primer-dimer interactions, and off-target amplification, enabling robust primer selection even in repetitive or homologous regions. Efficient algorithms generate ultra-high-plex panels within seconds to minutes on standard hardware, up to 15 min when the full human reference genome is screened, and both gene-specific and universal designs are supported across homologous gene families. We validated the approach by designing 237 primer pairs (474 primers) targeting all exonic regions of the three established Alport syndrome genes COL4A3, COL4A4 and COL4A5, together with the adjacent COL4A6. Experimental validation on four clinical Alport syndrome DNA samples confirmed successful amplification of all 237 primer pairs. As this was a pilot run, sequencing of one representative patient sample, in which each primer pool was amplified separately and sequenced at two template dilutions as 14 libraries (pool 1.1 as a 50-primer sub-pool during protocol optimisation), was used to evaluate analytical performance: read alignment rates were 94.2–99.3
Powdery mildew is a group of plant diseases caused by fungi of the Erysiphaceae family and threatening virtually all important crop plants. The main causative agents of powdery mildew in tomato are the fungi Leveilula taurica and Erysiphe neolycopersici. The genetic basis of tomato resistance to powdery mildew includes a series of six Ol-loci conferring resistance to E. neolycopersici and a single Lv locus of resistance to L. taurica. S-genes such as SlMlo1, SlPMR4, and SlDMR1 confer resistance against multiple powdery mildew species and thus are considered more promising targets for molecular breeding using methods of genome editing by CRISPR/Cas system. Specific resistance to powdery mildew is realized through specific leucine-rich repeat (LRR) proteins, as a general mechanism of pathogen-specific plant immunity. The present review focuses on the current advances in the tomato breeding for resistance to powdery mildew, based on the known molecular mechanisms of susceptibility and resistance.
Recent advances in bioinstrumentation, such as the development of long-read sequencing, have reignited interest in methods for extracting long, intact nucleic acids from complex samples. One traditional method for this purpose is gel electrophoresis followed by electroelution from gel slices into a salt cushion. However, this method has become largely overlooked because its standard implementation is laborious, time-consuming, and incompatible with high-throughput workflows. In our recent work, we revisited this experimental approach and developed a simple, fast, and efficient method for purifying intact nucleic acids of varying lengths from complex samples. The method is available in both horizontal and vertical electrophoresis configurations, has the potential for automation and scalability, and is suitable for purifying high molecular weight (HMW) DNA for long-read sequencing. In this paper, we discuss the origins of the method, the stages of its development, and its advantages. The successful implementation of the method demonstrates how looking at a traditional technique from a new perspective can help meet the demands of next-generation technologies such as long-read sequencing. Our results highlight the importance of rethinking the applications of well-established yet underutilized methodological approaches in rapidly evolving fields such as biotechnology and bioinstrumentation.
Background: DNA isolation is the first step for many molecular workflows, and success depends on removing contaminants such as polysaccharides, polyphenols, lipids, pigments, humic substances, and other inhibitors. Sequencing is especially sensitive to DNA purity. With third-generation long-read platforms, purity alone is not enough high molecular weight and intact DNA are essential to realize reads spanning tens to hundreds of kilobases. Materials and methods: We present a DNA purification method combining agarose gel electrophoresis with electroelution. It exploits the reduced electrophoretic mobility of DNA in high-salt conditions. After DNA separation in a standard gel, a high-salt gel block is placed ahead of the DNA path, leaving a gap (sample collection reservoir). Reapplying current causes DNA to migrate into the gap, where it slows and accumulates. DNA is then easily collected by pipetting and used directly or after desalting. This cost-effective method requires no specialized equipment and is ideal for challenging samples with complex biomolecular mixtures. Results: Isolating high molecular weight (HMW) DNA from complex samples like plants and soil is typically difficult, yielding low purity and quantity. Plant cells contain polysaccharides and polyphenols that interfere with extraction, while soil samples pose greater challenges due to fragmented DNA and humic substances. Despite these obstacles, our method consistently produced HMW DNA with high yield and purity. Conventional methods often recover less than 10% of starting DNA due to poor retention or partitioning losses. In contrast, our approach based on a distinct principle achieved yields up to 50% for plant samples and 30% for soil. Conclusion: This on-gel electroelution method yields HMW DNA for long-read sequencing and other demanding applications from eukaryotes, prokaryotes, organelles (mitochondria, chloroplasts), large DNA viruses, and more. It is particularly useful when other approaches are impractical enabling single step recovery from complex, inhibitor-rich samples with low target abundance. The workflow is low-cost, scalable, and readily automatable. Acknowledgement: This study was funded by the Committee of Science of the Ministry of Science and Higher Education of the Republic of Kazakhstan (grant AP23483529). Key words: nucleic acid, DNA purification and extraction, electroelution, DNA sequencing, long-read sequencing References: Kalendar, R., Ivanov, K.I., Akhmetollayev, I., Kairov, U., Samuilova, O., Burster, T., Zamyatnin, A. An improved method and device for nucleic acid isolation using a high-salt gel electroelution trap. Analytical Chemistry 96(39), 15526-15530 (2024). Kalendar, R., Ivanov, K.I., Samuilova, O., Kairov, U., Zamyatnin, A. Isolation of high- molecular-weight DNA for long-read sequencing using a high-salt gel electroelution trap. Analytical Chemistry 95(48), 17818-17825 (2023).
Background: Esophageal squamous cell carcinoma (ESCC) is one of the most lethal cancers in Central Asia, yet its molecular landscape in Kazakhstan remains poorly explored. However, the fusion-gene landscape in Kazakhstan remains undercharacterized. Fusion transcripts can serve as diagnostic biomarkers or therapeutic targets, motivating a focused survey in a local cohort. Materials and methods: We analyzed paired tumor and adjacent normal tissues from ESCC patients collected in Kazakhstan (34 normal, 33 tumor samples). Tumor specimens were obtained from patients who had undergone Ivor-Lewis esophagectomy without receiving prior chemotherapy or radiotherapy. Total RNA was sequenced (paired-end), reads aligned to GRCh38 (Gencode v44 / CTAT genome library). Fusion calling was performed with two orthogonal tools (STAR-Fusion v1.13.0 and Arriba v2.4.0). Callsets were intersected to prioritize concordant events; candidate junctions were manually inspected in IGV to remove artifacts. Primers for prioritized fusions were designed and PCR followed by Sanger sequencing is underway for orthogonal validation. Results: Intersection of STAR-Fusion and Arriba reduced noise and produced a concise set of high-confidence fusion candidates detected in multiple tumors and absent from matched normals. Manual IGV review confirmed strong split- and spanning-read support at exon boundaries for several recurrent events. Notable recurrent partners include MAP4K5-L2HGDH, ADAMTS2-ENSG00000253652, RAB8A-CIB3, BAZ2B-WDSUB1, HUWE1-SUPT3H, and PLEKHA5-KRAS. Conclusion: Using complementary fusion callers and rigorous manual curation, we generated an assay-ready shortlist of recurrent fusion transcripts in Kazakh ESCC. Pending PCR/Sanger confirmation, these candidates warrant follow-up as potential biomarkers or mechanistic leads for ESCC in Central Asia. Acknowledgement: We thank all the participants in this study. This research has been funded by the Science Committee of the Ministry of Science and Higher Education of the Republic of Kazakhstan (Grants No. AP23490594, BR18574184, BR24993023, BR24992841, BR27199879), Nazarbayev University funding CRP grants 021220CRP2222, 211123CRP1608. Key words: esophageal squamous cell carcinoma, gene fusions, RNA-seq, Kazakhstan
Polymerase chain reaction (PCR) is a quick and easy technique used to detect nucleotide polymorphisms and sequence variations in various fields such as medicine, agriculture, and basic research. There exists a set of PCR variants, collectively called allele-specific PCR, which use competitive reactions in the presence of allele-specific primers to amplify only specific alleles. In this chapter, we present protocols for a bioinformatics tool that can be used to design PCR-based genotyping assays for variable fragment length allele-specific genotyping. This assay is designed in both directions and can target single-nucleotide polymorphisms and insertion/deletions. This chapter provides a step-by-step guide to design a multiplexed primer assay for variable fragment length allele-specific genotyping.
We present a comprehensive web-based platform that integrates a suite of molecular biology tools tailored for PCR-based applications. Key features include custom multiplex tiling PCR panel design for amplicon sequencing, loop-mediated isothermal amplification (LAMP), allele-specific PCR genotyping assay development, Gibson assembly, oligonucleotide analysis and design, and the de novo identification, classification, and visualization of repetitive sequences. The platform supports a wide range of PCR primer design tasks, including standard, multiplex, reverse, long-range, quantitative fluorescence (TaqMan and MGB probe), and bisulfite PCR. In silico PCR analysis ensures primer and probe specificity across diverse applications such as gene discovery via homology analysis, molecular diagnostics, DNA profiling, and repeat sequence identification. The multiplex tiling PCR tool is optimized for both next-generation and third-generation sequencing technologies. LAMP primer design includes support for loop primers that target eight distinct regions of the DNA template. The allele-specific PCR module enables genotyping of single nucleotide polymorphisms, insertions/deletions, multi-nucleotide variants, and haplotypes at specific loci. Additionally, the platform facilitates genome-wide identification, masking, and clustering of repetitive elements, including interspersed repeats and low-complexity sequences such as simple sequence repeats and telomeres. All tools are freely accessible at https://primerdigital.com/tools/.
Изучение и сохранение биологического разнообразия является одной из важнейших проблем современного мира. Сокращение или исчезновение видового и генетического разнообразия представляет угрозу для воспроизводства природных экосистем, приводит к нарушению их стабильности и сохранности. Редкие и эндемичные виды растений обладают меньшим генетическим потенциалом, более подвержены угрозе исчезновения при изменении условий окружающей среды и воздействии антропогенного фактора. В этой связи необходим постоянный мониторинг генетического потенциала редких и исчезающих видов растений. В эти мероприятия входит анализ внутрипопуляционного полиморфизма, генетической дифференциации популяций, с комплексным изучением их морфолого-биологических особенностей. Применение существующих на сегодняшний день молекулярно-генетических методов исследований очень важно при выборе стратегии сохранения редких и исчезающих видов растений, так как данные методы позволяют выявить генетическое разнообразие в популяциях и предпринять меры для снижения генетического дрейфа и сохранения редких генотипов. Для исследования генетического разнообразия растений используются различные типы молекулярно-генетических маркеров, в том числе и ретротранспозоны, мобильные элементы, составляющие большую часть генома у всех эукариот. Метод Inter Primer Binding Sequence (iPBS), разработанный Kalendar et al. позволяет без знаний последовательностей геномов использовать высококонсервативный домен PBS, характерный для всех LTR-ретротранспозонов, для связывания с праймерами. Нуклеотидные последовательности PBS-участков универсальны для всех ретротранспозонов и относятся к высокоповторяющимся повторам, характерным для всех высших эукариот. Поэтому данный метод является универсальной и эффективной маркерной системой для исследования генетического разнообразия между и внутри различных популяций, а также видов и разновидностей. Исследования, проведенные на различных видах растений, таких как Allium sp., Rhodiola sp., Tulipa sp., Paeonia sp.. показали, что данный тип маркеров может быть успешно применен для исследования генетического полиморфизма популяций, а также для выявления сомаклональной изменчивости всех эукариотических организмов. В результате исследований были выявлены популяции с низким уровнем разнообразия, имеющие дефицит эффективных аллелей, что указывает на необходимость разработки стратегии их сохранения. Также были выявлены корреляции между PBS-полиморфизмом и адаптивным потенциалом растений, что подтверждает непосредственное участие мобильных элементов в адаптации растений к экологическим стрессам. Таким образом, анализ биоразнообразия с использованием в качестве маркеров ретротранспозонов позволяет повысить понимание структуры и генетических взаимоотношений популяций дикорастущих растений для последующей разработки стратегии их сохранения.
Background: Scots Pine is one of the main forest-forming species in boreal forests; it has great economic and ecological significance. This study aimed to develop and test primers for detecting nucleotide polymorphisms in genes that are promising for detecting adaptive genetic variability in populations of Pinus sylvestris in the Urals and adjacent territories. Objectives: The objects of the study were 13 populations of Scots Pine located in the Perm Territory, Chelyabinsk Region, and the Republic of Bashkortostan. Results: Sixteen pairs of primers to loci of potentially adaptively significant genes were developed, from which three pairs of primers were selected to detect the nucleotide diversity of the studied populations. The indicator of total haplotype diversity determined in the three studied loci varied from 0.620 (Pinus-12 locus) to 0.737 (Pinus-11 locus) and, on average, amounted to 0.662. The nucleotide diversity indicators in P. sylvestris in the study region were, on average, low (π = 0.004, θW = 0.013). Their highest values were found at the Pinus-12 locus (π = 0.005; θW = 0.032), and the lowest were found at the Pinus-15 locus (π = 0.003; θW = 0.002). This indicates that Pinus-15 is the most conserved of the three studied loci. In the three studied P. sylvestris loci associated with adaptation to environmental factors, 97 polymorphic positions were identified. The 13 populations of P. sylvestris are characterized by an average level of genetic diversity (Hd = 0.662; π = 0.004; θ = 0.013). Conclusions: The polymorphic loci of adaptively significant genes of P. sylvestris can help identify the adaptive potential of pine forests in conditions of increasing ambient temperatures.
BackgroundTuberculosis (TB) is a major public health emergency in many countries, including Kazakhstan. Despite the decline in the incidence rate and having one of the highest treatment effectiveness in the world, the incidence rate of TB remains high in Kazakhstan. Social and environmental factors along with host genetics contribute to pulmonary tuberculosis (PTB) incidence. Due to the high incidence rate of TB in Kazakhstan, our research aimed to study the epidemiology and genetics of PTB in Kazakhstan.Materials and methods1,555 participants were recruited to the case–control study. The epidemiology data was taken during an interview. Polymorphisms of selected genes were determined by real-time PCR using pre-designed TaqMan probes.ResultsEpidemiological risk factors like diabetes (χ2 = 57.71, p < 0.001), unemployment (χ2 = 81.1, p < 0.001), and underweight-ranged BMI (<18.49, χ2 = 206.39, p < 0.001) were significantly associated with PTB. VDR FokI (rs2228570) and VDR BsmI (rs1544410) polymorphisms were associated with an increased risk of PTB. A/A genotype of the TLR8 gene (rs3764880) showed a significant association with an increased risk of PTB in Asians and Asian males. The G allele of the rs2278589 polymorphism of the MARCO gene increases PTB susceptibility in Asians and Asian females. VDR BsmI (rs1544410) polymorphism was significantly associated with PTB in Asian females. A significant association between VDR ApaI polymorphism and PTB susceptibility in the Caucasian population of Kazakhstan was found.ConclusionThis is the first study that evaluated the epidemiology and genetics of PTB in Kazakhstan on a relatively large cohort. Social and environmental risk factors play a crucial role in TB incidence in Kazakhstan. Underweight BMI (<18.49 kg/m2), diabetes, and unemployment showed a statistically significant association with PTB in our study group. FokI (rs2228570) and BsmI (rs1544410) polymorphisms of the VDR gene can be used as possible biomarkers of PTB in Asian males. rs2278589 polymorphism of the MARCO gene may act as a potential biomarker of PTB in Kazakhs. BsmI polymorphism of the VDR gene and rs2278589 polymorphism of the MARCO gene can be used as possible biomarkers of PTB risk in Asian females as well as VDR ApaI polymorphism in Caucasians.
Nucleic acid amplification assays represent a pivotal category of methodologies for targeted sequence detection within contemporary biological research, boasting diverse utility in diagnostics, identification, and DNA sequencing. The foundational principles of these assays have been extrapolated to various simple and intricate nucleic acid amplification technologies. Concurrently, a burgeoning trend toward computational or virtual methodologies is exemplified by in silico PCR analysis. In silico PCR analysis is a valuable and productive adjunctive approach for ensuring primer or probe specificity across a broad spectrum of PCR applications encompassing gene discovery through homology analysis, molecular diagnostics, DNA profiling, and repeat sequence identification. The prediction of primer and probe sensitivity and specificity necessitates thorough database searches, accounting for an optimal balance of mismatch tolerance, sequence similarity, and thermal stability. This software facilitates in silico PCR analyses of both linear and circular DNA templates, including bisulfited treatment DNA, enabling multiple primer or probe searches within databases of varying scales alongside advanced search functionalities. This tool is suitable for processing batch files and is essential for automation when working with large amounts of data.
AbstractEnvironmental DNA (eDNA) technology is an essential tool for monitoring living organisms in ecological research. The combination of eDNA methods with traditional methods of ecological observation can significantly improve the study of the ecology of rare species. Here, we present the development and application of an eDNA approach to identify rare sturgeons in the lower reaches of the Ural River (Zhaiyk) (~1084 km). The presence of representatives of the genus Sturgeon was detected at all sites in spring (nine sites) and autumn (ten sites) while they were absent during the summer period, consistent with their semi‐anadromous ecology. Detection in spring and autumn indicates the passage of spring and winter forms to the lower and upper spawning grounds, respectively. This study confirms the difficulties of species‐specific identification of Eurasian sturgeon and provides the first documented eDNA detection of specimens of the genus Sturgeon in the Ural River. It also provides a biogeographic snapshot of their distribution, experimentally confirming their seasonal migrations in the lower reaches of the river. The successful detection of sturgeon motivates further eDNA surveys of this and other fish species for accurate species identification and population assessment, opening up prospects for the management of these threatened species.
The success of DNA analytical methods, including long-read sequencing, depends on the availability of high-quality, purified DNA. Previously, we developed a method and device for isolating high-molecular-weight (HMW) DNA for long-read sequencing using a high-salt gel electroelution trap. Here, we present an improved version of this method for purifying nucleic acids with high yield and purity from even the most challenging biological samples. The proposed method is a significant improvement over the previously published procedure, offering a simple, fast, and efficient solution for isolating HMW DNA and smaller DNA and RNA molecules. The method utilizes vertical gel electrophoresis in two nested, partially overlapping electrophoretic columns. The upper, smaller-diameter column has a thin layer of agarose gel at the bottom, which separates nucleic acids from impurities, and an electrophoresis buffer on top. After the target nucleic acid has been gel-purified on the upper column, a larger-diameter column with a layer of high-salt gel overlaid with electrophoresis buffer is inserted from below. The purified nucleic acid is then electroeluted into the buffer-filled gap between the separating gel and the high-salt gel, where excess counterions from the high-salt gel slow its migration and cause it to accumulate. The proposed vertical purification system outperforms the previously described horizontal system in terms of ease of use, speed, scalability, and compatibility with high-throughput workflows. Furthermore, the vertical system allows for the sequential purification of several nucleic acid species from the same sample using interchangeable salt-gel columns.
EDITORIAL article Front. Plant Sci., 05 September 2022Sec. Plant Breeding https://doi.org/10.3389/fpls.2022.1026492
Background/Objectives: The spruces of the Picea abies–P. obovata complex have a total range that is the most extensive in the world flora of woody conifers. Hybridization between the nominative species has led to the formation of a wide introgression zone, which probably increases the adaptive potential of the entire species complex. This study aimed to search the genes associated with drought resistance, develop primers for the informative loci of these genes, identify and analyze SNPs, and establish the parameters of nucleotide diversity in the studied populations. Methods: The objects of this study were eight natural populations of the spruce complex in the Urals. Nucleotide sequences related to drought resistance spruce genes with pronounced single-nucleotide substitutions were selected, based on which 16 pairs of primers to their loci were developed and tested. Results: Based on the developed primers, six pairs of primers were chosen to identify SNPs and assess the nucleotide diversity of the studied populations. All selected loci were highly polymorphic (6 to 27 SNPs per locus). It was found that the Pic01 locus is the most variable (Hd = 0.947; π = 0.011) and selectively neutral, and the Pic06 locus is the most conservative (Hd = 0.516; π = 0.002) and has the most significant adaptive value. Conclusions: The nucleotide diversity data for the studied populations reveal similar values among the populations and are consistent with the literature data. The discovered SNPs can be used to identify adaptive genetic changes in spruce populations, which is essential for predicting the effects of climate change.
Epilepsy is one of the most common neurological disorders affecting approximately 50 million people worldwide. It impacts people of all genders and ages, but evidence suggests a higher incidence rate in children and the elderly. Given that childhood epilepsy has the risk of causing developmental epileptic encephalopathy, which is associated with intellectual, behavioral, and/or motor disabilities, proper assessment of children with new-onset epilepsy at an early stage is essential to prevent threats affecting neurodevelopmental processes. The aim of this study was to investigate whole genome sequencing data of children diagnosed with epilepsy. Our results revealed an identification of a novel mutation in a 2-year-old male patient who suffered from recurrent epileptic seizures of unknown etiology. The detected variant is heterozygous and located in gene CHRNA2 (chr8:27321348, NM_000742, c.612G > A, p.Trp204∗) in exon 6. The databases such as Varsome, GeneCards, and NCBI did not reveal any matches with previously identified variants, implying the novelty of the finding. Moreover, according to various prediction tools (MutationTaster, SIFT, CADD, FATHMM-MKL, LRT, DANN, Eigen, and BayesDel), the mutation is characterized as pathogenic, which corresponds to the American College of Medical Genetics and Genomics (ACMG) classification. According to the findings, mutation of the CHRNA2 gene is closely associated with two disorders known as autosomal dominant nocturnal frontal epilepsy (ADNFLE), and benign familial infantile epilepsy (BFIS). Comparison of proband's clinical manifestations showed that it is difficult to attribute precisely the patient's symptoms to either of the conditions, however the evidence suggests that the patient's symptoms are more consistent with those of ADNFLE. In this report, we expanded the spectrum of existing variations in the CHRNA2 gene contributing and associated with the development of epilepsy with the important and novel causative genetic variant.
Genomic repeats are functionally ubiquitous structural units found in all genomes. These repeating patterns have a manifold signature and structure, making identification challenging. To address this challenge, we developed software that can rapidly and accurately detect any type of repeated sequences de novo in genomic sequences in the form of interspersed or clustered repeats. Numerous forms of repeated sequences and “repeat within repeat” patterns can be identified even for very complex sequence variants and for implicit or mixed types of repeat blocks. Direct and inverted-repeat elements, perfect and imperfect microsatellite repeats, and any type of short- or long-tandem repeats belonging to a wide range of organized into higher-order repeat structures of telomers or large satellite sequences can be detected. By combining precision and versatility, our tool significantly contributes to elucidating the intricate landscape of genomic repeats.