Background Large-scale genome projects have advanced the characterization of human genetic diversity; however, the lack of deeply sequenced, multiomic resources representing the Korean population has limited systematic evaluation of population-specific variation and its functional and translational impact. Results We present Korea10K, a population-scale genomic and multiomic resource comprising 10,239 high-depth whole genomes (mean coverage 30×) with integrated molecular and phenotypic data. Analysis of 9,000 unrelated individuals enabled near-complete discovery of rare and ultra-rare variants and supported the construction of a high-resolution, population-specific imputation panel. Koreans exhibit pronounced autosomal genetic homogeneity despite substantial diversity in Y-chromosomal, mitochondrial, and HLA lineages, reflecting long-term demographic continuity. We further identified 16.4 million variants that alter CpG dinucleotide context, revealing widespread sequence-driven modulation of the genomic CG landscape. Notably, population-specific CG-eliminating variants disrupt CpG probe targets in widely used methylation arrays, introducing a systematic source of bias in epigenome-wide association studies and epigenetic clock estimation. Conclusion Korea10K establishes a high-resolution genomic and multiomic reference for the Korean population and reveals a previously unrecognized interaction between genetic variation and epigenomic measurement. This resource provides a foundation for precision medicine and highlights the need for ancestry-aware interpretation of molecular data.
We present Korea10K, the largest genomic dataset of the Korean population, comprising 10,239 high-coverage whole genomes (mean depth 30×) with matched multi-omic profiles and phenotype data. Korea10K achieves complete and near-complete discovery of very rare and ultra-rare alleles, respectively, at 9,000 Korean genomes. This dataset provides the high-quality population-specific imputation panel, enabling accurate inference of low-frequency variants. Admixture analyses confirm the overall genetic homogeneity of the Korean population, despite its diverse Y-chromosomal, mitochondrial, and HLA repertoires. This pattern reflects a long and continuous lineage history characterized by persistent internal admixture and genomic homogenization over thousands of years on the Korean peninsula. We also identified 16.8 million genomic variants that directly modify CG sites by creating or abolishing CG dinucleotides, providing the population-scale evidence of coordinated genomic-epigenomic regulatory mechanism in Koreans. ### Competing Interest Statement S. Jeon is CEO and H.R. is an employee of AgingLab Inc. Y. Cho is an employee of CG Invites Co., LTD and Invites Genomics Co., LTD. B.L. is CEO and J.L. is an employee of nSAGE. G.C. is co-founder of Nebula Genomics and Glottatech.com. The remaining authors declare no competing interests. * ABHet : Heterozygotic Allele Balance AC : Allele Count AF : Allele Frequency AFR : African AI : Artificial Intelligence Alt : Alternative allele AMR : Ad Mixed American BEB : Bengali in Bangladesh CDX : Chinese Dai in Xishuangbanna, China CGV : CG context-associated Variant CHB : Han Chinese in Beijing, China CHS : Han Chinese South, China CLM : Colombian in Medellin, Colombia COVID-19 : Coronavirus Disease 2019 EAS : East Asian EBI : The European Bioinformatics Institute EUR : European ESN : Esan in Nigeria FDR : False Discovery Rate FIN : Finnish in Finland GATK : Genome Analysis Tool Kit GBR : British from England and Scotland GeDiPNet : Genes, Diseases and Pathway Networks GIH : Gujarati Indians in Houston, Texas, USA gnomAD : Genome Aggregation Database GO : Gene Ontology GVCF : Genomic Variant Call Format GWAS : Genome-Wide Association Study GWD : Gambian in Western Division – Mandinka HC : Health Check-up HLA : Histocompatibility Leukocyte Antigen HWE : Hardy-Weinburg Equilibrium IBS : Iberian populations in Spain ISOGG : International Society of Genetic Genealogy ITU : Indian Telugu in the UK JPT : Japanese in Tokyo, Japan KGP : Korean Genome Project KHV : Kinh in Ho Chi Minh City, Vietnam KOR : Korean in Korea10K KOREF : Korean Reference KPGP : Korean Personal Genome Project KEGG : Kyoto Encyclopedia of Gene and Genome LD : Linkage Disequilibrium LEPR : Leptin Receptor LQ : Lifestyle Questionnaire LWK : Luhya in Webuye, Kenya MAF : Minor Allele Frequency MFP : ModelFinder MHC : Major Histocompatibility Complex MSL : Mende in Sierra Leone MTA : Material Transfer Agreement MXL : Mexican Ancestry in Los Angeles CA USA PBMC : Peripheral Blood Mononuclear Cells PCA : Principal Component Analysis PEL : Peruvian in Lima, Peru PUR : Punjabi in Lahore, Pakistan\ QC : Quality Control QTL : Quantitative Trait Loci R² : Squared Correlation Coefficients rCRS : revised Cambridge Reference Sequence RSRS : Reconstructed Sapiens Reference Sequence SAS : South Asian SD : Standard Deviation SLC15A5 : Solute Carrier Family 15 Member 5 SNV : Single Nucleotide Variant TOPMed : Trans-Omics for Precision Medicine TSI : Toscani in Italia TWAS : Transcriptome-Wide Association Study VEP : Variants Effect Predictor VQSR : Variant Quality Score Recalibration WGS : Whole-Genome Sequencing YRI : Yoruba in Ibadan, Nigeria 1KGP : The 1000 Genome Project Ulsan National Institute of Science and Technology, https://ror.org/017cjz748, 1.200108.01, 1.200047.01 Ministry of SMEs and Startups, 1425157301, 1425156792, 1425157253 Ministry of Trade, Industry and Energy, 20016225, RS-2024-00435468
BACKGROUND:Phenome-wide association studies (PheWASs) have been conducted on Asian populations, including Koreans, but many were based on chip or exome genotyping data. Such studies have limitations regarding whole genome-wide association analysis, making it crucial to have genome-to-phenome association information with the largest possible whole genome and matched phenome data to conduct further population-genome studies and develop health care services based on population genomics. RESULTS:Here, we present 4,157 whole genome sequences (Korea4K) coupled with 107 health check-up parameters as the largest genomic resource of the Korean Genome Project. It encompasses most of the variants with allele frequency >0.001 in Koreans, indicating that it sufficiently covered most of the common and rare genetic variants with commonly measured phenotypes for Koreans. Korea4K provides 45,537,252 variants, and half of them were not present in Korea1K (1,094 samples). We also identified 1,356 new genotype-phenotype associations that were not found by the Korea1K dataset. Phenomics analyses further revealed 24 significant genetic correlations, 14 pleiotropic associations, and 127 causal relationships based on Mendelian randomization among 37 traits. In addition, the Korea4K imputation reference panel, the largest Korean variants reference to date, showed a superior imputation performance to Korea1K across all allele frequency categories. CONCLUSIONS:Collectively, Korea4K provides not only the largest Korean genome data but also corresponding health check-up parameters and novel genome-phenome associations. The large-scale pathological whole genome-wide omics data will become a powerful set for genome-phenome level association studies to discover causal markers for the prediction and diagnosis of health conditions in future studies.
The DNA Features pipeline is the analysis pipeline at EMBL-EBI that annotates repeat elements, including transposable elements. With Ensembl’s goal to stay at the cutting edge of genome annotation, we proved that this pipeline needed an update. We then created a new analysis that allowed the Ensembl database to store the repeat classification from the PGSB repeat classification (Recat). This new dataset was then fetched using Perl scripts and used to prove that the pipeline modification induced a gain in sensitivity. Finally, we performed a comparative analysis of transposable element distribution in all plant species available, raising new questions about transposable elements in certain branches of the taxonomic tree.
We present LT1, the first high-quality human reference genome from the Baltic States. LT1 is a female de novo human reference genome assembly, constructed using 57× nanopore long reads and polished using 47× short paired-end reads. We utilized 72 GB of Hi-C chromosomal mapping data for scaffolding, to maximize assembly contiguity and accuracy. The contig assembly of LT1 was 2.73 Gbp in length, comprising 4490 contigs with an NG50 value of 12.0 Mbp. After scaffolding with Hi-C data and manual curation, the final assembly has an NG50 value of 137 Mbp and 4699 scaffolds. Assessment of gene prediction quality using Benchmarking Universal Single-Copy Orthologs (BUSCO) identified 89.3% of the single-copy orthologous genes included in the benchmark. Detailed characterization of LT1 suggests it has 73,744 predicted transcripts, 4.2 million autosomal SNPs, 974,616 short indels, and 12,079 large structural variants. These data may be used as a benchmark for further in-depth genomic analyses of Baltic populations.
We present 4,157 whole-genome sequences (Korea4K) coupled with 107 health check-up parameters as the largest whole genomic resource of Koreans. Korea4K provides 45,537,252 variants and encompasses most of the common and rare variants in Koreans. We identified 1,356 new geno-phenotype associations which were not found by the previous Korea1K dataset. Phenomics analyses revealed 24 genetic correlations, 1,131 pleiotropic variants, and 127 causal relationships from Mendelian randomization. Moreover, the Korea4K imputation reference panel showed a superior imputation performance to Korea1K. Collectively, Korea4K provides the most extensive genomic and phenomic data resources for discovering clinically relevant novel genome-phenome associations in Koreans.
Coronavirus disease, COVID-19 (coronavirus disease 2019), caused by SARS-CoV-2 (severe acute respiratory syndrome coronavirus 2), has a higher case fatality rate in European countries than in others, especially East Asian ones. One potential explanation for this regional difference is the diversity of the viral infection efficiency. Here, we analyzed the allele frequencies of a nonsynonymous variant rs12329760 (V197M) in the TMPRSS2 gene, a key enzyme essential for viral infection and found a significant association between the COVID-19 case fatality rate and the V197M allele frequencies, using over 200,000 present-day and ancient genomic samples. East Asian countries have higher V197M allele frequencies than other regions, including European countries which correlates to their lower case fatality rates. Structural and energy calculation analysis of the V197M amino acid change showed that it destabilizes the TMPRSS2 protein, possibly negatively affecting its ACE2 and viral spike protein processing.
BACKGROUND:DNBSEQ-T7 is a new whole-genome sequencer developed by Complete Genomics and MGI using DNA nanoball and combinatorial probe anchor synthesis technologies to generate short reads at a very large scale-up to 60 human genomes per day. However, it has not been objectively and systematically compared against Illumina short-read sequencers. FINDINGS:By using the same KOREF sample, the Korean Reference Genome, we have compared 7 sequencing platforms including BGISEQ-500, DNBSEQ-T7, HiSeq2000, HiSeq2500, HiSeq4000, HiSeqX10, and NovaSeq6000. We measured sequencing quality by comparing sequencing statistics (base quality, duplication rate, and random error rate), mapping statistics (mapping rate, depth distribution, and percent GC coverage), and variant statistics (transition/transversion ratio, dbSNP annotation rate, and concordance rate with single-nucleotide polymorphism [SNP] genotyping chip) across the 7 sequencing platforms. We found that MGI platforms showed a higher concordance rate for SNP genotyping than HiSeq2000 and HiSeq4000. The similarity matrix of variant calls confirmed that the 2 MGI platforms have the most similar characteristics to the HiSeq2500 platform. CONCLUSIONS:Overall, MGI and Illumina sequencing platforms showed comparable levels of sequencing quality, uniformity of coverage, percent GC coverage, and variant accuracy; thus we conclude that the MGI platforms can be used for a wide range of genomics research fields at a lower cost than the Illumina platforms.
The Welfare Genome Project (WGP) provided 1,000 healthy Korean volunteers with detailed genetic and health reports to test the social perception of integrating personal genetic and healthcare data at a large-scale. WGP was launched in 2016 in the Ulsan Metropolitan City as the first large-scale genome project with public participation in Korea. The project produced a set of genetic materials, genotype information, clinical data, and lifestyle survey answers from participants aged 20–96. As compensation, the participants received a free general health check-up on 110 clinical traits, accompanied by a genetic report of their genotypes followed by genetic counseling. In a follow-up survey, 91.0% of the participants indicated that their genetic reports motivated them to improve their health. Overall, WGP expanded not only the general awareness of genomics, DNA sequencing technologies, bioinformatics, and bioethics regulations among all the parties involved, but also the general public’s understanding of how genome projects can indirectly benefit their health and lifestyle management. WGP established a data construction framework for not only scientific research but also the welfare of participants. In the future, the WGP framework can help lay the groundwork for a new personalized healthcare system that is seamlessly integrated with existing public medical infrastructure.
Cymbidium goeringii, commonly known as the spring orchid, has long been favoured for horticultural purposes in Asian countries. It is a popular orchid with much demand for improvement and development for its valuable varieties. Until now, its reference genome has not been published despite its popularity and conservation efforts. Here, we report the de novo assembly of the C. goeringii genome, which is the largest among the orchids published to date, using a strategy that combines short- and long-read sequencing and chromosome conformation capture (Hi-C) information. The total length of all scaffolds is 3.99 Gb, with an N50 scaffold size of 178.2 Mb. A total of 29,556 protein-coding genes were annotated and 3.55 Gb (88.87% of genome) repetitive sequences were identified. We constructed pseudomolecular chromosomes using Hi-C, incorporating 89.4% of the scaffolds in 20 chromosomes. We identified 220 expanded and 106 contracted genes families in C. goeringii after divergence from its close relative. We also identified new gene families, resistance gene analogues and changes within the MADS-box genes, which control a diverse set of developmental processes during orchid evolution. Our high quality chromosomal-level assembly of C. goeringii can provide a platform for elucidating the genomic evolution of orchids, mining functional genes for agronomic traits and for developing molecular markers for accelerated breeding as well as accelerating conservation efforts.
We present LT1, the first high-quality human reference genome from the Baltic States. LT1 is a female de novo human reference genome assembly constructed using 57× of ultra-long nanopore reads and 47× of short paired-end reads. We also utilized 72 Gb of Hi-C chromosomal mapping data to maximize the assembly’s contiguity and accuracy. LT1’s contig assembly was 2.73 Gbp in length comprising of 4,490 contigs with an N50 value of 13.4 Mbp. After scaffolding with Hi-C data and extensive manual curation, we produced a chromosome-scale assembly with an N50 value of 138 Mbp and 4,699 scaffolds. Our gene prediction quality assessment using BUSCO identify 89.3% of the single-copy orthologous genes included in the benchmarking set. Detailed characterization of LT1 suggested it has 73,744 predicted transcripts, 4.2 million autosomal SNPs, 974,000 short indels, and 12,330 large structural variants. These data are shared as a public resource without any restrictions and can be used as a benchmark for further in-depth genomic analyses of the Baltic populations.
Coronavirus disease (COVID-19), caused by SARS-CoV-2, has a higher case fatality rate (CFR) in European ethnic groups than in others, especially East Asians. One explanation to this phenomenon might be TMPRSS2, a key processing enzyme essential for viral infection. Here, we analyzed the allele frequencies of two nonsynonymous variants rs12329760 (V197M) and rs75603675 (G8V) in the TMPRSS2 gene using over 200,000 present-day and ancient genomic samples. We found a significant association between the CFR of COVID-19 and the allele frequencies of the two variants. Interestingly, they had opposing effects on the CFR: inverse correlation by V197, proportional correlation by G8V. East Asians have higher V197M and lower G8V allele frequencies than Europeans, possibly endowing resistance against SARS-CoV-2. Structural and energy calculation analysis of the V197M amino acid change showed that it destabilizes the TMPRSS2 protein, possibly affecting its ACE2 and viral spike protein processing negatively, ultimately resulting in reduced SARS-CoV-2 infection efficiency and CFR in East Asian ethnic groups.
We present the initial phase of the Korean Genome Project (Korea1K), including 1094 whole genomes (sequenced at an average depth of 31×), along with data of 79 quantitative clinical traits. We identified 39 million single-nucleotide variants and indels of which half were singleton or doubleton and detected Korean-specific patterns based on several types of genomic variations. A genome-wide association study illustrated the power of whole-genome sequences for analyzing clinical traits, identifying nine more significant candidate alleles than previously reported from the same linkage disequilibrium blocks. Also, Korea1K, as a reference, showed better imputation accuracy for Koreans than the 1KGP panel. As proof of utility, germline variants in cancer samples could be filtered out more effectively when the Korea1K variome was used as a panel of normals compared to non-Korean variome sets. Overall, this study shows that Korea1K can be a useful genotypic and phenotypic resource for clinical and ethnogenetic studies.
Background Early diagnosis and continuous monitoring are necessary for an efficient management of cervical cancers (CC). Liquid biopsy, such as detecting circulating tumor DNA (ctDNA) from blood, is a simple, non-invasive method for testing and monitoring cancer markers. However, tumor-specific alterations in ctDNA have not been extensively investigated or compared to other circulating biomarkers in the diagnosis and monitoring of the CC. Therfore, Next-generation sequencing (NGS) analysis with blood samples can be a new approach for highly accurate diagnosis and monitoring of the CC. Method Using a bioinformatics approach, we designed a panel of 24 genes associated with CC to detect and characterize patterns of somatic single-nucleotide variations, indels, and copy number variations. Our NGS CC panel covers most of the genes in The Cancer Genome Atlas (TCGA) as well as additional cancer driver and tumor suppressor genes. We profiled the variants in ctDNA from 24 CC patients who were being treated with systemic chemotherapy and local radiotherapy at the Jeonbuk National University Hospital, Korea. Result Eighteen out of 24 genes in our NGS CC panel had mutations across the 24 CC patients, including somatic alterations of mutated genes ( ZFHX3– 83%, KMT2C- 79% , KMT2D- 79%, NSD1–67%, ATM- 38% and RNF213 –27%). We demonstrated that the RNF213 mutation could be used potentially used as a monitoring marker for response to chemo- and radiotherapy. Conclusion We developed our NGS CC panel and demostrated that our NGS panel can be useful for the diagnosis and monitoring of the CC, since the panel detected the common somatic variations in CC patients and we observed how these genetic variations change according to the treatment pattern of the patient.
Ensembl Genomes (http://www.ensemblgenomes.org) is an integrating resource for genome-scale data from non-vertebrate species, complementing the resources for vertebrate genomics developed in the context of the Ensembl project (http://www.ensembl.org). Together, the two resources provide a consistent set of interfaces to genomic data across the tree of life, including reference genome sequence, gene models, transcriptional data, genetic variation and comparative analysis. Data may be accessed via our website, online tools platform and programmatic interfaces, with updates made four times per year (in synchrony with Ensembl). Here, we provide an overview of Ensembl Genomes, with a focus on recent developments. These include the continued growth, more robust and reproducible sets of orthologues and paralogues, and enriched views of gene expression and gene function in plants. Finally, we report on our continued deeper integration with the Ensembl project, which forms a key part of our future strategy for dealing with the increasing quantity of available genome-scale data across the tree of life.
We provide a Kazakh whole genome sequence (MJS) and analyses with the largest comparative Kazakh genomic data available to date. We found 102,240 novel SNVs and a high level of heterozygosity. ADMIXTURE analysis confirmed a significant proportion of variations in this individual coming from all continents except Africa and Oceania. A principal component analysis showed neighboring Kalmyk, Uzbek, and Kyrgyz populations to have the strongest resemblance to the MJS genome which reflects fairly recent Kazakh history. MJS's mitochondrial haplogroup, J1c2, probably represents an early European and Near Eastern influence to Central Asia. This was also supported by the heterozygous SNPs associated with European phenotypic features and strikingly similar Kazakh ancestral composition inferred by ADMIXTURE. Admixture (f3) analysis showed that MJS's genomic signature is best described as a cross between the Neolithic East Asian (Devil's Gate1) and the Bronze Age European (Halberstadt_LBA1) components rather than a contemporary admixture.
Ensembl Genomes (http://www.ensemblgenomes.org) is an integrating resource for genome-scale data from non-vertebrate species, complementing the resources for vertebrate genomics developed in the Ensembl project (http://www.ensembl.org). Together, the two resources provide a consistent set of programmatic and interactive interfaces to a rich range of data including genome sequence, gene models, transcript sequence, genetic variation, and comparative analysis. This paper provides an update to the previous publications about the resource, with a focus on recent developments and expansions. These include the incorporation of almost 20 000 additional genome sequences and over 35 000 tracks of RNA-Seq data, which have been aligned to genomic sequence and made available for visualization. Other advances since 2015 include the release of the database in Resource Description Framework (RDF) format, a large increase in community-derived curation, a new high-performance protein sequence search, additional cross-references, improved annotation of non-protein-coding genes, and the launch of pre-release and archival sites. Collectively, these changes are part of a continuing response to the increasing quantity of publicly-available genome-scale data, and the consequent need to archive, integrate, annotate and disseminate these using automated, scalable methods.
Advances in genome sequencing and assembly technologies are generating many high-quality genome sequences, but assemblies of large, repeat-rich polyploid genomes, such as that of bread wheat, remain fragmented and incomplete. We have generated a new wheat whole-genome shotgun sequence assembly using a combination of optimized data types and an assembly algorithm designed to deal with large and complex genomes. The new assembly represents >78% of the genome with a scaffold N50 of 88.8 kb that has a high fidelity to the input data. Our new annotation combines strand-specific Illumina RNA-seq and Pacific Biosciences (PacBio) full-length cDNAs to identify 104,091 high-confidence protein-coding genes and 10,156 noncoding RNA genes. We confirmed three known and identified one novel genome rearrangements. Our approach enables the rapid and scalable assembly of wheat genomes, the identification of structural variants, and the definition of complete gene models, all powerful resources for trait analysis and breeding of this key global crop.