This study investigates the genomic basis of immune adaptation in the transparent glass catfish (Kv: Kryptopterus vitreolus), focusing on the loss of the Toll-like receptor 21 (TLR21) gene. Comparative genomic analysis with closely related non-transparent North African catfish (Cg: Clarias gariepinus) revealed 11 TLR genes in the latter, while only 8 TLR genes (KvTLR1, 2, 3, 5, 7, 9, 13, and 20) were retained in the glass catfish, with TLR21 specifically absent. Collinearity analysis confirmed that the genomic region containing TLR21 is conserved across eight siluriform species, with loss exclusively in the glass catfish, supporting its lineage-specific absence. Structural expansion was notable in KvTLR5, KvTLR7, and KvTLR20. Molecular docking indicated that binding stability between CpG oligonucleotides and TLR21 varies significantly, with CpG-B 1681 showing the strongest interaction, which highlights sequence-dependent ligand recognition. Interestingly, absence of the TLR1 gene in another transparent teleost, the X-ray tetra (Pristella maxillaris), suggests that transparent fishes may share an evolutionary trend of lineage-specific TLR gene loss. Together, these findings reveal a distinctive evolutionary trajectory in the innate immune receptor family of transparent fishes and provide new molecular insights into their adaptive immune strategies. These insights will benefit the academic community by improving comparative frameworks for fish innate immunity, and they may inform disease prevention and health management strategies in aquaculture and the ornamental fish trade.
As a protandrous hermaphrodite with natural male-to-female sex change, maroon clownfish (Premnas biaculeatus) serves as an ideal model organism for investigating sequential hermaphroditism. However, genomic resources for this interesting species remain scarce, thereby limiting in-depth research on its unique biological traits. In this study, we generated the first telomere-to-telomere (T2T) gap-free genome assembly of maroon clownfish by integrating multi-platform sequencing data, including MGI short reads, PacBio HiFi long reads, ONT ultra-long reads, and Hi-C sequencing data. The final haplotypic genome spans 884.39 Mb, with all sequences successfully anchored onto 24 chromosomes. This assembly is highly contiguous, with a contig N50 of 37.98 Mb. Comprehensive genomic characterization revealed the precise localization of telomeric repeats and centromeric region within each chromosome. Independent quality assessments, such as QV of 71.01, CRAQ score of 98.98%, and BUSCO completeness of 99.98%, confirmed good assembly accuracy. Additionally, alignment of ONT ultra-long and PacBio HiFi reads to the assembly yielded a high mapping rate exceeding 99%, further validating a good assembly integrity. Repetitive elements constituted 33.51% (296.37 Mb) of the assembled genome, and a total of 24,556 protein-coding genes were annotated. This high-quality T2T genome assembly will not only provide a valuable genetic resource to advance related research in comparative genomics, population genetics, molecular breeding, and functional genomics of maroon clownfish, but also lay a solid foundation for resolving molecular mechanisms underlying its protandrous reproductive strategy.
As a protandrous hermaphroditic fish species with natural sex change from male to female, Asian seabass (Lates calcarifer) represents an attractive model for studying sequential hermaphroditism. In this study, we constructed the first telomere-to-telomere (T2T) gap-free genome assembly of Asian seabass, by integration of MGI short-read, PacBio HiFi long-read, ONT ultra-long and Hi-C sequencing technologies. The haplotypic 614.19 Mb genome sequences were successfully anchored onto 24 chromosomes, demonstrating exceptional contiguity with a contig N50 of 26.57 Mb. Comprehensive annotation revealed precise localization of telomeric repeats and centromeric regions across various chromosomes. Good results from Merqury (QV: 57.8), CRAQ (99.45%) and BUSCO (100%) indicate a high level of accuracy for the assembled genome. ONT ultra-long and PacBio HiFi sequencing data were aligned with the assembly using minimap2, resulting in a mapping rate over 98%. Repetitive elements accounted for 18.18% (111.64 Mb) of the entire genome, and a total of 25,093 protein-coding genes were annotated. This high-quality T2T genome assembly provides a valuable genetic resource for in-depth comparative genomics, population genetics, molecular breeding, and functional studies of this economically important marine species. This reference assembly also facilitates investigations into the detailed molecular mechanisms underlying its unique reproductive strategy of the protandrous hermaphrodite Asian seabass.
The well-known striped catfish (Pangasianodon hypophthalmus), belonging to the order Siluriformes and the family Pangasiidae, has become an important freshwater economic fish species due to its outstanding growth performance, environmental adaptability and disease resistance. In this study, we reported the first telomere-to-telomere (T2T) gap-free genome assembly of striped catfish, generated by integrating MGI short-read, ONT ultra-long read, PacBio HiFi long read and Hi-C data. A total of 772.03 Mb genome sequences were successfully anchored onto 30 chromosomes, demonstrating exceptional contiguity with a high contig N50 of 26.47 Mb. A comprehensive annotation revealed precise structure of telomeric repeats and centromeric region within each chromosome. In the assembled genome, repetitive elements accounted for 39.53% (305.21 Mb), with DNA transposons (20.77%) as the predominant repeat type. A total of 24,596 protein-coding genes were predicted (BUSCO: 96.21%), and 99.22% of these genes were functionally annotated. Our combined results about identification of telomeric repeats and centromeric regions, BUSCO assessment (99.45%), mapping coverage (99.55%), and Clipping information for Revealing Assembly Quality (CRAQ: 99.62%) support the high quality of this genome assembly. These complete chromosomal sequences will serve as a valuable genetic resource for in-depth biological investigations, evolutionary studies, comparative genomics, and molecular breeding to improve economic value of the striped catfish.
Three-spotted seahorse (Hippocampi trimaculata) is a unique fish with important economic and medicinal values, and its total chromosome number is potentially quite different from other seahorse species. Herein, we constructed a chromosome-level genome assembly for this special seahorse by integration of MGI short-read, PacBio HiFi long-read and Hi-C sequencing techniques. A 416.57-Mb haplotypic genome assembly was obtained. Subsequently, 99.38% of its scaffold sequences were anchored onto 18 chromosomes, with identification of 29.1% repeat sequences in the assembled genome. Additional karyotype analysis validated the diploid chromosomes of 2n = 36, which are remarkably different from other seahorses’ 2n = 42 or 44. The genome completeness (BUSCO score: 96.5%, CEGMA score: 97.87%) confirmed that this chromosome-scale assembly is indeed of high quality. Moreover, a total of 18,712 protein-coding genes were annotated, of which 96.36% could be predicted with functions. Based on construction of a phylogenetic tree, we estimated that Hippocampus and Syngnathoides diverged approximately 50.1 million years ago (Mya). Taken together, our genome data presented in this study provide a valuable genetic resource for numerical chromosome changes and in-depth evolutionary and functional investigations, as well as conservation and molecular breeding of this endangered teleost.
As an economically important species endemic to the upper tributaries of Yangtze River in China, long-finned gudgeon fish (Rhinogobio ventralis) has been classified as endangered due to habitat destruction and population decline. In this study, we constructed a chromosome-level genome assembly of R. ventralis by integration of MGI, PacBio and Hi-C sequencing technologies. The final genome assembly was 1015.9 Mb in length (contig N50: 25.91 Mb; scaffold N50: 39.99 Mb), and 97.19% of the haplotypic genome sequences were anchored onto 25 chromosomes. Repetitive elements accounted for 51.00% of the entire genome assembly. A total of 23,220 protein-coding genes were predicted for the assembled genome, of which 99.79% were functionally annotated. Genome evaluation revealed 99.72% completeness for the genome assembly. Through genome-wide prediction of antimicrobial peptides (AMPs), we identified and localized 561 putative AMP-containing genes in the R. ventralis genome. These genes were further classified into 185 distinct functional categories based on public databases, with the top ten components of Penetratin (21.74%), Histone (5.70%), E6AP (4.09%), Scolopendin 1 (2.67%), D38 (2.31%), WBp-1 (2.13%), Defensin (2.13%), Claudin 1 (1.96%), Azurocidin (AZU1, 1.78%), and Ubiquitin (1.60%). Our data presented here provide a potential genetic resource for promoting fundamental research and wild population conservation of this endangered fish species.
IntroductionCompared to mammals and birds, sex-determining genes differ in most fish species. Largemouth bass (Micropterus Salmoides) is one of the most important cultured fish species in China, and there are growth differences between males and females. However, its sex-determining genes and mechanisms currently remain unknown.MethodsWe explored the sex-determination mechanism by integrating whole-genome sequencing, resequencing and comparative genomics approaches.ResultsIn this study, we employed HiFi and Hi-C sequencing technologies to construct a chromosome-level haplotypic genome assembly for male largemouth bass, with a genome size of 875.69 Mb. The assembled genome contains 23 chromosomes, covering 95.31% of the complete sequences with a high scaffold N50 of 35.93 Mb. A genome-wide association study (GWAS) of sex was performed with four populations consisting of 62 males and 58 females. For the sex trait, a total of 3,838 SNP loci were identified to be significantly associated with sexual discrepancy. Interestingly, almost all these significant SNPs (3,825) were clustered on chromosome 10 (Chr10), within a 3.5-Mb sex-determination region (SDR). They were homozygous in females while heterozygous in males. We therefore speculate that largemouth bass owns a XX/XY sex determination system. By comparing genomics data and examining coverage depth of resequencing reads, we revealed a ~51-kb male-specific region (MSR) on Chr10. Gene annotation discovered a coding sequence (msy) within MSR-1, which may contribute to sex determination of largemouth bass. By differential expression analysis, two candidate sex-determining genes (ccdc103 and jockey) were predicted within the target SDR. Moreover, we applied two male-specific non-coding fragments (within MSR-2 and MSR-3) to design specific sex markers, successfully obtaining universal gender identity in examined largemouth bass.DiscussionOverall, our findings improve our understanding of the molecular basis for sex determination in largemouth bass, which will thereby promote the mono-sexual breeding progress in the aquaculture industry.
As personalized cancer vaccines advance, precise modeling of antigen presentation by MHC class I and II is crucial. High-quality training data is essential for clinical models. Existing deep learning models focus on prediction performance but lack interpretability. We introduce Pep2Vec, a modular, transformer-based model trained on MHC I and II ligandome data, transforming input sequences into interpretable vectors. This approach integrates source protein features and elucidates the source of its performance gains, revealing regions that correlate with gene expression and protein-protein interactions. Pep2Vec's peptide latent space shows relationships between peptides of varying MHC class, allotype, lengths, and submotifs. This enables identifying four major contaminant types, constituting 5.0% of our data. Pep2Vec enhances MHC presentation prediction, achieving higher average precision on our presentation test set and immunogenicity datasets than existing models, and reducing contaminant-like peptide recommendations. Pep2Vec addresses a critical need for the development of more precise and effective applications of peptide MHC models, such as for cancer vaccines and antibody deimmunization. ### Competing Interest Statement All authors are employees of Genentech
The Chinese sturgeon (Acipenser sinensis) is an ancient, complex autooctoploid fish species that is currently facing conservation challenges throughout its distribution. To comprehensively characterize the expression profiles of genes and their associated biological functions across different tissues, we performed a transcriptome-scale gene expression analysis, focusing on housekeeping genes (HKGs), tissue-specific genes (TSGs), and co-expressed gene modules in various tissues. We collected eleven tissues to establish a transcriptomic repository, including data from Pacific Biosciences isoform sequencing (PacBio Iso-seq) and RNA sequencing (RNA-seq), and then obtained 25,434 full-length transcripts, with lengths from 307 to 9515 bp and an N50 of 3195 bp. Additionally, 20,887 transcripts were effectively identified and classified as known homologous genes. We also identified 787 HKGs, and the number of TSGs varied from 25 in the liver to 2073 in the brain. TSG functions were mainly enriched in certain signaling pathways involved in specific physiological processes, such as voltage-gated potassium channel activity, nervous system development, glial cell differentiation in the brain, and leukocyte transendothelial migration in the spleen and pronephros. Meanwhile, HKGs were highly enriched in some pathways involved in ribosome biogenesis, proteasome core complex, spliceosome activation, elongation factor activity, and translation initiation factor activity, which have been strongly implicated in fundamental biological tissue functions. We also predicted five modules, with eight hub genes in the brown module, most of which (such as rps3a, rps7, rps23, rpl11, rpl17, rpl27, and rpl28) were linked to ribosome biogenesis. Our results offer insights into ribosomal proteins that are indispensable in ribosome biogenesis and protein synthesis, which are crucial in various cell developmental processes and neural development of Chinese sturgeon. Overall, these findings will not only advance the understanding of fundamental biological functions in Chinese sturgeon but also supply a valuable genetic resource for characterizing this extremely important species.
Antigen presentation on MHC class II (pMHCII presentation) plays an essential role in the adaptive immune response to extracellular pathogens and cancerous cells. But it can also reduce the efficacy of large-molecule drugs by triggering an anti-drug response. Significant progress has been made in pMHCII presentation modeling due to the collection of large-scale pMHC mass spectrometry datasets (ligandomes) and advances in machine learning. Here, we develop graph-pMHC, a graph neural network approach to predict pMHCII presentation. We derive adjacency matrices for pMHCII using Alphafold2-multimer and address the peptide-MHC binding groove alignment problem with a simple graph enumeration strategy. We demonstrate that graph-pMHC dramatically outperforms methods with suboptimal inductive biases, such as the multilayer-perceptron-based NetMHCIIpan-4.0 (+20.17% absolute average precision). Finally, we create an antibody drug immunogenicity dataset from clinical trial data and develop a method for measuring anti-antibody immunogenicity risk using pMHCII presentation models. Our model increases receiver operating characteristic curve (ROC)-area under the ROC curve (AUC) by 2.57% compared to just filtering peptides by hits in OASis alone for predicting antibody drug immunogenicity.
Based on the success of cancer immunotherapy, personalized cancer vaccines have emerged as a leading oncology treatment. Antigen presentation on MHC class I (MHC-I) is crucial for the adaptive immune response to cancer cells, necessitating highly predictive computational methods to model this phenomenon. Here, we introduce HLApollo, a transformer-based model for peptide-MHC-I (pMHC-I) presentation prediction, leveraging the language of peptides, MHC, and source proteins. HLApollo provides end-to-end treatment of MHC-I sequences and deconvolution of multi-allelic data, using a negative-set switching strategy to mitigate misassigned negatives in unlabelled ligandome data. HLApollo shows a 12.65% increase in average precision (AP) on ligandome data and a 4.1% AP increase on immunogenicity test data compared to next-best models. Incorporating protein features from protein language models yields further gains and reduces the need for gene expression measurements. Guided by clinical use, we demonstrate pan-allelic generalization which effectively captures rare alleles in underrepresented ancestries.
Endemic to the upper and middle reaches of the Yangtze River in China, elongate loach (Leptobotia elongata) has become a vulnerable species mainly due to overfishing and habitat destruction. Thus far, no genome data of this species are reported. As a result, lacking of such genomic information has restricted practical conservation and utilization of this economic fish. Here, we constructed chromosome-level genome assemblies for both male and female elongate loach by integration of MGI, PacBio HiFi and Hi-C sequencing technologies. Two primary genome assemblies (586-Mb and 589-Mb) were obtained for female and male fishes, respectively. Indeed, 98.22% and 98.61% of the contig sequences were anchored onto 25 chromosomes, with identification of 26.22% and 25.92% repeat contents in both assembled genomes. Meanwhile, a total of 25,215 and 25,253 protein-coding genes were annotated, of which 97.41% and 98.8% could be predicted with functions. Taken together, our genome data presented here provide a valuable genomic resource for in-depth evolutionary and functional research, as well as molecular breeding and conservation of this economic fish species.
Introduction: Mudskippers are a large group of amphibious fishes that have developed many morphological and physiological capacities to live on land. Genomics comparisons of chromosome-level genome assemblies of three representative mudskippers, Boleophthalmus pectinirostris (BP), Periophthalmus magnuspinnatus (PM) and P. modestus (PMO), may be able to provide novel insights into the water-to-land evolution and adaptation. Methods: Two chromosome-level genome assemblies for BP and PM were respectively sequenced by an integration of PacBio, Nanopore and Hi-C sequencing. A series of standard assembly and annotation pipelines were subsequently performed for both mudskippers. We also re-annotated the PMO genome, downloaded from NCBI, to obtain a redundancy-reduced annotation. Three-way comparative analyses of the three mudskipper genomes in a large scale were carried out to discover detailed genomic differences, such as different gene sizes, and potential chromosomal fission and fusion events. Comparisons of several representative gene families among the three amphibious mudskippers and some other teleosts were also performed to find some molecular clues for terrestrial adaptation. Results: We obtained two high-quality haplotype genome assemblies with 23 and 25 chromosomes for BP and PM respectively. We also found two specific chromosome fission events in PM. Ancestor chromosome analysis has discovered a common fusion event in mudskipper ancestor. This fusion was then retained in all the three mudskipper species. A loss of some SCPP (secretory calcium-binding phosphoprotein) genes were identified in the three mudskipper genomes, which could lead to reduction of scales for a part-time terrestrial residence. The loss of aanat1a gene, encoding an important enzyme (arylalkylamine N-acetyltransferase 1a, AANAT1a) for dopamine metabolism and melatonin biosynthesis, was confirmed in PM but not in PMO (as previously reported existence in BP), suggesting a better air vision of PM than both PMO and BP. Such a tiny variation within the genus Periophthalmus exemplifies to prove a step-by-step evolution for the mudskippers’ water-to-land adaptation. Conclusion: These high-quality mudskipper genome assemblies will become valuable genetic resources for in-depth discovery of genomic evolution for the terrestrial adaptation of amphibious fishes.
Orange spotted grouper (Epinephelus coioides) is an important mariculture fish, and genomic breeding of this grouper species has been hindered due to lack of efficient genotyping tools. Here, we developed a single nucleotide polymorphism (SNP) genotyping technology based on multiplex PCR enrichment capture sequencing, which mainly aims at target area for high-throughput sequencing, and 741 SNPs were designed for genomic selection (GS) of growth and ammonia tolerance traits at the same time. The multiplex PCR enrichment capture sequencing assay showed that the genotyping efficiency was more than 99% in the orange-spotted grouper and the predictive accuracy of body weight and ammonia tolerance traits was 82% and 96%, respectively. More importantly, the average identity of the sequences with these SNPs aligned to the genomes of giant grouper (E. lanceolatus) and brown-marbled grouper (E. fuscoguttatus) were both over 96%. Test data showed that the SNP genotyping efficiency was more than 94% in both giant grouper and brown-marbled grouper. In summary, these results indicated that the development of SNP loci and genotyping approach based on the multiple PCR enrichment capture sequencing are suitable for GS of growth and ammonia tolerance traits in various grouper species, and it would provide technical support for practical grouper breeding.
Antigen presentation on MHC class I (MHC-I) is key to the adaptive immune response to cancerous cells. Computational prediction of peptide presentation by MHC-I has enabled individualized cancer immunotherapies. Here, we introduce HLApollo, a transformer-based approach with end-to-end modeling of MHC-I sequence, deconvolution, and flanking sequences. To achieve this, we develop a novel training strategy, negative set switching, which greatly reduces overfitting to falsely presumed negatives that are necessarily found in presentation datasets. HLApollo shows a meaningful improvement compared to recent MHC-I models on peptide presentation (20.19% average precision (AP)) and immunogenicity (4.1% AP). As expected, adding gene expression boosts the performance of HLApollo. More interestingly, we show that introduction of features from a protein language model, ESM 1b, remarkably recoups much of the benefits of gene expression in absence of true expression measurements. Finally, we demonstrate excellent pan-allelic generalization, and introduce a framework for estimating the expected accuracy of HLApollo for untrained alleles. This guides the use of HLApollo in a clinical setting, where rare alleles may be observed in some subjects, particularly for underrepresented minorities.
The economically important Southern bluefin tuna (Thunnus maccoyii) is a world-famous fast-swimming fish, but its genomic information is limited. Here, we performed whole genome sequencing and assembled a draft genome for Southern bluefin tuna, aiming to generate useful genetic data for comparative functional prediction. The final genome assembly is 806.54 Mb, with scaffold and contig N50 values of 3.31 Mb and 67.38 kb, respectively. Genome completeness was evaluated to be 95.8%. The assembled genome contained 23,403 protein-coding genes and 236.1 Mb of repeat sequences (accounting for 29.27% of the entire assembly). Comparative genomics analyses of this fast-swimming tuna revealed that it had more than twice as many hemoglobin genes (18) as other relatively slow-moving fishes (such as seahorse, sunfish, and tongue sole). These hemoglobin genes are mainly localized in two big clusters (termed as “MNˮ and “LAˮ respectively), which is consistent with other reported fishes. However, Thr39 of beta-hemoglobin in the MN cluster, conserved in other fishes, was mutated as cysteine in tunas including the Southern bluefin tuna. Since hemoglobins are reported to transport oxygen efficiently for aerobic respiration, our genomic data suggest that both high copy numbers of hemoglobin genes and an adjusted function of the beta-hemoglobin may support the fast-swimming activity of tunas. In summary, we produced a primary genome assembly and predicted hemoglobin-related roles for the fast-swimming Southern bluefin tuna.
DATA REPORT article Front. Genet., 07 July 2021Sec. Livestock Genomics https://doi.org/10.3389/fgene.2021.695700
A high concentration of ammonia is toxic to intensively cultured groupers, with a potential to affect their survival and production. It is therefore useful to select groupers with a high tolerance to ammonia. Based on genome-wide loci, to predict the genomic estimated breeding value (GEBV) of individuals, genomic selection (GS) can expedite genetic improvement over traditional pedigree-based approaches. However, GS has not previously been applied in grouper breeding. In this study, 600 orange-spotted groupers (Epinephelus coioides) were genotyped by whole-genome resequencing and over 3 million single nucleotide polymorphisms (SNPs) were identified. A moderate heritability of 0.36 +/- 0.12 was estimated using these SNPs. Four methods, i.e. BayesA, BayesB, BayesC, and rrBLUP, were used for GS, and the average AUC (Area under receiver operating characteristic) values of these methods were 0.641, 0.643, 0.642, and 0.640, respectively. The predictive accuracies of a genome-wide association study (GWAS) informative SNPs and those of the randomly selected SNPs were compared. The accuracy of GWAS-informative SNPs reached 100%, which was higher than that of the random SNPs. These results suggest that GWAS can be used to improve the cost-efficiency of GS, due to a substantial reduction in genotyping cost and time. GS is therefore suitable for improving the ammonia tolerance of groupers.
Although there are various Conus species with publicly available transcriptome and proteome data, no genome assembly has been reported yet. Here, using Chinese tubular cone snail ( C. betulinus ) as a representative, we sequenced and assembled the first Conus genome with original identification of 133 genome-widely distributed conopeptide genes. After integration of our genomics, transcriptomics, and peptidomics data in the same species, we established a primary genetic central dogma of diverse conopeptides, assuming a rough number ratio of ~1:1:1:10s for the total genes: transcripts: proteins: post-translationally modified peptides. This ratio may be special for this worm-hunting Conus species, due to the high diversity of various Conus genomes and the big number ranges of conopeptide genes, transcripts, and peptides in previous reports of diverse Conus species. Only a fraction (45.9%) of the identified conotopeptide genes from our achieved genome assembly are transcribed with transcriptomic evidence, and few genes individually correspond to multiple transcripts possibly due to intraspecies or mutation-based variances. Variable peptide processing at the proteomic level, generating a big diversity of venom conopeptides with alternative cleavage sites, post-translational modifications, and N-/C-terminal truncations, may explain how the 133 genes and ~123 transcripts can generate thousands of conopeptides in the venom of individual C. betulinus . We also predicted many conopeptides with high stereostructural similarities to the putative analgesic ω-MVIIA, addiction therapy AuIB and insecticide ImI, suggesting that our current genome assembly for C. betulinus is a valuable genetic resource for high-throughput prediction and development of potential pharmaceuticals.