The pan-genome represents the complete genomic diversity of specific species, serving as a valuable resource for studying species evolution, crop domestication, and guiding crop breeding and improvement. While there are several single-species-specific plant pan-genome databases, the availability of multi-species pan-genome databases is limited. Additionally, variations in methods and data types used for plant pan-genome analysis across different databases hinder the comparison and integration of pan-genome information from various projects at multi-species or single-species levels. To tackle this challenge, we introduce PlantPan, a comprehensive database housing the results of pan-genome analysis for 195 genomes from 11 plant species. PlantPan aims to provide extensive information, including gene-centric and sequence-centric pan-genome information, graph-based pan-genome, pan-genome openness profiles, gene functions and its variation characteristics, homologous genes, and gene clusters across different species. Statistically, PlantPan incorporates 9 163 011 genes, 694 191 gene clusters, 526 973 370 genome variations, and 1 616 089 non-redundant genome variation groups at the species level, 33 455,098 genome synteny, and 177 827 non-redundant genome synteny groups at the species level. Regarding functional genes, PlantPan contains 5 222 720 genes related to transcription factors, 395 247 literature-reported resistance genes, 455 748 predicted microbial/disease resistance genes, and 1 612 112 genes related to molecular pathways. In summary, PlantPan is a vital platform for advancing the application of pan-genomes in molecular breeding for crops and evolutionary research for plants.
Cassava is a highly resilient tropical crop that produces large, starchy storage roots and high biomass. However, how did cassava's remarkable environmental adaptability and key economic traits evolve from its wild species remain unclear. In this study, we obtained near complete telomere-to-telomere genome assemblies and their haplotype forms for the cultivar AM560, the wild ancestors FLA4047 and W14, constructed a graphic pan-genome of 30 representatives with a size of 1.15 Gb, and built a clarified evolutionary tree of 486 accessions. A comparison of structural variations and single-nucleotide variations between the ancestors and cultivated cassavas reveals predominant expansions and contractions of numbers of genes and gene families, which are mainly driven by transposons. Significant selective sweeping occurred in 122 footprints of genomes and affects 1,519 domesticated genes. We identify selective mutations in MeCSK and MeFNR2 that could promote photoreactions associated with MeNADP-ME in C4 photosynthesis in modern cassava. Coevolution of retard floral primordia and initiation of storage roots may arise from MeCOL5 variants with altered bindings to MeFT1, MeFT2, and MeTFL2. Mutations in MeMATE1 and MeGTR occur in sweet cassava, and MeAHL19 has evolved to regulate the biosynthesis, transport, and endogenous remobilization of cyanogenic glucosides in cassava. These extensive genomic and gene resources provided here, along with the findings on the evolutionary mechanisms responsible for beneficial traits in modern cultivars, lay a strong foundation for future breeding improvements of cassava.
BackgroundThe immune system is of paramount importance in maintaining human health and defending against pathogens. Among them, the adaptive immune system is a crucial component of the immune system, as it is responsible for generating and modulating the immune repertoire, which is vital for immune responses.MethodsWe conducted a comprehensive analysis of T cell receptor (TCR) and B cell receptor (BCR) clonotypes in the peripheral blood immune repertoire of 20 patients with benign and malignant ovarian tumors. The analysis elucidates the differences between the two immune repertoires in various aspects and constructs an early screening machine learning model for ovarian tumors based on the characteristics of the immune repertoire.ResultThe finding revealed that patients with malignant ovarian tumors exhibited a reduction in balance, richness, and diversity in their immune repertoires compared to those with benign tumors. Additionally, there was a negative correlation between patient age and immune repertoire diversity, and the immune repertoire of patients with malignant tumors displayed high heterogeneity. By employing machine learning techniques, we have developed an early screening model based on 16 TCR V-J genes and 11 BCR V-J genes, which achieved an average AUC of 0.93 (TCR) and 0.958 (BCR) on the ovarian tumor test dataset. Moreover, a comparison of the spatial distributions of TCR and BCR revealed, for the first time, that TCR was more significantly associated with the benign-to-malignant transformation of ovarian tumors.ConclusionsOur study highlights the critical role of the adaptive immune repertoire in distinguishing between benign and malignant ovarian tumors. TCR demonstrated more distinct spatial distribution patterns between benign and malignant states, suggesting its potential as a more sensitive biomarker for ovarian tumor detection. These findings provide new insights into the immunological landscape of ovarian tumors and offer a promising avenue for early diagnosis and prognosis assessment.
Nanobelts are a rapidly developing family of macrocycles with appealing features. However, their hostguest chemistry is currently limited to the recognition of fullerenes via Tr-Tr interactions. Herein, we report two heteroatom-bridged [8]cyclophenoxathiin nanobelts ([8]CP-Me and [8]CP) encapsulate corannulene (Cora) to form bowl-in-bowl supramolecular structures stabilized mainly through CH-Tr interactions in solid-state. The convex surface of corannulene is oriented towards the cavity due to geometry complementarity. The complex Cora subset of[8]CP exhibits a unique 2:2 capsule-like structure in crystal packing, in which corannulene adopts a concave-to-concave assembling fashion. This work enriches the molecular recognition of nanobelts and demonstrates that CH-Tr interactions can act as the main driving force for nanobelts host-guest complexes. (c) 2025 Published by Elsevier B.V. on behalf of Chinese Chemical Society and Institute of Materia Medica, Chinese Academy of Medical Sciences.
Tropical crops are vital for tropical agriculture, with resource scarcity, functional diversity and extensive market demand, providing considerable economic benefits for the world's tropical agriculture-producing countries. The rapid development of sequencing technology has promoted a milestone in tropical crop research, resulting in the generation of massive amount of data, which urgently needs an effective platform for data integration and sharing. However, the existing databases cannot fully satisfy researchers' requirements due to the relatively limited integration level and untimely update. Here, we present the Tropical Crop Omics Database (TCOD, https://ngdc.cncb.ac.cn/tcod), a comprehensive multi-omics data platform for tropical crops. TCOD integrates diverse omics data from 15 species, encompassing 34 chromosome-level de novo assemblies, 1 255 004 genes with functional annotations, 282 436 992 unique variants from 2048 WGS samples, 88 transcriptomic profiles from 1997 RNA-Seq samples and 13 381 germplasm items. Additionally, TCOD not only employs genes as a bridge to interconnect multi-omics data, enabling cross-species comparisons based on homology relationships, but also offers user-friendly online tools for efficient data mining and visualization. In short, TCOD integrates multi-species, multi-omics data and online tools, which will facilitate the research on genomic selective breeding and trait biology of tropical crops.
A molecular belt [8]NCP incorporating naphthalene moieties was precisely engineered through bottom–up synthesis. It exhibits remarkable selective binding ability for C 70 compared to C 60 , forming a 1 : 1 complex.
Here we present a compact and precise [2]catenane rotary motor that functions with a single recognition site, capable of achieving a 360° directional rotation powered by chemical fuels. The motor is propelled by an acid-base fueled benzimidazolium pumping cassette and deemed the smallest (molecular weight ∼ 994 Da) catenane rotary motor to date. It can effectively undergo a 180° rotation by transitioning the [24]crown-6 ether (24C6) from the benzimidazolium site to the less favorable alkyl moiety through sequential deprotonation, slipping, and re-protonation operations, generating a meta stable co-conformer. Subsequently, a discharging phase, triggered by de-benzylation and re-benzylation, facilitates the other half-rotation of the motor, returning the 24C6 to its initial position and completing the full directional rotation of the [2]catenane rotary motor within 18 hours. The precision of the motor's operation enables further advances in artificial molecular machines.
When classifying breeds of dogs, the accuracy of classification significantly affects breed identification and dog research. Using images to classify dog breeds can improve classification efficiency; however, it is increasingly challenging due to the diversities and similarities among dog breeds. Traditional image classification methods primarily rely on extracting simple geometric features, while current convolutional neural networks (CNNs) are capable of learning high-level semantic features. However, the diversity of dog breeds continues to pose a challenge to classification accuracy. To address this, we developed a model that integrates multiple CNNs with a machine learning method, significantly improving the accuracy of dog images classification. We used the Stanford Dog Dataset, combined image features from four CNN models, filtered the features using principal component analysis (PCA) and gray wolf optimization algorithm (GWO), and then classified the features with support vector machine (SVM). The classification accuracy rate reached 95.24% for 120 breeds and 99.34% for 76 selected breeds, respectively, demonstrating a significant improvement over existing methods using the same Stanford Dog Dataset. It is expected that our proposed method will further serve as a fundamental framework for the accurate classification of a wider range of species.
This study explores the synthesis, structural characterization, and host‐guest interactions of heteroatom bridged nanobelts, focusing on a cyclothianthrene nanobelt and a fused nanobelt incorporating thianthrene and phenoxathiin. Utilizing a cyclization‐followed‐by‐bridging synthetic approach, both molecular belts were successfully synthesized, and their structures confirmed through NMR and MALDI‐TOF‐MS analysis. Crystallographic studies revealed that the cyclothianthrene nanobelt adopts an octagonal column‐like conformation, while the hybrid belt forms an oval tub‐shaped shape, both exhibiting distinct assembly motifs. The host‐guest chemistry of these nanobelts was investigated with fullerenes (C60, C70, and PC61BM). The cyclothianthrene belt showed no interaction with these fullerenes, whereas the other belt demonstrated adaptive binding capabilities, forming stable complexes with C60 and C70 through π−π interactions and C‒H∙∙∙S hydrogen bonds. The binding constants indicated that the hybrid belt has a stronger affinity for C70 due to better size complementarity. Additionally, its interaction with PC61BM showcased a specific 1:1 binding mode despite exhibiting a smaller binding constant. This study underscores the impact of heteroatom incorporation on the structural and functional properties of nanobelts, offering insights for future molecular design strategies.
Crown ethers (CEs), known for their exceptional host–guest complexation, offer potential as linkers in covalent organic frameworks (COFs) for enhanced performance in catalysis and host–guest binding. However, their highly flexible conformation and low symmetry limit the diversity of CE-derived COFs. Here, we introduce a novel C 3 -symmetrical azacrown ether (ACE) building block, tris(pyrido)[18]crown-6 ( TPy18C6 ), for COF fabrication ( ACE-COF-1 and ACE-COF-2 ) via reticular synthesis. This approach enables precise integration of CEs into COFs, enhancing Ni 2+ ion immobilization while maintaining crystallinity. The resulting Ni 2+ -doped COFs ( Ni@ACE-COF-1 and Ni@ACE-COF-2 ) exhibit high discharge capacity (up to 1.27 mAh ⋅ cm −2 at 8 mA ⋅ cm −2 ) and exceptional cycling stability (>1000 cycles) as cathode materials in aqueous alkaline nickel-zinc batteries. This study serves as an exemplar of the seamless integration of macrocyclic chemistry and reticular chemistry, laying the groundwork for extending the macrocyclic-synthon driven strategy to a diverse array of COF building blocks, ultimately yielding advanced materials tailored for specific applications.
Whole-exome sequencing (WES) data are frequently used for cancer diagnosis and genome-wide association studies (GWAS), based on high-coverage read mapping, informative variant calling, and high-quality reference genomes. The center position of the currently used genome assembly, GRCh38, is now challenged by two newly published telomere-to-telomere (T2T) genomes, T2T-CHM13 and T2T-YAO, and it becomes urgent to have a comparative study to test population specificity using the three reference genomes based on real case WES data. Here, we report our analysis along this line for 19 tumor samples collected from Chinese patients. The primary comparison of the exon regions among the three references reveals that the sequences in up to ∼ 1% of target regions in T2T-YAO are widely diversified from GRCh38 and may lead to off-target in sequence capture. However, T2T-YAO still outperforms GRCh38 by obtaining 7.41% of more mapped reads. Due to more reliable read-mapping and closer phylogenetic relationship with the samples than GRCh38, T2T-YAO reduces half of variant calls of clinical significance which are mostly benign, while maintaining sensitivity in identifying pathogenic variants. T2T-YAO also outperforms T2T-CHM13 in reducing calls of Chinese-specific variants. Our findings highlight the critical need for employing population-specific reference genomes in genomic analysis to ensure accurate variant analysis and the significant benefits of tailoring these approaches to the unique genetic background of each ethnic group.
ABSTRACT To analyze the characteristics of Mycoplasma pneumoniae as well as macrolide antibiotic resistance through whole-genome sequencing and comparative genomics. Thirteen clinical strains isolated from 2003 to 2019 were selected, 10 of which were resistant to erythromycin (MIC >64 µg/mL), including 8 P1-type I and 2 P1-type II. Three were sensitive (<1 µg/mL) and P1-type II. One resistant strain had an A→G point mutation at position 2064 in region V of the 23S rRNA, the others had it at position 2063, while the three sensitive strains had no mutation here. Genome assembly and comparative genome analysis revealed a high level of genome consistency within the P1 type, and the primary differences in genome sequences concentrated in the region encoding the P1 protein. In P1-type II strains, three specific gene mutations were identified: C162A and A430G in L4 gene and T1112G mutation in the CARDS gene. Clinical information showed seven cases were diagnosed with severe pneumonia, all of which were infected with drug-resistant strains. Notably, BS610A4 and CYM219A1 exhibited a gene multi-copy phenomenon and shared a conserved functional domain with the DUF31 protein family. Clinically, the patients had severe refractory pneumonia, with pleural effusion, necessitating treatment with glucocorticoids and bronchoalveolar lavage. The primary variations between strains occur among different P1-types, while there is a high level of genomic consistency within P1-types. Three mutation loci associated with specific types were identified, and no specific genetic alterations directly related to clinical presentation were observed. IMPORTANCE Mycoplasma pneumoniae is an important pathogen of community-acquired pneumonia, and macrolide resistance brings difficulties to clinical treatment. We analyzed the characteristics of M. pneumoniae as well as macrolide antibiotic resistance through whole-genome sequencing and comparative genomics. The work addressed primary variations between strains that occur among different P1-types, while there is a high level of genomic consistency within P1-types. In P1-type II strains, three specific gene mutations were identified: C162A and A430G in L4 gene and T1112G mutation in the CARDS gene. All the strains isolated from severe pneumonia cases were drug-resistant, two of which exhibited a gene multi-copy phenomenon, sharing a conserved functional domain with the DUF31 protein family. Three mutation loci associated with specific types were identified, and no specific genetic alterations directly related to clinical presentation were observed.
Pan-genome studies are important for understanding plant evolution and guiding the breeding of crops by containing all genomic diversity of a certain species. Three short-read-based strategies for plant pan-genome construction include iterative individual, iteration pooling, and map-to-pan. Their performance is very different under various conditions, while comprehensive evaluations have yet to be conducted nowadays. Here, we evaluate the performance of these three pan-genome construction strategies for plants under different sequencing depths and sample sizes. Also, we indicate the influence of length and repeat content percentage of novel sequences on three pan-genome construction strategies. Besides, we compare the computational resource consumption among the three strategies. Our findings indicate that map-to-pan has the greatest recall but the lowest precision. In contrast, both two iterative strategies have superior precision but lower recall. Factors of sample numbers, novel sequence length, and the percentage of novel sequences’ repeat content adversely affect the performance of all three strategies. Increased sequencing depth improves map-to-pan’s performance, while not affecting the other two iterative strategies. For computational resource consumption, map-to-pan demands considerably more than the other two iterative strategies. Overall, the iterative strategy, especially the iterative pooling strategy, is optimal when the sequencing depth is less than 20X. Map-to-pan is preferable when the sequencing depth exceeds 20X despite its higher computational resource consumption.
Convolutional neural network (CNN) has been widely used for fine-grained image classification, which has proven to be an effective approach for the classification and identification of specific species. For breed classification of dog, there are several proposed methods based on dog images, however, the highest accuracy rate for dogs (about 93%) is still below expectations compared to other animals or plants (more than 95% on birds and more than 97% on flowers). In this study, we used the Stanford Dog Dataset, combined image features from four CNN models, filtered the features using principal component analysis (PCA) and gray wolf optimization algorithm (GWO), and then classified the features with support vector machine (SVM). Eventually, the classification accuracy rate reached 95.24% for 120 breeds and 99.34% for 76 selected breeds, respectively, demonstrating a significant improvement over existing methods using the same Stanford Dog Dataset. It is expected that our proposed method will further serve as a fundamental framework for accurate classification of a wider range of species.
Homology is fundamental to infer genes' evolutionary processes and relationships with shared ancestry. Existing homolog gene resources vary in terms of inferring methods, homologous relationship and identifiers, posing inevitable difficulties for choosing and mapping homology results from one to another. Here, we present HGD (Homologous Gene Database, https://ngdc.cncb.ac.cn/hgd), a comprehensive homologs resource integrating multi-species, multi-resources and multi-omics, as a complement to existing resources providing public and one-stop data service. Currently, HGD houses a total of 112 383 644 homologous pairs for 37 species, including 19 animals, 16 plants and 2 microorganisms. Meanwhile, HGD integrates various annotations from public resources, including 16 909 homologs with traits, 276 670 homologs with variants, 398 573 homologs with expression and 536 852 homologs with gene ontology (GO) annotations. HGD provides a wide range of omics gene function annotations to help users gain a deeper understanding of gene function.
SUMMARY Cassava is the most important starch sources, a tropical model crop. We constructed nearly T2T genomes of cultivar AM560, wild ancestors FLA4047 and W14, pan-genome of 24 representatives and a clarified evolutionary tree with 486 accessions. Comparison of SVs and SNVs between the ancestors and cultivated cassavas revealed predominant expansion, contraction of genes and gene families. Significantly selective sweeping occurred in the cassava genomes in 122 footprints with 1,519 candidate domestication genes. We identified selective mutations in MeCSK and MeFNR3 promoting photoreaction associated with MeNADP-ME of C 4 assimilation in modern cassava. Co-evolved retardation of floral primordia and initiation of storage roots arose from MeCOL5 mutants with altered bindings to MeFT1, MeFT2 and MeTFL2. MebHLHs evolved to regulate the biosynthesis, transport and endogenous remobilization of cyanogenic glucosides, with new functionalities of MeMATE1, MeGTR in selected sweet cassava. These findings enhanced comprehensive knowledge and database on the evolution and breeding of cassava. HIGHLIGHTS Three nearly T2T cassava genomes of cultivar AM560 and its wild ancestors FLA4047 and W14. A species-level cassava panSV haplotype map across 346,322 structural variations over 31,362 gene families and 96,032,008 SNPs and InDels variations globally and a clarified evolutionary tree with 486 accessions. Selective mutations in MeCSK and MeFNR3 promoted photoreaction associated with MeNADP-ME of C 4 assimilation shaped the C 3 -C 4 intermediate photosynthesis of modern cassava. Coevolution of floral primordia contrary to initiated storage root is pivotal for the domestication of cassava, and arose from MeCOL5 mutants altered the binding with MeFT1, MeFT2 (SP6A), and MeTFL2 .
Xanthomonas oryzae pv. oryzae ( Xoo) is the causal bacterium of rice bacterial blight (BB), which is one of the most destructive bacterial diseases of rice worldwide. Although more than 40 BB resistance genes ( Xa genes) have been identified, only a few conferring broad resistance (e.g., Xa4 and Xa21) have been widely used for rice breeding in Asia. The narrow genetic basis for the BB resistance of rice cultivars has resulted in the emergence of new pathogenic strains that can overcome the resistance of commonly grown rice cultivars derived from breeding programs. For example, Xoo strain HN2011, which was isolated in Hainan (China), can overcome the resistance mediated by multiple BB resistance genes, including Xa4 and Xa21. We herein present the complete genome sequence of HN2011 generated by Nanopore sequencing. The genomic data will be useful for elucidating the genomic variations of Xoo and will contribute to clarifying the interaction between Xoo and rice. [Formula: see text] Copyright © 2023 The Author(s). This is an open access article distributed under the CC BY-NC-ND 4.0 International license .
Hydrogen-bonded capsules have been widely employed as supramolecular hosts for organic molecular guests. Encapsulation of fullerenes by capsules is relatively scarce, especially those that utilize sulfur atoms as hydrogen-bond acceptors. Herein, we describe, in both solution and solid state, a bowl-shaped nanobelt [8]cyclophenoxathiin la and its tetra-methylated derivative lb that can form C-H center dot center dot center dot S hydrogen-bonded capsules induced by complexation with suitable fullerenes. la strongly encapsulates C-60, C-70, or 6,6-phenyl-C-61-butyric acid methyl ester (PC61BM) to form a 2:1 ternary complex featuring 16 equatorial (sp(2))C-H center dot center dot center dot S hydrogen bonds. A pseudorotaxane structure was further obtained for the complex of la with PC61BM. Conversely, a 1:1 inclusion complex was observed for binding C-60 or PC61BM with lb indicating the reduced tendency to form capsules by introducing methyl groups into the belt. Surprisingly, the capsule-like structure was retained for the 1:2 complex of C(70 )with lb as observed by the presence of multiple (sp(3))C-H center dot center dot center dot S hydrogen bonds. The strong binding affinity and tailorable complexation mode enable further applications of nanobelts in fullerene chemistry. [GRAPHICS] .
Sequences that are homologous to the SARS-CoV-2 genome were found in numerous sequencing runs that were not associated with the SARS-CoV-2 studies in the public database. It is unclear whether they are derived from the possible progenitor of SARS-CoV-2 or contamination of more recent SARS-CoV-2 variants circulated in the population due to the lack of information on the collection, library preparation, and sequencing processes.
Since its initial release in 2001, the human reference genome has been continuously improved in both continuity and accuracy, and the recently-released telomere-to-telomere version—T2T-CHM13—reaches its top quality after 20 years of effort. However, T2T-CHM13 does not represent an authentic diploid human genome, but rather one derived from a simplified, nearly homozygous genome of a hydatidiform mole cell line. To address this limitation and provide an alternative pertinent to the Chinese population, the largest ethnic group in the world, we have assembled a complete diploid human genome of a male Han Chinese, T2T-YAO, which includes telomere-to-telomere assemblies for all the 22+X+M and 22+Y chromosomes in his two haploids inherited separately from his parents. Both haplotypes contain no artificial sequences or model nucleotides and possess a high quality comparable to CHM13, with fewer than one error per ∼14 Mb. Derived from the individual who lives in the aboriginal region of Han Chinese, T2T-YAO shows clear ancestry and potential genetic continuity from the ancient ancestors of the Han population. Each haplotype of T2T-YAO possesses ∼340 Mb exclusive sequences and ∼3100 unique genes as compared to CHM13, and their genome sequences show greater genetic distance to CHM13 than to each other in terms of nucleotide polymorphism and structural variations. The construction of T2T-YAO would serve as a high-quality diploid reference that enables precise delineation of genomic variations in a haplotype-sensitive manner, which could advance our understandings in human evolution, hereditability of diseases and phenotypes, especially within the context of the unique variations of the Chinese population.