Abstract Quantifying cross-species relationships among cell types from single-cell transcriptomic data can reveal both conserved and divergent patterns of cell-type hierarchies. However, existing cross-species integration methods can be limited in modeling genes beyond orthologs by leveraging cell-type-resolved transcriptional context, or in learning explicit type-level representations. Here we present CHORD, a cross-species integration framework that jointly learns representations of genes, cells and cell types. We demonstrate that CHORD can integrate cross-species single-cell atlases and support cell-type annotation with unknown cell-type detection. In the frog–zebrafish embryogenesis and mammalian motor cortex atlases, CHORD infers cell-type trees that place conserved cell types from different species in relative proximity and summarize hierarchical relationships among cell types. CHORD also supports cross-species comparison of continuous phenotypic variation by placing embryonic cells along an aligned developmental timeline. CHORD further yields gene embeddings that capture orthologous and functional relationships, and gene importance scores linking genes to cell types.
Human speech and language depend on hemispheric specialization across cortical regions and cortico-striatal circuits. We profiled 125 human cortical samples from 13 Brodmann areas (BAs), bilaterally, across five donors to generate a hemisphere-resolved transcriptomic atlas and quantify region-specific lateralization. Integrating genome-wide association signals for speech-, language-, and reading-related traits with brain cis-expression quantitative trait loci (eQTL) and enhancer maps prioritized a regulatory axis linking rs62060948 to MYC binding and WNT3 expression. WNT3 was higher in right BA44 than left, and cellular assays supported MYC occupancy and showed reduced WNT3 after MYC knockdown. In mice, unilateral Wnt3 overexpression (Wnt3 OE) in the dorsal striatum selectively altered ultrasonic vocalizations (USVs), locomotor activity, and myelin basic protein expression. These results connect regulatory variation to lateralized gene control and circuit function relevant to vocal communication, and provide a multiregional resource to support mechanistic studies in human tissue and animal models.
Vocal fold length is a primary anatomical determinant of pitch and fundamental frequency in terrestrial mammal vocalizations, playing a key role in social communication, mate attraction, and competition. However, its phenotypic diversification and genetic basis remain unclear. Here, using a phylogenetic genotype-to-phenotype mapping (PhyloG2P) framework to 75 species representing Primates, Carnivora, and Artiodactyla, we investigate the evolutionary dynamics and genetic architecture of relative vocal fold length (RVFL), accounting for allometric trends. Phylogenetic comparative analyses reveal similar patterns of RVFL variation among orders and strong adaptive responses supporting the acoustic size exaggeration hypothesis. Correlation analyses further reveal the potential driving role of social organization in RVFL evolution. Comparative genomics identifies 25 hearing-related genes showing rapid evolution, positive selection, lineage-specific mutations, and coevolution with RVFL, suggesting a reciprocal evolutionary interplay between vocal and auditory systems. This genetic link is supported by laryngeal MRI evidence from wild-type and Pjvk knock-out mice, where PJVK is identified as a candidate gene through our PhyloG2P analyses. Additionally, pathways related to peptide hormone secretion and neural regulation appear to mediate RVFL variation. Together, these findings advance our understanding of the genetic and evolutionary mechanisms driving vocal fold diversity in terrestrial mammals and highlight the links between vocal and auditory systems, offering new insights into the origins of vocal communication.
SIGNIFICANCE:Osteoarthritis (OA), one of the most prevalent joint diseases affecting more than 240 million people, strongly influences human health and reduces life quality. This review aims to fill the current research gap regarding the application and potential of mitochondrial quality control (MQC) based therapies in the treatment of OA, thereby providing guidance for future research and clinical practice. RECENT ADVANCES:Chondrocytes respond to the inflammatory microenvironment via an array of signaling pathways and thus are critical in cartilage degeneration and OA progression. Mitochondria, as an important metabolic center in chondrocytes, play a vital role in responding to inflammatory stimuli. Multiple MQC mechanisms, including mitochondrial antioxidant defense, mitochondrial protein quality control, mitochondrial DNA repair, mitochondrial dynamics, mitophagy, and mitochondrial biogenesis, sustain mitochondrial homeostasis under pathological conditions. CRITICAL ISSUES:Despite extensive OA research, effective therapies remain limited. Elucidating MQC mechanisms in disease progression and post-traumatic cartilage repair is crucial. While preclinical studies demonstrate potential, clinical translation requires addressing protocol standardization, patient stratification, and long-term efficacy, as well as safety validation. FUTURE DIRECTIONS:Future research should focus on developing personalized MQC-based OA therapies guided by biomarker profiling and signaling pathway modulation. However, translational challenges persist, particularly regarding pervasive off-target effects, inadequate OA-specific targeting capacity, interpatient heterogeneity, and reliable evaluation of long-term therapeutic efficacy. Strategic prioritization of OA-specific MQC targets coupled with delivery system optimization may significantly improve both clinical translatability and therapeutic outcomes.
Language evolution in the Gansu-Qinghai (GQ) region provides a key perspective for understanding cultural development along the eastern Silk Road. Previous genetic and archaeological studies have revealed complex, multi-ethnic interactions in this region, shaped by migration and sociocultural exchange. However, the lack of structured linguistic data and computational tools for studying GQ language contacts has limited rigorous analysis of sociocultural evolution. Here, we presented a new hybrid dataset of phonological and morpho-syntactic features from languages sampled across the GQ region. We introduced a computational framework to assess language contact and admixture, allowing us to quantify interaction among GQ languages and trace their origins. Our results showed that GQ languages exhibit distinct contact patterns in their phonological and morpho-syntactic systems, with some languages displaying clear evidence of mixture. Using a new statistical method, Trait Sharing Among Languages (TSAL), based on tree topology, we identified significant influences from Sinitic, Tibetan, Mongolic, and Turkic languages in shaping the GQ linguistic diversity. These findings highlight the GQ region as a linguistic convergence zone on the eastern Silk Road, providing a foundation for quantitative research on language contact and admixture. Our work enhances the linguistic perspective on cultural evolution in the GQ region and supports future interdisciplinary studies that integrate languages, genes, and material cultures.
BACKGROUND:Given the limitations of current diagnostic methods, there is a compelling need to develop a more accurate and efficient method for diagnosing Parkinson's disease (PD). As dysarthria is one of the most characteristic and core symptoms of PD, acoustic analysis of patients' voice recordings may be a new means to diagnose PD. OBJECTIVES:To assess the feasibility of employing acoustic analysis to diagnose PD in the Chinese population. METHODS:A total of 140 Chinese participants, 96 PD patients and 44 healthy controls respectively, provided their voice recordings on sustained vowels (/a:/, /i:/, /u:/). A deliberate acoustic analysis was performed to identify acoustic parameters with diagnostic values. Machine learning-based diagnostic models were also constructed, which can automatically output a possible diagnosis for reference based on one's voice recordings. RESULTS:Through acoustic analysis, we have identified 14 instances of significant differences across 8 unique acoustic parameters. Light gradient boosting machine model achieved an area under the curve of 0.96 along with the highest accuracy of 92.86%, the sensitivity of 89.47%, and the specificity of 100.00%. CONCLUSION:Our preliminary results in the Chinese population showed that PD patients presented noticeable phonatory deficits and proved the feasibility of applying acoustic analysis to PD diagnosis, which may have significance in guiding further clinical research.
BACKGROUND:Osteoarthritis (OA) is a prevalent degenerative disease, characterized by articular cartilage lesions, synovial inflammation, and osteophyte formation, which ultimately results in joint deformity and limited mobility. Polyphenols, a class of bioactive phytochemicals widely distributed in traditional Chinese medicine, have been extensively utilized for centuries in the management of OA. PURPOSE:Despite extensive research efforts, effective pharmacological interventions for treating OA have yet to be fully realized. This study aimed to identify anti-inflammatory polyphenols as potential therapeutic alternatives for osteoarthritis treatment. METHODS:The high-content screening (HCS) technique was utilized to identify active compounds from a polyphenol library. An interleukin-1β (IL-1β) stimulated inflammatory model, which significantly increased the expression of chondrocyte degradation markers such as matrix metallopeptidase 13 (MMP-13), was established in vitro. Polydatin (PD) was identified as a promising substance for attenuating MMP-13 expression and promoting type II collagen (COL2) expression. Subsequently, destabilization of the medial meniscus (DMM) and monosodium iodoacetate (MIA) mouse models of OA were utilized to evaluate the efficacy of PD in vivo. After model establishment and drug administration, knee joints were harvested at 4 and 12 weeks post-treatment and subjected to radiological and histological assessments. Finally, mitochondrial function assays were conducted to further explore the underlying therapeutic mechanisms. RESULTS:Within a library of 16 polyphenolic compounds, HCS-based phenotypic screening identified six active compounds that potently inhibited the expression of MMP-13. PD exhibited promising anti-inflammatory effects in a dose-dependent manner. Western blot and immunofluorescence analyses revealed that PD markedly suppressed MMP-13 synthesis and COL2 degradation in a dose-dependent manner. The RNA sequencing results revealed that the anti-inflammatory efficacy of PD may be associated with the Wnt signaling pathway. Three-dimensional micro-CT reconstructions revealed a reduction in osteophyte formation in the PD treatment groups relative to the OA model groups. Safranin O and fast green (SO&FG) staining, hematoxylin and eosin staining, and COL2 immunofluorescence staining revealed that PD treatment alleviated cartilage matrix breakdown and reduced osteoarthritis scores. Furthermore, PD upregulated the expression of sirtuin 3 (SIRT3) and superoxide dismutase 2 (SOD2), thereby increasing mitochondrial membrane potential and reducing mitochondrial superoxide (mtROS) levels, suggesting antioxidant-driven anti-inflammatory effects. CONCLUSION:This study successfully employed the HCS technique to identify polydatin from a polyphenol library as an active compound for treating OA. PD alleviates cartilage degradation and lesions by activating the SIRT3/SOD2/mtROS axis, increasing the mitochondrial membrane potential, and exerting antioxidant effects.
Understanding the mechanisms of language competition is crucial for mitigating language extinction and promoting cultural sustainability. Nevertheless, how to quantify the effects of socio-linguistic factors, such as social prestige and bilingualism, on language competition remains a critical challenge. Here, we present Markov-process-based language competition models to explore the interactions among monolingual and bilingual groups. Based on these models, we develop a Bayesian parametric estimation strategy, which enables quantifying the effects of socio-linguistic factors through rigorous statistical examinations. With six empirical cases worldwide, we observe a general trend of minority monolinguals shifting towards majority ones, where the presence of bilingualism can decelerate this shift. Typically, bilingualism can sometimes accelerate the reverse shift when the majority language possesses a higher social prestige than the minority one. Our findings emphasize the protective role of bilingualism in mitigating language extinction, particularly when competing languages exhibit distinct social prestige. We expect that our Bayesian computational framework could serve as a useful tool for assessing the roles of socio-linguistic factors in language competition and aiding language preservation and cultural sustainability.
Advancements in genetic correlation estimation have elucidated genome-wide pleiotropy's influence on phenotypic correlations among human complex traits and diseases. However, the role of proteomic domains in these correlations remains underexplored. Traditional genetic correlation analysis assumptions, including the minute effects of SNPs and their linkage disequilibrium, do not suit proteomic data. We present a novel method, Likelihood-based Estimation for Proteomic Correlation (LEAP), tailored to provide unbiased estimation of shared proteomic architectures between trait pairs. LEAP notably decreases computational demands by approximately 1000-fold compared to conventional bivariate linear mixed models. We applied LEAP to data from the UK Biobank Pharma Proteomics Project, identifying 585 significant proteomic correlations among 1,225 pairs of 50 biochemical, anthropometric, and behavioral traits. Furthermore, we quantified the distinct proteomic and genetic contributions to phenotypic correlations, highlighting significant gender differences. This study provides a comprehensive computational approach for proteomic correlation estimation, clarifying the specific roles of genomics and proteomics in complex trait correlations. Our findings not only advance the understanding of proteomic contributions to phenotypic traits but also suggest potential applications for evaluating shared omics architectures in other domains such as transcriptomics and metabolomics. ### Competing Interest Statement The authors have declared no competing interest.
People acquire concepts through rich physical and social experiences and use them to understand and navigate the world. In contrast, large language models (LLMs), trained solely through next-token prediction on text, exhibit strikingly human-like behaviors. Are these models developing concepts akin to those of humans? If so, how are such concepts represented, organized, and related to behavior? Here, we address these questions by investigating the representations formed by LLMs during an in-context concept inference task. We found that LLMs can flexibly derive concepts from linguistic descriptions in relation to contextual cues about other concepts. The derived representations converge toward a shared, context-independent structure, and alignment with this structure reliably predicts model performance across various understanding and reasoning tasks. Moreover, the convergent representations effectively capture human behavioral judgments and closely align with neural activity patterns in the human brain, providing evidence for biological plausibility. Together, these findings establish that structured, human-like conceptual representations can emerge purely from language prediction without real-world grounding, highlighting the role of conceptual structure in understanding intelligent behavior. More broadly, our work suggests that LLMs offer a tangible window into the nature of human concepts and lays the groundwork for advancing alignment between artificial and human intelligence.
Sulfate is the second most common nonmetallic ion in modern oceans, as its concentration dramatically increased alongside tectonic activity and atmospheric oxidation in the Proterozoic. Microbial sulfate/sulfite metabolism, involving organic carbon or hydrogen oxidation, is linked to sulfur and carbon biogeochemical cycles. However, the coevolution of microbial sulfate/sulfite metabolism and Earth's history remains unclear. Here, we conducted a comprehensive phylogenetic analysis to explore the evolutionary history of the dissimilatory sulfite reduction (Dsr) pathway. The phylogenies of the Dsr-related genes presented similar branching patterns but also some incongruencies, indicating the complex origin and evolution of Dsr. Among these genes, dsrAB is the hallmark of sulfur-metabolizing prokaryotes. Our detailed analyses suggested that the evolution of dsrAB was shaped by vertical inheritance and multiple horizontal gene transfer events and that selection pressure varied across distinct lineages. Dated phylogenetic trees indicated that key evolutionary events of dissimilatory sulfur-metabolizing prokaryotes were related to the Great Oxygenation Event (2.4-2.0 Ga) and several geological events in the "Boring Billion" (1.8-0.8 Ga), including the fragmentation of the Columbia supercontinent (approximately 1.6 Ga), the rapid increase in marine sulfate (1.3-1.2 Ga), and the Neoproterozoic glaciation event (approximately 1.0 Ga). We also proposed that the voluminous iron formations (approximately 1.88 Ga) might have induced the metabolic innovation of iron reduction. In summary, our study provides new insights into Dsr evolution and a systematic view of the coevolution of dissimilatory sulfur-metabolizing prokaryotes and the Earth's environment.
Modern humans have experienced explosive population growth in the past thousand years. We hypothesized that recent human populations have inhabited environments with relaxation of selective constraints, possibly due to the more abundant food supply after the Last Glacial Maximum. The ratio of nonsynonymous to synonymous mutations (N/S ratio) is a useful and common statistic for measuring selective constraints. In this study, we reconstructed a high-resolution phylogenetic tree using a total of 26,419 East Eurasian mitochondrial DNA genomes, which were further classified into expansion and nonexpansion groups on the basis of the frequencies of their founder lineages. We observed a much higher N/S ratio in the expansion group, especially for nonsynonymous mutations with moderately deleterious effects, indicating a weaker effect of purifying selection in the expanded clades. However, this observation on N/S ratio was unlikely in computer simulations where all individuals were under the same selective constraints. Thus, we argue that the expanded populations were subjected to weaker selective constraints than the nonexpanded populations were. The mildly deleterious mutations were retained during population expansion, which could have a profound impact on present-day disease patterns.
The Han Chinese history is shaped by substantial demographic activities and sociocultural transmissions. However, it remains challenging to assess the contributions of demic and cultural diffusion to Han culture and language, primarily due to the lack of rigorous examination of genetic-linguistic congruence. Here we digitized a large-scale linguistic inventory comprising 1,018 lexical traits across 926 dialect varieties. Using phylogenetic analysis and admixture inference, we revealed a north-south gradient of lexical differences that probably resulted from historical migrations. Furthermore, we quantified extensive horizontal language transfers and pinpointed central China as a dialectal melting pot. Integrating genetic data from 30,408 Han Chinese individuals, we compared the lexical and genetic landscapes across 26 provinces. Our results support a hybrid model where demic diffusion predominantly impacts central China, while cultural diffusion and language assimilation occur in southwestern and coastal regions, respectively. This interdisciplinary study sheds light on the complex social-genetic history of the Han Chinese.
The phenotype-first approach (PFA) and data-driven approach (DDA) have both greatly facilitated anthropological studies and the mapping of trait-associated genes. However, the pros and cons of the two approaches are poorly understood. Here, we systematically evaluated the two approaches and analyzed 14,838 facial traits in 2,379 Han Chinese individuals. Interestingly, the PFA explained more facial variation than the DDA in the top 100 and 1,000 except in the top 10 phenotypes. Accordingly, the ratio of heterogeneous traits extracted from the PFA was much greater, while more homogenous traits were found using the DDA for different sex, age, and BMI groups. Notably, our results demonstrated that the sex factor accounted for 30% of phenotypic variation in all traits extracted. Furthermore, we linked DDA phenotypes to PFA phenotypes with explicit biological explanations. These findings provide new insights into the analysis of multidimensional phenotypes and expand the understanding of phenotyping approaches.
A proper source of stem cells is key to muscle injury repair. Dental pulp stem cells (DPSCs) are an ideal source for the treatment of muscle injuries due to their high proliferative and differentiation capacities. However, the current myogenic induction efficiency of human DPSCs hinders their use in muscle regeneration due to the unknown induction mechanism. In this study, we treated human DPSCs with Noggin, a secreted antagonist of bone morphogenetic protein (BMP), and discovered that Noggin can effectively promote myotube formation. We also found that Noggin can accelerate the skeletal myogenic differentiation (MyoD) of DPSCs and promote the generation of Pax7+ satellite-like cells. Noggin increased the expression of myogenic markers and the transcriptional and translational abundance of satellite cell (SC) markers in DPSCs. Moreover, BMP4 inhibited Pax7 expression and activated p-Smad1/5/9, while Noggin eliminated BMP4-induced p-Smad1/5/9 in DPSCs. This finding suggests that Noggin antagonizes BMP by downregulating p-Smad and facilitates the MyoD of DPSCs. Then, we implanted Noggin-pretreated DPSCs combined with Matrigel into the mouse tibialis anterior muscle with volumetric muscle loss (VML) and observed a 73% reduction in the size of the defect and a 69% decrease in scar tissue. Noggin-treated DPSCs can benefit the Pax7+ SC pool and promote muscle regeneration. This work reveals that Noggin can enhance the production of satellite-like cells from the MyoD of DPSCs by regulating BMP/Smad signaling, and these satellite-like cell bioconstructs might possess a relatively fast capacity for muscle regeneration.
Obstructive sleep apnea (OSA) leads to chronic intermittent hypoxia (CIH) and is not well addressed by current therapies. The genioglossus (GG) is the largest upper airway dilator controlling OSA pathology, making its repair a potential treatment. This study investigates dental pulp stem cells (DPSCs) in repairing GG injury in a CIH mouse model. We induced DPSCs to myogenic lineage cells (iDPSCs) and transplanted them into GG of CIH mice. DPSCs/iDPSCs grafts improved EMGGG and muscle type transitions while reducing tumor necrosis factor α (TNF-α), alanine aminotransferase (ALT), lactate dehydrogenase (LDH), and creatine kinase (CK) levels, improving body weight. Moreover, iDPSCs increased Pax7+/Ki67+ and human-derived STEM121 cells in the GG compared with DPSCs. DPSCs/iDPSCs enhanced Desmin+ myotube formation in myoblasts under hypoxia in vitro, with iDPSCs increased human-derived myogenic markers and nuclei in myotubes. These results indicate that iDPSCs, beyond their paracrine effects like DPSCs, directly participate in myogenic differentiation, supporting the potential use of DPSCs for OSA treatment.
Abstract Background Hmong–Mien (HM) speakers are linguistically related and live primarily in China, but little is known about their ancestral origins or the evolutionary mechanism shaping their genomic diversity. In particular, the lack of whole-genome sequencing data on the Yao population has prevented a full investigation of the origins and evolutionary history of HM speakers. As such, their origins are debatable. Results Here, we made a deep sequencing effort of 80 Yao genomes, and our analysis together with 28 East Asian populations and 968 ancient Asian genomes suggested that there is a strong genetic basis for the formation of the HM language family. We estimated that the most recent common ancestor dates to 5800 years ago, while the genetic divergence between the HM and Tai–Kadai speakers was estimated to be 8200 years ago. We proposed that HM speakers originated from the Yangtze River Basin and spread with agricultural civilization. We identified highly differentiated variants between HM and Han Chinese, in particular, a deafness-related missense variant (rs72474224) in the GJB2 gene is in a higher frequency in HM speakers than in others. Conclusions Our results indicated complex gene flow and medically relevant variants involved in the HM speakers’ evolution history.
Reconstructing the spatial evolution of languages can deepen our understanding of the demic diffusion and cultural spread. However, the phylogeographic approach that is frequently used to infer language dispersal patterns has limitations, primarily because the phylogenetic tree cannot fully explain the language evolution induced by the horizontal contact among languages, such as borrowing and areal diffusion. Here, we introduce the language velocity field estimation, which does not rely on the phylogenetic tree, to infer language dispersal trajectories and centre. Its effectiveness and robustness are verified through both simulated and empirical validations. Using language velocity field estimation, we infer the dispersal patterns of four agricultural language families and groups, encompassing approximately 700 language samples. Our results show that the dispersal trajectories of these languages are primarily compatible with population movement routes inferred from ancient DNA and archaeological materials, and their dispersal centres are geographically proximate to ancient homelands of agricultural or Neolithic cultures. Our findings highlight that the agricultural languages dispersed alongside the demic diffusions and cultural spreads during the past 10,000 years. We expect that language velocity field estimation could aid the spatial analysis of language evolution and further branch out into the studies of demographic and cultural dynamics.
Large-scale genomic projects and ancient DNA innovations have ushered in a new paradigm for exploring human evolutionary history. However, the genetic legacy of spatiotemporally diverse ancient Eurasians within Chinese paternal lineages remains unresolved. Here, we report an integrated Y-chromosome genomic database encompassing 15,563 individuals from both modern and ancient Eurasians, including 919 newly reported individuals, to investigate the Chinese paternal genomic diversity. The high-resolution, time-stamped phylogeny reveals multiple diversification events and extensive expansions in the early and middle Neolithic. We identify four major ancient population movements, each associated with technological innovations that have shaped the Chinese paternal landscape. First, the expansion of early East Asians and millet farmers from the Yellow River Basin predominantly carrying O2/D subclades significantly influenced the formation of the Sino-Tibetan people and facilitated the permanent settlement of the Tibetan Plateau. Second, the dispersal of rice farmers from the Yangtze River Valley carrying O1 and certain O2 sublineages reshapes the genetic makeup of southern Han Chinese, as well as the Tai-Kadai, Austronesian, Hmong-Mien, and Austroasiatic people. Third, the Neolithic Siberian Q/C paternal lineages originated and proliferated among hunter-gatherers on the Mongolian Plateau and the Amur River Basin, leaving a significant imprint on the gene pools of northern China. Fourth, the J/G/R paternal lineages derived from western Eurasia, which were initially spread by Yamnaya-related steppe pastoralists, maintain their presence primarily in northwestern China. Overall, our research provides comprehensive genetic evidence elucidating the significant impact of interactions with culturally distinct ancient Eurasians on the patterns of paternal diversity in modern Chinese populations.