Background Effective patient education requires accurate communication aligned with patients’ emotional and semantical needs. Text-based large language models (LLMs) lack access to non-verbal cues, which may contribute to misaligned responses. Methods We evaluated emotional and semantic misalignment in a text-based LLM using 64,200 utterances from 16,583 patient education cases across six departments and three centers. Dolphin was developed integrating text and audio cues and evaluated through emotion recognition, semantic consistency assessment, branch-level ablations, and a double-blinded randomized trial against a matched text-based LLM comparator (Chinese Clinical Trial Registry: (ChiCTR2500095933). Findings The text-based LLM showed emotional misalignment in 36.7% of responses and semantic misalignment in 28.3% of cases, with higher misalignment under greater burden. Dolphin outperformed the text-based LLM in emotion recognition accuracy (0.886 vs. 0.713) and semantic consistency (84.9% vs. 82.1%; both adjusted p < 0.001). Ablations supported contribution of audio branches. Dolphin received higher expert ratings than the text-based LLM and human educators (all p < 0.001). In 555 patients, Dolphin was associated with greater patient satisfaction (98.6% vs. 93.8%), suggestion acceptance (76.1% vs. 58.9%; p < 0.001), proactive disclosure (44.6% vs. 26.5%; p < 0.001), and fewer 7-day unplanned recontact (12.9% vs. 22.9%; p = 0.002). No unsafe recommendations or safety events were identified. Conclusions Compared with text-based LLM, Dolphin improved emotional-semantic alignment and patient-education outcomes, supporting bimodal alignment as a strategy for reducing misalignment-driven communication failures. Funding National Natural Science Foundation of China, State Key Laboratory Special Fund, and Chinese Academy of Medical Sciences Innovation Fund.
Medical reasoning is fundamental to clinical decision-making, underpinning tasks such as patient communication, diagnosis, and treatment planning. Inspired by psychological findings that peer interaction promotes self-correction, we introduce model confrontation and collaboration (MCC), a debate intelligence framework that transcends static ensemble methods by integrating critique and self-reflection to iteratively refine reasoning through structured, multi-round confrontation and collaboration among diverse large language models (LLMs). In multiple-choice benchmarks, MCC achieved mean accuracy on MedQA (92.6%) and PubMedQA (84.8%) and demonstrated strong performance on medical subsets of MMLU. In long-form medical question answering, MCC outperformed all individual LLMs and the domain-specific LLM Med-PaLM 2 in both physician and layperson evaluations. In diagnostic dialog tasks, MCC further excelled in both history-taking and diagnostic accuracy, reaching a top-1 diagnosis rate of 80%. These results position MCC as a scalable, model-agnostic framework that advances medical reasoning through collaborative deliberation.
Genome-wide association studies identified a melanoma- and nevus count-associated locus on chromosome band 9q34.13. Fine-mapping and melanocyte expression data collectively suggest two potential causal genes with opposite association with risk: higher levels of Rap guanine nucleotide exchange factor 1 ( RAPGEF1 ) and lower levels of uridine-cytidine kinase 1 ( UCK1 ). Colocalization analyses and conditional TWAS suggest multiple causal cis -regulatory sequence variants in partial linkage disequilibrium (LD) to each other. Melanocyte capture-HiC and CRISPR-inhibition demonstrated regulatory interactions between fine-mapped variants and the RAPGEF1 and UCK1 promoters. Focusing on RAPGEF1 , we demonstrate RAPGEF1 expression promotes melanocyte growth and drives malignant transformation of human immortalized melanocytes. Following treatment with human EGF, RAPGEF1 overexpression activated both RAP1 and RAS. Further, we show RAPGEF1 expression is significantly enriched in melanomas lacking strongly activating RAS-MAPK mutations, suggesting that RAPGEF1 may promote oncogenic RAS-MAPK signaling in melanomas. Furthermore, in these tumors, we provide preliminary evidence to support the prognostic relevance of RAPGEF1 expression in patients lacking RAS or BRAF mutations. Together with other recent studies, these data suggest that germline variation influencing RAS activation may play a key role in nevus development and melanoma risk.
Large language models (LLMs) are increasingly applied in clinical communication, yet their reliability depends on high-quality conversational corpora. Real-world doctor-patient recordings are frequently degraded by noise, transcription errors, speaker overlap, and fragmented dialogue structure, limiting their usability for downstream model training. Here, we present an agent-based transcription framework that autonomously converts raw unstructured conversation transcriptions (RUCT) into structured conversation transcriptions (SCT) suitable for LLM fine-tuning. The system integrates three coordinated modules-Planner, Memory, and Executor-to orchestrate noise removal, content correction, speaker identification, and dialogue segmentation within a self-correcting workflow. Applied to 7197 minutes of Chinese clinical recordings across eight departments, with an additional 240 minutes of English-language dialogues used as a limited portability check, the agent achieved high reconstruction accuracy (94.7% denoising, 96.9% content correction, 88.6% speaker identification, 92.7% segmentation) and operated 3.6× faster than manual processing. In controlled comparisons against a cascaded deep-learning pipeline, a sequential non-agent execution, and an end-to-end large-context model, the agent achieved consistently higher performance across all four processing tasks. Architectural ablation further revealed marked degradation when Planner or Memory modules were removed (e.g., up to 47.6% reduction in speaker identification), supporting the contribution of coordinated task decomposition and cross-step state retention. To assess downstream impact, we fine-tuned an independent open-weight model (Qwen3-32B) on agent-generated SCT versus RUCT derived from an identical training set. Agent-generated SCT fine-tuning significantly improved overall quality scores (3.1 to 3.7; P < 0.001; Fleiss' κ = 0.82) in blinded expert evaluation across six clinically grounded dimensions, and also yielded higher scores on an external medical dialogue benchmark (HealthBench) than both RUCT fine-tuning and the non-fine-tuned baseline. These findings indicate that agent-structured clinical corpora enhance LLM fine-tuning performance and provide a scalable framework for reliable medical conversational AI development.
Genetic regulation of splicing uniquely contributes to trait-associated genome-wide association studies (GWAS) signals. However, quantitative trait loci (QTL) analysis using short-read sequencing of bulk tissues fails to capture full-length and cell-type-specific isoforms. Here, we present an isoform-level lung cell atlas from 129 never-smoking Korean women using single-cell long-read RNA-sequencing, identifying abundant unannotated and cell-type-specific isoforms. Isoform-level signatures of 37 lung cell types display a larger difference and therefore improve cell-type classification compared to gene-level expression. Notably, isoform-QTLs (isoQTLs) detect unannotated and/or cell-type-specific isoforms with independent genetic regulation from expression-QTL (eQTL), supported by enriched splicing functional elements. IsoQTLs nominate susceptibility isoforms from previously unexplained lung function and cancer GWAS loci, via eQTL-independent signals. We highlight a potentially functional novel variant of PPIL6 in multiciliated cells underlying lung cancer risk through alternative splicing. This isoform-level resource advances our understanding of cell-type-specific isoform regulation and its contribution to lung traits and diseases.
BACKGROUND:Chronic obstructive pulmonary disease (COPD) is the fourth leading cause of death worldwide. COPD is characterized by progressive airflow restriction and wide-spectrum heterogeneity across clinical manifestations, treatment responses, and underlying mechanisms. While recent omics approaches have advanced our understanding of COPD, in-depth multi-omics characterization remains scarce, leaving a critical gap in knowledge for the understudied population. METHODS:We conducted deep multi-omics profiling on a unique Chinese population from a region with high COPD prevalence, high-altitude residence, and widespread exposure to biomass fuels (74.8%). We recruited 159 COPD patients from 5 medical centers and integrated their radiomics, metabolomics, microbiomics, and genomics data, identifying three distinct molecular subtypes: "stable state" (SS), "restrained state" (RS), and "crumbly state" (CS). FINDINGS:The SS subtype is marked by the least acute exacerbation and mildest airflow limitation with potential eosinophilic inflammation. The RS subtype is characterized by extensive pulmonary structural damage, greatest airflow limitation, but mild clinical symptoms. The CS subtype is typified by disturbed microbial interaction, high triethanolamine, coal dust exposure, and nicotine dependence. CONCLUSIONS:This study provides a comprehensive multi-omics profile of COPD in a previously understudied population and reveals three molecular subtypes that enhance the understanding of COPD heterogeneity.
Single-cell expression quantitative trait loci (sc-eQTL) analyses are powerful in identifying context-specific susceptibility genes from genome-wide association studies (GWAS) loci. However, few studies have comprehensively investigated cells of lung cancer origin in non-European populations. Here, we built a lung sc-eQTL dataset from 129 Korean women never-smokers with epithelial cell enrichment. eQTL mapping identified 2,229 genes with an eQTL in 33 cell types, including East Asian-specific findings when compared to predominantly European datasets. Integration with single-cell chromatin accessibility data demonstrated an enrichment of cell-type specific eQTLs in cell-type matched candidate enhancers, while shared eQTLs were more frequently found near promoters. Colocalization and transcriptome-wide association study unveiled 36 susceptibility genes from 22 cell types in 22 lung cancer loci, including 10 loci not achieving genome-wide significance in prior GWAS. Around 47% of these genes were from cells of the alveoli, underscoring their importance, especially in lung adenocarcinoma (LUAD) susceptibility. Focusing on the trajectory of alveolar epithelial cell regeneration, we detected 785 cell-state-interacting QTLs, which overlapped with 28% (10) of the identified susceptibility genes. Finally, we experimentally validated East Asian- and alveolar type 2 cell-specific eQTL of TCF7L2 underlying East Asian LUAD locus, 10q25.2. Consistent with its role as a Wnt/β-catenin effector, TCF7L2 displayed significant effect on lung adenocarcinoma cell growth. Our data highlighted context-specific susceptibility genes, especially from alveolar cells of lung, contributing to lung cancer etiology.
Deep phenotyping data are essential for understanding disease mechanisms and individual health trajectories. However, its unprecedented scale and complexity pose significant analytical challenges that conventional approaches struggle to address. Here, we introduce ukbFound, a foundation model that encoded thousands of individual-level traits into language-like sequences. By incorporating domain-specific tokenization, position-free embedding, and interpretable reasoning, ukbFound effectively captures latent disease-trait relationships from 502,118 UK Biobank individuals. We demonstrate its versatility in three downstream applications. In disease stratification, ukbFound identifies distinct patient subgroups in 289 diseases, with 53/289 (18.3%) showing FDR-significant prognostic differences. Notably, ukbFound reveals two subgroups in chronic obstructive pulmonary disease distinguished by unique basophil count distributions, suggesting a novel indicator for lung function decline. In multimorbidity network analysis, ukbFound uncovers previously unreported associations (e.g., low platelet disorder and gout) and disease communities with shared etiological mechanisms. In disease prediction, ukbFound identified high-risk individuals based solely on lifestyle and dietary data, outperforming ten benchmark models by ΔAUC gains of +0.03 to +0.16. Notably, the highest-risk group showed 17.5-fold greater odds of developing gout up to 8 years in advance. Collectively, ukbFound provides a scalable and interpretable framework for modeling deep phenotyping data and advancing precision medicine.
Abstract Background: Lung cancer is one of the most prevalent and life-threatening cancers worldwide. Genetic factors contribute to lung cancer risk in smokers and non-smokers, and genome-wide association studies (GWAS) identified > 50 genomic loci from diverse populations. Expression quantitative trait loci (eQTL) and splicing QTL (sQTL) analyses can reveal distinct genetic effects on gene- or transcript isoform-level regulation contributing to GWAS signals to identify target genes and elucidate etiology. However, current sQTL studies using short-read sequencing of bulk tissues fail to capture full-length and novel isoforms or cell-type-specific splicing events contributing to tumorigenesis.Methods: We present an isoform-level lung cell atlas from 129 never-smoking Korean women using single-cell long-read RNA-sequencing. Epithelial cells, where lung cancer originate, were enriched by FACS sorting. Isoform signatures of each lung cell type were identified by differential analysis. We mapped isoform-level QTLs (isoQTLs) using jaxQTL negative binomial model of pseudo-bulk counts. Colocalization and transcriptome-wide association study (TWAS) with GWAS were performed to prioritize susceptibility isoforms.Results: We identified 325,865 full-length isoforms from 360,133 lung cells, where 83% are novel isoforms not annotated in GENCODE v32. We identified isoform-level signatures of 37 lung cell types, where 67.2% of differential isoforms display larger differences than in gene levels. Isoform-QTLs (isoQTLs) identified unreported genes in bulk tissue-based sQTL studies attributed to cell-type-specific and unannotated isoforms. Compared to the barcode-matched short-read expression data, 46% of isoQTLs did not colocalize with eQTLs in the same cell type. Consistently, isoQTLs were distinctly enriched in the functional elements of splicing and post-transcriptional regulation. Colocalization of isoQTLs with lung cancer and trait GWAS signals nominated candidate isoforms, where 69% were previously unreported at the gene level. Moreover, 71% of GWAS-colocalized isoforms were independent from eQTLs, including PPIL6-207 for lung cancer. TWAS of ancestry-matched lung adenocarcinoma identified 12 isoforms from the sub-threshold GWAS regions, which include lineage or cell-type-specific genes.Conclusions: We established an isoform-level lung cell atlas using single-cell long-read sequencing and detected isoQTLs that are cell-type-specific and independent from eQTLs. This isoform-level resource advances our understanding of cell-type-specific isoform regulation and its contribution to lung cancer and diseases. Citation Format: Bolun Li, Thong Luong, Elelta Sisay, Jinhu Yin, Zixuan Zhang, Ju Hye Shin, Jinyoung Byun, Yoon Soo Chang, Maria Teresa Landi, Nat Rothman, Erping Long, Qing Lan, Christopher I. Amos, Tongwu Zhang, Jianxin Shi, Nicholas Mancuso, Haoyu Zhang, Jin Gu Lee, Eon Young Kim, Jiyeon Choi. Single-cell full-length transcriptome of lung cells reveals genetic effects on isoform regulation beyond eQTL [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2026; Part 1 (Regular Abstracts); 2026 Apr 17-22; San Diego, CA. Philadelphia (PA): AACR; Cancer Res 2026;86(7 Suppl):Abstract nr 5924.
Genome-wide association studies (GWASs) of melanoma risk have identified 68 independent signals at 54 loci. For most loci, specific functional variants and their respective target genes remain to be established. Capture-HiC is an assay that links fine-mapped risk variants to candidate target genes by comprehensively mapping chromatin interactions. We performed a melanoma GWAS region-focused capture-HiC assay in human primary melanocytes to identify physical interactions between fine-mapped risk variants and potential causal melanoma-susceptibility genes. Overall, chromatin-interaction data alone nominated potential causal genes for 61 of the 68 melanoma risk signals, identifying many candidates beyond those reported by previous studies. We further integrated these data with epigenomic (chromatin state, accessibility), gene expression (expression quantitative trait locus [eQTL]/transcriptome-wide association study [TWAS]), DNA methylation (methylation QTL [meQTL]/methylome-wide association study [MWAS]), and massively parallel reporter assay (MPRA) data generated from melanoma-relevant cell types to prioritize potentially cis-regulatory variants and their respective candidate gene targets. From the set of fine-mapped variants across these loci, we identified 140 prioritized credible causal variants linked to 195 candidate genes at 42 risk signals. In addition, we developed an integrative scoring system to facilitate candidate gene prioritization, integrating melanocyte and melanoma datasets. Notably, at several GWAS risk signals, we observed long-range chromatin connections (500 kb to >1 Mb) with distant candidate target genes. We validated several such cis-regulatory interactions using CRISPR inhibition, providing evidence for known cancer driver genes MDM4 and CBL, as well as the SRY-box transcription factor SOX4, as likely melanoma risk genes.
Understanding how genetic variation contributes to phenotypic variation is a fundamental question in genetics. Genome-wide association studies (GWASs) have discovered numerous genetic associations with various human phenotypes, most of which contain co-inherited variants in strong linkage disequilibrium (LD) with indistinguishable statistical significance. The experimental and analytical difficulty in identifying the “causal variant” among the co-inherited variants has traditionally led mechanistic studies to focus on relatively simple loci, where a single functional variant is presumed to explain most of the association signal and affect a target gene. The notion that a single causal variant is responsible for an association signal, while other variants in LD are merely correlated, has often been assumed in functional studies. However, emerging evidence powered by high-throughput experimental tools and context-specific functional databases argues that even a single independent signal may involve multiple functional variants in strong LD, each contributing to the observed genetic association. In this perspective, we articulate this evolving understanding of causal variants through examples from both traditional locus-by-locus approaches and more recent high-throughput functional studies. We then discuss the implications and prospects of this notion in understanding the genetic architecture of complex traits and interpreting the variant-level causality in GWAS follow-up studies.
Despite lung cancer affecting all races and ethnicities, disparities are observed in incidence and mortality rates among different ethnic groups in the United States. Non-Hispanic African Americans had a high incidence rate of lung cancer at 55.8 per 100 000 people, as well as the highest death rate at 37.2 per 100 000 people from 2016 to 2020. While previous genome-wide association studies (GWAS) have identified over 45 susceptibility risk loci that influence lung cancer development, few GWAS have investigated the etiology of lung cancer in African Americans. To address this gap in knowledge, we conducted GWAS of lung cancer focused on studying African Americans, comprising 2267 lung cancer cases and 4264 controls. We identified three loci associated with lung cancer, one with lung adenocarcinoma, and four with lung squamous cell carcinoma in this population at the genomic-wide significance level. Among them, three novel loci were identified near VWF at 12p13.31 for overall lung cancer and GACAT3 at 2p24.3 and LMAN1L at 15q24.1 for lung squamous cell carcinoma. In addition, we confirmed previously reported risk loci with known or new lead variants near CHRNA5 at 15q25.1 and CYP2A6 at 19q13.2 associated with lung cancer and TRIP13 at 5p15.33 and ERC1 at 12p13.33 associated with lung squamous cell carcinoma. Further multi-step functional analyses shed light on biological mechanisms underlying these associations of lung cancer in this population. Our study highlights the importance of ancestry-specific studies for the potential alleviation of lung cancer burden in African Americans.
The evolution of surgical techniques aims to augment surgeons' capabilities through digital guidance and robotization for higher precision and consistency. Currently, surgeries heavily rely on surgeon's experience and visual judgment, causing operation variations. Artificial intelligence (AI) offers a solution by extracting and digitizing surgical trajectories and features from videos to provide digital guidance. In this study, we collected 17,538 videos of capsulorhexis, a crucial step in cataract surgery, to create an AI-driven system named Meta Surgery (MetaS). MetaS evaluates and identifies ideal cases, extracts their digital characteristics, and fits an optimal capsulorhexis path in real-time during surgery. Surgeons performed capsulorhexis benefited from MetaS's guidance and a lens caliper, which increased the rate of ideal capsulorhexis by ~40%. Additionally, these digital features enabled a surgical robot to perform precise capsulorhexis autonomously in porcine eyes. This approach augments surgeons' surgical skills and paves the way for the autonomous operation of surgical robots.
Even genetically identical cells in homogeneous environments exhibit heterogeneous mRNA abundance, typically called “gene expression noise,” which is involved in key cellular activities, evolutionary processes, and disease mechanisms. However, determinants of the gene expression noise and its functional role in variations of human complex traits remain largely unexplored. Here, we established an atlas of gene expression noise from 1.23 million human peripheral blood cells of 981 individuals, identifying its age- and gender-dependent pattern. We then identified 10,770 independent expression noise quantitative trait loci (enQTLs) for 6,743 unique enGenes across seven immune cell types. Most enQTLs were distinct from expression quantitative trait loci (eQTLs) and showed differential enrichment of functional elements across the genome. Colocalization of enQTLs with trait-associated genetic loci interpreted previously unexplained loci. Overall, this study unravels the genetic determinants of gene expression noise and implicates as a previously underappreciated mechanism underlying variation of human complex traits and diseases.
ABSTRACT Chronic obstructive pulmonary disease (COPD) associates with increased lung cancer incidence and shares genetic susceptibility, yet its independent causal role and driver mechanisms are poorly understood. We integrated data from the National Health and Nutrition Examination Survey (NHANES) cohort with genome‐wide association studies (GWAS) summary statistics and Mendelian randomization analyses to map genetic correlations and infer causality between COPD phenotypes and lung cancer. Post‐GWAS methods—including transcriptome‐wide association study, colocalization, partitioned heritability via heritability estimation from summary statistics (ρ‐HESS), and cross‐phenotype association (CPASSOC)—identified shared susceptibility loci, highlighting IREB2 and CD27⁺ B cells as potential mediators. Elevated IREB2 expression correlated with accelerated lung‐function decline in COPD but predicted improved prognosis in lung cancer B cells, whereas higher CD27⁺ B cell levels in COPD were associated with protumorigenic activity. Single‐cell transcriptomic analysis and in vitro knockdown experiments confirmed IREB2's role in modulating B‐cell activation and apoptosis pathways within tumors. These results support COPD as an independent lung cancer risk factor and implicate IREB2 and CD27⁺ B cells in COPD‐to‐cancer progression, laying groundwork for early detection and targeted intervention in high‐risk individuals.
The dynamics of chromatin conformation involve continuous and reversible changes within the nucleus of a cell, which participate in regulating processes such as gene expression, DNA replication, and damage repair. Here, SEE is introduced, an artificial intelligence (AI) method that utilizes autoencoder and transformer techniques to analyze chromatin dynamics using single-cell RNA sequencing data and a limited number of single-cell Hi-C maps. SEE is employed to investigate chromatin dynamics across different scales, enabling the detection of (i) rearrangements in topologically associating domains (TADs), and (ii) oscillations in chromatin interactions at gene loci. Additionally, SEE facilitates the interpretation of disease-associated single-nucleotide polymorphisms (SNPs) by leveraging the dynamic features of chromatin conformation. Overall, SEE offers a single-cell, high-resolution approach to analyzing chromatin dynamics in both developmental and disease contexts.
Heart failure with preserved ejection fraction (HFpEF) as a high heterogeneity clinical syndrome, is commonly associated with diastolic dysfunction, and has no effective therapy, which is obviously distinct from Heart failure with reduced ejection fraction (HFrEF). Currently the differences of cell type heterogeneity between HFpEF and HFrEF remain largely unknown. Here we illustrate an atlas consisting of 21,747 cardiac cells, including both HFpEF and HFrEF. Cell-cell communication analysis reveals cardiomyocytes rather than endothelium or fibroblasts were dominant communication "hub" in HFpEF. The subtypes of cardiomyocytes are highly heterogeneous between HFpEF and HFrEF. Notably, a specific subtype of cardiomyocytes shows significant gene expression associated with the metabolism of fatty acids. Additionally, regulon analysis reveals that Ppargc1a, Atf6, E2f6, and Mitf exhibited specific elevated regulation in the subtype of cardiomyocytes of HFpEF. Furthermore, we have identified 210 HF susceptibility genes from HF-associated GWAS data. After integrating scRNA-seq, GWAS, and eQTL data, the genetic susceptibility underlying HFpEF and HFrEF were discussed. In conclusion, this study not only comprehensively characterizes the differences of cardiomyocytes changes but also provides insights into potential targets for cell type- and subtype-specific molecules between HFpEF and HFrEF.
Genome-wide association studies (GWAS) identified over fifty loci associated with lung cancer risk. However, underlying mechanisms and target genes are largely unknown, as most risk-associated variants might regulate gene expression in a context-specific manner. Here, we generate a barcode-shared transcriptome and chromatin accessibility map of 117,911 human lung cells from age/sex-matched ever- and never-smokers to profile context-specific gene regulation. Identified candidate cis-regulatory elements (cCREs) are largely cell type-specific, with 37% detected in one cell type. Colocalization of lung cancer candidate causal variants (CCVs) with these cCREs combined with transcription factor footprinting prioritize the variants for 68% of the GWAS loci. CCV-colocalization and trait relevance score indicate that epithelial and immune cell categories, including rare cell types, contribute to lung cancer susceptibility the most. A multi-level cCRE-gene linking system identifies candidate susceptibility genes from 57% of the loci, where most loci display cell-category-specific target genes, suggesting context-specific susceptibility gene function. Multiple genetic loci are associated with lung cancer risk, but the underlying genetic mechanisms remain poorly understood. Here, the authors perform single-cell RNA-seq and ATAC-seq analyses of lung cells from ever- and never-smokers; they report candidate cis-regulatory elements that colocalise with candidate causal variants in lung cancer risk loci and potential susceptibility genes.