
Drug-target affinity (DTA) prediction is an important task in computer-aided drug discovery. Existing methods usually use a fixed drug-protein interaction strategy. This design is hard to adapt to heterogeneous samples. It is also limited in cold-start and low-similarity scenarios. This study proposes SIGMA-DTA, a similarity-guided adaptive interaction modeling framework for DTA prediction. The framework uses drug and protein similarities as explicit reasoning priors. It introduces a similarity-driven routing mechanism. The mechanism assigns weights to different interaction paths according to the distance between a sample and the training distribution. This enables sample-specific interaction modeling. Unlike conventional unified interaction schemes, SIGMA-DTA adjusts information propagation under different similarity levels. It can model complex drug-target relationships more effectively. Experiments on the Davis and KIBA datasets verify the effectiveness of the method. The results show that integrating similarity information into the interaction modeling process improves robustness and generalization in DTA prediction.
OBJECTIVE:Large-scale medical instruction datasets often contain substantial noise and redundancy, leading to high training costs without significant performance gains. Existing data selection and distillation methods rely on single static criteria, failing to capture the diversity and dynamic nature of medical samples during learning. This study proposes TriFuse, a three-stage data distillation framework that integrates multiple complementary learning signals for efficient training of medical question-answering models. METHODS:The framework comprises three stages: (1) Proxy-student filtering: utilizing masked language modeling loss to remove information-deficient samples; (2) Gradient-median filtering: under the guidance of a weak teacher model, retaining samples with stable and moderate gradient contributions to suppress noise and outliers; (3) Forgettable-count filtering: tracking sample stability across training epochs to identify and remove samples exhibiting unstable learning behaviors. The framework was applied to the large-scale medical instruction dataset MedS-Ins to construct a refined subset (MedCore) comprising less than 30% of the original data. RESULTS:Models trained on MedCore (28% of the original data) achieved comparable or superior performance across multiple tasks while substantially reducing training time and computational cost compared with full-data training. MedCore-Llama outperformed GPT-4 on French medical QA (65.6%) and HeadQA (63.2%), and achieved state-of-the-art results in Participant Extraction (82.61%) and medical concept explanation. Ablation studies confirmed the necessity of all three filtering stages. CONCLUSION:TriFuse effectively distills high-value medical training data through multi-signal progressive filtering, significantly reducing computational overhead while maintaining model performance. This framework provides a scalable, low-cost solution for efficient medical language model training in resource-constrained environments.
OBJECTIVE:To evaluate challenges in managing multisite, longitudinal, and multisystem clinical data for Down syndrome (DS) research and identify approaches to improve standardization and reuse. METHODS:This report summarizes the outcomes and lessons learned from a NIH workshop convened by the INCLUDE Project to align and modernize clinical data management (CDM) practices using common data elements (CDEs) and artificial intelligence (AI)-enabled tools. RESULTS:(1) consensus that shared CDEs, ontologies, and consistent data models are foundational for cross-study harmonization and interoperability; (2) identification of complementary electronic health record strategies, including standards-based exchange and research data models to support scalable extraction and analysis; (3) recognition that AI-enabled methods, including natural language processing with human review, can accelerate abstraction of information from clinical narratives and support data harmonization across heterogeneous sources; and (4) prioritization of governance needs for privacy protection, transparency, bias mitigation, and ongoing oversight when applying AI to sensitive health data. CONCLUSION:The major conclusion is that combining standardized CDE-driven design with appropriately governed AI-enabled workflows can reduce manual burden, improve data quality, and enable integrated multimodal research, with lessons that are not only applicable the INCLUDE Project but also to other complex clinical research programs.
Knowledge graphs are increasingly used to organize and retrieve complex medical information, yet existing graph-based retrieval systems often suffer from high construction costs, limited scalability as knowledge grows, and limited interpretability in clinical practice. These challenges are amplified in cardiovascular medicine, where data are heterogeneous, noisy, and linked by complex relationships. In this work, we present CGX, a domain-oriented GraphRAG framework that mirrors clinical reasoning for explainable heart failure analysis. CGX structures cardiovascular knowledge into a three-layer hierarchy that spans patient-level observations, guideline-based evidence, and standardized ontologies. An OCR-enhanced preprocessing pipeline combined with a zero-shot biomedical transformer converts PDF-based biomedical literature/guidelines and machine-readable clinical narratives into semantic triples, reducing error propagation compared with vanilla RAG. A Hybrid U-Retrieval mechanism then exploits the graph topology through top-down summary retrieval and bottom-up path refinement, producing explicit evidence chains that support each answer. Initial experiments on heart-failure-focused clinical question answering show that CGX improves evidence retrieval quality and perceived answer reliability over conventional retrieval methods, while reducing total graph construction time by 69.7% under the same input corpus and hardware setting. Under blinded expert clinical evaluation, CGX reduces the rate of clinically risky answers from 12.4%-14.0% to 8.3%, alongside substantially higher scores across all five expert-rated Likert criteria compared with both baselines. These results suggest that CGX offers a scalable and reusable GraphRAG architecture for integrating structured medical knowledge with large language models to support trustworthy clinical decision-making.
Objective Repeated measurements capture the progression of health over time and may inform survival prediction. The goal of this review is to identify and compare how the presented methodologies for survival prediction with repeated measurements have been applied to predict survival outcomes with multivariate structured repeated measurements and cluster them into families of methodologies. Methods We performed a search on MEDLINE via Ovid, PubMed, Web of Science, and Embase. We included original peer-reviewed publications where a Time-To-Event (TTE) prediction model was developed, based on multivariate repeated measurements of health data. The protocol was registered in PROSPERO (CRD42024529572). The main outcome of interest was the strategy used for dealing with repeated measurements. We assigned a score to each methodology in terms of temporal modelling, ease of implementation, interpretability and explainability, flexibility, computational power, scalability, ability to deal with irregular sampling and dynamic prediction, based on the data extracted from the corresponding studies. Results After screening 4029 records, we included 58 studies in the review. The studies were categorized based on the underlying survival models into traditional statistics and machine learning methodologies. In parallel, four families of methodologies to deal with repeated measurements were identified: joint models (statistics n = 18), deep learning based methods (machine learning n = 9), landmarking (statistics n = 13, machine learning n = 6) and data manipulation (statistics n = 3, machine learning n = 9). Statistical studies had a higher risk of bias, reported models trained on less longitudinal covariates and were based on smaller sample sizes, while machine learning could deal less often with irregular sampling of measurements. Joint models and deep learning were generally more suitable when dealing with complex temporal dependencies, while landmarking and data manipulation were simpler and computationally lighter options. Conclusions This systematic review identified and compared published studies that developed TTE prediction methodologies using repeated covariate measurements to predict clinical outcomes. Greater emphasis on external validation, transparent reporting, interpretability and explainability, and multimodal data integration is crucial to advance the field and enhance its impact on health outcomes.
The precise prediction of Antibody-Antigen Interaction (AAI) is a pivotal task for accelerating antibody drug discovery and virtual screening. To address the challenges of suboptimal multimodal fusion and the paucity of physical interpretability in existing approaches, this paper proposes PAC-Net, a physical prior-guided end-to-end deep learning framework. First, the model incorporates a Gated Multimodal Fusion mechanism that effectively integrates sequence semantics with structural features via dynamic weight allocation, thereby achieving adaptive alignment of multimodal information. Furthermore, the core Physics-Guided Hybrid Interaction Module encodes biophysical laws, including charge complementarity and hydrophobic interactions, directly as inductive biases for the attention mechanism. By integrating these biases with parallel depthwise separable convolutions within a unified architecture, the model synergistically captures both global long-range dependencies and local structural patterns among residues. Experimental results on two public datasets, HIV and CoV-AbDab, demonstrate that PAC-Net significantly outperforms current state-of-the-art methods in terms of prediction accuracy and robustness. Particularly in highly challenging antibody and antigen cold-start scenarios, the model exhibits exceptional cross-entity generalization performance, driven by its two innovative mechanisms: gated multimodal fusion and physical rule-guided attention. Consequently, PAC-Net provides a high-precision and interpretable computational tool for the virtual screening of antibody therapeutics. The source codes are publicly available at the following link https://github.com/WeiSongJian/PAC-Net.
OBJECTIVE:To evaluate whether small language models adapted on public medical benchmarks transfer reliably to institution-constrained open-answer clinical QA, and to identify which adaptation choices - supervision format, optimization objective, and backbone - determine local answer quality and evidence coverage. METHODS:We used DistilGPT2 (82M) as the primary student and Llama3 70B as the teacher. We compared four strategies: public-benchmark answer-only adaptation, public-benchmark chain-of-thought adaptation, in-domain teacher-supervised question-answer fine-tuning (QAFT), and direct preference optimization (DPO). We evaluated each strategy on cleaned public benchmarks and on an internal EHR-grounded open-answer task using token-F1, exact match, hallucinated evidence rate, and evidence coverage. We tested robustness via multi-seed reruns, a controlled DPO pair-construction ablation, a hard-negative DPO variant, and cross-backbone replication on TinyLlama 1.1B and a modern Qwen2.5-3B model. RESULTS:Public-benchmark adaptation did not produce stable gains in repeated-seed external evaluation. In the primary internal comparison, in-domain QAFT outperformed public-benchmark transfer (F1 0.1310 vs 0.1090). DPO did not improve answer quality uniformly; instead, it shifted models toward shorter, stricter responses with lower evidence coverage. In a controlled fixed-split three-run ablation, this shift varied with rejected-response construction. TinyLlama replication showed backbone-specific DPO effects. CONCLUSION:For institution-constrained clinical QA, in-domain teacher-supervised fine-tuning was the most reliable evaluated adaptation path. Public-benchmark transfer was not a dependable proxy for local utility in the primary internal comparison, and DPO should be interpreted as operating-point control with local validation before deployment. We propose a local evaluation protocol: adapt on the target task, validate DPO as an operating-point control, and interpret groundedness jointly with evidence coverage.
OBJECTIVE:Data extraction is among the most resource-intensive and error-prone stages of systematic review production. Large language models (LLMs) offer potential for automating or semi-automating this process, yet their performance characteristics remain incompletely characterised. This systematic review aimed to comprehensively evaluate LLM accuracy, reliability, and efficiency for data extraction in evidence synthesis, and to identify optimal implementation strategies. METHODS:We searched PubMed, Embase, Web of Science, and preprint servers (medRxiv, arXiv) through December 2025. Studies were eligible if they evaluated one or more LLMs for data extraction against a human reference standard and reported quantitative performance metrics. Two reviewers independently extracted data and assessed methodological quality using PROBAST + AI and reporting completeness using TRIPOD-LLM. Narrative synthesis was performed due to substantial heterogeneity precluding meta-analysis. RESULTS:Twenty-seven studies met inclusion criteria, evaluating models including GPT-4/4o (n = 15), Claude versions 2-3.5 (n = 10), Gemini (n = 3), and open-source alternatives including Llama, Mistral, Qwen, and DeepSeek. Overall accuracy ranged from 47% to 99.9%, with substantial heterogeneity by task type and data granularity. Categorical and string variables were extracted more reliably (74-96%) than numerical data (47-88%). Claude 3.5 Sonnet achieved high accuracy in an assistive workflow (91.0%; 95% CI: 90.4-91.6%), exceeding human-only extraction (89.0%). Claude models outperformed GPT in head-to-head comparisons (OR 1.70 for event counts). Omissions were the dominant error type (60-74%), with hallucination rates of only 0.08-6%, challenging widespread fabrication concerns. Time savings of 33% to 87% were reported, although most included studies did not quantitatively assess efficiency. Methodological quality was generally robust, with 74.1% of studies rated low risk of bias under PROBAST + AI. Mean TRIPOD-LLM compliance was 88.5%, though gaps in inference settings and model version documentation were common. CONCLUSION:LLMs demonstrate promising but variable performance for data extraction in evidence synthesis. Current evidence supports their integration as assistive tools within dual-extraction workflows requiring human verification, rather than as autonomous extractors. Categorical data is extracted more reliably than numerical outcomes, and few-shot prompting with structured output formats consistently improves performance. Standardised benchmarks and prospective comparative studies remain priorities for future research.
Objective Clinical abbreviations introduce semantic ambiguity that hinders automated understanding in healthcare informatics. While generative large language models (LLMs) show promise, direct generation often lacks clinical precision. We present S-ACAD, a task-specific framework that shifts the paradigm from generation to constrained selection using dual-pathway candidate construction to improve clinical abbreviation disambiguation. Methods The S-ACAD pipeline consists of four stages: (i) Anchor generation via preliminary LLM expansion; (ii) Phrase-level retrieval of the top-3 semantic candidates; (iii) Contextual prototyping using pseudo-texts to identify three additional candidates; and (iv) Discriminative selection from the resulting 7-option pool. We evaluated the framework on the original and Adams’s denoised CASI datasets using six open-source LLMs. Results S-ACAD demonstrates superior performance, with Gemma-2-9B-IT emerging as the best-performing model, achieving remarkable results on Adams’s denoised CASI dataset (accuracy: 0.9228, Macro-F1: 0.9279). Even on the challenging original dataset, it maintains high consistency (accuracy: 0.7585, Macro-F1: 0.7585). Ablation studies confirm that shifting the paradigm from generation to selection via dual-pathway candidate construction is the primary driver of performance, significantly outperforming direct generative baselines. Conclusion S-ACAD reframes clinical abbreviation disambiguation as a constrained selection task for generative LLMs, effectively unlocking the reasoning potential of small models through dual-pathway candidate construction. This offers a high-precision, economical, and privacy-preserving solution for standardizing electronic health records, thereby enhancing their downstream utility in clinical informatics.
Objective Selective classification improves reliability by allowing models to abstain on uncertain inputs, which is critical in safety-sensitive domains such as healthcare. However, commonly used evaluation metrics obscure fairness issues under class imbalance, leading to disproportionately high rejection of minority classes and misleadingly favorable assessments of selective strategies. Methods We introduce two imbalance-aware evaluation metrics, Class-Averaged AURC (CA-AURC) and Class-Averaged AUGRC (CA-AUGRC), which integrate risk against class-specific coverage and then average across classes, alongside the Area under the IAM-coverage curve (AUIC) as a complementary imbalance-aware performance metric. In addition, we propose a class-conditional coverage-matching selection strategy that enforces balanced rejection across diagnostic categories. The proposed framework is evaluated on a clinical complete blood count dataset comprising 3316 patient records, 11 laboratory features, and 9 diagnostic classes, and validated on three additional publicly available benchmark datasets covering different domains, class structures, and imbalance levels. Results While conventional metrics such as AURC and AUGRC favor classical selective strategies, the proposed class-averaged metrics reveal substantial disparities in rejection behavior under class imbalance. Using CA-AUGRC and AUIC, the class-conditional strategy consistently outperforms both the classical approach and LABEL, a theoretically principled set-valued baseline, across all four datasets and five uncertainty measures. Rank-based statistical comparisons confirm significant advantages of the class-conditional strategy on CA-AUGRC (p=0.009, r=0.638) and AUIC (p<0.001, r=0.957). Conclusion The results demonstrate that both evaluation and selection in selective classification must be class-aware to ensure fairness and clinical usefulness. Class-averaged metrics and class-conditional selection provide a more reliable basis for assessing selective classifiers in imbalanced medical data, with consistent generalizability across diverse datasets and uncertainty measures.
OBJECTIVE:To develop and evaluate a time-aware data mining pipeline that integrates accelerometry-based Physical Activity (PA), static dietary intake, and sociodemographic factors to predict Healthy Weight Status (HWS), and identify features with predictive importance for HWS. METHODS:We propose TimePAD, Time-Based Physical Activity and Dietary Intake data mining, a three-stage prediction pipeline that integrates time-based PA representations, static dietary intake, and sociodemographic variables to predict HWS categories. Stage 1 derives intensity-specific hourly PA sequences and learns time-of-day PA embeddings using a Transformer encoder trained with masked Self-Supervised Learning (SSL) reconstruction. Stage 2 combines the learnt PA embeddings with static dietary variables and participant characteristics derived from the Food Frequency Questionnaire (FFQ). Stage 3 trains prediction models and ranks predictors by their predictive importance for HWS. TimePAD was evaluated on a real-world dataset of 206 adolescents (10-16 years) with 7-day continuous wrist-worn accelerometer data, FFQ records, and sociodemographic attributes, using 10-fold cross-validation and comparison to an ARIMA-based feature engineering baseline. RESULTS:TimePAD achieves an accuracy of 82.90% and an F1-score of 67.92% for HWS prediction, outperforming the best baseline (ARIMA-based PA features + diet, 77.51% accuracy). Across experiments, time-of-day light PA (LPA) features were consistently ranked as important predictors. Other key features include time-of-day Moderate-to-Vigorous Activity (MVPA) and Sedentary (SED) levels, Tanner stage, total weekly sleep time, age (in months), weekly fruit intake, weekly vegetable and legume intake, and intake of sugar-sweetened beverages. DISCUSSION:TimePAD contributes a pipeline for learning from the time-of-day structure in wearable PA time series and integrating it with static dietary and contextual data for prediction and feature analysis. The findings suggest that LPA is likely to have a significant association with HWS, calling for further attention and investigation to better understand the role of LPA in overall health outcomes. This illustrates the potential benefits of TimePAD in modelling PA with dietary intake context in shaping healthy behaviours.
OBJECTIVE:The selective permeability of the blood-brain barrier (BBB) hinders the delivery of central nervous system (CNS) drugs to brain targets. It is therefore crucial to determine BBB permeability during CNS drug development. METHODS:We propose BiGranMolNet (Bi-Granularity Molecular Graph Network), a BBB permeability prediction model based on graph convolutional networks. We integrated datasets from multiple sources to construct comprehensive benchmarks containing 2148 regression samples and 16,904 classification samples. Molecular graphs were constructed at two granularities (atom-level and motif-level), and graph convolutional networks were used to learn representations for each granularity. A cross-attention mechanism fused the dual-granularity features, and a weighted loss function was introduced to mitigate class imbalance. RESULTS:In 5-fold cross-validation and structure-aware evaluations, BiGranMolNet achieved stable and competitive performance on both regression and classification tasks. CONCLUSIONS:By integrating molecular structural information at atomic and motif levels, BiGranMolNet provides a reliable computational tool for BBB permeability prediction. The proposed framework can support early-stage CNS drug screening by prioritizing compounds with favorable brain exposure potential and by offering structure-aware clues for molecular optimization.
OBJECTIVE:Most hallucination mitigation for large language models (LLMs) operates post-hoc, leaving safety-critical clinical deployment without real-time warning capability. We present a calibrated hidden-state probing pipeline that enables token-time hallucination detection under explicit false-positive-rate (FPR) constraints, making it suitable for streaming clinical deployment. METHODS:We formulate token-time hallucination detection as a constrained sequential decision problem separating non-circular supervision construction (using ROUGE-L, exact-choice, and exploratory NLI verifiers), FPR-constrained operating-point selection, and downstream intervention evaluation into a modular, reusable protocol. A lightweight two-layer hidden-state probe classifies each generated token into one of three risk states (Safe, AtRisk, or Hallucinating), where the intermediate AtRisk state captures pre-error instability. We validate across four medical QA benchmarks (Endoscopy, PubMedQA, MedHallu, MedQA-USMLE) spanning approximately 44,500 generation trajectories and three backbone LLMs (Qwen3-8B, Llama-3.1-8B-Instruct, BioMistral-7B). RESULTS:On the public biomedical QA benchmarks the probe yields viable detection across all backbone-dataset combinations: EDR@5 up to 0.390 (PubMedQA) and 0.352 (MedHallu) at the strict cap γ≤0.10, rising to 0.587 with 29%-59% hallucination reduction at the monitoring cap γ≤0.20. The trigger fires roughly 10-35 tokens before error onset, whereas a post-hoc check has zero lead time by construction. A single-institution Endoscopy corpus serves as a case study (EDR@5 0.613, Llama-3.1), treated as illustrative because its small confirmed-correct denominator (26-52 per backbone) yields wide confidence intervals. Probe latency is ∼10ms per answer versus ∼10s for multi-sample baselines. CONCLUSION:Hidden-state probing with FPR-constrained calibration provides a practical, low-latency solution for real-time hallucination monitoring in clinical LLM deployments. The modular pipeline separating supervision construction, probe training, and trigger calibration is directly reusable with alternative detectors or verifiers.
INTRODUCTION:Emergency care in the emergency department (ED) requires continuous, multifaceted decision-making based on evolving clinical information. This study aimed to develop and externally validate a comprehensive ED decision-support model based on a Temporal Fusion Transformer (TFT) with multi-modal time-series data to jointly predict testing, treatment, diagnosis, and disposition. METHODS:Adult ED visit data from one hospital were used for model development, and data from another hospital for external validation. The static inputs included patient characteristics, ED visit-related information, and triage notes, while the time-varying inputs included vital signs, laboratory results, management, and clinical notes. The primary outcomes were the areas under the receiver operating characteristic curves (AUCs) for predicting computed tomography, magnetic resonance imaging, echocardiography, gastrointestinal endoscopy, mechanical ventilation (MV), antibiotic administration, oxygen therapy, vasopressor use, transfusion, primary diagnosis, and ED disposition. A single TFT model was trained to predict all targets jointly, and separate Random Forest (RF) models were developed for comparison. RESULTS:The development and external validation datasets included 272,058 and 138,343 patients, respectively (median age, 61-62 years; females, 51.3%-51.7%). Across internal and external validation, the TFT demonstrated strong discrimination for tests (AUC 0.877-0.961), treatments (AUC 0.912-0.990), and disposition outcomes (AUC 0.811-0.905). The top-5 accuracy for diagnosis prediction was 73.7% and 65.4% in the internal and external validation, respectively. The TFT outperformed RF models for most targets and showed comparable performance in predicting MV, oxygen therapy, and vasopressor use. CONCLUSION:The TFT model achieved high accuracy across multiple ED decisions, demonstrating its potential as a comprehensive and temporally aware decision-support tool throughout the ED trajectory.
Paired biomedical assays increasingly measure different molecular or clinical views from the same sample. The statistical problem is simple to state but hard to solve: the views often have different dimensions, noise models, and dynamic ranges, yet downstream analysis requires a common representation. Single-cell CITE-seq is a useful example because transcript counts and surface-protein abundances are observed in the same cell. Existing paired-omics methods, including probabilistic models, matrix-factorization approaches, and neural fusion models, have addressed this setting with different assumptions. Fewer studies, however, have asked whether a symmetric contrastive objective can align the two views while retaining modality-specific signal through reconstruction. We present scCLIP, a contrastive masked-reconstruction framework for paired single-cell multi-omics integration. scCLIP trains RNA and ADT branches jointly with a bidirectional cross-modal contrastive loss and masked reconstruction losses. The branches use the same architectural template but retain separate input/output adapters, encoder-decoder parameters, and projection heads; the projected embeddings are compared in an ℓ2-normalized space with a learnable logit scale. We evaluate scCLIP on five paired RNA-protein datasets, where it is compared against TotalVI, BREMSC, jointDIMMSC, scMM, and SCOIT and achieves the highest ARI and FMI on every dataset (with TotalVI second-best overall and marginally higher NMI on three of the five datasets), and on the larger NeurIPS 2021 BMMC CITE-seq benchmark (90,261 cells, 134 proteins, 45 cell types) where scCLIP scales without architectural change and produces strong batch mixing on a 12-batch dataset. We additionally provide direct evidence of RNA-ADT alignment through retrieval and distance-based alignment metrics, and report standard batch-mixing scores on the NeurIPS embedding. The results support scCLIP as a reusable paired-view representation-learning template, with RNA-protein integration serving as the primary empirical testbed. The source code and an end-to-end tutorial for applying scCLIP to new CITE-seq data are publicly available at https://github.com/xubohao39-cyber/scCLIP.
The emergence of foundation models has marked a transformative shift in AI, enabling robust generalization across diverse downstream tasks through putative zero-shot learning. Large Language Models and Vision-Language Models have demonstrated strong capabilities in tasks such as image interpretation, report generation, and question answering by effectively learning from multimodal data - images paired with associated text - often with minimal supervision. In the healthcare domain, this ability to align visual and textual information reduces the reliance on extensive manual annotations, as models can leverage existing clinical reports and imaging data to learn meaningful representations. This integration holds promise for improving diagnostic support, treatment planning, and overall patient care, even in data-constrained settings. In this review, we provide a definitive taxonomy of the medical VLM landscape, tracing the evolution from early Contrastive Alignment and Generative MLLMs to the cutting-edge frontiers of Dense Pixel-Grounding, Sparse Mixture-of-Experts (MoE), and Reasoning-Incentivized (RL) architectures. We critically examine the "medical bottleneck"-identifying the persistent challenges of data scarcity, the "hallucination" risks in generative diagnostics, the computational strain of 3D volumetric processing, and the lack of standardized, clinically-grounded evaluation metrics.
OBJECTIVE:Single-cell foundation models such as scGPT have been promoted as representations of gene regulation, but their advertised regulatory signal has been judged largely from attention weights, which are correlational. We ask a deliberately bounded question: whether direct interventions on scGPT gene tokens-causal with respect to the model's own computation-recover model-internal transcription-factor (TF)-target dependencies that align with curated references, whether that signal is robust, and whether it transfers to real perturbation responses. We separate these into two explicit validation axes and report both, including where the model-internal signal does not correspond to biological causality. METHODS:On Tabula Sapiens kidney, lung, and immune subsets, plus an external Krasnow lung atlas and three CRISPR perturbation datasets (Adamson, Dixit, Shifrut), we ablate or swap TF token values inside scGPT and quantify changes in target-token readouts. Axis 1 (reference alignment): intervention scores are evaluated against curated TRRUST and DoRothEA references and against attention and coexpression baselines, with new robustness studies over cell count (120-500), positive/negative pair count, four negative-sampling designs, and four readout strategies (mean, L2, cosine, max). We trace component-level circuits via activation patching, and benchmark against twelve GRN inference methods including a neural-network baseline on a matched substrate, additionally giving the classical methods larger cell budgets to test the equal-information question. Axis 2 (perturbation transfer): scores are evaluated against CRISPR perturbation-derived edges under balanced labels and AUROC. RESULTS:On Axis 1, lung shows reproducible enrichment that is stable across cell count and survives all four negative-sampling designs (permutation p improving from 0.07 to 0.03 as cells grow from 120 to 500); kidney enrichment is significant at small samples but does not survive scaling (permutation p rising to 0.16 at 500 cells); immune is at baseline. Richer readouts (L2, cosine) recover signal the scalar mean compresses away, particularly in kidney. On the matched benchmark scGPT leads classical and neural baselines at the shared 120-cell budget, but the lead narrows as the classical methods are given more cells, confirming this is a signal-efficiency result, not general superiority. On Axis 2, balanced evaluation of all three CRISPR datasets, including the TF-screen Dixit data, yields AUROC ≈ 0.50: the model-internal signal does not predict real perturbation responses. CONCLUSION:scGPT encodes tissue-conditional, intervention-sensitive regulatory structure that is aligned with literature-curated TF-target edges (robustly in lung) but is representational rather than biologically causal: it does not transfer to perturbation outcomes. The pipeline is a practical mechanistic-audit toolkit for biological foundation models, and the gap between reference alignment and perturbation transfer is a concrete cautionary result for using such models in regulatory inference.
Accurate prediction of drug-target interactions (DTI) is essential for drug discovery. Despite the success of pre-trained language models (PLMs) in learning robust molecular and protein representations, a fundamental challenge remains in characterizing the fine-grained, localized biochemical interactions between drug substructures and protein binding sites. Such critical interaction patterns are often underrepresented in conventional global embedding approaches, thereby limiting both predictive accuracy and biological interpretability. To address this challenge, we propose CoAff-DTI, an end-to-end deep learning framework designed to enhance multi-scale interaction modeling for DTI prediction. The model introduces three key components. First, a token-level decomposition strategy is employed to transform global embeddings into pharmacophore- and residue-level representations, facilitating the capture of localized features. Second, an Affinity-Guided Cross-Attention (AGCA) module is designed to explicitly model fine-grained interactions between ligand substructures and protein residues. Third, an Affinity-Gating Fusion (AGF) module is proposed to enhance cross-modal feature integration by dynamically modeling element-wise interactions. Extensive experiments on multiple benchmark datasets demonstrate that CoAff-DTI consistently outperforms state-of-the-art methods. In addition, attention-based visualization results suggest improved interpretability, as the model's learned attention patterns align effectively with experimentally verified binding regions.
OBJECTIVE:This study introduces a set of metrics for evaluating temporal preservation in synthetic longitudinal patient data, defined as artificially generated data that mimic real patients' repeated measurements over time. METHODS:The proposed metrics assess how synthetic data reproduce key temporal characteristics, categorized into marginal, covariance, individual-level and measurement structures. RESULTS:Strong marginal-level resemblance may be observed even when the covariance structures and individual trajectories are substantially different. Temporal preservation is influenced by factors such as original data quality, measurement frequency, and preprocessing strategies, including binning, variable encoding and precision. Variables with sparse or highly irregular measurement times provide limited information for learning temporal dependencies, yielding reduced resemblance between the synthetic and original data. CONCLUSION:No single metric adequately captures temporal preservation; instead, a multidimensional evaluation across all characteristics provides a more comprehensive assessment of synthetic data quality. Overall, the proposed metrics elucidate how and why temporal structures are preserved or degraded, enabling more reliable evaluation and improvement of generative models and supporting the creation of temporally realistic synthetic longitudinal patient data.
OBJECTIVE:Cross-lingual biomedical named entity recognition (BioNER) remains challenging in low-resource settings due to scarce annotations, heterogeneous corpora, and severe label imbalance dominated by the O class. We aim to improve the robustness of BioNER fine-tuning under such conditions without modifying the backbone architecture. METHODS:We propose RouteB-v4, a validation-driven training controller that adaptively reweights token-level loss using validation-side entity-level feedback. At periodic validation checkpoints, the controller performs conservative, bounded small-step updates to entity-type loss weights and explicitly regulates the dominant O-label weight to reduce training instability. The method is designed as a drop-in module for standard token-classification pipelines. RESULTS:We evaluate RouteB-v4 on CRAFT-based English-Spanish transfer settings, PharmaCoNER, and the external Spanish clinical benchmark CANTEMIST. Within the current experimental scope of English-Spanish transfer and Spanish biomedical/clinical datasets, RouteB-v4 achieves more stable span-level F1 improvements over strong static, heuristic dynamic, and learning-based non-RL dynamic reweighting baselines, with clearer gains on minority entity types and external generalization. Using XLM-R, RouteB-v4 reaches 0.868 F1 on PharmaCoNER and improves external generalization on CANTEMIST under the same training budget. Under a strict split-isolation protocol, the main performance advantage is largely retained after separating controller feedback from checkpoint selection. CONCLUSION:Within the current experimental scope of English-Spanish transfer and Spanish biomedical/clinical datasets, validation-driven adaptive loss control provides an effective way to improve the robustness of BioNER fine-tuning under label imbalance and distribution shift without altering the backbone architecture.