MOTIVATION:Accurate TNM staging from lung cancer radiology reports is crucial for treatment planning and prognosis assessment. Manual staging processes are time-consuming and subject to inter-observer variability. Large language models (LLMs) offer opportunities to automate TNM staging with enhanced interpretability and clinical reasoning. RESULTS:We developed two complementary systems for automated TNM staging from English radiology reports. System I employs GPT-4o with reasoning-based few-shot learning and multi-step voting. System II integrates multiple LLMs (GPT-4o and Gemini-2) using DSPy framework with MIPROv2 optimization. In NTCIR-18 RadNLP 2024 English main task, our approaches achieved first (joint accuracy: 0.6543) and second place (joint accuracy: 0.6296), demonstrating superior performance in T, N, and M classification with accuracies of 0.7037/0.9136/0.8889 and 0.7284/0.9383/0.8395, respectively. AVAILABILITY AND IMPLEMENTATION:Source code freely available at https://github.com/nlptmu/multi-expert-tnm-staging under MIT license. An archival snapshot of the version used in this study is deposited on Zenodo at https://doi.org/10.5281/zenodo.20338561. Implemented in Python 3.12+ with PyTorch 2.6 and DSPY 3.0, supporting Linux.
Film inhomogeneity of Si and CIGS which plays an important role for solar cell scaling are studied. The inhomogeneity is due to the process variations. The Si and CIGS modules with different key material parameter like the lifetime, doping, and band gap are simulated. The simulation results show that CIGS have larger performance degradation from the variation comparing with the Si case. And this could explain why CIGS solar cell has larger efficiency gap between the module and the small cell.
Natural language processing (NLP) has become an essential technique in various fields, offering a wide range of possibilities for analyzing data and developing diverse NLP tasks. In the biomedical domain, understanding the complex relationships between compounds and proteins is critical, especially in the context of signal transduction and biochemical pathways. Among these relationships, protein-protein interactions (PPIs) are of particular interest, given their potential to trigger a variety of biological reactions. To improve the ability to predict PPI events, we propose the protein event detection dataset (PEDD), which comprises 6823 abstracts, 39 488 sentences and 182 937 gene pairs. Our PEDD dataset has been utilized in the AI CUP Biomedical Paper Analysis competition, where systems are challenged to predict 12 different relation types. In this paper, we review the state-of-the-art relation extraction research and provide an overview of the PEDD's compilation process. Furthermore, we present the results of the PPI extraction competition and evaluate several language models' performances on the PEDD. This paper's outcomes will provide a valuable roadmap for future studies on protein event detection in NLP. By addressing this critical challenge, we hope to enable breakthroughs in drug discovery and enhance our understanding of the molecular mechanisms underlying various diseases.
We propose a semantic template-based distributed representation for the convolutional neural network called Semantic Template-based Convolutional Neural Network (STCNN) for text categorization that imitates the perceptual behavior of human comprehension. STCNN is a highly automatic approach that learns semantic templates that characterize a domain from raw text and recognizes categories of documents using a semantic-infused convolutional neural network that allows a template to be partially matched through a statistical scoring system. Our experiment results show that STCNN effectively classifies documents in about 140,000 Chinese news articles into predefined categories by capturing the most prominent and expressive patterns and achieves the best performance among all compared methods for Chinese topic classification. Finally, the same knowledge can be directly used to perform a semantic analysis task.
AbstractIn this research, we explored various state-of-the-art biomedical-specific pre-trained Bidirectional Encoder Representations from Transformers (BERT) models for the National Library of Medicine - Chemistry (NLM CHEM) and LitCovid tracks in the BioCreative VII Challenge, and propose a BERT-based ensemble learning approach to integrate the advantages of various models to improve the system’s performance. The experimental results of the NLM-CHEM track demonstrate that our method can achieve remarkable performance, with F1-scores of 85% and 91.8% in strict and approximate evaluations, respectively. Moreover, the proposed Medical Subject Headings identifier (MeSH ID) normalization algorithm is effective in entity normalization, which achieved a F1-score of about 80% in both strict and approximate evaluations. For the LitCovid track, the proposed method is also effective in detecting topics in the Coronavirus disease 2019 (COVID-19) literature, which outperformed the compared methods and achieve state-of-the-art performance in the LitCovid corpus.Database URL: https://www.ncbi.nlm.nih.gov/research/coronavirus/.
Code-mixing is a phenomenon where at least two languages are combined in a hybrid manner in the context of a single conversation. The use of mixed language is widespread in multilingual and multicultural countries and poses significant challenges for the development of automated language processing tools. In Taiwan’s electronic health record (EHR) systems, unstructured EHR texts are usually represented in a mixture of English and Chinese which increases the difficulty for de-identification and synthetization of protected health information (PHI). We explored this problem by applying several state-of-the-art pre-trained mono- and multilingual language models and propose to exploit the principle-based approach (PBA) for the tasks of PHI recognition and resynthesis on a code-mixed EHR corpus annotated with 6 main categories and 25 subcategories of PHIs. A hierarchical principle slot schema is defined in the PBA to encode knowledge of code-mixed PHIs and utilize slots to learn from the training set to assemble principles for recognizing PHI mentions and synthesizing surrogates simultaneously. In addition, a semantic disambiguation process is implemented to disambiguate ambiguous PHI categories in the de-identification process and to dynamically extend the knowledge encoded in PBA during the knowledge augmentation process. The experiment results demonstrate that the proposed method can achieve the best micro- and macro-F-scores in comparison to the other mono- and multilingual language models fine-tuned on our code-mixed corpus.
Conventional rule‐based approaches use exact template matching to capture linguistic information and necessarily need to enumerate all variations. We propose a novel flexible template generation and matching scheme called the principle‐based approach (PBA) based on sequence alignment, and employ it for reference metadata extraction (RME) to demonstrate its effectiveness. The main contributions of this research are threefold. First, we propose an automatic template generation that can capture prominent patterns using the dominating set algorithm. Second, we devise an alignment‐based template‐matching technique that uses a logistic regression model, which makes it more general and flexible than pure rule‐based approaches. Last, we apply PBA to RME on extensive cross‐domain corpora and demonstrate its robustness and generality. Experiments reveal that the same set of templates produced by the PBA framework not only deliver consistent performance on various unseen domains, but also surpass hand‐crafted knowledge (templates). We use four independent journal style test sets and one conference style test set in the experiments. When compared to renowned machine learning methods, such as conditional random fields (CRF), as well as recent deep learning methods (i.e., bi‐directional long short‐term memory with a CRF layer, Bi‐LSTM‐CRF), PBA has the best performance for all datasets.
Background Coronavirus disease 19 (COVID-19) first appeared in the city of Wuhan, in the Hubei province of China. Since its emergence, the COVID-19-causing virus, SARS-CoV-2, has been rapidly transmitted around the globe, overwhelming the medical care systems in many countries and leading to more than 3.3 million deaths. Identification of immunological epitopes on the virus would be highly useful for the development of diagnostic tools and vaccines that will be critical to limiting further spread of COVID-19. Methods To find disease-specific B-cell epitopes that correspond to or mimic natural epitopes, we used phage display technology to determine the targets of specific antibodies present in the sera of immune-responsive COVID-19 patients. Enzyme-linked immunosorbent assays were further applied to assess competitive antibody binding and serological detection. VaxiJen, BepiPred-2.0 and DiscoTope 2.0 were utilized for B-cell epitope prediction. PyMOL was used for protein structural analysis. Results 36 enriched peptides were identified by biopanning with antibodies from two COVID-19 patients; the peptides 4 motifs with consensus residues corresponding to two potential B-cell epitopes on SARS-CoV-2 viral proteins. The putative epitopes and hit peptides were then synthesized for validation by competitive antibody binding and serological detection. Conclusions The identified B-cell epitopes on SARS-CoV-2 may aid investigations into COVID-19 pathogenesis and facilitate the development of epitope-based serological diagnostics and vaccines.
Pre-trained language models may reduce the amount of training data required. Among the models, PEGASUS, a recently proposed self-supervised approach, is trained to generate the pseudo-summary given the partially masked document. PEGASUS uses gap sentence generation for summarization. The most important sentences are masked, and then PEGASUS predicts the masked sentences as the output summary. In this study, however, we apply PEGASUS in a novel downstream task. We reformulate the task to generate the masked question part in a primary math word problem. In the past research, PEGASUS has shown good potentials on the few-shot datasets, so we try a smaller set of primary math text problems as well. The fine-tuning dataset sizes used in this study are 1000, 500, 50, 10 samples. Their performance are measured by a non-weighted average of the ROUGE-1, ROUGE-2, and ROUGE-L scores. The results show the outstanding performance of PEGASUS applied in our novel downstream task.
Mass spectrometry-based proteomics using isobaric labeling for multiplex quantitation has become a popular approach for proteomic studies. We present Multi-Q 2, an isobaric-labeling quantitation tool which can yield the largest quantitation coverage and improved quantitation accuracy compared to three state-of-the-art methods. Multi-Q 2 supports identification results from several popular proteomic data analysis platforms for quantitation, offering up to 12% improvement in quantitation coverage for accepting identification results from multiple search engines when compared with MaxQuant and PatternLab. It is equipped with various quantitation algorithms, including a ratio compression correction algorithm, and results in up to 336 algorithmic combinations. Systematic evaluation shows different algorithmic combinations have different strengths and are suitable for different situations. We also demonstrate that the flexibility of Multi-Q 2 in customizing algorithmic combination can lead to improved quantitation accuracy over existing tools. Moreover, the use of complementary algorithmic combinations can be an effective strategy to enhance sensitivity when searching for biomarkers from differentially expressed proteins in proteomic experiments. Multi-Q 2 provides interactive graphical interfaces to process quantitation and to display ratios at protein, peptide, and spectrum levels. It also supports a heatmap module, enabling users to cluster proteins based on their abundance ratios and to visualize the clustering results. Multi-Q 2 executable files, sample data sets, and user manual are freely available at http://ms.iis.sinica.edu.tw/COmics/Software_Multi-Q2.html .
Interleukin (IL)-10 is a homodimer cytokine that plays a crucial role in suppressing inflammatory responses and regulating the growth or differentiation of various immune cells. However, the molecular mechanism of IL-10 regulation is only partially understood because its regulation is environment or cell type-specific. In this study, we developed a computational approach, ILeukin10Pred (interleukin-10 prediction), by employing amino acid sequence-based features to predict and identify potential immunosuppressive IL-10-inducing peptides. The dataset comprises 394 experimentally validated IL-10-inducing and 848 non-inducing peptides. Furthermore, we split the dataset into a training set (80%) and a test set (20%). To train and validate the model, we applied a stratified five-fold cross-validation method. The final model was later evaluated using the holdout set. An extra tree classifier (ETC)-based model achieved an accuracy of 87.5% and Matthew’s correlation coefficient (MCC) of 0.755 on the hybrid feature types. It outperformed an existing state-of-the-art method based on dipeptide compositions that achieved an accuracy of 81.24% and an MCC value of 0.59. Our experimental results showed that the combination of various features achieved better predictive performance..
In this paper, we propose a principle-based approach, combining a corpus-based approach and a knowledge-based approach. There are two parts of a principle: (1) Structure: such as verb frames, address patterns; (2) Statistics: mostly related to collocation and frequency. The method is based on a pipeline of four main stages, allowing to contribute to automatically building ontologies from text corpora, aimed to extract relations between words from the unstructured text.
BACKGROUND:Personal genomics and comparative genomics are becoming more important in clinical practice and genome research. Both fields require sequence alignment to discover sequence conservation and variation. Though many methods have been developed, some are designed for small genome comparison while some are not efficient for large genome comparison. Moreover, most existing genome comparison tools have not been evaluated the correctness of sequence alignments systematically. A wrong sequence alignment would produce false sequence variants.RESULTS:In this study, we present GSAlign that handles large genome sequence alignment efficiently and identifies sequence variants from the alignment result. GSAlign is an efficient sequence alignment tool for intra-species genomes. It identifies sequence variations from the sequence alignments. We estimate performance by measuring the correctness of predicted sequence variations. The experiment results demonstrated that GSAlign is not only faster than most existing state-of-the-art methods, but also identifies sequence variants with high accuracy.CONCLUSIONS:As more genome sequences become available, the demand for genome comparison is increasing. Therefore an efficient and robust algorithm is most desirable. We believe GSAlign can be a useful tool. It exhibits the abilities of ultra-fast alignment as well as high accuracy and sensitivity for detecting sequence variations.
Natural language processing (NLP) is widely applied in biological domains to retrieve information from publications. Systems to address numerous applications exist, such as biomedical named entity recognition (BNER), named entity normalization (NEN) and protein-protein interaction extraction (PPIE). High-quality datasets can assist the development of robust and reliable systems; however, due to the endless applications and evolving techniques, the annotations of benchmark datasets may become outdated and inappropriate. In this study, we first review commonlyused BNER datasets and their potential annotation problems such as inconsistency and low portability. Then, we introduce a revised version of the JNLPBA dataset that solves potential problems in the original and use state-of-the-art named entity recognition systems to evaluate its portability to different kinds of biomedical literature, including protein-protein interaction and biology events. Lastly, we introduce an ensembled biomedical entity dataset (EBED) by extending the revised JNLPBA dataset with PubMed Central full-text paragraphs, figure captions and patent abstracts. This EBED is a multi-task dataset that covers annotations including gene, disease and chemical entities. In total, it contains 85000 entity mentions, 25000 entity mentions with database identifiers and 5000 attribute tags. To demonstrate the usage of the EBED, we review the BNER track from the AI CUP Biomedical Paper Analysis challenge. Availability: The revised JNLPBA dataset is available at https://iasl-btm.iis.sinica.edu.tw/BNER/Content/Re vised_JNLPBA.zip. The EBED dataset is available at https://iasl-btm.iis.sinica.edu.tw/BNER/Content/AICUP _EBED_dataset.rar. Contact: Email: thtsai@g.ncu.edu.tw, Tel. 886-3-4227151 ext. 35203, Fax: 886-3-422-2681 Email: hsu@iis.sinica.edu.tw, Tel. 886-2-2788-3799 ext. 2211, Fax: 886-2-2782-4814 Supplementary information: Supplementary data are available at Briefings in Bioinformatics online.
MOTIVATION:Natural Language Processing techniques are constantly being advanced to accommodate the influx of data as well as to provide exhaustive and structured knowledge dissemination. Within the biomedical domain, relation detection between bio-entities known as the Bio-Entity Relation Extraction (BRE) task has a critical function in knowledge structuring. Although recent advances in deep learning-based biomedical domain embedding have improved BRE predictive analytics, these works are often task selective or use external knowledge-based pre-/post-processing. In addition, deep learning-based models do not account for local syntactic contexts, which have improved data representation in many kernel classifier-based models. In this study, we propose a universal BRE model, i.e. LBERT, which is a Lexically aware Transformer-based Bidirectional Encoder Representation model, and which explores both local and global contexts representations for sentence-level classification tasks.RESULTS:This article presents one of the most exhaustive BRE studies ever conducted over five different bio-entity relation types. Our model outperforms state-of-the-art deep learning models in protein-protein interaction (PPI), drug-drug interaction and protein-bio-entity relation classification tasks by 0.02%, 11.2% and 41.4%, respectively. LBERT representations show a statistically significant improvement over BioBERT in detecting true bio-entity relation for large corpora like PPI. Our ablation studies clearly indicate the contribution of the lexical features and distance-adjusted attention in improving prediction performance by learning additional local semantic context along with bi-directionally learned global context.AVAILABILITY AND IMPLEMENTATION:Github. https://github.com/warikoone/LBERT.SUPPLEMENTARY INFORMATION:Supplementary data are available at Bioinformatics online.
Advancement of deep learning has improved performances on a wide variety of tasks. However, language reasoning and understanding remain difficult tasks in Natural Language Processing (NLP). In this work, we consider this problem and propose a novel Knowledge-Infused Document Embedding Representation (KIDER) for text categorization. We use knowledge patterns to generate high quality document representation. These patterns preserve categorical-distinctive semantic information, provide interpretability, and achieve superior performances at the same time. Experiments show that the KIDER model outperforms state-of-the-art methods on two important NLP tasks, i.e., emotion analysis and news topic detection, by 7% and 20%. In addition, we also demonstrate the potential of highlighting important information for each category and news using these patterns. These results show the value of knowledge-infused patterns in terms of interpretability and performance enhancement.
This article presents a general experimental protocol for programmable one-pot oligosaccharide synthesis and demonstrates how to use Auto-CHO software for generating potential synthetic solutions. The programmable one-pot oligosaccharide synthesis approach is designed to empower fast oligosaccharide synthesis of large amounts using thioglycoside building blocks (BBLs) with the appropriate sequential order of relative reactivity values (RRVs). Auto-CHO is a cross-platform software with a graphical user interface that provides possible synthetic solutions for programmable one-pot oligosaccharide synthesis by searching a BBL library (containing about 150 validated and >50,000 virtual BBLs) with accurately predicted RRVs by support vector regression. The algorithm for hierarchical one-pot synthesis has been implemented in Auto-CHO and uses fragments generated by one-pot reactions as new BBLs. In addition, Auto-CHO allows users to give feedback for virtual BBLs to keep valuable ones for further use. One-pot synthesis of stage-specific embryonic antigen 4 (SSEA-4), which is a pluripotent human embryonic stem cell marker, is demonstrated in this work.
ABSTRACTCharacterizing the taxonomic diversity of a microbial community is very important to understand the roles of microorganisms. Next generation sequencing (NGS) provides great potential for investigation of a microbial community and leads to Metagenomic studies. NGS generates DNA fragment sequences directly from microorganism samples, and it requires analysis tools to identify microbial species (or taxonomic composition) and estimate their relative abundance in the studied community. However, only a few tools could achieve strain-level identification and most tools estimate the microbial abundances simply according to the read counts. An evaluation study on metagenomic analysis tools concludes that the predicted abundance differed significantly from the true abundance. In this study, we present StrainPro, a novel metagenomic analysis tool which is highly accurate both at characterizing microorganisms at strain-level and estimating their relative abundances. A unique feature of StrainPro is it identifies representative sequence segments from reference genomes. We generate three simulated datasets using known strain sequences and another three simulated datasets using unknown strain sequences. We compare the performance of StrainPro with seven existing tools. The results show that StrainPro not only identifies metagenomes with high precision and recall, but it is also highly robust even when the metagenomes are not included in the reference database. Moreover, StrainPro estimates the relative abundance with high accuracy. We demonstrate that there is a strong positive linear relationship between observed and predicted abundances.
OBJECTIVEIn this era of digitized health records, there has been a marked interest in using de-identified patient records for conducting various health related surveys. To assist in this research effort, we developed a novel clinical data representation model entitled medical knowledge-infused convolutional neural network (MKCNN), which is used for learning the clinical trial criteria eligibility status of patients to participate in cohort studies.MATERIALS AND METHODSIn this study, we propose a clinical text representation infused with medical knowledge (MK). First, we isolate the noise from the relevant data using a medically relevant description extractor; then we utilize log-likelihood ratio based weights from selected sentences to highlight "met" and "not-met" knowledge-infused representations in bichannel setting for each instance. The combined medical knowledge-infused representation (MK) from these modules helps identify significant clinical criteria semantics, which in turn renders effective learning when used with a convolutional neural network architecture.RESULTSMKCNN outperforms other Medical Knowledge (MK) relevant learning architectures by approximately 3%; notably SVM and XGBoost implementations developed in this study. MKCNN scored 86.1% on F1metric, a gain of 6% above the average performance assessed from the submissions for n2c2 task. Although pattern/rule-based methods show a higher average performance for the n2c2 clinical data set, MKCNN significantly improves performance of machine learning implementations for clinical datasets.CONCLUSIONMKCNN scored 86.1% on the F1 score metric. In contrast to many of the rule-based systems introduced during the n2c2 challenge workshop, our system presents a model that heavily draws on machine-based learning. In addition, the MK representations add more value to clinical comprehension and interpretation of natural texts.
Ting-Yi Sung合作论文数Institute of Information Science
Academia Sinica, Taiwan64
Min-Yuh Day合作论文数Intelligent Agent Systems Lab, Institute of Information Science, Academia Sinica25
Hsu-Chun Yen合作论文数Department of Electrical Engineering;National Taiwan University9