DA-ATE: Data Augmentation for Automatic Term Extraction. | AMiner
DA-ATE: Data Augmentation for Automatic Term Extraction.
Shubhanker Banerjee,Bharathi Raja Chakravarthi,John P. Mccrae
LINKING MEANING SEMANTIC TECHNOLOGIES SHAPING THE FUTURE OF AI(2025)
Univ Galway
被引用0|浏览0
摘要
Automatic term extraction (ATE) identifies domain-specific concepts from specialized corpora, but suffers from limited annotated training data across diverse domains. We propose three novel LLM-based data augmentation schemes for ATE: context-level augmentation (generating diverse sentences using existing terms), term-level augmentation (replacing terms with domain-relevant alternatives), and combined augmentation (creating novel sentences with new terminology). Our approach leverages both ChatGPT-4o and Wikipedia-derived domain lexicons to generate synthetic training data. Experiments across four domains in the ACTER dataset demonstrate consistent improvements over state-of-the-art XLM-RoBERTa baselines, with gains of up to 28% F1-score in few-shot scenarios (5-10 samples) and 1-2% improvements in larger datasets (100-500 samples). Context-level and term-level augmentation consistently outperform combined augmentation, while LLM-based methods surpass Wikipedia-based augmentation. Our findings establish the effectiveness of targeted data augmentation for ATE across varying data availability scenarios, with performance gains extending beyond few-shot settings to practical dataset sizes.
更多
查看译文
关键词
automatic term extraction,large language models,domain-specific concepts,LLM-based data augmentation