KILM: Knowledge Injection into Encoder-Decoder Language Models

Yan Xu,Mahdi Namazifar,Devamanyu Hazarika,Aishwarya Padmakumar,Yang Liu,Dilek Hakkani-Tür

conf_acl（2023）

引用 0|浏览131

暂无评分

摘要

Large pre-trained language models (PLMs) have been shown to retain implicit knowledge within their parameters. To enhance this implicit knowledge, we propose Knowledge Injection into Language Models (KILM), a novel approach that injects entity-related knowledge into encoder-decoder PLMs, via a generative knowledge infilling objective through continued pre-training. This is done without architectural modifications to the PLMs or adding additional parameters. Experimental results over a suite of knowledge-intensive tasks spanning numerous datasets show that KILM enables models to retain more knowledge and hallucinate less, while preserving their original performance on general NLU and NLG tasks. KILM also demonstrates improved zero-shot performances on tasks such as entity disambiguation, outperforming state-of-the-art models having 30x more parameters.

查看译文

关键词

knowledge injection,models,encoder-decoder

AI 理解论文

溯源树

样例

生成溯源树，研究论文发展脉络

Chat Paper

正在生成论文摘要