Chrome Extension
WeChat Mini Program
Use on ChatGLM

The Impact of Word Splitting on the Semantic Content of Contextualized Word Representations

arXiv (Cornell University)(2024)

Cited 0|Views12
No score
Abstract
When deriving contextualized word representations from language models, adecision needs to be made on how to obtain one for out-of-vocabulary (OOV)words that are segmented into subwords. What is the best way to represent thesewords with a single vector, and are these representations of worse quality thanthose of in-vocabulary words? We carry out an intrinsic evaluation ofembeddings from different models on semantic similarity tasks involving OOVwords. Our analysis reveals, among other interesting findings, that the qualityof representations of words that are split is often, but not always, worse thanthat of the embeddings of known words. Their similarity values, however, mustbe interpreted with caution.
More
Translated text
Key words
Language Modeling,Part-of-Speech Tagging,Syntax-based Translation Models
AI Read Science
Must-Reading Tree
Example
Generate MRT to find the research sequence of this paper
Chat Paper
Summary is being generated by the instructions you defined