An Analysis of Euclidean vs. Graph-Based Framing for Bilingual Lexicon Induction from Word Embedding Spaces

التفاصيل البيبلوغرافية
العنوان: An Analysis of Euclidean vs. Graph-Based Framing for Bilingual Lexicon Induction from Word Embedding Spaces
المؤلفون: Marchisio, Kelly, Park, Youngser, Saad-Eldin, Ali, Alyakin, Anton, Duh, Kevin, Priebe, Carey, Koehn, Philipp
سنة النشر: 2021
المجموعة: Computer Science
مصطلحات موضوعية: Computer Science - Computation and Language
الوصف: Much recent work in bilingual lexicon induction (BLI) views word embeddings as vectors in Euclidean space. As such, BLI is typically solved by finding a linear transformation that maps embeddings to a common space. Alternatively, word embeddings may be understood as nodes in a weighted graph. This framing allows us to examine a node's graph neighborhood without assuming a linear transform, and exploits new techniques from the graph matching optimization literature. These contrasting approaches have not been compared in BLI so far. In this work, we study the behavior of Euclidean versus graph-based approaches to BLI under differing data conditions and show that they complement each other when combined. We release our code at https://github.com/kellymarchisio/euc-v-graph-bli.
Comment: EMNLP Findings 2021 Camera-Ready
نوع الوثيقة: Working Paper
URL الوصول: http://arxiv.org/abs/2109.12640
رقم الأكسشن: edsarx.2109.12640
قاعدة البيانات: arXiv