Abstract
We present a new method for sentiment lex- icon induction that is designed to be appli- cable to the entire range of typological di- versity of the world’s languages. We eval- uate our method on Parallel Bible Corpus+ (PBC+), a parallel corpus of 1593 languages. The key idea is to use Byte Pair Encodings (BPEs) as basic units for multilingual em- beddings. Through zero-shot transfer from English sentiment, we learn a seed lexicon for each language in the domain of PBC+. Through domain adaptation, we then gener- alize the domain-specific lexicon to a general one. We show – across typologically diverse languages in PBC+ – good quality of seed and general-domain sentiment lexicons by intrin- sic and extrinsic and by automatic and human evaluation. We make freely available our code, seed sentiment lexicons for all 1593 languages and induced general-domain sentiment lexi- cons for 200 languages
| Dokumententyp: | Konferenz |
|---|---|
| EU Funded Grant Agreement Number: | 740516 |
| EU-Projekte: | Horizon 2020 > ERC Grants > ERC Advanced Grant > ERC Grant 740516: NonSequeToR - Non-sequence models for tokenization replacement |
| Fakultätsübergreifende Einrichtungen: | Centrum für Informations- und Sprachverarbeitung (CIS) |
| Themengebiete: | 000 Informatik, Informationswissenschaft, allgemeine Werke > 000 Informatik, Wissen, Systeme
400 Sprache > 410 Linguistik |
| URN: | urn:nbn:de:bvb:19-epub-72191-0 |
| ISBN: | 978-1-950737-48-2 |
| Ort: | Stroudsburg, USA |
| Sprache: | Englisch |
| Dokumenten ID: | 72191 |
| Datum der Veröffentlichung auf Open Access LMU: | 20. Mai 2020 09:35 |
| Letzte Änderungen: | 04. Nov. 2020 13:53 |

