A Multilingual BPE Embedding Space for Universal Sentiment Lexicon Induction

www.lmu.de | UB | Blättern | Hilfe

Zur erweiterten Suche

English

Zur erweiterten Suche

Zhao, Mengjie und Schütze, Hinrich (28. Juli 2019): A Multilingual BPE Embedding Space for Universal Sentiment Lexicon Induction. , July 28 - August 2, 2019, Florence, Italy [PDF, 544kB]

Vorschau

Creative Commons: Namensnennung 4.0 (CC-BY)

DOI: 10.5282/ubm/epub.72191

Abstract

We present a new method for sentiment lex- icon induction that is designed to be appli- cable to the entire range of typological di- versity of the world’s languages. We eval- uate our method on Parallel Bible Corpus+ (PBC+), a parallel corpus of 1593 languages. The key idea is to use Byte Pair Encodings (BPEs) as basic units for multilingual em- beddings. Through zero-shot transfer from English sentiment, we learn a seed lexicon for each language in the domain of PBC+. Through domain adaptation, we then gener- alize the domain-specific lexicon to a general one. We show – across typologically diverse languages in PBC+ – good quality of seed and general-domain sentiment lexicons by intrin- sic and extrinsic and by automatic and human evaluation. We make freely available our code, seed sentiment lexicons for all 1593 languages and induced general-domain sentiment lexi- cons for 200 languages

Dokumententyp:	Konferenz
EU Funded Grant Agreement Number:	740516
EU-Projekte:	Horizon 2020 > ERC Grants > ERC Advanced Grant > ERC Grant 740516: NonSequeToR - Non-sequence models for tokenization replacement
Fakultätsübergreifende Einrichtungen:	Centrum für Informations- und Sprachverarbeitung (CIS)
Themengebiete:	000 Informatik, Informationswissenschaft, allgemeine Werke > 000 Informatik, Wissen, Systeme 400 Sprache > 410 Linguistik
URN:	urn:nbn:de:bvb:19-epub-72191-0
ISBN:	978-1-950737-48-2
Ort:	Stroudsburg, USA
Sprache:	Englisch
Dokumenten ID:	72191
Datum der Veröffentlichung auf Open Access LMU:	20. Mai 2020 09:35
Letzte Änderungen:	04. Nov. 2020 13:53

Dokument bearbeiten