Identifying Necessary Elements for BERT’s Multilinguality

www.lmu.de | UB | Blättern | Hilfe

Zur erweiterten Suche

English

Zur erweiterten Suche

Dufter, Philipp und Schütze, Hinrich (1. Mai 2020): Identifying Necessary Elements for BERT’s Multilinguality. [PDF, 631kB]

Vorschau

DOI: 10.5282/ubm/epub.72199

Externer Volltext: https://arxiv.org/abs/2005.00396

Abstract

It has been shown that multilingual BERT (mBERT) yields high quality multilingual rep- resentations and enables effective zero-shot transfer. This is suprising given that mBERT does not use any kind of crosslingual sig- nal during training. While recent literature has studied this effect, the exact reason for mBERT’s multilinguality is still unknown. We aim to identify architectural properties of BERT as well as linguistic properties of lan- guages that are necessary for BERT to become multilingual. To allow for fast experimenta- tion we propose an efficient setup with small BERT models and synthetic as well as natu- ral data. Overall, we identify six elements that are potentially necessary for BERT to be mul- tilingual. Architectural factors that contribute to multilinguality are underparameterization, shared special tokens (e.g., “[CLS]”), shared position embeddings and replacing masked to- kens with random tokens. Factors related to training data that are beneficial for multilin- guality are similar word order and comparabil- ity of corpora.

Dokumententyp:	Paper
EU Funded Grant Agreement Number:	740516
EU-Projekte:	Horizon 2020 > ERC Grants > ERC Advanced Grant > ERC Grant 740516: NonSequeToR - Non-sequence models for tokenization replacement
Fakultätsübergreifende Einrichtungen:	Centrum für Informations- und Sprachverarbeitung (CIS)
Themengebiete:	000 Informatik, Informationswissenschaft, allgemeine Werke > 000 Informatik, Wissen, Systeme 400 Sprache > 410 Linguistik
URN:	urn:nbn:de:bvb:19-epub-72199-8
Sprache:	Englisch
Dokumenten ID:	72199
Datum der Veröffentlichung auf Open Access LMU:	20. Mai 2020 07:46
Letzte Änderungen:	04. Nov. 2020 13:53

Dokument bearbeiten