An Information-Theoretic Approach and Dataset for Probing Gender Stereotypes in Multilingual Masked Language Models

www.lmu.de | UB | Blättern | Hilfe

Zur erweiterten Suche

English

Zur erweiterten Suche

Steinborn, Victor; Dufter, Philipp; Jabbar, Haris und Schütze, Hinrich (Juli 2022): An Information-Theoretic Approach and Dataset for Probing Gender Stereotypes in Multilingual Masked Language Models. NAACL 2022, Seattle, United States, July 2022. Carpuat, Marine; de Marneffe, Marie-Catherine und Meza Ruiz, Ivan Vladimir (Hrsg.): In: Findings of the Association for Computational Linguistics: NAACL 2022 - Findings : July 10-15, 2022 : NAACL 2022, Stroudsburg, PA: Association for Computational Linguistics (ACL). S. 921-932 [PDF, 2MB]

[thumbnail of 2022.findings-naacl.69.pdf]

Vorschau

Creative Commons: Namensnennung 4.0 (CC-BY)

DOI: 10.18653/v1/2022.findings-naacl.69

Externer Volltext: https://aclanthology.org/2022.findings-naacl.pdf

Abstract

Bias research in NLP is a rapidly growing and developing field. Similar to CrowS-Pairs (Nangia et al., 2020), we assess gender bias in masked-language models (MLMs) by studying pairs of sentences with gender swapped person references. Most bias research focuses on and often is specific to English.Using a novel methodology for creating sentence pairs that is applicable across languages, we create, based on CrowS-Pairs, a multilingual dataset for English, Finnish, German, Indonesian and Thai.Additionally, we propose SJSD, a new bias measure based on Jensen–Shannon divergence, which we argue retains more information from the model output probabilities than other previously proposed bias measures for MLMs.Using multilingual MLMs, we find that SJSD diagnoses the same systematic biased behavior for non-English that previous studies have found for monolingual English pre-trained MLMs. SJSD outperforms the CrowS-Pairs measure, which struggles to find such biases for smaller non-English datasets.

Dokumententyp:	Konferenzbeitrag (Paper)
EU Funded Grant Agreement Number:	740516
EU-Projekte:	Horizon 2020 > ERC Grants > ERC Advanced Grant > ERC Grant 740516: NonSequeToR - Non-sequence models for tokenization replacement
Fakultätsübergreifende Einrichtungen:	Centrum für Informations- und Sprachverarbeitung (CIS)
Themengebiete:	000 Informatik, Informationswissenschaft, allgemeine Werke > 000 Informatik, Wissen, Systeme 400 Sprache > 400 Sprache 400 Sprache > 410 Linguistik
URN:	urn:nbn:de:bvb:19-epub-107421-7
Ort:	Stroudsburg, PA
Bemerkung:	ISBN 978-1-955917-76-6
Sprache:	Englisch
Dokumenten ID:	107421
Datum der Veröffentlichung auf Open Access LMU:	20. Okt. 2023 06:04
Letzte Änderungen:	20. Okt. 2023 06:04

Dokument bearbeiten