Grounded Sequence to Sequence Transduction

www.lmu.de | UB | Blättern | Hilfe

Zur erweiterten Suche

English

Zur erweiterten Suche

Specia, Lucia; Barrault, Loic; Caglayan, Ozan; Duarte, Amanda; Elliott, Desmond; Gella, Spandana; Holzenberger, Nils; Lala, Chiraag; Lee, Sun Jae; Libovicky, Jindrich; Madhyastha, Pranava; Metze, Florian; Mulligan, Karl; Ostapenko, Alissa; Palaskar, Shruti; Sanabria, Ramon; Wang, Josiah und Arora, Raman (2020): Grounded Sequence to Sequence Transduction. In: IEEE Journal of Selected Topics in Signal Processing, Bd. 14, Nr. 3: S. 577-591

Volltext auf 'Open Access LMU' nicht verfügbar.

DOI: 10.1109/JSTSP.2020.2998415

Abstract

Speech recognition and machine translation have made major progress over the past decades, providing practical systems to map one language sequence to another. Although multiple modalities such as sound and video are becoming increasingly available, the state-of-the-art systems are inherently unimodal, in the sense that they take a single modality - either speech or text - as input. Evidence from human learning suggests that additional modalities can provide disambiguating signals crucial for many language tasks. In this article, we describe the How2 dataset , a large, open-domain collection of videos with transcriptions and their translations. We then show how this single dataset can be used to develop systems for a variety of language tasks and present a number of models meant as starting points. Across tasks, we find that building multimodal architectures that perform better than their unimodal counterpart remains a challenge. This leaves plenty of room for the exploration of more advanced solutions that fully exploit the multimodal nature of the How2 dataset , and the general direction of multimodal learning with other datasets as well.

Dokumententyp:	Zeitschriftenartikel
Fakultät:	Sprach- und Literaturwissenschaften > Department 2
Themengebiete:	400 Sprache > 400 Sprache
ISSN:	1932-4553
Sprache:	Englisch
Dokumenten ID:	88687
Datum der Veröffentlichung auf Open Access LMU:	25. Jan. 2022 09:28
Letzte Änderungen:	25. Jan. 2022 09:28

Dokument bearbeiten