Logo Logo
Hilfe
Hilfe
Switch Language to English

Schiel, Florian (2010): BAStat : New Statistical Resources at the Bavarian Archive for Speech Signals. 7th International Conference on Language Resources and Evaluation (LREC), Valetta, Malta, 19. - 21. Mai 2010. Calzolari, Nicoletta (Hrsg.): In: Proceedings of the Seventh International Conference on Language Resources and Evaluation (LREC'10), Paris: S. 1069-1076 [PDF, 177kB]

[thumbnail of schiel_13681.pdf]
Vorschau
Download (177kB)

Abstract

A new type of language resource ’BAStat’ has been released by the Bavarian Archive for Speech Signals. In contrast to primary resources like speech and text corpora BAStat comprises statistical estimates based on a number of primary resources: first and second order occurrence probability of phones, syllables and words, duration statistics, probabilities of pronunciation variants of words and probabilities of context information. Unlike other statistical speech resources BAStat is based solely on recordings of conversational German and therefore models spoken language. It consists of 7-bit ASCII tables and matrices to maximize inter-operability between different platforms and can be downloaded from the BAS web-site. This paper gives a detailed description about the empirical basis, the contained data types, some interesting interpretations and a brief comparison to the text-based statistical resource CELEX.

Dokument bearbeiten Dokument bearbeiten