Abstract
A new type of language resource ’BAStat’ has been released by the Bavarian Archive for Speech Signals. In contrast to primary resources like speech and text corpora BAStat comprises statistical estimates based on a number of primary resources: first and second order occurrence probability of phones, syllables and words, duration statistics, probabilities of pronunciation variants of words and probabilities of context information. Unlike other statistical speech resources BAStat is based solely on recordings of conversational German and therefore models spoken language. It consists of 7-bit ASCII tables and matrices to maximize inter-operability between different platforms and can be downloaded from the BAS web-site. This paper gives a detailed description about the empirical basis, the contained data types, some interesting interpretations and a brief comparison to the text-based statistical resource CELEX.
Item Type: | Conference or Workshop Item (Paper) |
---|---|
Faculties: | Languages and Literatures > Department 2 > Speech Science |
Subjects: | 400 Language > 400 Language |
URN: | urn:nbn:de:bvb:19-epub-13681-1 |
Place of Publication: | Paris |
Language: | English |
Item ID: | 13681 |
Date Deposited: | 19. Jul 2012, 09:25 |
Last Modified: | 04. Nov 2020, 12:54 |