Design, creation, and analysis of Czech corpora for structural metadata extraction from speech.

Saved in:
Bibliographic Details
Title: Design, creation, and analysis of Czech corpora for structural metadata extraction from speech.
Authors: Kolář, Jáchym1 jachym@kky.zcu.cz
Source: Language Resources & Evaluation. Dec2011, Vol. 45 Issue 4, p439-462. 24p.
Subjects: Metadatabases, Sentences (Grammar), Automatic speech recognition, Corpora, Spoken Czech, Statistics
Abstract: Structural metadata extraction (MDE) research aims to develop techniques for automatic conversion of raw speech recognition output to forms that are more useful to humans and downstream automatic processes. The MDE annotation includes inserting boundaries of sentence-like units to the flow of speech, labeling non-content words like filled pauses and discourse markers for optional removal, and identifying sections of disfluent speech. This paper describes design, creation, and analysis of data resources for structural MDE from spoken Czech. The annotation is based on the LDC's MDE annotation standard for English, with changes applied to accommodate specific phenomena of Czech. In addition to the necessary language-dependent modifications, we further proposed and applied several language-independent modifications slightly refining the original annotation scheme. We created two Czech MDE speech corpora-one in the domain of broadcast news and the other in the domain of broadcast conversations. Both corpora have already been published at LDC. The analysis section of this paper presents a variety of statistics about fillers, edit disfluencies, and sentence-like units. The two Czech corpora are not only compared with each other, but also with statistics relating to the available English MDE corpora. We also report the statistics indicating that edit disfluencies have a different part of speech (POS) distribution in comparison with the overall POS distribution. The findings from the corpus analysis should help guide strategies for developing automatic MDE systems. [ABSTRACT FROM AUTHOR]
Copyright of Language Resources & Evaluation is the property of Springer Nature and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Database: Engineering Source
FullText Links:
  – Type: pdflink
Text:
  Availability: 0
Header DbId: egs
DbLabel: Engineering Source
An: 67104764
AccessLevel: 6
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: Design, creation, and analysis of Czech corpora for structural metadata extraction from speech.
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Kolář%2C+Jáchym%22">Kolář, Jáchym</searchLink><relatesTo>1</relatesTo><i> jachym@kky.zcu.cz</i>
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="JN" term="%22Language+Resources+%26+Evaluation%22">Language Resources & Evaluation</searchLink>. Dec2011, Vol. 45 Issue 4, p439-462. 24p.
– Name: Subject
  Label: Subjects
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Metadatabases%22">Metadatabases</searchLink><br /><searchLink fieldCode="DE" term="%22Sentences+%28Grammar%29%22">Sentences (Grammar)</searchLink><br /><searchLink fieldCode="DE" term="%22Automatic+speech+recognition%22">Automatic speech recognition</searchLink><br /><searchLink fieldCode="DE" term="%22Corpora%22">Corpora</searchLink><br /><searchLink fieldCode="DE" term="%22Spoken+Czech%22">Spoken Czech</searchLink><br /><searchLink fieldCode="DE" term="%22Statistics%22">Statistics</searchLink>
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: Structural metadata extraction (MDE) research aims to develop techniques for automatic conversion of raw speech recognition output to forms that are more useful to humans and downstream automatic processes. The MDE annotation includes inserting boundaries of sentence-like units to the flow of speech, labeling non-content words like filled pauses and discourse markers for optional removal, and identifying sections of disfluent speech. This paper describes design, creation, and analysis of data resources for structural MDE from spoken Czech. The annotation is based on the LDC's MDE annotation standard for English, with changes applied to accommodate specific phenomena of Czech. In addition to the necessary language-dependent modifications, we further proposed and applied several language-independent modifications slightly refining the original annotation scheme. We created two Czech MDE speech corpora-one in the domain of broadcast news and the other in the domain of broadcast conversations. Both corpora have already been published at LDC. The analysis section of this paper presents a variety of statistics about fillers, edit disfluencies, and sentence-like units. The two Czech corpora are not only compared with each other, but also with statistics relating to the available English MDE corpora. We also report the statistics indicating that edit disfluencies have a different part of speech (POS) distribution in comparison with the overall POS distribution. The findings from the corpus analysis should help guide strategies for developing automatic MDE systems. [ABSTRACT FROM AUTHOR]
– Name: AbstractSuppliedCopyright
  Label:
  Group: Ab
  Data: <i>Copyright of Language Resources & Evaluation is the property of Springer Nature and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.)
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=egs&AN=67104764
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.1007/s10579-010-9126-8
    Languages:
      – Code: eng
        Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 24
        StartPage: 439
    Subjects:
      – SubjectFull: Metadatabases
        Type: general
      – SubjectFull: Sentences (Grammar)
        Type: general
      – SubjectFull: Automatic speech recognition
        Type: general
      – SubjectFull: Corpora
        Type: general
      – SubjectFull: Spoken Czech
        Type: general
      – SubjectFull: Statistics
        Type: general
    Titles:
      – TitleFull: Design, creation, and analysis of Czech corpora for structural metadata extraction from speech.
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Kolář, Jáchym
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 12
              Text: Dec2011
              Type: published
              Y: 2011
          Identifiers:
            – Type: issn-print
              Value: 1574020X
          Numbering:
            – Type: volume
              Value: 45
            – Type: issue
              Value: 4
          Titles:
            – TitleFull: Language Resources & Evaluation
              Type: main
ResultId 1