On Acoustic Diversification Front-End for Spoken Language Identification.

Saved in:
Bibliographic Details
Title: On Acoustic Diversification Front-End for Spoken Language Identification.
Authors: Khe Chai Sim1,2 kcsim@i2r.a-star.edu.sg, Haizhou Li2,3 hli@i2r.a-star.edu.sg
Source: IEEE Transactions on Audio, Speech & Language Processing. Jul2008, Vol. 16 Issue 5, p1029-1037. 9p. 2 Diagrams, 6 Charts.
Subjects: Acoustic models, Acoustical engineering, Telephone systems, Telecommunication systems, Phonotactics, Language & languages
Abstract: The parallel phone recognition followed by language model (PPRLM) architecture represents one of the state-of-the-art spoken language identification systems. A PPRLM system comprises multiple parallel subsystems, where each subsystem employs a phone recognizer with a different phone set for a particular language. The phone recognizer extracts phonotactic attributes from the speech input to characterize a language. The multiple parallel subsystems are devised to capture the phonetic diversification available in the speech input. Alternatively, this paper investigates a new approach for building a PPRLM system that aims at improving the acoustic diversification among its parallel subsystems by using multiple acoustic models. These acoustic models are trained on the same speech data with the same phone set but using different model structures and training paradigms. We examine the use of various structured precision (inverse covariance) matrix modeling techniques as well as the maximum likelihood and maximum mutual information training paradigms to produce complementary acoustic models. The results show that acoustic diversification, which requires only one set of phonetically transcribed speech data, yields similar performance improvements compared to phonetic diversification. In addition, further improvements were obtained by combining both diversification factors. The best performing system reported in this paper combined phonetic and acoustic diversifications to achieve EERs of 4.71% and 8.61% on the 2003 and 2005 NIST LRE sets, respectively, compared to 5.77% and 9.94% using phonetic diversification alone. [ABSTRACT FROM AUTHOR]
Copyright of IEEE Transactions on Audio, Speech & Language Processing is the property of IEEE and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Database: Engineering Source
FullText Text:
  Availability: 0
Header DbId: egs
DbLabel: Engineering Source
An: 32989467
AccessLevel: 6
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: On Acoustic Diversification Front-End for Spoken Language Identification.
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Khe+Chai+Sim%22">Khe Chai Sim</searchLink><relatesTo>1,2</relatesTo><i> kcsim@i2r.a-star.edu.sg</i><br /><searchLink fieldCode="AR" term="%22Haizhou+Li%22">Haizhou Li</searchLink><relatesTo>2,3</relatesTo><i> hli@i2r.a-star.edu.sg</i>
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="JN" term="%22IEEE+Transactions+on+Audio%2C+Speech+%26+Language+Processing%22">IEEE Transactions on Audio, Speech & Language Processing</searchLink>. Jul2008, Vol. 16 Issue 5, p1029-1037. 9p. 2 Diagrams, 6 Charts.
– Name: Subject
  Label: Subjects
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Acoustic+models%22">Acoustic models</searchLink><br /><searchLink fieldCode="DE" term="%22Acoustical+engineering%22">Acoustical engineering</searchLink><br /><searchLink fieldCode="DE" term="%22Telephone+systems%22">Telephone systems</searchLink><br /><searchLink fieldCode="DE" term="%22Telecommunication+systems%22">Telecommunication systems</searchLink><br /><searchLink fieldCode="DE" term="%22Phonotactics%22">Phonotactics</searchLink><br /><searchLink fieldCode="DE" term="%22Language+%26+languages%22">Language & languages</searchLink>
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: The parallel phone recognition followed by language model (PPRLM) architecture represents one of the state-of-the-art spoken language identification systems. A PPRLM system comprises multiple parallel subsystems, where each subsystem employs a phone recognizer with a different phone set for a particular language. The phone recognizer extracts phonotactic attributes from the speech input to characterize a language. The multiple parallel subsystems are devised to capture the phonetic diversification available in the speech input. Alternatively, this paper investigates a new approach for building a PPRLM system that aims at improving the acoustic diversification among its parallel subsystems by using multiple acoustic models. These acoustic models are trained on the same speech data with the same phone set but using different model structures and training paradigms. We examine the use of various structured precision (inverse covariance) matrix modeling techniques as well as the maximum likelihood and maximum mutual information training paradigms to produce complementary acoustic models. The results show that acoustic diversification, which requires only one set of phonetically transcribed speech data, yields similar performance improvements compared to phonetic diversification. In addition, further improvements were obtained by combining both diversification factors. The best performing system reported in this paper combined phonetic and acoustic diversifications to achieve EERs of 4.71% and 8.61% on the 2003 and 2005 NIST LRE sets, respectively, compared to 5.77% and 9.94% using phonetic diversification alone. [ABSTRACT FROM AUTHOR]
– Name: AbstractSuppliedCopyright
  Label:
  Group: Ab
  Data: <i>Copyright of IEEE Transactions on Audio, Speech & Language Processing is the property of IEEE and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.)
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=egs&AN=32989467
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.1109/TASL.2008.924150
    Languages:
      – Code: eng
        Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 9
        StartPage: 1029
    Subjects:
      – SubjectFull: Acoustic models
        Type: general
      – SubjectFull: Acoustical engineering
        Type: general
      – SubjectFull: Telephone systems
        Type: general
      – SubjectFull: Telecommunication systems
        Type: general
      – SubjectFull: Phonotactics
        Type: general
      – SubjectFull: Language & languages
        Type: general
    Titles:
      – TitleFull: On Acoustic Diversification Front-End for Spoken Language Identification.
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Khe Chai Sim
      – PersonEntity:
          Name:
            NameFull: Haizhou Li
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 07
              Text: Jul2008
              Type: published
              Y: 2008
          Identifiers:
            – Type: issn-print
              Value: 15587916
          Numbering:
            – Type: volume
              Value: 16
            – Type: issue
              Value: 5
          Titles:
            – TitleFull: IEEE Transactions on Audio, Speech & Language Processing
              Type: main
ResultId 1