Minimum Phone Error Training of Precision Matrix Models.

Saved in:
Bibliographic Details
Title: Minimum Phone Error Training of Precision Matrix Models.
Authors: Khe Chai Sim1,2 kcs23@eng.cam.ac.uk, Gales, Mark J. F.2,3 mjfg@eng.cam.ac.uk
Source: IEEE Transactions on Audio, Speech & Language Processing. May2006, Vol. 14 Issue 3, p882-889. 8p. 4 Charts.
Subjects: Speech processing systems, Speech perception, Gaussian processes, Density functionals, Sound recording & reproducing, Speech
Abstract: Gaussian mixture models (GMMs) are commonly used as the output density function for large-vocabulary continuous speech recognition (LVCSR) systems. A standard problem when using multivariate GMMs to classify data is how to accurately represent the correlations in the feature vector. Full covarianee matrices yield a good model, but dramatically increase the number of model parameters. Hence, diagonal covariance matrices are commonly used. Structured precision matrix approximations provide an alternative, flexible, and compact representation. Schemes in this category include the extended maximum likelihood linear transform and subspace for precision and mean models. This paper examines how these precision matrix models can be discriminatively trained and used on state-of-the-art speech recognition tasks. In particular, the use of the minimum phone error criterion is investigated. Implementation issues associated with building LVCSR systems are also addressed. These models are evaluated and compared using large vocabulary continuous telephone speech and broadcast news English tasks. [ABSTRACT FROM AUTHOR]
Copyright of IEEE Transactions on Audio, Speech & Language Processing is the property of IEEE and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Database: Engineering Source
FullText Text:
  Availability: 0
Header DbId: egs
DbLabel: Engineering Source
An: 23119635
AccessLevel: 6
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: Minimum Phone Error Training of Precision Matrix Models.
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Khe+Chai+Sim%22">Khe Chai Sim</searchLink><relatesTo>1,2</relatesTo><i> kcs23@eng.cam.ac.uk</i><br /><searchLink fieldCode="AR" term="%22Gales%2C+Mark+J%2E+F%2E%22">Gales, Mark J. F.</searchLink><relatesTo>2,3</relatesTo><i> mjfg@eng.cam.ac.uk</i>
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="JN" term="%22IEEE+Transactions+on+Audio%2C+Speech+%26+Language+Processing%22">IEEE Transactions on Audio, Speech & Language Processing</searchLink>. May2006, Vol. 14 Issue 3, p882-889. 8p. 4 Charts.
– Name: Subject
  Label: Subjects
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Speech+processing+systems%22">Speech processing systems</searchLink><br /><searchLink fieldCode="DE" term="%22Speech+perception%22">Speech perception</searchLink><br /><searchLink fieldCode="DE" term="%22Gaussian+processes%22">Gaussian processes</searchLink><br /><searchLink fieldCode="DE" term="%22Density+functionals%22">Density functionals</searchLink><br /><searchLink fieldCode="DE" term="%22Sound+recording+%26+reproducing%22">Sound recording & reproducing</searchLink><br /><searchLink fieldCode="DE" term="%22Speech%22">Speech</searchLink>
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: Gaussian mixture models (GMMs) are commonly used as the output density function for large-vocabulary continuous speech recognition (LVCSR) systems. A standard problem when using multivariate GMMs to classify data is how to accurately represent the correlations in the feature vector. Full covarianee matrices yield a good model, but dramatically increase the number of model parameters. Hence, diagonal covariance matrices are commonly used. Structured precision matrix approximations provide an alternative, flexible, and compact representation. Schemes in this category include the extended maximum likelihood linear transform and subspace for precision and mean models. This paper examines how these precision matrix models can be discriminatively trained and used on state-of-the-art speech recognition tasks. In particular, the use of the minimum phone error criterion is investigated. Implementation issues associated with building LVCSR systems are also addressed. These models are evaluated and compared using large vocabulary continuous telephone speech and broadcast news English tasks. [ABSTRACT FROM AUTHOR]
– Name: AbstractSuppliedCopyright
  Label:
  Group: Ab
  Data: <i>Copyright of IEEE Transactions on Audio, Speech & Language Processing is the property of IEEE and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.)
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=egs&AN=23119635
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.1109/TSA.2005.858062
    Languages:
      – Code: eng
        Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 8
        StartPage: 882
    Subjects:
      – SubjectFull: Speech processing systems
        Type: general
      – SubjectFull: Speech perception
        Type: general
      – SubjectFull: Gaussian processes
        Type: general
      – SubjectFull: Density functionals
        Type: general
      – SubjectFull: Sound recording & reproducing
        Type: general
      – SubjectFull: Speech
        Type: general
    Titles:
      – TitleFull: Minimum Phone Error Training of Precision Matrix Models.
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Khe Chai Sim
      – PersonEntity:
          Name:
            NameFull: Gales, Mark J. F.
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 05
              Text: May2006
              Type: published
              Y: 2006
          Identifiers:
            – Type: issn-print
              Value: 15587916
          Numbering:
            – Type: volume
              Value: 14
            – Type: issue
              Value: 3
          Titles:
            – TitleFull: IEEE Transactions on Audio, Speech & Language Processing
              Type: main
ResultId 1