On Acoustic Diversification Front-End for Spoken Language Identification.
Saved in:
| Title: | On Acoustic Diversification Front-End for Spoken Language Identification. |
|---|---|
| Authors: | Khe Chai Sim1,2 kcsim@i2r.a-star.edu.sg, Haizhou Li2,3 hli@i2r.a-star.edu.sg |
| Source: | IEEE Transactions on Audio, Speech & Language Processing. Jul2008, Vol. 16 Issue 5, p1029-1037. 9p. 2 Diagrams, 6 Charts. |
| Subjects: | Acoustic models, Acoustical engineering, Telephone systems, Telecommunication systems, Phonotactics, Language & languages |
| Abstract: | The parallel phone recognition followed by language model (PPRLM) architecture represents one of the state-of-the-art spoken language identification systems. A PPRLM system comprises multiple parallel subsystems, where each subsystem employs a phone recognizer with a different phone set for a particular language. The phone recognizer extracts phonotactic attributes from the speech input to characterize a language. The multiple parallel subsystems are devised to capture the phonetic diversification available in the speech input. Alternatively, this paper investigates a new approach for building a PPRLM system that aims at improving the acoustic diversification among its parallel subsystems by using multiple acoustic models. These acoustic models are trained on the same speech data with the same phone set but using different model structures and training paradigms. We examine the use of various structured precision (inverse covariance) matrix modeling techniques as well as the maximum likelihood and maximum mutual information training paradigms to produce complementary acoustic models. The results show that acoustic diversification, which requires only one set of phonetically transcribed speech data, yields similar performance improvements compared to phonetic diversification. In addition, further improvements were obtained by combining both diversification factors. The best performing system reported in this paper combined phonetic and acoustic diversifications to achieve EERs of 4.71% and 8.61% on the 2003 and 2005 NIST LRE sets, respectively, compared to 5.77% and 9.94% using phonetic diversification alone. [ABSTRACT FROM AUTHOR] |
| Copyright of IEEE Transactions on Audio, Speech & Language Processing is the property of IEEE and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.) | |
| Database: | Engineering Source |
| FullText | Text: Availability: 0 |
|---|---|
| Header | DbId: egs DbLabel: Engineering Source An: 32989467 AccessLevel: 6 PubType: Academic Journal PubTypeId: academicJournal PreciseRelevancyScore: 0 |
| IllustrationInfo | |
| Items | – Name: Title Label: Title Group: Ti Data: On Acoustic Diversification Front-End for Spoken Language Identification. – Name: Author Label: Authors Group: Au Data: <searchLink fieldCode="AR" term="%22Khe+Chai+Sim%22">Khe Chai Sim</searchLink><relatesTo>1,2</relatesTo><i> kcsim@i2r.a-star.edu.sg</i><br /><searchLink fieldCode="AR" term="%22Haizhou+Li%22">Haizhou Li</searchLink><relatesTo>2,3</relatesTo><i> hli@i2r.a-star.edu.sg</i> – Name: TitleSource Label: Source Group: Src Data: <searchLink fieldCode="JN" term="%22IEEE+Transactions+on+Audio%2C+Speech+%26+Language+Processing%22">IEEE Transactions on Audio, Speech & Language Processing</searchLink>. Jul2008, Vol. 16 Issue 5, p1029-1037. 9p. 2 Diagrams, 6 Charts. – Name: Subject Label: Subjects Group: Su Data: <searchLink fieldCode="DE" term="%22Acoustic+models%22">Acoustic models</searchLink><br /><searchLink fieldCode="DE" term="%22Acoustical+engineering%22">Acoustical engineering</searchLink><br /><searchLink fieldCode="DE" term="%22Telephone+systems%22">Telephone systems</searchLink><br /><searchLink fieldCode="DE" term="%22Telecommunication+systems%22">Telecommunication systems</searchLink><br /><searchLink fieldCode="DE" term="%22Phonotactics%22">Phonotactics</searchLink><br /><searchLink fieldCode="DE" term="%22Language+%26+languages%22">Language & languages</searchLink> – Name: Abstract Label: Abstract Group: Ab Data: The parallel phone recognition followed by language model (PPRLM) architecture represents one of the state-of-the-art spoken language identification systems. A PPRLM system comprises multiple parallel subsystems, where each subsystem employs a phone recognizer with a different phone set for a particular language. The phone recognizer extracts phonotactic attributes from the speech input to characterize a language. The multiple parallel subsystems are devised to capture the phonetic diversification available in the speech input. Alternatively, this paper investigates a new approach for building a PPRLM system that aims at improving the acoustic diversification among its parallel subsystems by using multiple acoustic models. These acoustic models are trained on the same speech data with the same phone set but using different model structures and training paradigms. We examine the use of various structured precision (inverse covariance) matrix modeling techniques as well as the maximum likelihood and maximum mutual information training paradigms to produce complementary acoustic models. The results show that acoustic diversification, which requires only one set of phonetically transcribed speech data, yields similar performance improvements compared to phonetic diversification. In addition, further improvements were obtained by combining both diversification factors. The best performing system reported in this paper combined phonetic and acoustic diversifications to achieve EERs of 4.71% and 8.61% on the 2003 and 2005 NIST LRE sets, respectively, compared to 5.77% and 9.94% using phonetic diversification alone. [ABSTRACT FROM AUTHOR] – Name: AbstractSuppliedCopyright Label: Group: Ab Data: <i>Copyright of IEEE Transactions on Audio, Speech & Language Processing is the property of IEEE and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.) |
| PLink | https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=egs&AN=32989467 |
| RecordInfo | BibRecord: BibEntity: Identifiers: – Type: doi Value: 10.1109/TASL.2008.924150 Languages: – Code: eng Text: English PhysicalDescription: Pagination: PageCount: 9 StartPage: 1029 Subjects: – SubjectFull: Acoustic models Type: general – SubjectFull: Acoustical engineering Type: general – SubjectFull: Telephone systems Type: general – SubjectFull: Telecommunication systems Type: general – SubjectFull: Phonotactics Type: general – SubjectFull: Language & languages Type: general Titles: – TitleFull: On Acoustic Diversification Front-End for Spoken Language Identification. Type: main BibRelationships: HasContributorRelationships: – PersonEntity: Name: NameFull: Khe Chai Sim – PersonEntity: Name: NameFull: Haizhou Li IsPartOfRelationships: – BibEntity: Dates: – D: 01 M: 07 Text: Jul2008 Type: published Y: 2008 Identifiers: – Type: issn-print Value: 15587916 Numbering: – Type: volume Value: 16 – Type: issue Value: 5 Titles: – TitleFull: IEEE Transactions on Audio, Speech & Language Processing Type: main |
| ResultId | 1 |